Vertically integrated neural network computing systems and associated systems and methods

By vertically stacking volatile and flash memory dies in high-bandwidth memory (HBM) devices and using through silicon holes (TSVs) to achieve high-bandwidth communication, the problems of power consumption and heat generation in neural network computing of HBM devices are solved, and efficient neural network computing operations are achieved.

CN119990206APending Publication Date: 2025-05-13MICRON TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411580793.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-10
Filing Date
2024-11-07
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing high-bandwidth memory (HBM) devices have problems with power consumption and heat generation when implementing neural network computing operations, especially during machine learning computing, frequent data transmission leads to higher power demand and heat generation.

Method used

A functional high bandwidth memory (HBM) device is designed, which includes a controller die, a volatile memory die and a flash memory die. High bandwidth communication is achieved through vertical stacking structures and through silicon through-silicon holes (TSVs), and the weights and inputs of neural network computing operations are programmed on the flash memory die to complete neural network computing operations.

Benefits of technology

By completing neural network computing operations in the HBM device, power consumption and heat generation are significantly reduced, computing efficiency and system energy efficiency performance are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990206A_ABST
    Figure CN119990206A_ABST
Patent Text Reader

Abstract

The invention relates to a vertically integrated neural network computing system and associated systems and methods. A system-in-package (SiP) with a functional high bandwidth memory (HBM) device, and associated systems and methods, are disclosed herein. In some embodiments, the functional HBM device may include a controller die, one or more volatile memory dies, a flash memory die, and an HBM bus communicatively coupled to each of the controller, volatile memory, and flash memory dies. The flash memory die may include one or more word lines and a plurality of bit lines each having a plurality of programmable memory cells. Each of the bit lines is coupled from each of the one or more word lines to a corresponding programmable memory cell. During operation, the controller die is configured to control the volatile memory die and the flash memory die through a shared bus between the volatile memory die and the flash memory die to perform one or more neural network computing operations within the functional HBM device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology generally relates to vertically stacked semiconductor devices, and more specifically to stacking volatile and functional dies in semiconductor packages for neural network processing. Background Art

[0002] Microelectronic devices, such as memory devices, microprocessors, and other electronic devices, typically include one or more semiconductor dies mounted to a substrate and enclosed in a protective cover layer. The semiconductor die includes functional features, such as memory cells, processor circuits, imager devices, interconnect circuitry, etc. In order to meet the demand for ever-decreasing size, wafers, individual semiconductor dies, and / or active components are typically manufactured in batches, singulated, etc., and then stacked on a supporting substrate, such as a printed circuit board (PCB) or other suitable substrate. The stacked dies can then be coupled to a supporting substrate (sometimes also referred to as a packaging substrate) by bonding wires in shingled stacked dies (e.g., stacked dies with an offset for each die) and / or by through-substrate vias (TSVs) between the die and the supporting substrate. Summary of the invention

[0003] In one aspect, the present disclosure provides a method of operating a high bandwidth memory (HBM) device, the method comprising: reading a plurality of weights for a trained neural network from one or more memory dies in the HBM device; programming a threshold voltage of each of a plurality of memory cells on a flash memory die in the HBM device based on the plurality of weights of the trained neural network; reading a plurality of inputs for a neural network computing operation from the one or more memory dies in the HBM device; and performing the neural network computing operation using the plurality of inputs.

[0004] In another aspect, the present disclosure provides a functional high bandwidth memory (HBM) device comprising: a controller die; one or more volatile memory dies carried by the controller die; a flash memory die carried by the one or more volatile memory dies, the flash memory die comprising: one or more rows each having a word line and two or more programmable memory cells coupled to the word line; and two or more bit lines, wherein each of the two or more bit lines is coupled to a corresponding programmable memory cell from each of the one or more rows; and a shared bus electrically coupled to the controller die, the one or more volatile memory dies, and each of the flash memory dies, wherein the controller die is configured to control the one or more volatile memory dies and the flash memory die through the shared bus to implement one or more neural network computing operations within the functional HBM device.

[0005] On the other hand, the present disclosure provides a system-in-package (HBM) device, comprising: a base substrate; a processing unit carried by the base substrate; and a functional high-bandwidth memory HBM device carried by the base substrate and electrically coupled to the processing unit through a (SiP) bus, wherein the functional HBM device comprises: a controller die; one or more memory dies; a neural network computing die; and a shared bus electrically coupled to each of the controller die, the one or more memory dies, and the neural network computing die, wherein the controller die is configured to control the one or more memory dies and the neural network computing die through the shared bus to implement one or more neural network computing operations within the functional HBM device. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 is a schematic diagram illustrating an environment incorporating a high bandwidth memory architecture.

[0007] Figure 2 is a schematic diagram illustrating an environment in which a neural network computing device is incorporated into a high bandwidth memory architecture according to some embodiments of the present technology.

[0008] Figure 3 is a partial cross-sectional schematic diagram of a system-in-package having a functional high-bandwidth memory device configured in accordance with some embodiments of the present technology.

[0009] Figure 4 is a partially exploded schematic diagram of a functional high bandwidth memory device configured in accordance with some embodiments of the present technology.

[0010] Figure 5A is a top plan schematic diagram of components of a functional high bandwidth memory device configured in accordance with some embodiments of the present technology.

[0011] Figure 5B is a schematic diagram for routing signals through a functional high bandwidth memory device according to some embodiments of the present technology.

[0012] Figure 6 is a schematic diagram of a neural network computing operation according to some embodiments of the present technology.

[0013] Figure 7 is a partial circuit diagram of a flash memory cell configured according to some embodiments of the present technology.

[0014] Figure 8 is a partial circuit diagram of a cell on a flash memory device configured according to some embodiments of the present technology.

[0015] Fig. 9is a flow chart of a process for implementing neural network computing operations within a functional high bandwidth memory device according to some embodiments of the present technology.

[0016] Fig.10 is a flow chart of a process for programming a flash memory device for neural network computing operations in accordance with some embodiments of the present technology.

[0017] Fig.11 is a partial circuit diagram of a cell on a flash memory device configured according to a further embodiment of the present technology.

[0018] Fig.12 is a partially exploded schematic diagram of a functional high bandwidth memory device configured in accordance with further embodiments of the present technology.

[0019] The drawings are not necessarily drawn to scale. In addition, it will be understood that several drawings have been drawn schematically and / or partially schematically. Similarly, for the purpose of discussing some embodiments of the present technology, some components and / or operations may be separated into different blocks or combined into a single block. In addition, although the present technology may be subjected to various modifications and alternative forms, specific embodiments have been shown by examples in the drawings and described in detail below. However, it is not intended to limit the technology to the specific embodiments described. DETAILED DESCRIPTION

[0020] High data reliability, high-speed memory access, lower power consumption, and reduced chip size are required features of semiconductor memory. In recent years, three-dimensional (3D) memory devices have been introduced. Some 3D memory devices are formed by vertically stacking memory dies and interconnecting the dies using through-silicon (or through-substrate) vias (TSVs). The benefits of 3D memory devices include: shorter interconnects (which reduce signal delays and power consumption); a larger number of vertical vias between layers (which allow wide bandwidth buses between functional blocks in different layers, such as memory dies); and a relatively small footprint. Therefore, 3D memory devices contribute to higher memory access speeds, lower power consumption, and chip size reduction. Example 3D memory devices include hybrid memory cubes (HMCs) and high-bandwidth memories (HBMs). For example, HBM is a type of memory that includes a vertical stack of dynamic random access memory (DRAM) dies and an interface die (which, for example, provides an interface between the DRAM die of the HBM device and a host device).

[0021] In a system-in-package (SiP) configuration, the HBM device can be integrated with a host device (e.g., a graphics processing unit (GPU) and / or a computer processing unit (CPU)) using a base substrate (e.g., a silicon interposer, a substrate of organic material, a substrate of inorganic material, and / or any other suitable material that provides interconnection between the GPU / CPU and the HBM device and / or provides mechanical support for the components of the SiP device), through which the HBM device and the host communicate. Because the traffic between the HBM device and the host device resides within the SiP (e.g., using signals routed through the silicon interposer), a higher bandwidth can be achieved between the HBM device and the host device than in conventional systems. In other words, the TSVs that interconnect the DRAM dies within the HBM device and the silicon interposer that integrates the HBM device and the host device enable routing of a greater number of signals (e.g., a wider data bus) than is typically found between a packaged memory device and a host device (e.g., through a printed circuit board (PCB)). The high-bandwidth interface within the SiP enables large amounts of data to be quickly moved between the host device (e.g., GPU / CPU) and the HBM device during operation. For example, the high bandwidth channel may be approximately 1000 gigabytes per second (GB / s, sometimes also referred to as gigabits (Gb)). It should be appreciated that such high bandwidth data transfer between the GPU / CPU and the memory of the HBM device may be advantageous in various high performance computing applications, such as video rendering, high resolution graphics applications, artificial intelligence and / or machine learning (AI / ML) computing systems and other complex computing systems and / or various other computing applications.

[0022] Figure 1 is a schematic diagram illustrating an environment 100 incorporating a high bandwidth memory architecture. Figure 1 As shown in FIG. 1 , the environment 100 includes a SiP device 110 having one or more processing devices 120 (in Figure 1 , sometimes referred to herein as one or more "hosts"), and one or more high bandwidth memory (HBM) devices 130 (described in Figure 11 ), which is integrated with a silicon interposer 112 (or any other suitable base substrate). The environment 100 additionally includes a storage device 140 coupled to the SiP device 110. The processing device 120 may include one or more CPUs and / or one or more GPUs, referred to as CPU / GPU 122, each of which may include registers 124 and a first level cache 126. The first level cache 126 (also referred to herein as an "L1 cache") is communicatively coupled to a second level cache 128 (also referred to herein as an "L2 cache") via a first communication path 152. In the illustrated embodiment, the L2 cache 128 is incorporated into the processing device 120. However, it should be understood that the L2 cache 128 may be integrated into the SiP device 110 separately from the processing device 120. Purely by way of example, a processing device 120 may be carried by a base substrate adjacent to an L2 cache 128 (e.g., an interposer which is itself carried by a package substrate) and communicate with the L2 cache 128 via one or more signal lines therein (or other suitable signal routing lines). The L2 cache 128 may be shared by one or more processing devices 120 (and the CPU / GPU 122 therein). During operation of the SiP device 110, the CPU / GPU 122 may use registers 124 and the L1 cache 126 to complete processing operations, and whenever a cache miss occurs in the L1 cache 126, attempt to retrieve data from the larger L2 cache 128. Thus, the multiple levels of cache may help speed up the average time it takes a processing device 120 to access data, thereby speeding up the overall processing rate.

[0023] like Figure 1 As further described in the figure, the L2 cache 128 is communicatively coupled to the HBM device 130 via a second communication channel 154. As described, the processing device 120 (and the L2 cache 128 therein) and the HBM device 130 are carried by and electrically coupled to the silicon interposer 112 (e.g., integrated by the silicon interposer 112). The second communication channel 154 is provided by the silicon interposer 112 (e.g., the silicon interposer includes and routes interface signals that form the second communication channel, such as through one or more redistribution layers (RDLs)). Figure 1140 , the L2 cache 128 is also communicatively coupled to the memory device 140 via a third communication channel 156. As illustrated, the memory device 140 is outside the SiP device 110 and utilizes signal routing components that are not contained within the silicon interposer 112 (e.g., between the packaged SiP device 110 and the packaged memory device 140). For example, the third communication channel 156 can be a peripheral bus, such as a peripheral component interconnect express (PCIe) bus, for connecting components on a motherboard or PCB. Thus, during operation of the SiP device 110, the processing device 120 can read data from and / or write data to the HBM device 130 and / or the memory device 140 via the L2 cache 128.

[0024] In the illustrated environment 100, the HBM device 130 includes one or more stacked volatile memory dies 132 (e.g., DRAM dies, Figure 1 1). As explained above, the HBM device 130 may be positioned on the silicon interposer 112, and the processing device 120 may also be positioned on the silicon interposer 112. Thus, the second communication channel 154 may provide a high bandwidth (e.g., approximately 1000 GB / s) channel through the silicon interposer 112. Furthermore, as explained above, each HBM device 130 may provide a high bandwidth channel (not shown) between the volatile memory dies 132 therein. Thus, data may be communicated between the processing device 120 and the HBM device 130 (and the volatile memory dies 132 therein) at high speeds, which may be advantageous for data-intensive processing operations.

[0025] Although the HBM device 130 of the SiP device 110 provides relatively high bandwidth communications, there are certain disadvantages to its integration on the silicon interposer 112. For example, communicating data via the second communication channel 154 may require a relatively large amount of power. During machine learning calculations, such as applying a trained neural network to a data set, the calculations may require several rounds of back-and-forth communication (e.g., loading each set of inputs from the HBM device 130 to the processing device 120). As a result, the calculations may consume a large amount of power and / or generate a large amount of heat.

[0026] Disclosed herein are HBM devices and associated systems and methods that address the shortcomings discussed above. For example, an HBM device may include an interface die, one or more volatile memory dies (e.g., DRAM dies), and one or more computing dies (e.g., programmable NOR flash dies, programmable NAND dies, and / or the like). An HBM device may also include one or more TSVs that electrically couple an interface die, one or more volatile memory dies, and one or more computing dies to establish a communication path therebetween. As described herein, the TSVs may provide a wide communication path (e.g., approximately 1024 I / Os) between the interface die, volatile memory die, and computing die of the HBM device, thereby enabling high bandwidth therebetween. In other words, the disclosed HBM device combines both volatile memory and computing die (referred to herein as "functional HBM devices") with high bandwidth communication between the dies of the functional HBM device. However, because the communication channels between the dies are significantly shorter, communicating data within a functional HBM device may require significantly less power than communicating data to a separate processing device. Furthermore, as explained herein, a compute die may be programmed to perform neural network computations, thereby allowing a functional HBM device to implement neural network computation operations entirely within the functional HBM device. Thus, a functional HBM device may significantly reduce the power required (and the heat generated thereby) to implement neural network computation operations.

[0027] Additional details about functional HBM devices and associated systems and methods are set forth below. For ease of reference, semiconductor packages (and components thereof) are sometimes described herein with reference to the front and back, top and bottom, up and down, upward and downward, and / or horizontal planes, xy planes, vertical or z directions relative to the spatial orientation of the embodiments shown in the figures. However, it should be understood that the semiconductor assembly (and components thereof) can be moved to and used in different spatial orientations without changing the structure and / or function of the disclosed embodiments of the present technology. In addition, signals within semiconductor packages (and components thereof) are sometimes described herein with reference to downstream and upstream, forward and backward, and / or read and write relative to the embodiments shown in the figures. However, it should be understood that signal flows can be described in various other terms without changing the structure and / or function of the disclosed embodiments of the present technology.

[0028] Furthermore, although the memory device architecture disclosed herein is primarily discussed in the context of implementing neural network computing operations entirely within HBM devices, those skilled in the art will appreciate that the scope of the present technology is not limited thereto. For example, the systems and methods disclosed herein may also be deployed to implement various other computing operations entirely within HBM devices (e.g., for various other machine learning applications).

[0029] Figure 22 is a schematic diagram illustrating an environment 200 incorporating an HBM architecture according to some embodiments of the present technology. Similar to the environment 100 discussed above, the environment 200 includes one or more processing devices 220 ( Figure 2 ) and one or more storage devices 240 ( Figure 2 1 ) is a SiP device 210. However, Figure 1 In contrast to the SiP device 110 described in Figure 2 The embodiment of the SiP device 210 illustrated in FIG. 1 includes one or more functional HBM devices 230 ( Figure 2 2 (illustrated in one of the figures), as further described below. The processing device 220 and the functional HBM device 230 are each integrated on an interposer 212 (e.g., a silicon interposer, another organic interposer, an inorganic interposer, and / or any other suitable base substrate) that may include one or more signal routing lines. The processing device 220 is driven by a CPU / GPU 222 that includes registers 224 and an L1 cache 226. The L1 cache 226 is communicatively coupled to the L2 cache 228 via a first communication channel 252. In addition, the L2 cache 228 is communicatively coupled to the functional HBM device 230 through a second communication channel 254, and is communicatively coupled to the storage device 240 through a third communication channel 256. Further, the second communication channel 254 may have a relatively high bandwidth (e.g., approximately 1000 GB / s), while the third communication channel 256 may have a relatively low bandwidth (e.g., approximately 8 GB / s).

[0030] exist Figure 2In the embodiment illustrated in FIG. 1 , the functional HBM devices 230 each include a stack of one or more volatile memory dies 232 (e.g., DRAM dies) and one or more functional dies 234 (e.g., programmable NOR flash dies, programmable NAND dies, or other suitable functional dies for neural network computing). That is, the one or more volatile memory dies 232 and the one or more functional dies 234 may be stacked vertically in the functional HBM device 230. Within the stack, the memory dies 232 and the functional dies 234 may be communicatively coupled via a fourth communication channel 258 (e.g., an HBM bus), which may have a relatively high bandwidth (e.g., about 1000 GB / s). Thus, data may be communicated quickly and efficiently within the functional HBM device 230. As discussed in more detail below, the functional die 234 may be used to complete one or more complex computations within the functional HBM device 230, such as various neural network computations, machine learning computations, and / or other suitable computations. As also discussed in more detail below, each of the volatile memory and functional dies 232, 234 may be coupled to a high-bandwidth bus in the functional HBM device 230. For example, each of the functional HBM devices 230 may include a plurality of TSVs interconnecting the volatile memory die 232 and the functional die 234 within the functional HBM device 230, thereby providing a high-bandwidth bus.

[0031] The combination of volatile memory and functional die within each functional HBM device 230 can provide certain advantages. For example, one or more volatile memory dies 232 can store weights associated with a trained neural network and / or values ​​associated with inputs to the neural network. In this example, as discussed in more detail below, the functional HBM device 230 can program weights into one or more functional dies 234 via TSVs, then write inputs into one or more functional dies 234 and read the results to complete one or more neural network computing operations. That is, the functional HBM device 230 can complete any number of neural network computing operations without communicating data to the processing device 220 via the second communication channel 254. Communicating data via TSVs can be faster than communicating data via the second communication channel 254 and / or can require less energy than communicating data via the second communication channel 254. Therefore, by completing neural network computing operations internally, the functional HBM device can speed up neural network computing operations and / or reduce the power requirements of the operating environment 200. Additionally, reduced power consumption may reduce heat generated by environment 200 which may have detrimental effects (eg, data loss).

[0032] Environment 200 may be configured to perform any of a wide variety of suitable computing, processing, storage, sensing, imaging, and / or other functions for various electronic devices. For example, representative examples of systems that include environment 200 (and / or its components, such as SiP device 210) include, but are not limited to, computers and / or other data processors, such as desktop computers, laptop computers, Internet appliances, handheld devices (e.g., palmtop computers, wearable computers, cellular or mobile phones, automotive electronics, personal digital assistants, music players, etc.), tablet computers, multi-processor systems, processor-based or programmable consumer electronics, network computers, and minicomputers. Additional representative examples of systems that include environment 200 (and / or its components) include lights, cameras, vehicles, and the like. With respect to these and other examples, environment 200 may be housed in a single unit or distributed over multiple interconnected units, such as via a communications network, in various locations on a motherboard, and the like. In addition, components of environment 200 (and / or any of its components) may be coupled to various other local and / or remote memory storage devices, processing devices, computer-readable storage media, and the like. Reference is made below to Figures 3 to 12 Additional details are set forth regarding the architecture of the environment 200 , SiP device 210 , functional HBM device 230 , and their operational processes.

[0033] Figure 3 FIG. 2 is a schematic diagram of a partial cross-section of a SiP device 300 having a functional HBM device 330 configured in accordance with some embodiments of the present technology. Figure 3 , the SiP device 300 includes a base substrate 310 (e.g., a silicon interposer, another suitable organic substrate, an inorganic substrate, and / or any other suitable material) and a CPU / GPU 320 and a functional HBM device 330 each integrated with an upper surface 312 of the base substrate 310. In the illustrated embodiment, the CPU / GPU 320 and associated components (e.g., registers, L1 cache, and the like) are illustrated as a single package, and the functional HBM device 330 includes a semiconductor die stack. The semiconductor die stack in the functional HBM device 330 includes an interface die 332, one or more volatile memory dies 334 ( Figure 3 ) and one or more functional dies 336 ( Figure 3 The CPU / GPU 320 is coupled to the functional HBM device 330 via a high bandwidth bus 340, which includes one or more routing lines 344 ( Figure 3 In various embodiments, routing line 344 may include one or more metallization layers formed in one or more RDL layers of base substrate 310 and / or one or more vias interconnecting metallization layers and / or traces. Figure 3 Not illustrated, but to be understood, the CPU / GPU 320 and the functional HBM device 330 may each be coupled to the routing line 344 via solder structures (eg, solder balls), metal-to-metal bonds, and / or any other suitable conductive bonds.

[0034] As discussed in more detail below, the high-bandwidth bus 340 may also include a plurality of through-substrate vias (TSVs) 342 extending from the interface die 332 through the volatile memory die 334 to the functional die 336. The TSVs 342 allow each die to communicate data within the functional HBM device 330 (e.g., between the volatile memory die 334 (e.g., DRAM die) and the functional die 336 (e.g., programmable NOR die)) at a relatively high rate (e.g., about 1000 GB / s or more). Additionally, the functional HBM device 330 may include one or more signal routing lines 341 (e.g., additional TSVs extending through the interface die 332) that couple the interface die 332 and / or the TSVs 342 to the routing lines 344. The signal routing lines 341, TSVs 342, and routing lines 344, in turn, allow the functional HBM device 330 and the dies in the CPU / GPU 320 to communicate data at high bandwidth.

[0035] Figure 4 is a partially exploded schematic diagram of a functional HBM device 400 configured according to some embodiments of the present technology. For example, the functional HBM device 400 can be used as the above reference Figure 3 The functional HBM device 330 discussed. In the illustrated embodiment, the functional HBM device 400 is an interface die 410, one or more volatile memory dies 420 ( Figure 4 4) and one or more functional dies 430 ( Figure 4 In addition, the functional HBM device 400 includes a shared HBM bus 440 that communicatively couples the interface die 410, the volatile memory die 420, and the functional die 430.

[0036] The interface die 410 may be a component that is connected between the shared HBM bus 440 and an external component (eg, Figure 3The interface die 410 may include one or more active components, such as a static random access memory (SRAM) cache, a memory controller, and / or any other suitable components. The volatile memory die 420 (also sometimes collectively referred to herein as "main memory") may be a DRAM memory die that provides low-latency memory access to the functional HBM device 400. The functional die 430 (sometimes referred to herein as a "programmable memory die," "flash memory die," "compute die," and the like) may provide programmable computing functionality for the functional HBM device 400. In a specific non-limiting example, the functional die may be a NOR flash die that includes one or more word lines, each word line coupled to one or more memory cells having a programmable threshold voltage. Each memory cell is coupled to a corresponding bit line and a common source line (also sometimes referred to herein as a "shared source line"). In some embodiments, at least some (up to all) memory cells coupled to the same word line are also coupled to the same shared source line. In this example, a constant voltage (e.g., a few volts, tens of volts, and / or the like) may be applied to the word line, while an input voltage (e.g., a negative voltage having an absolute value of a few volts, tens of volts, and / or the like) is applied to each bit line of the individual memory cells coupled to the shared word line. Thus, each of the individual memory cells drives a current to the common source line such that the current on the common source line reflects the sum of the currents output by each of the individual memory cells (i.e., the common source line sums or accumulates the currents from the memory cells coupled to it). As discussed in more detail below, this summing function (and / or an activation function applied to the summed current) may be used to implement a neural network computing operation. For example, a programmable threshold voltage at each of the memory cells coupled to the word line may correspond to an individual weight from a trained neural network, while the voltage applied to each bit line corresponds to an individual input value. In this example, the above operation multiplies the weight vector and the input vector and provides the result as an output current on the source line, thereby implementing a neural network calculation operation.

[0037] In the illustrated embodiment, the shared HBM bus 440 includes a plurality of TSVs 442 ( 444B ) extending from the interface die 410 through the volatile memory die 420 to the functional die 430. Figure 44, but any suitable number of TSVs is possible). Each of the TSVs 442 can support independent bidirectional read / write operations to communicate data between dies in the functional HBM device 400 (e.g., between the interface die 410 and the volatile memory die 420, between the functional die 430 and the volatile memory die 420, and / or the like). Because the TSVs 442 establish a shared HBM bus 440 between each of the dies in the functional HBM device 400, the shared HBM bus 440 can reduce (or minimize) the footprint required to establish a high-bandwidth communication route through the functional HBM device 400. Therefore, the shared HBM bus 440 can reduce (or minimize) the overall footprint of the functional HBM device 400.

[0038] Figure 5A is a top plan view schematic diagram of components of a functional HBM device 500 configured in accordance with some embodiments of the present technology. Figure 5A , the functional HBM device 500 is substantially similar to the above reference Figure 4 For example, the functional HBM device 500 includes an interface die 510, one or more volatile memory dies 520 ( Figure 5A ) and one or more functional dies 530 ( Figure 5A ), and a shared HBM bus 540 that communicatively couples the interface die 510, the volatile memory die 520, and the functional die 530.

[0039] Interface die 510 includes one or more read / write components 512 ( Figure 5A In various embodiments, read / write component 512 may couple interface die 510 to external components (e.g., to Figure 3 ), presenting information about the functional HBM device 500 and / or the die of the functional HBM device 500 to external components (e.g., Figure 3The interface die 510 includes a CPU / GPU 320 of the functional die 530, and / or includes memory control functions to control the movement of data between the volatile memory die 520 and / or the functional die 530. In some embodiments, the interface die 510 includes one or more controller components that are coupled to the volatile memory die 520 and the functional die 530 to implement various computing operations within the functional HBM device 500 (e.g., reading weights of neural network processing operations from the volatile memory die 520, programming the functional die 530 with weights, writing one or more inputs from the volatile memory die 520 to the functional die 530, and reading results from the functional die 530). The volatile memory die 520 each includes a memory circuit 522 that can store data in a volatile array (e.g., a line of capacitors and / or transistors). The functional die 530 includes a memory circuit 532 (e.g., a NOR flash memory circuit) that includes a plurality of programmable memory cells. As described herein, the programmable memory cells can be programmed with programmable threshold voltages (e.g., by a write operation).

[0040] like Figure 5A As further described in the above, the shared HBM bus 540 may include a plurality of TSVs 542 ( Figure 5A 32 (illustrated in FIG. 32 ), the TSVs 542 extend between each of the interface die 510, the volatile memory die 520, and the functional die 530. The TSVs 542 can be organized into subgroups (e.g., rows, columns, and / or any other suitable subgroupings) that are selectively coupled to the die in the functional HBM device 500 to simplify signal routing. For example, in Figure 5A In the embodiment illustrated in FIG. 1 , the first volatile memory die 520a may be selectively coupled to the first subgroup 542a of the TSVs 542 (e.g., the rightmost column of the TSVs 542). Therefore, read / write operations to the first volatile memory die 520a must be performed through the first subgroup 542a of the TSVs 542. Similarly, Figure 5A540, the second volatile memory die 520b may be selectively coupled to the second subgroup 542b of the TSVs 542, the third volatile memory die 520c may be selectively coupled to the third subgroup 542c of the TSVs 542, and the fourth volatile memory die 520d may be selectively coupled to the fourth subgroup 542d of the TSVs 542. In the illustrated embodiment, each of the first to fourth subgroups 542a to 542d is completely separate from the other subgroups. Therefore, despite being coupled to the shared HBM bus 540, each of the volatile memory dies 520 is completely individually addressed. However, it should be understood that in some embodiments, the first through fourth subgroupings 542a through 542d may share one or more of the TSVs 542, each of the first through fourth volatile memory dies 520a through 520d may be coupled to the shared subgrouping of the TSVs 542, and / or the TSVs 542 may include a fifth subgrouping coupled to each of the first through fourth volatile memory dies 520a through 520d to allow one or more read / write operations to send data to multiple first through fourth volatile memory dies 520a through 520d at one time (e.g., allowing the second volatile memory die 520b to store a copy of the data of the first volatile memory die 520a).

[0041] exist Figure 5A In the embodiment illustrated in FIG. 5 , the interface die 510 is coupled to each of the TSVs 542. Thus, the interface die 510 can clock and help route the read / write signals to any suitable destination. Similarly, the functional die 530 is coupled to each of the TSVs 542. Thus, the functional die 530 can use any available subgrouping of the TSVs 542 (and / or all of the TSVs 542) to send and / or receive read / write signals.

[0042] Figure 5B According to some embodiments of the present technology, Figure 5A Schematic diagram of the signal routing of the functional HBM device 500. Figure 5B , TSVs 542 are schematically represented by horizontal lines, and connections to TSVs 542 (e.g., by volatile memory die 520 and functional die 530) are illustrated by vertical lines that intersect the horizontal lines. It should be understood that each intersection point may represent a connection to one or more of TSVs 542 (e.g., Figure 5A 542a through 542d, eight of the TSVs 542 illustrated in each of the first through fourth subgroups 542a through 542d, a single TSV, two TSVs, and / or any other suitable number of connections).

[0043] exist Figure 5BIn the embodiment illustrated in FIG. 5 , the volatile memory die 520 is selectively coupled to the first to fourth subgroups 542a to 542d of the TSVs 542, and the interface die 510 is coupled to each of the TSVs 542. For example, the first volatile link V0 (corresponding to Figure 5A The first volatile memory die 520a of the first subgroup 542a is coupled to the first subgroup 542a, the second volatile link V1 (corresponding to the second volatile memory die 520b) is coupled to the second subgroup 542b, and the third volatile link V2 (corresponding to Figure 5A The third volatile memory die 520c of FIG. 5 is coupled to the third subgroup 542c, and the fourth volatile link V3 (corresponding to Figure 5A 542d). In addition, the functional die 530 is coupled to each of the first to fourth subgroups 542a to 542d of the TSVs 542 (e.g., at non-volatile link NV0). Thus, for example, a signal (e.g., a read request) from the interface die 510 may be written to the second subgroup 542b (via interface link I0). Thus, the signal may only be received by the second volatile link V1 (e.g., Figure 5A The second volatile memory die 520b may then write the requested data onto the second subpacket 542b, which may then be received only by the interface die 510 and the functional die 530.

[0044] like Figure 5B As further described in FIG. 5 , signals in the functional HBM device 500 can move along any of three bidirectional paths between the interface die 510, one of the volatile memory dies 520, and any two of the functional die 530. For example, a first signal path P1 extends between the interface die 510 and the volatile memory die 520. The first signal path P1 can be used during normal operation of the functional HBM device 500 to communicate between the interface die 510 (and any suitable components beyond the interface die 510, such as Figure 3 CPU / GPU 320 and / or Figure 2The first signal path P1 may be used to perform any number of read / write operations between the interface die 510 and the functional die 530. For example, the first signal path P1 may be used to load weights for a trained neural network and / or inputs for various neural network calculations onto the volatile memory die 520. The second signal path P2 extends between the volatile memory die 520 and the functional die 530. The second signal path P2 may be used to write a subset of the data in the volatile memory die 520 (e.g., weights and / or input values ​​for neural network processing operations) from the volatile memory die 520 to the functional die 530, write the results of some computer processing at the functional die 530 to the volatile memory die 520, and / or perform any other suitable operations. The third signal path P3 extends between the interface die 510 and the functional die 530. The third signal path P3 may be used to write the results of one or more neural network processing operations to the interface die 510 (e.g., communicated externally to the functional HBM device 500) and / or perform any other suitable operations. Because the operations described above with reference to the first to third signal travel paths P1 to P3 use a high-bandwidth channel (i.e., the TSVs 542 in the shared HBM bus 540), the operations may be completed at a relatively fast rate, with relatively low power requirements, and / or while generating a relatively small amount of heat (e.g., compared to performing the same read / write operations outside of the functional HBM device 500).

[0045] The bidirectional restriction of each of the TSVs 542 in the illustrated embodiment prevents any subset of the TSVs 542 from being used for multiple operations simultaneously (e.g., simultaneously along the first travel path P1 and the third travel path P3). However, it should be understood that in some embodiments, one or more of the signal travel paths may have multiple destinations. For example, a write operation to one of the volatile memory dies 520 along the first signal travel path P1 may simultaneously write data to the functional die 530 along the third signal travel path P3. In addition, the first subgroup 542a may be used for a first operation (e.g., writing data from one of the volatile memory dies 520 to the functional die 530 along the third signal travel path P3) while the second subgroup 542b is used for a second operation (e.g., writing data from the interface die 510 to one of the volatile memory dies 520 along the first signal travel path P1). Furthermore, as discussed in more detail below, the shared HBM bus 540 may include additional TSVs and / or additional subgroupings of TSVs to allow subgroups to be dedicated to various signal travel paths at the expense of the shared HBM bus 540 having a larger footprint.

[0046] Figure 66 is a schematic diagram of a neural network computing operation 600 according to some embodiments of the present technology. The neural network computing operation 600 (sometimes also referred to as an "artificial neuron") receives one or more inputs in an input vector 610, multiplies each of the inputs by a corresponding weight via a weight vector 620, and then applies a summing function 630 (sometimes also referred to as a transfer function) to sum each of the weighted inputs in an initial output 632. The result may then be passed through an activation function 640 to produce a final output 642 (sometimes also referred to as an "activation"). In some embodiments, the activation function 640 produces a binary output based on the sum of the weighted inputs in the initial output 632 (e.g., 0 if the sum is below a predetermined threshold, and 1 if the sum is equal to or above a predetermined threshold). In some embodiments, if the weighted sum is at or above a predetermined threshold, the activation function 640 passes the value of the weighted sum in the initial output 632 forward, otherwise outputs a predetermined value (e.g., 0). In some embodiments, the activation function 640 filters the weighted sum in the initial output 632 to a value within a predetermined range (e.g., producing a number between 0 and 1).

[0047] In a specific non-limiting example, the input may be an image value of a pixel in an image, and the weight may be based on the importance of the value in detecting an individual face in the image. In various embodiments, the weight may be based on the learned importance of the value (e.g., learned during the training process of the neural network computing operation 600) and / or loaded from a database (e.g., based on a previously trained model). When the weighted sum is above a predetermined threshold, the activation function 640 may set the final output 642 to 1 (corresponding to the detection of the individual face). Otherwise, the activation function 640 may set the final output 642 to 0 (corresponding to the non-detection of the individual face). In some related examples, the weight corresponds to the importance of the value in detecting the individual face from a first angle. In such examples, the process may be repeated with a second weight vector corresponding to the importance of the value in detecting the individual face from a second angle. In some related examples, the weight corresponds to the importance in calculating one or more features of the individual face. In such examples, the process may be repeated for multiple features to produce multiple final outputs 642. The final output 642 may then be input into another neural network computing operation that produces a final output corresponding to the detection or non-detection of the individual face.

[0048] Because the neural network computing operation 600 is based on logic operations applied to the input vector 610 and the weight vector 620 (e.g., generating a sum of weighted inputs), the neural network computing operation 600 can be implemented via multiple flash memory cells coupled to a shared source line, where each flash memory cell generates a current corresponding to the weighted input, and the shared source line captures a total current corresponding to the sum of the weighted inputs. The operation of individual flash memory cells and flash memory cell arrays that provide neural network computing operations is further described below.

[0049] Figure 7 700 is a partial circuit schematic of an individual flash memory cell 700 configured in accordance with some embodiments of the present technology. The flash memory cell 700 may be part of a functional die of a functional HBM device (e.g., a programmable NOR flash die for neural network computations). In the illustrated embodiment, the flash memory cell 700 is implemented with NMOS transistors (e.g., as a single NOR flash cell). However, it should be appreciated that in embodiments, the flash memory cell 700 may be implemented differently (e.g., with PMOS transistors as a NAND flash cell). The flash memory cell 700 includes a gate terminal 710, a drain terminal 712, and a source terminal 714. The gate terminal 710 may be coupled to a word line 702, the drain terminal 712 may be coupled to a bit line 704, and the source terminal 714 may be coupled to a source line 708. As described herein, word lines 702 and / or source lines 708 can be shared by other flash memory cells 700 in a row (or column and / or other suitable lines (linear or non-linear)), and / or bit lines 704 can be shared by other flash memory cells 700 in a column (or row and / or other suitable lines (linear or non-linear)). Flash memory cells 700 can also be associated with threshold values ​​706. As described herein, threshold values ​​706 can be programmed differently for individual flash memory cells 700.

[0050] During operation, when the bit line voltage V applied to the bit line 704 is bl is much lower than the word line voltage V applied to word line 702 gs and threshold 706 (ie, threshold voltage V th ), the flash memory cell 700 generates an output current (I ds ):

[0051] I ds = k(V gs -V th )*V bl (1)

[0052] where k is a process-dependent constant. As discussed in more detail below, the threshold voltage V thcan be programmed into the flash memory cell 700 .

[0053] To perform one or more operations associated with neural network computations, the weight vector 620 ( Figure 6 ) can be translated from the result of k*(Vgs-Vth), which can then be used to adjust the threshold voltage V of individual flash memory cells (e.g., flash memory cell 700). th Programming. The input vector 610 ( Figure 6 ) can then be translated into a bit line voltage V bl and applied to the flash memory cell, resulting in an output current I that can be translated into a weighted input ds Then, the output current I ds Can be driven on a shared source line that sums the output currents from multiple similar flash memory cells, thereby simulating Figure 6 The process can then use the bit line voltage V bl Repeat any number of times with different inputs to perform additional neural network computations with the same weight vector.

[0054] Figure 8 8 is a partial circuit diagram of a flash memory device 800 configured in accordance with some embodiments of the present technology. The flash memory device 800 may be part of a functional die of a functional HBM device (e.g., a programmable NOR flash die for neural network computing). In the illustrated embodiment, the flash memory device 800 includes a plurality of first flash memory cells 812 ( Figure 8 The first row 810 (n) includes a plurality of second flash memory cells 822 ( Figure 8 Each of the first flash memory cell 812 and the second flash memory cell 822 is substantially similar to the above reference numeral 820. Figure 7 The flash memory cell 700 discussed. For example, the first-first flash memory cell 812 in the first row 810 1 including coupling to the first word line 802 1 The gate terminal is coupled to the first bit line 804 1 The drain terminal, the first-first threshold voltage 814 1 and a source terminal coupled to the first source line. In the illustrated embodiment, the first word line 802 1 Each of the first flash memory cells 812 in the first row 810 (eg, 812 1 To 812 n ) to share the same (eg, constant) first word line voltage V gs1is applied to each of the first flash memory cells 812. In addition, each of the first flash memory cells 812 outputs current to a first source line 818 in an additive manner (e.g., the first source line 818 is a shared and / or common source line that sums the output currents from the first flash memory cells 812 in the first row 810). The first source line 818 then passes through a first filter circuit 819.

[0055] To simulate artificial neurons (e.g., Figure 6 type of discussion), based on the weight vector (e.g. Figure 6 The weight vector 620 of FIG. 61 is used to program each of the first threshold voltages 814. More specifically, the first-first threshold voltage 814 1 The first-first threshold voltage V th11 and the first word line voltage V gs1 The difference between corresponds to the first weight in the weight vector of the artificial neuron; the second-first threshold 814 2 The second-first threshold voltage V th12 and the first word line voltage V gs1 The difference between corresponds to the second weight in the weight vector; and so on to the nth - first threshold 814 n In other words, each of the first flash memory cells 812 associated with the row 810 can be programmed with an individual threshold voltage corresponding to a different weight in the weight vector.

[0056] Then the vectors from the input (e.g., Figure 6 The input vector 610) is translated into the bit line voltage V bl1-n , bit line voltage V bl1-n is loaded into bit line 804 (eg, V bl1 Load to 804 1 Up, V bl2 Load to 804 2 and so on), thereby causing a plurality of first output currents to be output to the first source line 818. For example, from the first-first flash memory cell 812 1 The first output current i 11 is output to the first source line 818; from the second-first flash memory cell 812 2 The second-first output current i 12 is output to the first source line 818; and so on to the nth-first flash memory cell 812 n The nth-first output current i 1n .

[0057] The first filter circuit 819 may then act as an activation function for the artificial neuron. For example, as discussed above, if the sum on the first source line 818 is below a predetermined threshold, the first filter circuit 819 may set the value of the final output from the first row 810 to 0, otherwise the first filter circuit 819 may set the value of the final output to 1. In various other examples, the first filter circuit 819 may adjust the final output to a value within a range (e.g., from 0 to 1); if the sum on the first source line 818 is below a predetermined threshold, set the value of the final output to 0, otherwise the final output is set to the sum; if the sum on the first source line 818 is above a predetermined threshold, set the value of the final output to 1, otherwise the final output is set to the sum; scale the final output; and / or apply various other suitable activation function filters.

[0058] like Figure 8 As further described in the second row 820, the second flash memory cell 822 can include a second threshold voltage 824 that can be programmed independently of the first threshold voltage 814. In addition, the bit line 804 can be shared by the corresponding first flash memory cell 812 and the second flash memory cell 822. For example, the first bit line 804 1 The first-first flash memory cell 812 of the first row 810 1 and the first-second flash memory cell 822 on the second row 820 1 Shared; second bit line 804 2 The second-first flash memory cell 812 of the first row 810 2 and the second-second flash memory cell 822 on the second row 820 2 Share; and so on to share the nth bit line 804 n The n-th first flash memory unit 812 n and the nth-second flash memory unit 822 n Thus, the first row 810 may act as a first artificial neuron associated with a first set of weights for a set of inputs, while the second row 820 acts as a second artificial neuron associated with a second set of weights for the same set of inputs to perform separate neural network computation operations simultaneously (or substantially simultaneously).

[0059] In a specific non-limiting example, the first row 810 may be programmed with weights that allow a first artificial neuron to detect an individual face in an image at a first angle, while the second row 820 may be programmed with weights that allow a second artificial neuron to detect an individual face in an image at a second angle. In this example, input values ​​from an image may be provided on a bit line 804 shared by the first row 810 and the second row 820, thereby allowing the flash memory device 800 to quickly perform multiple neural network computation operations to detect individual faces in an image.

[0060] It should be understood that although the operation of the flash memory device 800 has been discussed with reference to two rows (e.g., the first row 810 and the second row 820), the flash memory device 800 may include any suitable number of rows, each of which may be programmed based on a different weight vector. For example, in various embodiments, the flash memory device 800 may include ten rows, one hundred rows, one thousand rows, ten thousand rows, and / or any other suitable number of rows. As described above, each row may include a plurality of flash memory cells coupled to a shared word line and / or a shared source line associated with the row. In addition, the n flash memory cells in any of the rows may be one memory cell, two memory cells, ten memory cells, one hundred memory cells, one thousand memory cells, ten thousand memory cells, and / or any other suitable number of memory cells. Further, different rows in the flash memory device 800 may have different numbers of flash memory cells. Purely by example, the first row 810 may include ten thousand memory cells, while the second row 820 may include one thousand memory cells. The additional memory cells in the first row 810 may allow for more complex neural network computing operations. The smaller number of memory cells in the second row may simplify and / or speed up neural network computing operations (e.g., by requiring fewer flash memory cells to be programmed). In some embodiments, only some of the flash memory cells within a row are programmed.

[0061] It should be understood that although the flash memory device 800 has been described herein as having flash memory cells arranged in horizontal rows, the techniques described herein are not limited thereto. For example, in various embodiments, groups of memory cells may be coupled to shared word lines and / or source lines in columns (e.g., bit lines arranged in rows), coupled to shared word lines and / or source lines along non-linear lines, and / or the like. Thus, as used herein, a "row" should be understood to refer to a group of one or more memory cells that are communicatively coupled to a shared word line and / or shared source line, regardless of the location distribution of the memory cells and / or the orientation of the shared word lines and / or shared source lines.

[0062] Fig. 9900 is a flow chart of a process 900 for implementing one or more neural network computing operations within a functional HBM device according to some embodiments of the present technology. Process 900 may be performed, for example, by an interface die (e.g., Figure 3 An interface die 332; Figure 4 Interface die 410; and / or Figure 5A Thus, process 900 may be performed entirely within a functional HBM device to reduce the power and time required for process 900 and / or reduce the heat generated by process 900.

[0063] Process 900 begins at block 902 by reading one or more weights from a memory die in a functional HBM device. The weights may correspond to a weight vector (e.g., Figure 6 The read at block 902 may use an HBM bus (e.g., one or more TSVs (e.g., Figure 4 TSV 442, Figure 5A TSV 542 and / or the like)) is accomplished.

[0064] At block 904, process 900 includes programming weights into a functional die in a functional HBM device. For example, as discussed above, the threshold voltage of a NOR flash memory cell may be programmed based on the weights of a desired neural network computation operation. Fig.10 Additional details regarding an exemplary process for programming threshold voltages are discussed. In some embodiments, process 900 may loop through blocks 902, 904 to read and program weights for multiple neural network computation operations (e.g., to program multiple word lines with different weights). In some embodiments, process 900 may program each different row in one loop through blocks 902, 904 (e.g., by reading multiple weight vectors at block 902 and then programming the flash memory cells belonging to the row at block 904).

[0065] At block 906, process 900 includes reading input (e.g., one or more input vectors) from the memory die. The input may correspond to data to be processed by the neural network computing operation, such as one or more images, one or more videos, one or more text files, and / or any other suitable input. The reading of block 902 may be accomplished using an HBM bus (e.g., one or more TSVs coupled between an interface die, a memory die, and / or one or more functional dies).

[0066] At block 908, process 900 includes performing one or more neural network computing operations at the functional die. In the illustrated embodiment, performing the one or more neural network computing operations includes performing various sub-processes. Returning to the example of the NOR flash memory die for illustration, performing the neural network computing operations may include, at block 910, loading inputs onto one or more bit lines in the NOR flash memory die (e.g., loading inputs onto Figure 8 804) while applying a constant voltage to each word line. In various examples, the constant voltage may be between about 1 volt and about 100 volts (e.g., a few volts to tens of volts), between about 5 volts and about 50 volts, and / or any other suitable voltage. Thus, each flash memory cell (e.g., coupled to Figure 8 The first word line 802 1 Each first flash memory cell 812 in the first row 810 of FIG. 8 outputs current to the source line (eg, Figure 8 The output currents from each of the flash memory cells are summed on the source line, thereby performing the summing required for the neural network computing operation. Then, at block 912, process 900 applies a filter to the current on the source line. For example, the filter may be Figure 8 The filter simulates an activation function for a neural network computation operation to produce a final output, which may be binary (mimicking a firing neuron), a varying value between an upper and lower limit (e.g., between 0 and 1), and / or various other suitable outputs from the activation function discussed in more detail above. The process 900 then reads the output at block 914 and writes the output to a suitable destination (e.g., to a memory die, an interface die, and / or a suitable external component (e.g., Figure 3 CPU / GPU 320) for external use).

[0067] In some embodiments, the process 900 reuses the same weights for multiple neural network computation operations (e.g., to process multiple different inputs). In such embodiments, the process 900 may return to block 906 to read a second set of inputs from the memory die without reprogramming the weights into the functional die, and then perform the second neural network computation operation at block 908. In other words, once the functional die has been programmed with the relevant weights for the neural network computation operation, the process 900 may iteratively run through blocks 906, 908 for multiple different input vectors. Thus, in a specific non-limiting example, the process may run a face detection neural network computation operation on multiple different images without reprogramming the functional die.

[0068] Fig.10 1 is a flow chart of a process 1000 for programming a threshold voltage at one or more memory cells on a flash memory device according to some embodiments of the present technology. The process 1000 may be performed, for example, by an interface die (e.g., Figure 3 An interface die 332; Figure 4 Interface die 410; and / or Figure 5A The process 1000 may be performed by one or more controllers on the interface die 510 of the HBM device. Thus, the process 1000 may be completed entirely within a functional HBM device to reduce the power and time required and / or reduce the heat generated by the process 1000. The process 1000 may be performed to program the threshold voltage of one or more memory cells based on corresponding weights used in one or more neural network calculations. For example, the process 1000 may be performed as follows: Fig. 9 900 (e.g., as part of block 904 of process 900) to program one or more memory cells coupled to a word line (e.g., within a row). The one or more memory cells may be the ones described above with reference to Figure 8 Therefore, for the purpose of illustration, the following reference is made to Figure 8 discuss Fig.10 However, one skilled in the art will appreciate that process 1000 can be performed to program threshold voltages in various other flash memory devices.

[0069] Process 1000 begins at block 1002 by selecting a row to be programmed (e.g., Figure 8 In some embodiments, process 1000 selects multiple rows at block 1002 to program the same weights into multiple selected rows (e.g., to provide one or more backup rows to be able to average the outputs from the rows to account for minor variations in processes 900, 1000, and / or the like). Similarly, at block 1004, the process selects memory cells on the selected rows to be programmed, e.g. Figure 8 The first-first flash memory unit 812 1 .

[0070] At block 1006, process 1000 includes setting the word line voltage (eg, Figure 8 The first word line 802 1 V gs1 ) is set to a first voltage. More specifically, the first voltage may be based on the first-first flash memory unit 812 1 ( Figure 8), and various process-related constants (e.g., loss in programming the threshold voltage). As discussed above, programming is performed to set the first-first threshold voltage V th11 So that k*(V gs1 -V th11 ) is equal to the weight, where V gs1 is the constant word line voltage that will be applied to the word lines during neural network computing operations.

[0071] At block 1008, the process includes providing a memory cell coupled to a selected memory cell (eg, Figure 8 The first-first flash memory unit 812 1 ) is set to a second voltage. The second voltage can be zero (or approximately zero) and / or a negative value to create a voltage difference across the selected memory cell that is large enough to program the threshold voltage. At block 1010, process 1000 includes programming each other memory cell (e.g., Figure 8 The second to nth first flash memory cells 812 2-n ) is set to a third voltage. That is, the third voltage should be sufficiently similar to (or equal to) the first voltage, which is set based on the weights as described above, so as to avoid affecting the threshold voltage at each of the other memory cells (i.e., causing the other memory cells not to be activated and / or programmed). Returning to the example above, at block 1008, Figure 8 The first bit line 804 1 The first line voltage V bl1 will be set to zero (or a negative value), and at block 1010, the second-nth bit line voltage V bl2-n will be set at block 1006 for V gs1 It should be understood that in some embodiments, blocks 1008 and 1010 are performed simultaneously (or substantially simultaneously) to set each of the bit line voltages at one time. In some embodiments, block 1010 is performed before block 1006 (e.g., substantially simultaneously with or immediately after block 1004) to set the bit line voltages of the non-target memory cells to the target voltage before setting the bit line voltage of the first memory cell.

[0072] After each of the bit line voltages has been set at block 1010, process 1000 includes programming the selected memory cells at block 1012. Programming the selected memory cells may include applying the set voltages (if not already applied) to the word line and the bit line. Once applied, each of the memory cells in the selected row will see a voltage difference between the word line voltage and the bit line voltage. For a selected memory cell (e.g., Figure 8 The first-first flash memory unit 812 1 ), the voltage difference is equal to (or substantially equal to) the first voltage (or the first voltage minus the second voltage). Therefore, the voltage difference programs the first voltage to the desired threshold voltage at the selected memory cell (e.g., sets the first-first threshold voltage V th11 ). In each of the other non-target memory cells (e.g., Figure 8 The second to nth first flash memory cells 812 2-n ), there is no voltage difference (or almost no voltage difference). Therefore, process 1000 does not change the threshold voltage at any of the other memory cells (i.e., does not program them). In some embodiments, block 1012 is performed simultaneously (or substantially simultaneously) with block 1006 to set the bit line voltage V for each of the bit lines at blocks 1008 and 1010. bl Then set the word line voltage V gs1 and applies it to the word line.

[0073] At decision block 1014, process 1000 checks whether the selected memory cell is the last memory cell on the word line of the selected row that needs to be programmed. If the selected memory cell is the last memory cell, then process 1000 proceeds to decision block 1014, otherwise process 1000 returns to block 1004 to select the next memory cell on the word line in the selected row and program the threshold voltage at the next memory cell. For example, process 1000 may be based on a second memory cell (e.g., Figure 8 The second-first flash memory unit 812 2 ) The desired weight will be the word line voltage V gs1 Set to the first voltage; set the second bit line voltage V bl2 Set to a second voltage (eg, zero); set the first and third to nth bit line voltages V bl1,3-n is set to a third voltage (eg, equal to or substantially equal to the first voltage); and the word line voltage V gs1 Applied to the word line to the second-first threshold voltage V th12 Process 1000 may then continue looping through blocks 1004 through 1012 to sequentially program the voltage for each of the third-nth memory cells in the selected row.

[0074] At decision block 1016, process 1000 checks whether the selected row is the last row that needs to be programmed on the memory device. If the selected row is the last row, process 1000 proceeds to block 1018 to complete, otherwise process 1000 returns to block 1002 to select the next row. Process 1000 can then loop through blocks 1004 to 1014 to program the voltage of each of the memory cells in the newly selected row before returning to decision block 1016 to continue selecting rows until each row has been programmed.

[0075] At block 1018, process 1000 ends programming with the threshold voltage of each of the memory cells being set to a desired value. In some embodiments, process 1000 at block 1018 includes setting the word line voltage V gs The word lines are set to a predetermined constant (e.g., a number of voltages, tens of volts, and / or the like) to prepare each of the word lines for neural network computing operations. In some embodiments, process 1000 returns a completion signal at block 1018 (e.g., to a component in an interface die), thereby allowing another process (e.g., Fig. 9 Process 900) continues with the neural network computing operation.

[0076] Fig.11 is a partial circuit schematic of a flash memory device 1100 configured in accordance with further embodiments of the present technology. In the illustrated embodiment, the flash memory device 1100 includes a row 1110 including a plurality of flash memory cells 1112 (n flash memory cells are schematically illustrated). Fig.11 The flash memory cell 1112 illustrated in FIG. 1 is substantially similar to the flash memory cell 1112 illustrated in FIG. Figure 8 The first flash memory cell 812 in question. For example, the first flash memory cell 1112 1 including coupling to the first bit line 1104 1 The drain terminal, the first threshold voltage 1114 1 and a source terminal coupled to a source line 1118; a second flash memory cell 1112 2 includes coupling to the second bit line 1104 2 The drain terminal, the second threshold voltage 1114 2 and coupled to the source terminal of source line 1118; and so on to the nth flash memory cell 1112 n , which includes coupling to the nth bit line 1104 n The drain terminal, the nth threshold voltage 1114 nand a source terminal coupled to source line 1118. In addition, each of the flash memory cells 1112 shares source line 1118 and each outputs current to source line 1118 in an additive manner (e.g., source line 1118 sums the output currents from the flash memory cells 1112). Source line 1118 then passes through filter circuit 1119, which acts as an activation function for row 1110.

[0077] However, in the illustrated embodiment, each of the flash memory cells 1112 in the row 1110 is coupled to an independent word line 1102. For example, the first flash memory cell 1112 1 including coupling to the first word line 1102 1 The gate terminal is used to apply the first word line voltage V gs1 To the first flash memory cell 1112 1 ; Second flash memory unit 1112 2 Including coupling to the second word line 1102 2 The gate terminal is used to set the second word line voltage V gs2 Applied to the second flash memory cell 1112 2 ; and so on to the nth flash memory unit 1112 n , which includes coupling to the nth word line 1102 n The gate terminal of the nth word line voltage V gsn Applied to the nth flash memory cell 1112 n In other words, the flash memory device 1100 includes a plurality of word lines 1102 that are individually coupled to corresponding ones of the flash memory cells 1112 in the row 1110. Although the illustrated embodiment may require a larger footprint of the row 1110 (e.g., Figure 8 8 (compared to the first row 810 of the flash memory device 1100) to accommodate multiple word lines 1102, but the independent word line for each of the flash memory cells 1112 can speed up various operations of the flash memory device 1100.

[0078] For example, since each of the flash memory cells 1112 includes an independent word line 1102 and a bit line 1104, each of the thresholds 1114 (eg, the first to nth thresholds 1114) is 1-n ) Available threshold voltage V th More specifically, the programming process (for example, implemented by the interface die) may be performed by increasing each word line voltage V gs1-n Set to the target voltage at the corresponding one in the flash memory cell 1112; Set each bit line voltage V bl1-n Set to zero; and set the first to nth word line voltages V gs1-n Applied to the first to nth word lines 1102 1-n.in other words, Fig.11 The programming process of the flash memory device 1100 illustrated in FIG. 1 can program each of the threshold values ​​1114 in one pass without requiring iterations through blocks 1002 to 1008 in process 1000. Thus, the programming process can be completed more quickly (e.g., compared to a conventional method applied to a flash memory device 1100). Figure 8 1000 ) and thereby helps accelerate neural network computing operations.

[0079] Fig.12 is a partially exploded schematic diagram of a functional HBM device 1200 configured in accordance with a further embodiment of the present technology. Fig.12 , the functional HBM device 1200 is substantially similar to the above reference Figure 4 For example, the functional HBM device 1200 includes an interface die 1210, one or more volatile memory dies 1220 ( Fig.12 4) and one or more functional dies 1230 ( Fig.12 The functional HBM device 1200 also includes a shared HBM bus 1240 having a plurality of TSVs 1242 extending communicatively between and coupled to the interface die 1210, the volatile memory die 1220, and the functional die 1230.

[0080] However, in the illustrated embodiment, the functional die 1230 may be positioned below the interface die 1210. This positioning may reduce (or minimize) the distance between the interface die 1210 and the functional die 1230 to speed up control signals therebetween (e.g., thereby speeding up the above reference Fig.10 However, it should be understood that in various other embodiments, the functional die 1230 can be positioned at any suitable location in the functional HBM device 1200 (eg, directly above the interface die 1210 and / or any other suitable location).

[0081] like Fig.12 As further described in the figure, the functional HBM device 1200 may also include one or more additional dies 1250 ( Fig.12 1 ). The additional die 1250 may include a static random access memory (SRAM) die (e.g., to provide cache to the functional HBM device 1200 to support computing operations thereon, to act as an L3 (or higher) cache, and / or the like), a controller die or other suitable processing unit, a non-volatile memory die, a logic die, and / or any other suitable component.

[0082] In a specific non-limiting example, the additional die 1250 may be a non-volatile memory die, such as a NAND memory die. Figure 2 The non-volatile memory die may provide a relatively large memory capacity that is "closer" to the functional die 1230 and / or the volatile memory die 1220 (e.g., accessible within the functional HBM device 1200 via the relatively high bandwidth of the shared HBM bus 1240) compared to the storage device 240 discussed above. Thus, for example, a relatively large data set may be communicated from the non-SiP storage device to the non-volatile memory die for use during a processing operation (e.g., various trained weight vectors for a neural network computing operation, a large input data set for a neural network computing operation, and / or the like). The functional HBM device 1200 may then iteratively communicate subsets of the data to the volatile memory die 1220 for access, use, and / or processing in the functional die 1230.

[0083] in conclusion

[0084] According to the above, it will be understood that specific embodiments of the present technology have been described herein for the purpose of illustration, but the well-known structures and functions have not been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments of the present technology. To the extent that any material incorporated herein by reference conflicts with the present disclosure, the present disclosure shall prevail. Where the context permits, singular or plural items may also include plural or singular items, respectively. Furthermore, unless the word "or" is explicitly limited to a single item excluded from other items in a list that only means reference to two or more items, the use of "or" in this list will be interpreted as including any single item in the (a) list, all items in the (b) list, or any combination of items in the (c) list. In addition, as used herein, the phrase "and / or" in "A and / or B" refers to A alone, B alone, and both A and B. In addition, the terms "include", "comprise", "have", and "have" are always used to mean at least including the stated features, so that any larger number of the same features and / or other features of additional types are not excluded. In addition, the terms "about" and "approximately" are used herein to mean within at least 10% of a given value or limit. Purely by way of example, a similar ratio means within 10% of a given ratio.

[0085] Several implementations of the disclosed technology are described above with reference to the figures. A computing device on which the described technology can be implemented may include one or more central processing units, memory, input devices (e.g., keyboards and pointing devices), output devices (e.g., display devices), storage devices (e.g., disk drives), and network devices (e.g., network interfaces). Memory and storage devices are computer-readable storage media that can store instructions for implementing at least part of the described technology. In addition, data structures and message structures can be stored or transmitted via data transmission media (e.g., signals on a communication link). Various communication links can be used, such as the Internet, a local area network, a wide area network, or a point-to-point dial-up connection. Therefore, computer-readable media may include computer-readable storage media (e.g., "non-transitory" media) and computer-readable transmission media.

[0086] Based on the above, it will also be appreciated that various modifications may be made without departing from the present disclosure or the present technology. For example, those skilled in the art will appreciate that the various components of the present technology may be further divided into subcomponents, or the various components and functions of the present technology may be combined and integrated. In addition, specific aspects of the technology described in the context of a particular embodiment may also be combined or eliminated in other embodiments.

[0087] In addition, although the advantages associated with the embodiments have been described in the context of specific embodiments of the present invention, other embodiments may also exhibit such advantages, and not all embodiments must exhibit such advantages to fall within the scope of the present invention. Therefore, the present disclosure and associated technology may cover other embodiments not explicitly shown or described herein.

Claims

1. A method of operating a high bandwidth memory (HBM) device, the method comprising: reading a plurality of weights for a trained neural network from one or more memory dies in the HBM device; programming a threshold voltage of each of a plurality of memory cells on a flash memory die in the HBM device based on the plurality of weights of the trained neural network; reading a plurality of inputs for a neural network computation operation from the one or more memory dies in the HBM device; and The neural network computation operation is performed using the plurality of inputs.

2. The method of claim 1 , wherein programming the threshold voltage of each of the plurality of memory cells comprises: For a first memory cell of a plurality of memory cells coupled to a word line of the flash memory die, setting a word line voltage of the word line to a first voltage based on a first weight from the plurality of weights; setting a first bit line coupled to the first memory cell to a second voltage different from the first voltage; and For each other memory cell from the plurality of memory cells on the word line, a corresponding bit line is set to the first voltage.

3. The method of claim 2, wherein programming the threshold voltage of each of the plurality of memory cells further comprises: for a second memory cell of the plurality of memory cells, setting the word line voltage to a third voltage based on a second weight from the plurality of weights; setting a second bit line coupled to the second memory cell to a fourth voltage different from the third voltage; and For each other memory cell from the plurality of memory cells on the word line, the corresponding bit line is set to the third voltage.

4. The method of claim 2, wherein the word line is a first word line of a plurality of word lines, wherein the word line voltage is a first word line voltage, wherein the plurality of memory cells are a plurality of first memory cells, and wherein programming the threshold voltages of the plurality of memory cells further comprises: For a second memory cell from a plurality of second memory cells coupled to a second word line, setting a second word line voltage of the second word line to a third voltage based on a second weight from the plurality of weights; setting a second bit line coupled to the second memory cell to a fourth voltage different from the third voltage; and For each other memory cell from the second plurality of memory cells, a corresponding bit line is set to the third voltage.

5. The method according to claim 1, wherein: The flash memory die includes a plurality of word lines, a plurality of bit lines, and a plurality of source lines, wherein each of the plurality of word lines is coupled to two or more memory cells from the plurality of memory cells, and wherein each of the two or more memory cells is coupled to one of the plurality of bit lines and output from the plurality of source lines to a common source line; and Performing the neural network computation operation using the plurality of inputs includes: loading a bit line voltage to each of the plurality of bit lines based on individual values ​​from the plurality of inputs; applying a filter to the resulting current on each of the plurality of source lines; and An output is read from each of the plurality of source lines.

6. The method of claim 5, wherein the output from each source line of the plurality of source lines is based on a sum of output currents from the two or more memory cells.

7. The method of claim 1 , wherein the neural network computation operation is a first neural network computation operation and the plurality of inputs is a first plurality of inputs, and wherein the method further comprises: reading a second plurality of inputs for a second neural network computation operation from the one or more memory dies in the HBM device; and A second neural network computation operation is performed using the second plurality of inputs.

8. The method of claim 1 , wherein each memory cell from the plurality of memory cells is individually coupled to a word line from a plurality of word lines and a bit line from a plurality of bit lines, and wherein programming the threshold voltages of the plurality of memory cells comprises: For each individual memory cell in the plurality of memory cells, setting a word line voltage of a corresponding word line from the plurality of word lines to a first voltage based on a corresponding weight from the plurality of weights; and A corresponding bit line coupled to the individual memory cell is set to a second voltage different than the first voltage.

9. The method of claim 1, further comprising writing the plurality of weights for the trained neural network and the plurality of inputs for the neural network computation operations from a storage die peripheral to the HBM device into the one or more memory dies.

10. A functional high bandwidth memory HBM device, comprising: Controller die; one or more volatile memory dies carried by the controller die; a flash memory die carried by the one or more volatile memory dies, the flash memory die comprising: one or more rows, each having a word line and two or more programmable memory cells coupled to the word line; and two or more bit lines, wherein each of the two or more bit lines is coupled to a corresponding programmable memory cell from each of the one or more rows; and a shared bus electrically coupled to each of the controller die, the one or more volatile memory dies, and the flash memory die, Wherein the controller die is configured to control the one or more volatile memory dies and the flash memory die through the shared bus to implement one or more neural network computing operations within the functional HBM device.

11. The functional HBM device of claim 10, wherein performing the one or more neural network computing operations comprises: reading a plurality of weights for a trained neural network from one or more volatile memory dies; programming a threshold voltage of each of the two or more programmable memory cells in each of the one or more rows in the flash memory die based on the plurality of weights of the trained neural network; reading a plurality of inputs for the one or more neural network computation operations from the one or more volatile memory dies; and The one or more neural network computing operations are performed on the two or more programmable memory cells in each of the one or more rows using the plurality of inputs.

12. The functional HBM device of claim 11 , wherein programming the threshold voltage of a first individual programmable memory cell of the two or more programmable memory cells in a first row comprises: setting a word line voltage of a first word line in the first row to a first voltage based on a first weight from the plurality of weights; setting a bit line coupled to the first individually programmable memory cell to a second voltage different than the first voltage; and For each other programmable memory cell in the first row, the corresponding bit line is set to a third voltage different from the second voltage.

13. The functional HBM device of claim 12, wherein the second voltage is zero or negative.

14. The functional HBM device of claim 11 , wherein performing one of the one or more neural network computing operations using the input comprises: for each individual bit line from the two or more bit lines, applying a bit line voltage to the individual bit line based on an individual value from the plurality of inputs; for each individual row, applying an activation function to a current on a source line coupled to each of the two or more programmable memory cells in the individual row; and Read the output from the activation function.

15. The functional HBM device of claim 14, wherein applying the activation function comprises comparing the current on the source line to a reference current.

16. A system-in-package (SiP) device, comprising: base substrate; a processing unit carried by the base substrate; and a functional high bandwidth memory HBM device carried by the base substrate and electrically coupled to the processing unit through a SiP bus, wherein the functional HBM device comprises: Controller die; one or more memory dies; Neural network computing chips; and a shared bus electrically coupled to each of the controller die, the one or more memory dies, and the neural network computing die, Wherein the controller die is configured to control the one or more memory dies and the neural network computing die through the shared bus to implement one or more neural network computing operations within the functional HBM device.

17. The SiP device of claim 16, wherein the neural network computing die comprises a NOR flash array having a plurality of rows, wherein each of the plurality of rows comprises a shared word line and a plurality of memory cells communicatively coupled to the shared word line, wherein each of the plurality of memory cells has a programmable threshold voltage.

18. The SiP device of claim 16, wherein implementing the one or more neural network computing operations comprises: reading a plurality of weights for the one or more neural network computation operations from one or more memory dies; writing the plurality of weights into threshold voltages of a plurality of memory cells in the neural network compute die; reading from the one or more memory dies a plurality of inputs for each of the one or more neural network computation operations; and Each of the one or more neural network computation operations is performed.

19. The SiP device of claim 18, wherein writing the plurality of weights into threshold voltages of a plurality of memory cells comprises sequentially programming individual threshold voltages of each of the plurality of memory cells based on individual ones of the plurality of weights.

20. The SiP device of claim 18, wherein: The neural network die includes (1) a plurality of word lines, each communicatively coupled to a subset of the plurality of memory cells, and (2) a plurality of bit lines, each coupled to a memory cell on each of the plurality of word lines; and Executing each of the one or more neural network computing operations includes applying a bit line voltage to each of the plurality of bit lines based on an individual input from the plurality of inputs.