Nvm-based high-capacity neural network inference engine

CN117063148BActive Publication Date: 2026-10-09INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280022919.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-31
Filing Date
2022-03-30
Publication Date
2026-10-09
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

片外存储器访问经常是耗能、耗时的,并且需要大的封装形状因数

Benefits of technology

[0004] The above overview is not intended to describe every illustrated embodiment or implementation of this disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117063148B_ABST
    Figure CN117063148B_ABST
Patent Text Reader

Abstract

A system, method, and computer program product for a neural network inference engine are disclosed. The inference engine system can include a first memory and a processor in communication with the first memory. The processor can be configured to perform operations. The operations that the processor is configured to perform can include retrieving a first task with the first memory and delivering the first task to the processor for processing the first task. The operations can also include pre-fetching a second task with the first memory while the processor processes the first task. The operations can further include the first memory delivering the second task to the processor upon completion of processing the first task. The operations can further include the processor processing the second task.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This disclosure generally pertains to the memory field, and more specifically to retrieving data from memory.

[0002] Neural networks place increasingly significant demands on memory subsystems. This is especially true for deep neural networks (DNNs) with their ever-growing model sizes and datasets. Off-chip memory access is often power-intensive, time-consuming, and requires a large package form factor. Thermal challenges can also impose limitations on memory systems and the memory technologies used within them. Summary of the Invention

[0003] Embodiments of this disclosure include systems, methods, and computer program products for neural network inference engines. In some embodiments of this disclosure, the inference engine system may include a first memory and a processor communicating with the first memory. The processor may be configured to perform operations. Operations configured to be performed by the processor may include retrieving a first task using the first memory and delivering the first task to the processor for processing. Operations may also include prefetching a second task from the first memory while the processor is processing the first task. The operations may further include delivering the second task to the processor from the first memory upon completion of processing the first task. The operations may further include the processor processing the second task.

[0004] The above overview is not intended to describe every illustrated embodiment or implementation of this disclosure. Attached Figure Description

[0005] The accompanying drawings, which are included in and form a part of this disclosure, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings are merely illustrative of certain embodiments and do not limit the scope of the disclosure.

[0006] Figure 1 The memory stack according to this disclosure is shown.

[0007] Figure 2 A memory stack with integrated artificial intelligence according to an embodiment of the present disclosure is described.

[0008] Figure 3a A memory system according to an embodiment of the present disclosure is shown.

[0009] Figure 3b A timeline of tasks performed by the memory system according to an embodiment of the present disclosure is shown.

[0010] Figure 4a A memory system according to an embodiment of the present disclosure is shown.

[0011] Figure 4b A timeline of tasks performed by the memory system according to embodiments of the present disclosure is shown.

[0012] Figure 5a A memory system according to an embodiment of the present disclosure is shown.

[0013] Figure 5b A timeline of tasks performed by a memory system according to embodiments of the present disclosure is shown.

[0014] Figure 6 A memory system according to an embodiment of the present disclosure is described.

[0015] Figure 7 A storage system according to an embodiment of the present disclosure is shown.

[0016] Figure 8 A cloud computing environment according to an embodiment of the present disclosure is shown.

[0017] Figure 9 An abstract model layer according to an embodiment of this disclosure is described.

[0018] Figure 10 A high-level block diagram of an exemplary computer system, according to embodiments of the present disclosure, that can be used to implement one or more of the methods, tools, modules, and any associated functions described herein.

[0019] While the embodiments described herein are subject to various modifications and alternatives, their details have been illustrated by way of example in the accompanying drawings and will be described in detail. However, it should be understood that the specific embodiments described are not intended to be limiting. Rather, it is intended to cover all modifications, equivalents, and substitutions that fall within the scope of this disclosure. Detailed Implementation

[0020] This disclosure generally relates to the field of digital memory, and more specifically, to the field of retrieving data from memory. Additional aspects of this disclosure will be apparent to those skilled in the art. Some of these aspects are further described below.

[0021] Embodiments of this disclosure include systems, methods, and computer program products for high-capacity neural network inference engines based on non-volatile memory. Some embodiments may be particularly useful in deep neural network applications. In some embodiments of this disclosure, the inference engine system may include a first memory and a processor communicating with the first memory. The processor may be configured to perform operations. Operations configured to be performed by the processor may include retrieving a first task using the first memory and delivering the first task to the processor for processing. Operations may also include prefetching a second task using the first memory while the processor is processing the first task. The operations may further include delivering the second task to the processor from the first memory upon completion of processing the first task. The operations may further include the processor processing the second task.

[0022] To help understand this disclosure, Figure 1 A memory stack 100 according to an embodiment of the present disclosure is shown. The memory stack 100 includes memory dies 110a, 110b, 110c, 110d, 110e, 110f, and 110g stacked in a memory layer 110. Vertical interconnects 112 connect memory dies 110a, 110b, 110c, 110d, 110e, 110f, and 110g to buffer dies 120. Vertical interconnects 112 may be, for example, microbumps, pillars, or direct pad-to-pad bonding. Buffer dies 120 may have one or more buffer segments 120a, 120b, and 120c. The memory stack 100 may be connected to another system (e.g., a memory chip) via a controlled crash chip connection (C4) 130.

[0023] In some embodiments, memory modules 110a, 110b, 110c, 110d, 110e, 110f, and 110g may include one or more types of non-volatile memory. Memory modules 110a, 110b, 110c, 110d, 110e, 110f, and 110g may, for example, include high-density memory. Memory modules 110a, 110b, 110c, 110d, 110e, 110f, and 110g may include, for example, phase-change memory and / or magnetoresistive random access memory (MRAM).

[0024] In some embodiments, memory modules 110a, 110b, 110c, 110d, 110e, 110f, and 110g may preferably use memory types with high durability, the ability to maintain data at high temperatures, and low latency. In some embodiments, memory layer 110 may contain read-only data, or memory layer 110 may contain data that is rarely written to or changed. For example, in some embodiments, a fully developed artificial intelligence (AI) model may be stored on memory layer 110; the AI ​​model may be fully or partially retrieved from memory layer 110 through one or more buffer segments 120a, 120b, and 120c to be used by the processor. In this way, large AI models can be stored in memory layer 110, which consists of slow, high-density memory, and the latency of using the AI ​​model can be reduced by using fast, low-density memory as buffer segments 120a, 120b, and 120c.

[0025] Buffer module 120 may include one or more buffers. In some embodiments, a buffer may be segmented into multiple buffer segments 120a, 120b, and 120c. In some embodiments, multiple buffers may be used, and each buffer may be either unsegmented or segmented into multiple buffer segments 120a, 120b, and 120c. Multiple buffer segments 120a, 120b, and 120c may perform multiple acquisitions simultaneously, in series, or in some combination thereof.

[0026] Buffer segments 120a, 120b, and 120c may be memory. In some embodiments, buffer segments 120a, 120b, and 120c may be fast, low-density memory, such as, for example, static random access memory (SRAM) or dynamic random access memory (DRAM). In some embodiments, buffer segments 120a, 120b, and 120c may hold data while the processor is processing other data, may hold the data being processed by the processor, and / or may hold data while the processor processes and rewrites data (e.g., overwriting pulled data with computation). In some embodiments, buffer segments 120a, 120b, and 120c may deliver processed data (e.g., task computation) to memory layer 110 for storage; in some embodiments, buffer segments 120a, 120b, and 120c may deliver processed data to memory layer 110 for storage and then fetch other data to deliver to the processor for processing and / or computation.

[0027] In some embodiments of this disclosure, memory stack 100 may generate a first task computation in response to a processor processing a first task. The processor may send the first task computation to a first memory (e.g., a buffer or buffer segment 120a, 120b, or 120c); the first memory may accept the first task computation. The first memory may deliver the first task computation to a second memory (e.g., memory layer 110 or memory modules 110a, 110b, 110c, 110d, 110e, 110f, or 110g). In some embodiments, the first memory may be a low-density memory, and the second memory may be a high-density memory. In some embodiments, the second memory may be integrated into a three-dimensional stack of memories (e.g., a memory stack), wherein the three-dimensional stack of memories includes multiple memory layers (e.g., memory modules 110a, 110b, 110c, 110d, 110e, 110f, and 110g).

[0028] In some embodiments of this disclosure, the first memory is a buffer, the buffer is in a buffer module, and the buffer module has an artificial intelligence core. Figure 2 An AI integrated memory stack 204 and its components are shown according to an embodiment of this disclosure.

[0029] The memory stack 200 may have a memory module layer 210 including one or more memory modules 210a, 210b, and 210c. The memory stack 200 may also have a buffer module 220 having one or more buffer segments 220a, 220b, and 220c. The memory stack 200 may be combined with an AI unit 202. The AI ​​unit 202 may include an AI accelerator 232 having an AI core group 234. The AI ​​core group may have multiple AI cores 234a, 234b, 234c, and 234d. One or more of the AI ​​cores may include, for example, a notepad (e.g., a digital notepad and / or notepad memory); in some embodiments, the notepad may include multiple buffers (e.g., double buffering) to reduce or hide latency. The AI ​​accelerator 232 may be integrated into the buffer module 220 to form an AI-enabled buffer module 250. In some embodiments, the AI-enabled buffer module 250 may include AI acceleration and computing power (e.g., it may include a processor).

[0030] In some embodiments, a buffer module and a computation module can be combined. Figure 3a A memory system 300 according to this embodiment of the present disclosure is shown. Memory 310 is shown in communication with computing module 320. Buffer 326 in computing module 320 has acquired task A 324a to deliver it to computing core 328. Simultaneously, prefetch controller 322 guides task B 324b from prefetching from memory 310.

[0031] Figure 3b A graphical timeline 340 is shown illustrating tasks performed by the memory system 300 according to an embodiment of the present disclosure. The graphical timeline 340 shows the tasks performed over time period 342 (x-axis). The operation of each component of the memory system 300 is identified by the component's name (listed on the y-axis), and the tasks performed by the component are shown on the graph.

[0032] Buffer 326 performs buffer operation 352, which can also be referred to as operation performed by the buffer. Buffer operation 352 may include fetching (or prefetching) and containing data for task 352a. Buffer operation 352 may also include prefetching (or fetching) and containing data for task B 352b. Computation core 328 performs core operation 356, which can also be referred to as operation performed by the computation core. Core operation 356 includes executing task A 356a and executing task B 356b. Prefetch controller 322 performs prefetch controller operation 358, which can also be referred to as operation performed by the prefetch controller. Prefetch controller operation 358 may include prefetching data that triggers task B 358a and prefetching data that triggers task C 358b.

[0033] Data for task A 352a may be stored in buffer 326 before, during, and / or after computation core 328 executes task A 356a. In some embodiments, computation core 328 may receive data for task A 356a from buffer 326 for computation, thereby making buffer 326 available for performing other tasks. In such embodiments, computation core 328 may deliver the results of completed task 356a (e.g., processed data) to the same buffer 326 (or buffer segment) or different buffers (or buffer segments) for transfer to storage memory (e.g., memory 310 or a different memory where the processed data can be used for another task).

[0034] In some embodiments, both the buffer 326 and the AI ​​accelerator can be integrated into the computing module 320. The AI ​​accelerator can have its own sub-unit within the computing module 320, or it can be integrated into another component (e.g., it can be integrated into the buffer 326).

[0035] In some embodiments, alternating operation between a first buffer (or segment) and a second buffer (or segment) can be used (e.g., for data acquisition). Buffer units that alternate operation between buffers (or buffer segments) can be referred to as ping-pong buffers. Figure 4a A memory system 400 according to this embodiment of the present disclosure is shown.

[0036] Memory 410 is shown communicating with compute module 420. A first buffer 426a (or buffer segment) in compute module 420 has acquired task A 424a to deliver it to compute core 428, and a second buffer 426b in compute module 420 has acquired task B 424b to deliver it to compute core 428. Multiplexer (MUX) 430 can direct data traffic, and prefetch controller 422 can direct prefetching task C 424c from memory 410.

[0037] MUX 430 can sort and direct data from a first buffer 426a and a second buffer 424b. MUX 430 can, for example, first submit task A 424a from the first buffer 426a to computation core 428, and task B 424b from the second buffer 426b to computation core 428. MUX 430 can wait until a task is completed before submitting another task. MUX 430 can also direct data generated by tasks (e.g., processed data, such as computations generated by tasks) to buffers (e.g., the first buffer 426a and / or the second buffer 426b), which can then deliver it to storage memory (e.g., memory 410 or external memory). MUX 430 can alternatively or additionally direct buffers to deliver data generated by completed tasks to different processors (e.g., computation cores in a connected or separate system). Different processors can use the data generated by tasks to compute other data; for example, processed data can be forwarded to another system as input data for another computation.

[0038] A ping-pong buffer can be a buffer (or buffer segment) that fetches a task while the processor is working on another task and is therefore not yet ready to accept a new task. Fetching a task before the processor is ready to compute it can be called prefetching.

[0039] Figure 4b A graphical timeline 440 is shown illustrating the tasks performed by the memory system 400 according to an embodiment of the present disclosure. The graphical timeline 440 shows the tasks performed over time period 442 (x-axis). The operation of each component of the memory system 400 is identified by the component's name (listed on the y-axis), and the tasks performed by the component are shown on the graph.

[0040] The first buffer 426a performs the first buffer operation 452, which may also be referred to as the operation performed by the first buffer. The first buffer operation 452 may include fetching (or prefetching) and containing data for task 452a. The first buffer operation 452 may also include prefetching (or fetching) and containing data for task C 452c. The second buffer 426b performs the second buffer operation 454, which may also be referred to as the operation performed by the second buffer. The second buffer operation 454 may include fetching (or prefetching) and containing data for task B 452a.

[0041] Computational core 428 executes core task 456, which can also be referred to as the task performed by the computational core. Core task 456 includes executing task A 456a, executing task B 456b, and executing task C 456c. Prefetch controller 422 executes prefetch controller task 458, which can also be referred to as the task performed by the prefetch controller. Prefetch controller task 458 may include prefetching data that triggers task B 458a and prefetching data that triggers task C 458b.

[0042] Data for task A 452a may be stored in the first buffer 426a before, during, and / or after the computation core 428 executes task A 456a. In some embodiments, the computation core 428 may receive data for task A 456a from the first buffer 426 for computation, thereby making the buffer 426 available for performing other tasks. In such embodiments, the computation core 428 may deliver the result of completed task A 456a (e.g., processed data) to the same first buffer 426a (or buffer segment) or a different buffer (or buffer segment) for transfer to a storage memory (e.g., memory 410 or a different memory where the processed data can be used for another task).

[0043] In some embodiments, a ping-pong buffer may be preferred to reduce latency implemented by the end user. For example, memory 410 may be a slow, high-density memory, which could result in data retrieval taking several seconds; alternating operation allows computation core 428 to perform work on task A 452a in a first buffer 426a while a second buffer 426b prefetches task B 452b, such that task B 452b is retrieved and prepared for computation core 428 to work on immediately after completing task A 452a. In such an embodiment, the first buffer 426a can deliver computation to its destination and prefetch task 424c.

[0044] In some embodiments of this disclosure, the buffer may be a first memory having a first buffer segment and a second buffer segment. The first buffer segment can fetch a first task, and the second buffer segment can prefetch a second task. When the processor processes the first task, the second buffer segment can fetch the second task.

[0045] In some embodiments, the buffer and the computation core can be on separate modules. Figure 5a A memory system 500 according to this embodiment of the present disclosure is shown. A memory 510 is shown communicating with a buffer module 526 and a computation module 520. The buffer module 526 communicates with the computation module 520.

[0046] Buffer module 526 has a first buffer segment 526a and a second buffer segment 526b. Tasks fetched by each buffer segment can be delivered to the processor or compute core 528 for processing and / or computation. Task A 524a fetched by the first buffer segment 526a is delivered to compute core 528, and task B 524b fetched by the second buffer segment 526b is delivered to compute core 528. Simultaneously, prefetch controller 522 directs task C 524c to be prefetched from memory 510.

[0047] A MUX (not shown) can be used to order data traffic to and / or direct data from buffer segments to compute core 528. In some embodiments, the MUX can direct buffers to return processed data (e.g., computation results) to the same memory 510; in some embodiments, the MUX can direct processed data to different memories, another compute core, another processor, and / or other locations for storage and / or use.

[0048] Figure 5b A timeline of tasks performed by the memory system according to embodiments of the present disclosure is shown. Graphical timeline 540 illustrates the tasks performed over time period 542 (x-axis). The operation of each component of the memory system 500 is identified by the component's name (listed on the y-axis), and the tasks performed by the component are shown graphically.

[0049] Buffer module 526 performs buffer operation 552, which can also be referred to as operation performed by the buffer. Buffer operation 552 may include fetching (or prefetching) and containing data for task 552a. Buffer operation 552 may also include prefetching (or fetching) and containing data 552b for task B. Computation core 528 performs core operation 556, which can also be referred to as operation performed by the computation core. Core operation 556 includes executing task A 556a and executing task B 556b. Prefetch controller 522 performs prefetch controller operation 558, which can also be referred to as operation performed by the prefetch controller. Prefetch controller operation 558 may include prefetching data that triggers task B 558a and prefetching data that triggers task C 558b.

[0050] Data for task A 552a may be stored in buffer module 526 before, during, and / or after computation core 528 executes task A 556a. In some embodiments, computation core 528 may receive data for task A 556a from buffer module 526 for computation, thereby making buffer module 526 available for performing other tasks. In such embodiments, computation core 528 may deliver the results of completed task 556a (e.g., processed data) to the same buffer module 526 (or buffer segment) or a different buffer module (or buffer segment) for transfer to storage memory (e.g., memory 510 or a different memory where the processed data can be used for another task).

[0051] In some embodiments, both the buffer module 526 and the AI ​​accelerator can be integrated into the computing module 520. The AI ​​accelerator can have its own sub-unit within the computing module 520, or it can be integrated into another component (e.g., it can be integrated into the buffer module 526).

[0052] In some embodiments of this disclosure, the system may further include an error correction engine that communicates with the memory and the processor. Figure 6 A memory system 600 according to such an embodiment of the present disclosure is depicted.

[0053] Memory 610 is shown communicating with buffer 626 and computation module 620. Buffer 626 communicates with computation module 620. Memory 610 can transfer data bit 622a to computation module 620. This communication can be done directly (as shown) or through a buffer (not shown). Memory 610 can also transfer parity bit 622b to computation module 620. This communication can be done directly (not shown) or through a buffer (as shown). Figure 6 The parity bit 622b, which travels to the computation module 620 via buffer 626, and the data bit 622a, which is delivered directly from memory 610 to the computation module 620, are shown.

[0054] Data bits 622a and parity bits 622b are fed to the error correction engine 624. Data can be verified and corrected, if desired, before being submitted to the computation core. The error correction engine 624 can reside on a memory module, a buffer module, a computation module 620, or a different module (e.g., a module with a different memory or a dedicated error correction module). In some embodiments, data bits 622a, parity bits 622b, and the error correction engine 624 can be co-located (e.g., on the same memory module, the same buffer module, the same computation module, or a different error correction module).

[0055] In some embodiments of this disclosure, the sensor may be implemented as a protection system, data contained within the system, and / or one or more auxiliary systems (e.g., a data collection system). For example, in some embodiments, the system may include a first temperature sensor for sensing a first temperature; the first temperature sensor may communicate with a power gate, wherein the power gate throttles power if a first temperature threshold is reached. In some embodiments, the system may further include a second temperature sensor for sensing a second temperature communicating with the power gate. If the second temperature threshold is reached, the power gate may throttle. The power gate may be connected to any number of components and may throttle power to one or more of them (e.g., a buffer only or the entire system) based on the temperature threshold. Figure 7 A memory system 700 according to this embodiment of the present disclosure is shown.

[0056] Memory 710 is shown communicating with prefetch controller 722, buffer 726, and power gate 730 on buffer module 720. Buffer 726 communicates with compute core 728. Memory 710 can submit programs to compute core 728 via buffer 726. Figure 7 The diagram shows a computational core 728 operating on program A 724a and a buffer 726 including program B 724b. A prefetch controller 722 is submitting a prefetch request 724c for program C to memory.

[0057] Memory 710 embeds temperature sensors 732a, 732b, 732c, 732d, and 732e that communicate with power gate 730. Similarly, buffer module 720 is embedded with temperature sensors 734a and 734b that communicate with power gate 730. Temperature sensors 732a, 732b, 732c, 732d, and 732e in memory 710 may be uniformly distributed throughout memory 710 or concentrated in one or more bell-shaped regions (e.g., regions expected to reach a certain temperature or threshold most quickly). Similarly, temperature sensors 734a and 734b in buffer module 720 may be uniformly distributed throughout buffer module 720 or concentrated in one or more bell-shaped weather regions (e.g., regions expected to reach a specific temperature or threshold most quickly). The number and placement of temperature sensors 732a, 732b, 732c, 732d, 732e, 734a, and 734b in system 700 may vary depending on the system, its components, risk tolerance, and user preferences.

[0058] Power gate 730 can suppress (e.g., limit, increase, or eliminate) the power supply to any, some, or all of the components in system 700. Power throttling can be the result of manual (e.g., user command) or automatic (e.g., threshold reached) commands. An example of a manual command is that a user can identify the reason for shutting down system 700 and manually instruct power gate 730 to interrupt power to system 700. An example of an automatic command is that temperature sensor 734b can indicate to power gate 730 that the temperature near buffer 726 exceeds a safe temperature threshold, and power gate 730 can reduce the power supplied to buffer 726.

[0059] In some embodiments, the threshold temperature can be consistent throughout the system. For example, if any of the temperature sensors 732a, 732b, 732c, 732d, 732e, 734a, or 734b exceeds a threshold of 90°C, the power gate 730 can reduce the power of the system or any of its components. In some embodiments, the threshold temperature can be set specific to the relevant component. For example, the memory 710 can tolerate (e.g., operate at a temperature higher than that of the buffer 726) higher temperatures; in this case, the temperature sensors 732a, 732b, 732c, 732d, and 732e in the memory 710 can have a threshold temperature of 120°C, while the temperature sensors 734a and 734b in the buffer module 720 can have a temperature threshold of 80°C. If different components are held in different regions of the buffer module 720 and each component has a different temperature tolerance, the temperature thresholds of the temperature sensors 734a and 734b can be different from each other. Similarly, if different components are held in different regions of memory 710 with different temperature tolerances, the temperature thresholds of temperature sensors 732a, 732b, 732c, 732d, 732e, 734a, or 734b may vary depending on the sensor.

[0060] Thresholds can be set based on the materials included in system 700 and how system 700 is constructed. For example, if the memory cells used in memory 710 operate safely from 5°C to 75°C, then power gate 730 can trigger power throttling if one or more of the temperature sensors 732a, 732b, 732c, 732d, and 732e in memory 710 reaches 76°C or drops below 5°C. In some embodiments, the memory cells used in memory 710 can be selected for enhanced heat resistance (e.g., the ability to operate safely at elevated temperatures); if the memory cells used in memory 710 operate safely from 5°C to 125°C, then power gate 730 can trigger power throttling if one or more of the temperature sensors 732a, 732b, 732c, 732d, and 732e in memory 710 reaches 126°C or drops below 5°C.

[0061] Similarly, temperature thresholds can be set for the buffer module 720 and other components. For example, the buffer 726, which operates safely from 0°C to 50°C, may have a throttling threshold set to stop power supply if either of the temperature sensors 734a or 734b receives a temperature reading exceeding 50°C or dropping below 0°C. Likewise, temperatures can be sensed on or near the computing core 728, prefetch controller 722, power gate 730, and other components of the system 700, and the throttling thresholds can reflect the specific operating tolerances of different components.

[0062] In some embodiments of this disclosure, temperature can be sensed and / or tracked in or near different components of system 700, and the throttling threshold temperature can be the same for all components. For example, a user can set a uniform throttling threshold temperature of 45°C for any sensor in the system. The temperature threshold can be set automatically (e.g., the preset threshold is the same for any system), semi-autonomous (e.g., system specifications are used to automatically load temperature thresholds for a specific system and / or its components), manually (e.g., a user can input one or more temperature thresholds), or some combination thereof.

[0063] In some embodiments, the threshold may be set based on the materials included in system 700, how system 700 is constructed, and the geometry of the system. For example, some memory cells may be able to operate at temperatures above a standard temperature for up to three seconds, such that if the temperature falls within the standard temperature range within three seconds, the integrity of the memory cell is preserved. In this case, memory cells located in a well-ventilated, flat memory geometry may cool faster than memory cells in a memory stack (see [link to relevant documentation]). Figure 1 Therefore, the temperature thresholds associated with temperature sensors 732a, 732b, 732c, 732d, and 732e in the flat memory geometry can be triggered only after an excessively high temperature is reached for more than 3 seconds, while temperature sensors 732a, 732b, 732c, 732d, and 732e in the memory stack can be triggered immediately.

[0064] This disclosure can be implemented in a variety of systems, including but not limited to field-hardwired memory storage, cloud-accessible memory storage, analog storage, and digital storage. This disclosure enables faster memory access to any system that can use memory and computing, which can communicate integrally and directly (e.g., as part of a physical computer system) via a local connection (e.g., a local area network), a dedicated connection (e.g., a virtual private network), or some other connection (e.g., a wide area network or the Internet).

[0065] It should be understood that while this disclosure includes a detailed description of cloud computing, the implementation of the teachings cited herein is not limited to cloud computing environments. Rather, embodiments of this disclosure can be implemented in conjunction with any other type of computing environment now known or developed hereafter.

[0066] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0067] The features are as follows:

[0068] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring human interaction with the service provider.

[0069] Extensive network access: Capabilities are available through networks and accessed via standard mechanisms that facilitate the use of heterogeneous thin client or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0070] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, which features different physical and virtual resources that are dynamically allocated and reallocated based on demand. There is a sense of partial independence because consumers typically do not have control or knowledge of the exact portions of the resources provided, but may be able to specify portions at a higher level of abstraction (e.g., country, state, or data center).

[0071] Rapid flexibility: The ability to provide capacity quickly and flexibly, automatically scaling down and up rapidly in some situations to scale up rapidly. For consumers, the available supply capacity often appears unlimited and can be purchased in any quantity at any time.

[0072] Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the providers and consumers of the services being utilized.

[0073] The business model is as follows:

[0074] Software as a Service (SaaS): This provides consumers with the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from different client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.

[0075] Platform as a Service (PaaS): This provides consumers with the ability to deploy applications created by the consumer or acquired using programming languages ​​and tools supported by the provider onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and, if any, the configuration of the application hosting environment.

[0076] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, networking, and other basic computing resources that enable consumers to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather have control over the operating system, storage, and deployed applications, and may have limited control over selected networking components (e.g., host firewalls).

[0077] The deployment model is as follows:

[0078] Private cloud: A cloud infrastructure that operates solely for an organization. It can be managed by the organization or a third party and can exist on-site or off-site.

[0079] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and / or compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.

[0080] Public cloud: Makes cloud infrastructure available to the public or large industry groups and is owned by an organization that sells cloud services.

[0081] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported (e.g., cloud bursting for load balancing between clouds).

[0082] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure comprising a network of interconnected nodes.

[0083] Figure 8A cloud computing environment 810 according to an embodiment of the present disclosure is illustrated. As shown, the cloud computing environment 810 includes one or more cloud computing nodes 800, with local computing devices used by cloud consumers capable of communicating with these nodes. These local computing devices include, for example, personal digital assistants (PDAs) or cellular phones 800A, desktop computers 800B, laptop computers 800C, and / or automotive computer systems 800N. The nodes 800 can communicate with each other. They can be physically or virtually grouped in one or more networks (not shown), such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above.

[0084] This allows cloud computing environments 810 to provide infrastructure, platforms, and / or software services to cloud consumers without requiring them to maintain resources on their local computing devices. It should be understood that... Figure 8 The types of computing devices 800A-N shown are intended to be illustrative only, and the computing node 800 and cloud computing environment 810 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0085] Figure 9 The invention illustrates a method based on an embodiment of the invention. Figure 8 The abstract model layer 900 provided by the cloud computing environment 810. This should be understood in advance. Figure 9 The components, layers, and functions shown are intended to be illustrative only, and embodiments of this disclosure are not limited thereto. The following layers and corresponding functions are provided.

[0086] The hardware and software layer 915 includes hardware and software components. Examples of hardware components include: a host 902; a server 904 based on a RISC (Reduced Instruction Set Computer) architecture; a server 906; a blade server 908; a storage device 911; and a network and network components 912. In some embodiments, software components include network application server software 914 and database software 916.

[0087] The virtualization layer 920 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 922; virtual storage 924; virtual network 926, including virtual private network; virtual application and operating system 928; and virtual client 930.

[0088] In one example, management layer 940 can provide the functions described below. Resource provisioning 942 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 944 provides cost tracking of resources used within the cloud computing environment, and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, and protection for data and other resources. User portal 946 provides access to the cloud computing environment for consumers and system administrators. Service level management 948 provides cloud resource allocation and management to ensure that required service levels are met. Service level agreement (SLA) planning and fulfillment 950 provides pre-scheduling and procurement of cloud resources, anticipating future requirements for those resources according to the SLA.

[0089] The workload layer 960 provides examples of functionalities that can leverage a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 962; software development and lifecycle management 974; virtual classroom education delivery 966; data analytics and processing 968; transaction processing 970; and one or more high-capacity neural network inference engines based on NVM 972.

[0090] It should be understood that while this disclosure includes a detailed description of cloud computing, the implementation of the teachings cited herein is not limited to cloud computing environments. Rather, embodiments of this disclosure can be implemented in conjunction with any other type of computing environment currently known or developed in the future.

[0091] Figure 10 A high-level block diagram of an exemplary computer system 1001, according to embodiments of the present disclosure, is shown that can be used to implement one or more of the methods, tools, and modules described herein, as well as any associated functions (e.g., using one or more processor circuits or a computer processor). In some embodiments, the main components of the computer system 1001 may include a processor 1002 having one or more central processing units (CPUs) 1002A, 1002B, 1002C, and 1002D, a memory 1004, a terminal interface 1012, a storage interface 1017, an I / O (input / output) device interface 1014, and a network interface 1018, all of which may be directly or indirectly communicatively coupled to enable inter-component communication via a memory bus 1003, an I / O bus 1008, and an I / O bus interface 1010.

[0092] Computer system 1001 may include one or more general-purpose programmable CPUs 1002A, 1002B, 1002C, and 1002D, collectively referred to herein as CPU 1002. In some embodiments, computer system 1001 may include multiple processors typical of a relatively large system; however, in other embodiments, computer system 1001 may alternatively be a single CPU system. Each CPU 1002 may execute instructions stored in memory 1004 and may include one or more levels of on-board cache.

[0093] Memory 1004 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 1022 or cache memory 1024. Computer system 1001 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 1027 may be provided for reading from and writing to non-removable, non-volatile magnetic media, such as a “hard disk drive”. Although not shown, a disk drive may be provided for reading from or writing to a removable non-volatile disk (e.g., a “floppy disk”), or an optical disk drive may be provided for reading from or writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media). Furthermore, memory 1004 may include flash memory, such as a flash stick drive or a flash drive. The memory device may be connected to memory bus 1003 via one or more data media interfaces. Memory 1004 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of different embodiments.

[0094] One or more programs / utilities 1028 (each program having at least one set of program modules 830) may be stored in memory 1004. Programs / utilities 1028 may include a hypervisor (also called a virtual machine monitor), one or more operating systems, one or more applications, other program modules, and program data. Each or some combination of the operating system, one or more applications, other program modules, and program data may include an implementation of a network environment. Programs 1028 and / or program modules 1030 generally perform the functions or methods of different embodiments.

[0095] Although memory bus 1003 is Figure 10The diagram illustrates a single bus structure providing a direct communication path between CPU 1002, memory 1004, and I / O bus interface 1010. However, in some embodiments, memory bus 1003 may include multiple different buses or communication paths, which may be arranged in any of a variety of forms, such as point-to-point links in hierarchical, star, or network configurations, multiple hierarchical buses, parallel and redundant paths, or any other suitable type of configuration. Furthermore, while I / O bus interface 1010 and I / O bus 1008 are shown as a single corresponding unit, in some embodiments, computer system 1001 may include multiple I / O bus interfaces 1010, multiple I / O buses 1008, or both. Further, although multiple I / O interfaces 1010 are shown separating I / O bus 1008 from individual communication paths running to individual I / O devices, in other embodiments, some or all of the I / O devices may be directly connected to one or more system I / O buses 1008.

[0096] In some embodiments, computer system 1001 may be a multi-user mainframe computer system, a single-user system, a server computer, or a similar device that has little or no direct user interface but receives requests from other computer systems (clients). Further, in some embodiments, computer system 1001 may be implemented as a desktop computer, portable computer, laptop or notebook computer, tablet computer, pocket computer, telephone, smartphone, network switch or router, or any other suitable type of electronic device.

[0097] It is important to note that Figure 10 This description aims to depict representative major components of an exemplary computer system 1001. However, in some embodiments, a single component may have more than Figure 10 The greater or lesser complexity represented therein can exist differently from... Figure 10 The components shown or excluding Figure 10 Components other than those shown, and the number, type, and configuration of such components can vary.

[0098] This disclosure can be a system, method, and / or computer program product with any possible level of technical detail integration. The computer program product may comprise a computer-readable storage medium (or medium) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.

[0099] Computer-readable storage media can be tangible means for retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), SRAM, portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0100] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.

[0101] Computer-readable program instructions used to perform the operations disclosed herein may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may execute entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing the status information of the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this disclosure.

[0102] This document describes aspects of the disclosure with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0103] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0104] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, such that the instructions executed on the computer, other programmable apparatus, or other device perform the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, two blocks shown consecutively may actually be completed as a single step, executed simultaneously, substantially simultaneously, in a manner that partially or completely overlaps in time, or these blocks may sometimes be executed in reverse order depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0106] While this disclosure has been described with reference to specific embodiments, variations and modifications thereof are expected to become apparent to those skilled in the art. Different embodiments of this disclosure have been described for illustrative purposes but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements relative to technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein. Therefore, the following claims are intended to be construed as covering all such changes and modifications that fall within the scope of this disclosure.

Claims

1. An inference engine system, the system comprising: A memory stack including a buffer module and multiple memory modules, each of the buffer module and multiple memory modules being connected by a vertical interconnect, and the buffer module including an artificial intelligence core and a buffer, the buffer including a first buffer segment and a second buffer segment; A processor, configured to perform operations including: The first task is obtained from the first buffer segment; The first task is delivered to the processor by the first buffer segment for processing the first task; While the processor is processing the first task, the second task is prefetched by the second buffer segment; Upon completion of processing the first task, the second buffer segment transfers the second task to the processor; as well as The processor processes the second task.

2. The system according to claim 1, wherein, The operation further includes: In response to the processor processing the first task, a first task calculation is generated; The processor calculates and sends the first task to the second buffer segment; The first task is calculated by the second buffer segment; and The first task computation is delivered to the second memory; The buffer is a low-density memory, and the second memory is a high-density memory.

3. The system according to claim 2, wherein: The second memory is integrated in a three-dimensional stack of memory, wherein the three-dimensional stack of memory includes multiple memory layers.

4. The system according to claim 1, wherein, The system further includes: A first temperature sensor is used to sense a first temperature in communication with a power gate, wherein if a first temperature threshold is reached, the power gate throttles the power of the first memory.

5. The system according to claim 4, wherein, The system further includes: A second temperature sensor is used to sense a second temperature in communication with the power gate, wherein if the second temperature threshold is reached, the power gate throttles the power of the first memory.

6. The system according to claim 1, wherein, The system further includes: An error correction engine communicates with a first memory and the processor, wherein the data bits and check bits used by the error correction engine are located in the same position.

7. A method for storing and retrieving data in a memory, the method comprising: A memory stack is provided, the memory stack including a buffer module and a plurality of memory modules, each of the buffer module and the plurality of memory modules being connected by a vertical interconnect, and the buffer module including an artificial intelligence core and a buffer, the buffer including a first buffer segment and a second buffer segment; The first task is obtained from the first buffer segment; The first task is delivered to the processor by the first buffer segment for processing. While the processor is processing the first task, the second task is prefetched by the second buffer segment; Upon completion of processing the first task, the second task is delivered to the processor by the second buffer segment; as well as The processor processes the second task.

8. The method of claim 7, further comprising: In response to the processor processing the first task, a first task calculation is generated; The processor calculates and sends the first task to the second buffer segment; The second buffer segment receives the calculations for the first task; as well as The first task computation is delivered to the second memory; The buffer is a low-density memory, and the second memory is a high-density memory.

9. The method according to claim 8, wherein: The second memory is integrated in a three-dimensional stack of memory, wherein the three-dimensional stack of memory includes multiple memory layers.

10. The method of claim 7, further comprising: A first temperature is sensed using a first temperature sensor, wherein the first temperature sensor communicates with a power gate, wherein if a first temperature threshold is reached, the power gate throttles the power of the first memory.

11. The method of claim 10, further comprising: A second temperature is sensed using a second temperature sensor, wherein the second temperature sensor communicates with the power gate, and wherein if the second temperature threshold is reached, the power gate throttles the power of the first memory.

12. The method of claim 7, further comprising: Communication occurs between the error correction engine, the first memory, and the processor, wherein the data bits and check bits of the error correction engine are located in the same position.

13. A computer program product for memory storage and retrieval, the computer program product comprising program instructions executable by a processor to cause the processor to perform functions, the functions including: A memory stack is provided, the memory stack including a buffer module and a plurality of memory modules, each of the buffer module and the plurality of memory modules being connected by a vertical interconnect, and the buffer module including an artificial intelligence core and a buffer, the buffer including a first buffer segment and a second buffer segment; The first task is obtained from the first buffer segment; The first task is delivered to the processor by the first buffer segment for processing the first task; While the processor is processing the first task, the second task is prefetched by the second buffer segment; Upon completion of processing the first task, the second buffer segment transfers the second task to the processor; as well as The processor processes the second task.

14. The computer program product according to claim 13, wherein the function further includes: In response to the processor processing the first task, a first task calculation is generated; The processor calculates and sends the first task to the second buffer segment; The second buffer segment receives the calculations for the first task; as well as The first task computation is delivered to the second memory; The buffer is a low-density memory, and the second memory is a high-density memory.

15. The computer program product according to claim 13, wherein the function further includes: A first temperature is sensed using a first temperature sensor, wherein the first temperature sensor communicates with a power gate, wherein if a first temperature threshold is reached, the power gate throttles the power of the first memory.

Citation Information

Patent Citations

  • A machine learning reasoning coprocessor

    CN109814927A