Inference in memory
By integrating the computing core and dual-mode memory elements into the memory module, the problem of the impact on other components when expanding machine learning and artificial intelligence computing hardware in the prior art is solved, and flexible computing power enhancement is achieved.
Patent Information
- Application Number
- CN202180033985.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-26
- Filing Date
- 2021-05-20
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-05-20
AI Technical Summary
When adding machine learning and artificial intelligence computing hardware, existing computing systems often sacrifice the capabilities of other computing components and lack flexible interface expansion methods.
By integrating the computing core and dual-mode memory elements in the memory module, a portion of the memory is allocated to the host processing system and a portion to the computing core processing, enabling machine learning and artificial intelligence tasks.
Without affecting other computing components, it enhances the system's machine learning and artificial intelligence computing capabilities and provides flexible configuration options to optimize system performance.
Smart Images

Figure CN115516436B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims the benefit of and priority to U.S. Patent Application No. 16 / 883,869, filed May 26, 2020, the entire contents of which are incorporated herein by reference. BACKGROUND
[0003] Aspects of the present disclosure relate to performing machine learning and artificial intelligence tasks in non-traditional computing hardware, and in particular to performing such tasks in memory hardware.
[0004] In the era of big data, there is a sharp increase in demand for machine learning and artificial intelligence capabilities. Conventionally, machine learning has been used to generate models that can then generate inferences for artificial intelligence tasks, such as predictions, classifications, and the like. As the demand for inference capabilities increases, computing hardware manufacturers are seeking to expand the density of inference capabilities in existing computing platforms, such as desktop computers and servers, as well as other emerging types of processing systems, such as mobile devices and edge processing devices.
[0005] Conventionally, machine learning and artificial intelligence “accelerators” have been added to systems through hardware expansion interfaces (e.g., PCIe slots on a motherboard) to expand the capabilities of the underlying computing infrastructure. Unfortunately, using such hardware expansion interfaces for accelerators represents that these same interfaces cannot be used for other purposes, such as networking, graphics rendering, security, sound processing, and other common computing tasks. Thus, adding additional machine learning and artificial intelligence-optimized computing hardware to existing processing systems often comes at the expense of other important computing components.
[0006] Accordingly, what is needed are systems and methods that can add additional machine learning and artificial intelligence computing capabilities to existing processing systems without sacrificing other essential components. SUMMARY
[0007] Certain aspects of the present disclosure provide an enhanced memory module comprising: a computing core; and one or more dual-mode memory elements, wherein the enhanced memory module is configured to: allocate a first subset of memory in the one or more dual-mode memory elements as host processing system addressable memory and a second subset of memory in the one or more dual-mode memory elements as computing core addressable memory; receive data from the host processing system; process the data with the computing core to generate processed data; provide the processed data to the host processing system via the first subset of memory.
[0008] Further aspects provide a method for processing data with an enhanced memory module including a compute core, the method comprising: initializing the enhanced memory module by allocating a first subset of memory of the enhanced memory module as host processing system addressable memory and a second subset of memory of the enhanced memory module as compute core addressable memory; receiving, at the enhanced memory module, data from the host processing system; processing, on the enhanced memory module, the data with the compute core to generate processed data; and providing the processed data to the host processing system via the first subset of memory.
[0009] Further aspects provide a non-transitory computer readable medium comprising instructions that, when executed by one or more processors of a host processing system, cause the processing system to perform a method for processing data with an enhanced memory module including a compute core, the method comprising: initializing the enhanced memory module by allocating a first subset of memory of the enhanced memory module as host processing system addressable memory and a second subset of memory of the enhanced memory module as compute core addressable memory; receiving, at the enhanced memory module, data from the host processing system; processing, on the enhanced memory module, the data with the compute core to generate processed data; and providing the processed data to the host processing system via the first subset of memory.
[0010] The following description and associated drawings set forth certain illustrative features of one or more embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings illustrate aspects of one or more embodiments and thus are not to be considered limiting in scope.
[0012] Figure 1 An example of an enhanced memory module including a memory integrated accelerator is depicted.
[0013] Figure 2 An example processing system described with respect to an enhanced memory module is depicted. Figure 1 An example processing system described with respect to an enhanced memory module is depicted.
[0014] Figures 3A-3C An example configuration of an enhanced memory module is depicted.
[0015] Figure 4 An example method for transferring data to a compute core memory element on an enhanced memory module is depicted.
[0016] Figure 5 An example method for transferring data from a compute core memory element on an enhanced memory module is depicted.
[0017] Figure 6 An example method for processing data with a compute core on an enhanced memory module is depicted.
[0018] Figure 7 An example method for constructing an enhanced memory module is depicted.
[0019] Figure 8 An example memory map of a processing system including an enhanced memory module is depicted.
[0020] Figure 9 An example method for processing data with an enhanced memory module including a compute core is depicted.
[0021] Figure 10 An example electronic device that can be configured to perform data processing with an enhanced memory module is depicted.
[0022] For ease of understanding, the same reference numbers will be used in different drawings to designate the same elements common to the drawings. It can be contemplated that elements and features of one embodiment can be beneficially incorporated into other embodiments without further description. DETAILED DESCRIPTION
[0023] Aspects of the present disclosure provide systems and methods for adding additional machine learning and artificial intelligence computing capabilities to memory hardware in processing systems, such as computers, servers, and other computer processing devices.
[0024] Conventional computer processing systems can include a processor connected to memory modules, such as DRAM modules, via an interface, such as a DIMM interface, and connected to other peripheral devices via other interfaces, such as a PCIe interface. As the demand for machine learning and artificial intelligence “accelerators” increases, one solution for existing computer processing systems is to connect such accelerators to available system interfaces, such as available PCIe slots. However, the number of such peripheral interface slots can be limited on any given system, and in particular, conventional computer processing systems can typically include more DIMM slots than PCIe slots.
[0025] Described herein are enhanced memory modules that include integrated accelerators configured to perform machine learning and artificial intelligence tasks, such as training and inference. For example, memory integrated accelerators can be implemented in industry standard dual in-line memory modules (DIMMs), such as load-reduced DIMMs (LRDIMMs) or registered DIMMs (RDIMMs), using standard memory interfaces for data and instruction communication. In some embodiments, the enhanced memory modules can be referred to as, for example, “ML-DIMMs,” “AI-DIMMs,” or “inference DIMMs.”
[0026] In various embodiments, the memory integrated accelerators include one or more processing cores and can also be configured as a system on a chip (SoC) integrated into the enhanced memory modules. The processing cores enable on-memory processing of data in addition to conventional storage and retrieval. In particular, the processing cores can be configured to perform on-memory machine learning and artificial intelligence tasks.
[0027] Further, the enhanced memory modules described herein can be configured to partition and allocate the memory elements and / or memory space in each enhanced memory module between the host processing system (as in standard memory module configurations) and the memory integrated accelerators for on-memory machine learning and artificial intelligence tasks. In some embodiments, firmware elements can configure the memory allocation between host processing and on-memory processing at boot time.
[0028] The memory partitioning in the enhanced memory modules can also be dynamically configured by programs, applications installed on devices that include the enhanced memory modules, such that the performance of the overall processing system can be tuned for different tasks and performance requirements. Then, beneficially, the enhanced memory modules can be configured to act as ordinary memory modules to maximize the capacity of available memory for the host processing system, or can be configured to allocate some or all of that memory to the memory integrated accelerators to maximize the machine learning and artificial intelligence computing capabilities of the host processing system.
[0029] Accordingly, the embodiments described herein beneficially enable processing systems to utilize additional computing resources for computing tasks, such as machine learning and artificial intelligence tasks, without removing or preventing the use of other essential components due to a lack of available interfaces (e.g., PCIe), without adding or modifying additional physical interfaces, and without implementing entirely new interfaces. The additional computing resources greatly increase the capacity of existing processing systems, as well as new processing systems. Further, the embodiments described herein utilize a dual operating mode architecture that beneficially allows users to select different operating modes to configure in different situations to maximize the utility of the processing system.
[0030] The embodiments described herein can be used in many contexts. For example, data centers employing servers can implement the embodiments described herein to procure less servers while having the same amount of computing power, which saves device cost, power and cooling cost, space cost, and also provides more energy efficient processing centers. As another example, desktop and / or consumer grade computing systems can implement the embodiments described herein to bring improved data processing capabilities to home computers, rather than relying on external data processing capabilities (e.g., provided by cloud processing services). As yet another example, Internet of Things (IoT) devices can implement the embodiments described herein to extend data processing capabilities to entirely new types of devices, and further enable processing (e.g., machine learning and artificial intelligence inference) at the “edge.”
[0031] Example enhanced memory module with memory-integrated processing core
[0032] Figure 1 An example enhanced memory module 102 including a memory-integrated accelerator 104 is depicted.
[0033] In some embodiments, the enhanced memory module 102 is built according to a standardized form factor, such as an LR-DIMM form factor or another standardized form factor for memory modules. Utilizing a standardized form factor for the enhanced memory module 102 allows it to be integrated into existing processing systems without needing to modify the underlying system interface or memory access protocol. In this example, the memory module 102 includes pins 116 for interfacing with a DIMM slot in a processing system.
[0034] The memory module 102 includes a plurality of memory elements 112A-F, which can be, for example, dynamic random access memory (DRAM), synchronous DRAM (SDRAM), low power double data rate (LPDDR) DRAM, high bandwidth memory (HBM), graphics DDR (GDDR) memory, etc. In some examples, the memory elements 112A-F can include a mix of different types of memory elements.
[0035] Each of the memory elements 112A-F is connected to an adjacent memory buffer 114A-F, which is used to buffer data read from and written to the memory elements 112A-F by a host processing system (not shown).
[0036] Memory elements 112A-F can generally be referred to as dual-mode or dual-purpose memory elements, as their memory can be allocated in whole or in part to memory-integrated accelerator 104 to perform on-storage processing, and / or to the host processing system to perform regular system memory functions. Figures 3A-3C Several examples of different memory allocations are described.
[0037] Memory-integrated accelerator 104 includes flash memory 106, which is configured to store firmware that can be configured to initialize memory module 102. For example, the firmware can be configured to initialize computing core(s) 108 and any other logic on the enhanced memory module. In addition, a memory controller (not shown) can be configured to initialize memory elements 112A-F for use by computing core(s) 108 and / or the host processing system.
[0038] In some embodiments, the initialization of memory elements 112A-F can be handled by accelerator 104, which can improve the boot time of the host processing system (e.g., by reducing the boot time). In this case, the memory driver of the host processing system will only need to initialize the memory interface.
[0039] In some embodiments, the firmware in flash memory module 106 can configure memory module 102 at boot time and allocate (or map) some or all of the memory capacity (e.g., including memory elements and / or memory ranges, addresses, etc.) to accelerator 104. The firmware in flash memory module 106 can be updated so that the amount of memory allocated to accelerator 104 can change between runtimes. In embodiments where memory (e.g., ranges of memory addresses) from memory elements 112A-F is allocated at boot time, the consolidated processing system will only "see" the remaining available memory (not including the memory allocated to accelerator 104) and will not attempt to write to the memory allocated to accelerator 104, thereby avoiding memory conflicts.
[0040] In some embodiments, the firmware in flash memory module 106 can configure an initial memory allocation at boot time, which can then be dynamically changed by instructions from the consolidated processing system. In this case, the operating system of the consolidated processing system can be configured to dynamically reallocate memory for use by the processing system, thereby avoiding conflicts.
[0041] The accelerator 104 also includes computing core(s) 108, which can be configured to perform various types of processing tasks, including machine learning and artificial intelligence tasks, such as training and inference. In some embodiments, the computing core(s) 108 can be based on an ARM Cortex-M3, Cortex-M4, Cortex-M7, Cortex-M23, Cortex-R4, Cortex-R5, Cortex-R7, Cortex-R8, Cortex-A5, Cortex-A7, Cortex-A8, Cortex-A9, Cortex-A12, Cortex-A15, Cortex-A17, Cortex-A53, Cortex-A55, Cortex-A57, Cortex-A72, Cortex-A73, Cortex-A76, Cortex-A77, Cortex-A78, Cortex-A8x, Cortex-A12x, Cortex-A17x, Cortex-A53x, Cortex-A55x, Cortex-A57x, Cortex-A72x, Cortex-A73x, Cortex-A76x, Cortex-A77x, Cortex-A78x, Cortex-R4x, Cortex-R5x, Cortex-R7x, Cortex-R8x, RISC, or Complex Instruction Set Computer (CISC), to name a few examples. TM
[0042] In some embodiments, the computing core(s) 108 replace and integrate the functionality of a registry clock driver chip or circuit on a conventional memory module.
[0043] The accelerator 104 also includes a plurality of memory elements 110A-B that are dedicated to the computing core(s) 108, or in other words, the accelerator 104 also includes single mode or single purpose memory elements. Because the memory elements 110A-B are dedicated to the computing core(s) 108, the memory associated with the memory elements 110A to 110B is not addressable by the host processing system. The dedicated memory elements 110A-B allow the accelerator 104 to be able to perform computing tasks at all times. In some examples, the memory elements 110A-B can be configured to load model parameters, such as weights, biases, and the like, of a machine learning model that is configured to perform a machine learning or artificial intelligence task, such as training or inference.
[0044] In some embodiments, the accelerator 104 can be implemented as a System on a Chip (SOC) that is integrated with the memory module 102.
[0045] Notably, Figure 1 The number of each type of element (e.g., memory, buffer, flash, computing core, etc.) in the memory module 102 is merely one example, and many other configurations are possible. Moreover, although all of the elements are shown on a single side of the memory module 102, in other embodiments, the memory module 102 can include additional elements on the opposite side, such as additional memory elements, buffers, computing cores, flash memory modules, and the like. An example processing system including an enhanced memory module
[0046] Figure 2 An example processing system 200 is depicted that includes an enhanced memory module, such as the 102 described with respect to Figure 1
[0047] In this example, the processing system 200 implements a software stack including an operating system (OS) user space 202, an OS kernel space 212, and a CPU 218. The software stack is configured to interface with enhanced memory modules 224A and 224B.
[0048] Generally, an operating system can be configured to segregate virtual memory between an OS user space (such as 202) and an OS kernel space (such as 212) to provide memory protection and hardware protection against malicious or erroneous software behavior. In this example, the OS user space 202 is a memory region in which application software executes, such as the ML / AI application 208 and the accelerator runtime application 210. However, the kernel space 212 is a memory space reserved for running, for example, privileged operating system kernels, kernel extensions, and device drivers, such as the memory driver 214 and the accelerator driver 216.
[0049] Further, in this example, the OS user space 202 includes a machine learning and / or artificial intelligence application 208, which can implement model(s) 204 (e.g., machine learning models, such as artificial neural network models) configured to process data 206. The model(s) 204 can include, for example, weights, biases, and other parameters. The data 206 can include, for example, video data, image data, audio data, textual data, and other types of data, which the model(s) 204 can operate on to perform various machine learning tasks, such as image recognition and segmentation, speech recognition and translation, value prediction, and the like.
[0050] The ML / AI application 208 is also configured to interface with an accelerator runtime application 210, which is a user space process configured to direct data processing requests to processing accelerators, such as those found in the enhanced memory modules 224A-B. Notably, while the ML / AI application 208 is depicted as a single application, this is merely an example, and other types of processing applications can exist in the OS user space 202 and be configured to utilize the enhanced memory modules 224A-B. Figure 2 A single ML / AI application 208 is depicted in the middle, but this is merely an example, and other types of processing applications can exist in the OS user space 202 and be configured to utilize the enhanced memory modules 224A-B.
[0051] The accelerator runtime application 210 is configured to interface with an accelerator driver 216 in the OS kernel space 212. Generally, a driver in the OS kernel space 212 is software configured to control a particular type of device attached to a processing system, such as the processing system 200. For example, a driver can provide a software interface to a hardware device, such as the memory management unit 220, which enables an operating system and other computer programs to access the functionality of the hardware device without needing to know precise details about the hardware device.
[0052] In this example, the accelerator driver 216 is configured to provide access to the enhanced memory modules 224A-B through the memory driver 214. In some embodiments, the accelerator driver 216 is configured to perform the protocol used by the enhanced memory modules 224A-B. To do so, the accelerator driver 216 can interact with the memory driver 214 to send commands to the enhanced memory modules 224A-B. For example, such commands can be used to move data from memory allocated to the host processing system to memory allocated to an accelerator (e.g., to a compute core) of the enhanced memory modules 224A-B, to begin processing data, or to move data from memory allocated to the accelerator to memory allocated to the host processing system.
[0053] In some embodiments, the accelerator driver 216 can use a particular memory address to implement the protocol for writing to or reading from the memory of the enhanced memory modules 224A-B. In some cases, the memory address can be physically mapped to the enhanced memory modules. The accelerator driver 216 can also access memory addresses on the enhanced memory modules 224A-B that are not accessible to the OS user space 202, which can be used to copy data into memory allocated to an accelerator of the enhanced memory modules 224A-B and then to compute cores of those enhanced memory modules.
[0054] Further, the accelerator driver 216 is configured to interface with the memory driver 214. The memory driver 214 is configured to provide low-level access to the memory elements of the enhanced memory modules 224A-B. For example, the memory driver 214 can be configured to pass commands generated by the accelerator driver 216 directly to the enhanced memory modules 224A-B. The memory driver 214 can be part of the memory management of the operating system, but can be modified to implement the enhanced memory module functionality described herein. For example, the memory driver 214 can be configured to block access to the enhanced memory modules 224A-B during the period in which data is copied between memory allocated to an accelerator (e.g., a compute core) and memory allocated to the host processing system within the enhanced memory modules 224A-B.
[0055] The memory driver 214 is configured to interface with a memory management unit (MMU) 220, which is a hardware component configured to manage memory and cache operations associated with a processor, such as the CPU 218.
[0056] The memory management unit 220 is configured to interface with a memory controller 222, which is a hardware component configured to manage the flow of data to and from the main memory of the host processing system. In some embodiments, the memory controller 222 is a separate chip or is integrated into another chip, such as a component of a microprocessor, as with the CPU 218 (in which case it can be referred to as an integrated memory controller (IMC)). The memory controller 222 can alternatively be referred to as a memory chip controller (MCC) or a memory controller unit (MCU).
[0057] In this example, the memory controller 222 is configured to control the flow of data to and from the enhanced memory modules 224A-B. In some embodiments, as noted above, the enhanced memory modules 224A-B can be enhanced DIMMs.
[0058] The enhanced memory modules 224A-B include drivers 226A-B, respectively, that are capable of communicating with elements of the processing system 200. For example, the drivers 226A-B can communicate with other elements of the processing system 100 via a DIMM interface with the memory controller 222.
[0059] In general, the ML / AI application 208 can request processing of data 206 via the model 204 to be performed by the enhanced memory modules through the accelerator runtime application 210, which communicates the request to the accelerator driver 216, and then to the memory driver 214, which can be a regular memory driver. The memory driver 214 can then route the processing instructions and data to one or more of the enhanced memory modules 224A-B via the memory management unit 220 and the memory controller 222. This is just one example of a processing flow, and other processing flows are possible.
[0060] In some embodiments, if multiple enhanced memory modules (e.g., 224A-B) are used for the same application (e.g., ML / AI application 208), the accelerator runtime application 210 and the accelerator driver 216 can cooperate to manage data for the entire system. For example, the accelerator runtime application 210 can request the total available accelerator-allocated memory (and number of enhanced memory modules) from the accelerator driver 216, and then split the workload as efficiently as possible among the enhanced memory modules if a single enhanced memory module cannot handle the entire workload. As another example, for multiple workloads and multiple enhanced memory modules, the accelerator runtime application 210 can be configured to split the workloads to the available enhanced memory modules to maximize performance.
[0061] CPU 218 can generally be a central processing unit that includes one or more of its own processing cores. While in this example, the CPU 218 does not perform specific processing requests from the ML / AI application 208 that are routed to the enhanced memory modules 224A-B, in other examples, the CPU 218 can process requests from the ML / AI application 208 (and other applications) in parallel with the enhanced memory modules 224A-B.
[0062] Furthermore, while certain aspects of the processing system 200 are described with respect to this example, the processing system 200 can include other components. For example, the processing system 200 can include additional applications, additional hardware elements (e.g., reference Figure 10 other kinds of processing units described) and the like. Figure 2 Certain elements of the processing system 200 are of particular interest with respect to the described embodiments. Example operating modes of enhanced memory modules
[0063] Figures 3A-3C An example memory allocation configuration of an enhanced memory module (e.g., 102 in Figure 1 and 224A-B in Figure 2 is depicted.
[0064] In Figure 3A , the enhanced memory module 302 is configured by firmware in the flash memory 306 to use only its dedicated memory elements 310A-B for any processing by the computing core(s) 308, for example at startup time. In this configuration, the enhanced memory module 302 is set to provide maximum capacity to the system memory.
[0065] In Figure 3BIn this configuration, the enhanced memory module 302 is configured to utilize its dedicated memory elements 310A-B, in addition to the memory space provided by memory elements 312A-C for any processing of computing core(s) 308. On the other hand, the memory space provided by memory elements 312D-F is reserved for the host processing system, just as it would be for conventional system memory. Therefore, in this configuration, the enhanced memory module 302 is set to balance the computing power on the module with the memory requirements of the host processing system.
[0066] exist Figure 3C In this configuration, the enhanced memory module 302 is configured to utilize its dedicated memory elements 310A-B, in addition to the memory space provided by memory elements 312A-F for any processing of the computing core(s) 308. Therefore, in this configuration, the enhanced memory module 302 is set to maximize the computing power on the module. It is worth noting that in this configuration, another enhanced memory module or another conventional memory module can provide host-processor-addressable memory for host processing system operation.
[0067] It is worth noting that, although in this example, various memory elements are depicted as being allocated to compute core(s) 308 or host system memory, this representation is for convenience. For example, virtual memory address spaces can be allocated to on-memory compute or host system processing, regardless of how these address spaces are allocated to physical memory addresses within physical memory elements (e.g., via page tables).
[0068] also, Figures 3A-3C The example configurations depicted are merely some exemplary possibilities. The amount of host processing system memory space provided by the enhanced memory module 302 and allocated to compute core(s) 308 can be any ratio between 0 and 100% of the available host processing system addressable memory elements. However, in some embodiments, the ratio of memory space allocated to compute core(s) 308 can be limited such that at least some memory space on each enhanced memory module is reserved for the host system.
[0069] Example method for transferring data to the computing core memory element on the enhanced memory module.
[0070] Figure 4 A method for transferring data to an enhanced memory module (e.g., Figure 1 102 in Figure 2 224A-B and Figures 3A-3C Example method for computing core memory elements on (302) in the example.
[0071] Method 400 begins with step 402, in which a data command is received via the enhanced memory protocol to move data into a compute core memory element on the enhanced memory module. In some embodiments, an accelerator driver (such as discussed above with respect to Figure 2 ) can generate the command via the enhanced memory protocol.
[0072] Note that in this example, the compute core memory element can be a designated compute core memory unit, such as memory elements 310A-B in FIG. 3, or a dual-mode memory element assigned to a compute core, such as memory units 312A-C in FIG. 3. Further, method 400 is described with respect to moving data into and out of a memory element, but this is for ease of demonstration. Memory assigned to a compute core can be a range of memory addresses, a memory space, or any other logical and physical memory subdivision.
[0073] Method 400 then proceeds to step 404, in which the host processing system command is prevented from accessing the memory element on the enhanced memory module. This is to prevent any conflicts from occurring while the compute core on the enhanced memory module processes the data command.
[0074] In some embodiments, a memory driver (such as discussed above with respect to Figure 2 ) can be configured to prevent the host processing system command.
[0075] Method 400 then proceeds to step 406, in which a busy state is set on the data bus to prevent read requests from the host processing system while data is moved on the enhanced memory module. This is to prevent any interruptions from occurring while the compute core on the enhanced memory module processes the data command.
[0076] Method 400 then proceeds to step 408, in which data is transferred from a memory element assigned to the host processing system to a memory element assigned to a compute core on the enhanced memory module. As described above, the memory element assigned to the compute core can include a designated compute core memory element, such as memory elements 310A-B in FIG. 3, or can include a dual-mode memory element that has been assigned to a compute core, such as memory elements 312A-C in FIG. 3. Here, the data transferred into the compute core memory element can include data for processing as well as instructions for how to process the data. For example, the data can include data for processing (e.g., 206 in FIG. 2), model data (e.g., 204 in FIG. 2), and processing commands. Figure 2 Figure 2
[0077] Then, method 400 proceeds to step 410, where it is determined whether the data transfer is complete. If not, the method returns to step 408 and continues the data transfer.
[0078] If the transfer is completed in step 410, method 400 proceeds to step 412, where the busy state is cleared from the data bus.
[0079] Method 400 proceeds to step 414, wherein host processing system commands that are blocked from being allocated to memory elements of the host processing system are released, so that normal memory operations with the host processing system can continue.
[0080] Then, method 400 proceeds to step 416, in which data is processed using one or more computing cores of the enhanced memory module, such as the following regarding... Figure 6 In more detail. For example, processing data can include performing machine learning tasks, such as training or inference.
[0081] Example method for transferring data from the computing core memory element on the enhanced memory module.
[0082] Figure 5 Depicting the use of enhanced memory modules (e.g., Figure 1 102 in Figure 2 224A-B and Figures 3A-3C Example method 500 for computing core transfer data on (302) in the example. For example, method 500 can be used in... Figure 4 Step 416 is executed after data processing is complete.
[0083] Method 500 begins at step 502, where a result command is received via the enhanced memory protocol to move data out of the compute core memory element on the enhanced memory module. Note that in this example, the compute core memory element can be a designated compute core memory element, such as memory elements 310A-B in FIG3, or a dual-mode memory element already allocated to the compute core, such as memory elements 312A-C in FIG3.
[0084] Then, method 500 proceeds to step 504, in which host processing system commands are prevented from accessing memory elements on the enhanced memory module. This is to prevent any conflicts when data commands are processed on the enhanced memory module.
[0085] Then, method 500 proceeds to step 506, in which a busy state is set on the data bus to prevent read requests from the host processing system when data is moved on the enhanced memory module.
[0086] Method 500 then proceeds to step 508, where data is transferred from the compute core memory elements to the memory elements on the enhanced memory module that are allocated to the host processing system. As described above, the memory elements allocated to the compute core can include designated compute core memory elements, such as memory elements 310A-B in FIG. 3, or can include dual-mode memory elements that have been allocated to the compute core, such as memory elements 312A-C in FIG. 3.
[0087] Method 500 then proceeds to step 510, where it is determined whether the transfer is complete. If not, the method returns to step 508 and continues the data transfer.
[0088] If the transfer is complete at step 510, method 500 proceeds to step 512, where the busy status is cleared from the data bus.
[0089] Finally, method 500 proceeds to step 514, where the host processing system is unblocked from host processing system commands to the memory elements allocated to the host processing system, so that normal memory operations with the host processing system can continue.
[0090] Example method for processing data with compute cores on enhanced memory modules
[0091] Figure 6 An example method 600 is depicted for processing data with compute cores on enhanced memory modules. For example, processing data with compute cores on enhanced memory modules (e.g., Figure 1 102 in FIG. 1, Figure 2 224A-B in FIG. 2, and Figures 3A-3C 302 in FIG. 3) can proceed after transferring data to compute core memory elements on enhanced memory, such as described above with respect to Figure 4 The data for processing can be generated by an application, a sensor, other processor, etc. For example, the data for processing can include image, video, or sound data.
[0092] Method 600 begins at step 602, where a command is received from a requester to process data on an enhanced memory module. For example, the requester can be a data processing application, such as Figure 2 ML / AI application 208 in FIG. 2.
[0093] In some embodiments, the requester queries an accelerator runtime application (e.g., Figure 2 210 in FIG. 2) to determine how many enhanced memory modules are configured in the system, and how much memory on those enhanced memory modules is allocated to running accelerator-based workloads. The accelerator runtime application can then query an accelerator driver (e.g., Figure 2The requestor can then determine whether the workload is suitable within the configured amount of memory and execute the workload accordingly, with the information. In the event that the amount of available memory is insufficient to meet the workload, the requestor can reconfigure the workload, such as reducing the size of the workload or chunking the workload so that the available memory can be utilized.
[0094] Method 600 then proceeds to step 604, in which data is loaded into the compute core memory element(s) in the enhanced memory module. As described above, the compute core memory element(s) can be dedicated memory elements or dynamically allocated memory elements. In some embodiments, step 604 is performed in accordance with reference Figure 4 Method 400 described above.
[0095] In some embodiments, the requestor (e.g., ML / AI application 208 in FIG. 1) sends data to the accelerator runtime application (e.g., 210 in FIG. 1), which then forwards the data to the accelerator driver (e.g., 216 in FIG. 1). The accelerator driver can then send a protocol command to load the data to the memory driver (e.g., 214 in FIG. 1), and the memory driver then sends the command to the enhanced memory module (e.g., 224A-B in FIG. 1). Figure 2 Figure 2 In some embodiments, the requestor (e.g., ML / AI application 208 in FIG. 1) sends data to the accelerator runtime application (e.g., 210 in FIG. 1), which then forwards the data to the accelerator driver (e.g., 216 in FIG. 1). The accelerator driver can then send a protocol command to load the data to the memory driver (e.g., 214 in FIG. 1), and the memory driver then sends the command to the enhanced memory module (e.g., 224A-B in FIG. 1). Figure 2 Figure 2 In some embodiments, the requestor (e.g., ML / AI application 208 in FIG. 1) sends data to the accelerator runtime application (e.g., 210 in FIG. 1), which then forwards the data to the accelerator driver (e.g., 216 in FIG. 1). The accelerator driver can then send a protocol command to load the data to the memory driver (e.g., 214 in FIG. 1), and the memory driver then sends the command to the enhanced memory module (e.g., 224A-B in FIG. 1). Figure 2
[0096] Method 600 then proceeds to step 606, in which a processing command is received via the enhanced memory protocol, such as described in more detail below. The processing command can be configured to cause the processing core(s) of the enhanced memory module to perform data processing on the data stored in the compute core memory element.
[0097] In some embodiments, after the data is loaded, the accelerator driver (e.g., 216 in FIG. 1) sends a protocol command to the memory driver (e.g., 214 in FIG. 1) to process the data, which then forwards the command to the enhanced memory module (e.g., 224A-B in FIG. 1) for execution. Figure 2 Figure 2 In some embodiments, after the data is loaded, the accelerator driver (e.g., 216 in FIG. 1) sends a protocol command to the memory driver (e.g., 214 in FIG. 1) to process the data, which then forwards the command to the enhanced memory module (e.g., 224A-B in FIG. 1) for execution. Figure 2
[0098] Method 600 then proceeds to step 608, where data stored in the computing core memory element is processed by one or more computing cores on the enhanced memory module. In some embodiments, the processing may involve machine learning and / or artificial intelligence tasks, such as training or inference. Furthermore, in some embodiments, machine learning tasks may be processed in parallel across multiple computing cores on a single enhanced memory module, and across multiple enhanced memory modules.
[0099] Then, method 600 proceeds to step 610, where if the processing of the workload has not yet been completed, method 600 returns to step 608.
[0100] If the workload processing is completed in step 610, then method 600 proceeds to step 612, in which a result signal is sent via an enhanced memory protocol.
[0101] In some embodiments, the accelerator driver (e.g., Figure 2 216 in the middle) can periodically query the enhanced memory module (e.g., Figure 3B The enhanced memory module (224A-B) determines when processing is complete. When processing is complete, the enhanced memory module can output a signal on the bus indicating that processing is complete. The accelerator driver will then observe the result signal the next time it queries the enhanced memory module.
[0102] Method 600 then proceeds to step 612, in which the processed data is moved from the computing core memory element to the memory element allocated to the host processing system on the enhanced memory module. For example, regarding Figure 2 Data can be moved from one or more of the memory elements 310A-B and 312A-C allocated to the computing core to one or more of the memory cells 312D-F allocated to the host processing system.
[0103] In some embodiments, after processing is complete, the accelerator driver (e.g., Figure 2 216 in the middle) will send to the memory drive (e.g., Figure 2 (214) Sends a protocol command to retrieve data. The memory driver then passes the command to the enhanced memory module, which in turn copies the processed data to the host processing system's addressable memory.
[0104] Method 600 then ends in step 614, wherein the processed data is provided to the requester from the memory element on the enhanced memory module allocated to the host processing system.
[0105] In some embodiments, the accelerator driver (e.g., Figure 2 216 in the middle) will send to the memory drive (e.g.,Figure 2 (214) sends a command to read the processed data. The memory driver then forwards these read requests to the enhanced memory module. The accelerator driver will then use the accelerator runtime application (e.g., Figure 1 (210) These results are sent back to the requester.
[0106] Example aspects of enhanced memory protocols
[0107] Enhanced memory protocols can be implemented to work with enhanced memory modules (e.g., Figure 2 102 or Figures 3A-3C 224A-B or Figure 1 The 302) interaction. In some embodiments, the enhanced memory protocol may be based on the Joint Electronic Devices Engineering Committee (JEDEC) protocol or its extensions.
[0108] In some embodiments, the enhanced memory protocol can be implemented for use with memory-integrated accelerators (such as reference accelerators). Figure 1 The description includes various commands for the accelerator 104 interface connection.
[0109] For example, an enhanced memory protocol may include an "initialization" signal (or command) configured to initiate an initialization from, for example, a flash memory module (such as those mentioned above). Figure 1 The described module 106) loads the firmware and initializes the memory integrated accelerator.
[0110] The enhanced memory protocol may also include a “data” signal (or command) configured to initiate the transfer of data from the enhanced memory module (e.g., Figure 2 102 in Figures 3A-3C 224A-B and Figure 3B The allocation of system memory (e.g., in 302) on the memory (e.g., Figure 3B The memory element (312D-F) is moved to the computing core memory element (e.g., Figure 4 The process of 312A-C in the middle, such as through the process of 312A-C in the middle, Figure 1 The method described can be static or dynamic, as described above.
[0111] The “data” signal can have variations such as: “move all data”, where all data between the start and end addresses is moved between system-allocated memory and compute core-allocated memory; “move x bytes”, where x data bytes are moved from system-allocated memory to compute core-allocated memory; and “move x cycles”, where x cycles will be used to move data from system-allocated memory to compute core-allocated memory.
[0112] The enhanced memory protocol can also include a "process" signal (or command) configured to cause a compute core of the enhanced memory module (e.g., 102 in Figure 2 , Figures 3A-3C 224A-B in Figure 6 302) to begin executing instructions that have been stored in memory allocated to the compute core, such as by the methods described with respect to Figure 3B .
[0113] The enhanced memory protocol can also include a "result" signal (or command) configured to initiate processing to move data from memory allocated to the compute core (e.g., 312A-C in Figure 3B to memory allocated to the system (e.g., 312D-F in Figure 2 ), so that it can be accessed by other parts of the host processing system (e.g., system 200 in Figure 5 ), such as by the methods described with respect to Figure 2 .
[0114] Accordingly, in one example, the accelerator driver (e.g., 216 in Figure 2 ) sends a "data" signal, and the memory driver (e.g., 214 in Figure 2 ) forwards that signal to the enhanced memory module (e.g., 224A-B in Figure 7 ). The accelerator driver informs the memory driver to block traffic, which can be a single command with the "data" command, or a separate command. The memory driver then blocks all traffic that is not from the accelerator driver. In some cases, traffic can be held or sent back for retry. The enhanced memory module can also set a value in a buffer to output a data pattern indicating the status of the memory transfer. For example, the accelerator driver can request data from the enhanced memory module (the memory driver forwards the data), and if the data pattern is "busy", the accelerator driver waits and retries, and if the data pattern is "complete", the accelerator driver informs the memory driver to again allow traffic.
[0115] Example method for constructing an enhanced memory module
[0116] Figure 1 An example method 700 is depicted for constructing an enhanced memory module (e.g., 102 in Figure 2 , Figures 3A-3C 224A-B in Figure 1 302).
[0117] Method 700 begins with step 702, where dual mode memory elements are placed on an enhanced memory module, such as memory elements 112A-F on enhanced memory module 102 in Figure 1 In this example, "dual mode" refers to the ability of each memory element to be used in a conventional manner for system memory as well as for local memory of a compute core(s) on the module, such as 108 in Figure 1 The memory elements do not require physical modification to work in either mode.
[0118] Method 700 then proceeds to step 704, where compute cores are placed on the enhanced memory module, such as compute core(s) 108 in Figure 1
[0119] Method 700 then proceeds to step 706, where command / address lines are connected between the compute core(s) and each dual mode memory element. The command / address lines can be configured to transmit signals that determine which portions of memory are accessed and what the command is.
[0120] Method 700 then proceeds to step 708, where data lines are connected between each dual mode memory element and the compute core(s) (for local processing mode) as well as the memory buffer (for system memory mode).
[0121] Method 700 then proceeds to step 710, where compute core memory elements are placed on the enhanced memory module, such as memory elements 110A-B of enhanced memory module 102 in Figure 1
[0122] Method 700 then proceeds to step 712, where command / address lines and data lines are connected between the compute core(s) and the compute core memory elements.
[0123] Method 700 then proceeds to step 714, where a flash memory module is placed on the enhanced memory module, such as flash memory module 106 of enhanced memory module 102 in Figure 1
[0124] Method 700 then proceeds to step 716, where the flash memory module is connected to the compute core(s), such as compute core(s) 108 of enhanced memory module 102 in Figure 1
[0125] Method 700 then proceeds to step 718, where firmware is written to the flash memory module to initialize the enhanced memory module, such as Figure 1 enhanced memory module 102.
[0126] Notably, the method 700 is just one high-level example of how to construct an enhanced memory module (such as Figure 8 enhanced memory module 102) and there are many other examples. Discussion of specific manufacturing techniques is omitted for clarity.
[0127] Example memory map of a processing system including an enhanced memory module
[0128] Figure 1 An example memory map of a processing system including an enhanced memory module (e.g., 102 in Figure 2 224A-B in Figures 3A-3C 302 in Figure 1 is depicted.
[0129] In the depicted example, the processing system has six memory module slots (e.g., 6 DIMM slots) and each memory module in each DIMM slot has 8 GB of system addressable memory, for a total of 48 GB of system memory. The memory modules in each DIMM slot can be regular memory modules or enhanced memory modules, such that the system as a whole can be a “hybrid configuration” with a mix of regular and enhanced memory modules, or an “enhanced configuration” with all enhanced memory modules. Because enhanced memory modules can be configured to allocate memory between the host processing system and the on-die processing elements (e.g., to computing cores such as 108 in Figure 1 the more enhanced memory that is present in the system, the more total memory that is available to be allocated to on-die processing.
[0130] Further, in this example, a memory allocation threshold is implemented such that up to half of the memory in any enhanced memory module can be allocated to on-die processing (e.g., to computing cores such as 108 in Figure 1 and the remaining memory is allocated to the host processing system. In other embodiments, the memory allocation threshold can not be implemented, or can be implemented with different values. The memory threshold can generally be established to ensure that the host processing system retains enough allocated memory to execute properly.
[0131] Memory maps 802 or 804 depict example memory allocations in a processing system with a hybrid configuration of memory modules. Memory maps 806 and 808 depict example memory allocations in a processing system with an enhanced configuration, where all memory modules are enhanced memory modules. In the depicted examples, the memory allocations in the enhanced configurations of 806 and 808 include additional memory allocated to the compute cores for on-memory processing.
[0132] Furthermore, the memory allocation in memory maps 806 and 808 is specific to the computational cores of each enhanced memory module (e.g., enhanced memory modules 0-5). However, in other examples, memory may generally be allocated to computational cores as a group, such as in example memory maps 802 and 804, rather than specifically, such as in example memory maps 806 and 808.
[0133] Generally speaking, having more enhanced memory modules increases the amount of memory that can be allocated to on-memory processing, thereby improving the host processing system's ability to perform AI and ML-specific tasks such as training and inference, and increasing the flexibility in configuring the total system memory.
[0134] In the case of dynamic memory reallocation between the host processing system and on-memory processing (e.g., via... Figure 8 The accelerator 104 can update the memory mapping accordingly. For example, memory mapping 802 can be updated to memory mapping 806 or other combinations.
[0135] Figure 9 The examples depicted are intended to demonstrate several examples of various memory configurations using the enhanced memory module, and other examples are also possible.
[0136] Example methods for processing data using enhanced memory modules
[0137] Figure 1 Describing the use of computing cores (e.g., Figure 1 Enhanced memory modules (e.g., 108) Figure 2 102 in Figures 3A-3C 224A-B in Figure 1 Example method 900 for processing data (302) in the example. In some embodiments, the computing core may be part of a memory-integrated accelerator, such as Figures 3A-3C 104. In some implementations, the computing core may include multiple processing elements, such as multiple individual processing cores. In some embodiments, the enhanced memory module includes a load-reduced dual in-line memory module (LRDIMM).
[0138] Method 900 begins at step 902 with initializing the enhanced memory module by allocating a first subset of the memory of the enhanced memory module as host processing system addressable memory and a second subset of the memory of the enhanced memory module as compute core addressable memory, such as described with reference to Figure 1 .
[0139] In some embodiments of method 900, the initialization of the enhanced memory module can be performed at startup time. In other embodiments, the initialization of the enhanced memory module can be performed dynamically during operation of the host processing system.
[0140] In some embodiments, the allocation of memory can include allocating memory of one or more dual-mode memory elements (such as memory elements 112A-F in FIG. 1) between the compute core and the host processing system. For example, the allocation can include one or more memory ranges or memory address spaces, such as virtual address spaces or physical address spaces. Figure 1
[0141] Method 900 proceeds to step 904 with receiving, at the enhanced memory module, data from the host processing system. The data can be provided by the host processing system to the enhanced memory module for processing on the enhanced memory module, such as by the compute core.
[0142] In some embodiments, the received data can include data for processing by a machine learning model and machine learning model parameters. In some embodiments, the received data can also include processing commands or instructions.
[0143] Method 900 then proceeds to step 906 with storing the received data in the first subset of memory (host processing system addressable memory). In some embodiments, the host processing system addressable memory includes the memory space allocated to the host processing system in one or more dual-mode memory elements (such as memory elements 112A-F in FIG. 1) on the enhanced memory module. Figure 1
[0144] Method 900 then proceeds to step 908 with transferring the received data from the first subset of memory (host processing system addressable memory) to the second subset of memory (compute core addressable memory).
[0145] In some embodiments, the compute core addressable memory includes non-host processing system addressable memory, such as memory elements 110A-B in FIG. 1. In some embodiments, the compute core addressable memory includes only non-host processing system addressable memory, such as in memory elements 110A-B in FIG. 1. Figure 3A Figure 2 in the example of FIG. 1).
[0146] The method 900 then proceeds to step 910, where the received data is processed using the compute core on the enhanced memory module to generate processed data.
[0147] The method 900 then proceeds to step 912, where the processed data is transferred from the second subset of memory (compute core addressable memory) to the first subset of memory (host processing system addressable memory).
[0148] The method 900 then proceeds to step 914, where the processed data is provided to the host processing system via the first subset of memory (host processing system addressable memory), such as described above with respect to Figure 1 .
[0149] In some embodiments, initializing the enhanced memory module includes processing firmware instructions stored in a flash memory module on the enhanced memory, such as described above with respect to Figure 4 .
[0150] In some embodiments, the method 900 further includes setting a busy state on a data bus of the host processing system prior to transferring the data from the first subset of memory (host processing system addressable memory) to the second subset of memory (compute core addressable memory) to indicate that the data transfer is in progress and to block read requests to the host processing system addressable memory, such as described above with respect to Figure 4 .
[0151] In some embodiments, the method 900 further includes setting an available state on a data bus of the host processing system after transferring the data from the first subset of memory (host processing system addressable memory) to the second subset of memory (compute core addressable memory) to allow read or write requests from the host processing system to the host processing system addressable memory, such as described above with respect to Figure 4 .
[0152] In some embodiments, the method 900 further includes enabling a host processing system memory command block for the enhanced memory module prior to transferring the data from the first subset of memory (host processing system addressable memory) to the second subset of memory (compute core addressable memory), such as described above with respect to Figure 4 .
[0153] In some embodiments, the method 900 further includes disabling a host processing system memory command block for the enhanced memory module after transferring the data from the first subset of memory (host processing system addressable memory) to the second subset of memory (compute core addressable memory), such as described above with respect toFigure 3A The.
[0154] In some embodiments, the method 900 further includes reallocating memory of the enhanced memory module by de-allocating the first subset of memory and the second subset of memory and allocating a third subset of memory of the one or more dual-mode memory elements as host processing system addressable memory and a fourth subset of memory of the one or more dual-mode memory elements as compute core addressable memory, where the third subset of memory is different than the first subset, and where the fourth subset of memory is different than the second subset. In some embodiments, the reallocating is performed after a boot time, such as when the host processing system is running. Thus, the reallocating can be referred to as a dynamic or "hot" reallocation of memory.
[0155] For example, in some embodiments, the first allocation (such as Figure 3A the allocation) can be dynamically reallocated to a second allocation (such as Figure 3B , Figure 3C or Figure 8 the allocation) or other unillustrated allocation. In some embodiments, during the reallocation process, the host processing system can be prevented from sending memory commands to the enhanced memory module. In some embodiments, a page table or memory map can be updated after the reallocation so that the host processing system is properly configured for physical memory that is no longer accessible, or newly accessible. For example, Figure 10 the memory map example in may be updated after the reallocation of memory on the enhanced memory module.
[0156] Example electronic device for data processing with an enhanced memory module
[0157] Figures 4-6 An example electronic device 1000 is depicted, which can be configured to perform data processing with an enhanced memory module, such as described herein with respect to Figure 9 and Figure 1 In some embodiments, the electronic device 1000 can comprise a server computer.
[0158] The electronic device 1000 includes a central processing unit (CPU) 1002, which in some examples can be a multi-core CPU. Instructions executed at the CPU 1002 can be loaded, for example, from a program memory associated with the CPU 1002, or can be loaded from a memory partition 1024.
[0159] The electronic device 1000 also includes additional processing components that are customized for particular functions, such as a graphics processing unit (GPU) 1004, a digital signal processor (DSP) 1006, and a neural processing unit (NPU) 1008.
[0160] An NPU, such as 1008, is generally a specialized circuit that is configured to implement all of the control and arithmetic logic needed to perform machine learning algorithms, such as algorithms for processing artificial neural networks (ANN), deep neural networks (DNN), random forests (RF), and the like. An NPU can also sometimes be alternatively referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligent processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.
[0161] An NPU can be optimized for training or inference, or in some cases, configured to balance performance between the two. For an NPU capable of performing both training and inference, the two tasks can still typically be performed independently.
[0162] For example, an NPU, such as 1008, can be configured to accelerate the performance of common machine learning inference tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, multiple NPUs can be instantiated on a single chip, such as a system on a chip (SoC), while in other examples, they can be part of a specialized neural network accelerator. An NPU designed to accelerate inference is typically configured to operate on a complete model. Such an NPU can thus be configured to input new data and quickly process it through an already trained model to generate a model output (e.g., inference).
[0163] As another example, an NPU can be configured to accelerate common machine learning training tasks, such as processing a test dataset (typically labeled or tagged), iterating over the dataset, and then adjusting model parameters, such as weights and biases, to improve model performance. Typically, the optimization that is performed based on the error predictions involves propagating back through the layers of the model and determining gradients to reduce the prediction error.
[0164] In one implementation, the NPU 1008 is part of one or more of the CPU 1002, GPU 1004, and / or DSP 1006.
[0165] The electronic device 1000 can also include one or more input and / or output devices 1022, such as a screen, a network interface, physical buttons, and the like.
[0166] The electronic device 1000 also includes a memory 1024, which represents one or more static and / or dynamic memories, such as dynamic random access memory, static memory based on flash, and the like. In this example, the memory 1024 includes computer-executable components that can be executed by one or more of the above-described processors of the electronic device 1000.
[0167] The memory 1024 can represent one or more memory modules, such as the enhanced memory modules described herein (e.g., with respect to Figure 2 , Figure 10 and FIG. 3). For example, the memory 1024 can include computing core(s) 1024I and flash memory components 1024J. In some cases, the computing core(s) 1024I can include a NPU or other kind of processor described herein. Moreover, although not shown, some of the memory 1024 can be accessible only by the computing core(s) 1024J, while some memory can be configured to be accessible by the processing system 1000 or the computing core 1024J.
[0168] In this example, the memory 1024 includes a receiving component 1024A, a storing component 1024B, a transferring component 1024C, a processing component 1024D, a sending (or providing) component 1024E, an initializing component 1024F, an allocating component 1024G, and a splitting component 1024H. The depicted components and other, non-depicted components can be configured to perform various aspects of the methods described herein.
[0169] Although not shown in , the electronic device 1000 can also include one or more data buses for transferring data between various aspects of the electronic device 1000.
[0170] Example Clauses
[0171] Implementation is described in the following numbered clauses:
[0172] Clause 1 : An enhanced memory module comprising: a computing core; and one or more dual-mode memory elements, wherein the enhanced memory module is configured to: allocate a first subset of memory in the one or more dual-mode memory elements as host processing system addressable memory and a second subset of memory in the one or more dual-mode memory elements as computing core addressable memory; receive data from a host processing system; process the data with the computing core to generate processed data; and provide the processed data to the host processing system via the first subset of memory.
[0173] Clause 2: The enhanced memory module of clause 1, wherein the enhanced memory module is further configured to: store the received data in the first subset of memory; transfer the received data from the first subset of memory to the second subset of memory prior to processing the data with the compute core; and transfer the processed data from the second subset of memory to the first subset of memory after processing the data with the compute core.
[0174] Clause 3: The enhanced memory module of any of clauses 1-2, wherein the second subset of memory further comprises one or more single mode memories configured to be used by the compute core.
[0175] Clause 4: The enhanced memory module of any of clauses 1-3, further comprising: a flash memory module comprising firmware instructions for allocating the first subset of memory and the second subset of memory.
[0176] Clause 5: The enhanced memory module of any of clauses 1-4, wherein the enhanced memory module is further configured to: set a busy state on a data bus of the host processing system prior to transferring the data from the first subset of memory to the second subset of memory to indicate that a data transfer is in progress.
[0177] Clause 6: The enhanced memory module of clause 5, wherein the enhanced memory module is further configured to: set an available state on the data bus of the host processing system after transferring the data from the first subset of memory to the second subset of memory.
[0178] Clause 7: The enhanced memory module of any of clauses 1-6, wherein the enhanced memory module is further configured to: enable a host processing system memory command block for the enhanced memory module prior to transferring the data from the first subset of memory to the second subset of memory.
[0179] Clause 8: The enhanced memory module of clause 7, wherein the enhanced memory module is further configured to: disable the host processing system memory command block for the enhanced memory module after transferring the data from the first subset of memory to the second subset of memory.
[0180] Clause 9: The enhanced memory module of any of clauses 1-8, wherein the data comprises: data for processing by a machine learning model; and machine learning model parameters.
[0181] Clause 10: The enhanced memory module of any one of clauses 1-9, wherein the enhanced memory module comprises a dual in-line memory module (DIMM).
[0182] Clause 11 : The enhanced memory module of any one of clauses 1-10, wherein the enhanced memory module is further configured to: de-allocate the first subset of memory and the second subset of memory; and allocate a third subset of memory in the one or more dual-mode memory elements as host processing system addressable memory and a fourth subset of memory in the one or more dual-mode memory elements as compute core addressable memory, wherein the third subset of memory is different from the first subset, and wherein the fourth subset of memory is different from the second subset.
[0183] Clause 12: A method for processing data with an enhanced memory module comprising a compute core, comprising: initializing the enhanced memory module by allocating a first subset of memory of the enhanced memory module as host processing system addressable memory and a second subset of memory of the enhanced memory module as compute core addressable memory; receiving, at the enhanced memory module, data from a host processing system; processing the data with the compute core on the enhanced memory module to generate processed data; and providing the processed data to the host processing system via the first subset of memory.
[0184] Clause 13: The method of clause 12, further comprising: storing the received data in the first subset of memory; transferring the received data from the first subset of memory to the second subset of memory prior to processing the data with the compute core; and transferring the processed data from the second subset of memory to the first subset of memory after processing the data with the compute core.
[0185] Clause 14: The method of any one of clauses 12-13, wherein the first subset of memory and the second subset of memory are associated with one or more dual-mode memory elements.
[0186] Clause 15: The method of any one of clauses 12-14, wherein initializing the enhanced memory module comprises processing firmware instructions stored in a flash memory module on the enhanced memory module.
[0187] Clause 16: The method of any of clauses 12-15, wherein the second subset of memory is also associated with one or more single-mode memory elements configured to be used by the compute core.
[0188] Clause 17: The method of clause 13, further comprising setting a busy state on a data bus of the host processing system prior to transferring the data from the first subset of memory to the second subset of memory to indicate that a data transfer is in progress.
[0189] Clause 18: The method of clause 17, further comprising setting an available state on the data bus of the host processing system after transferring the data from the first subset of memory to the second subset of memory.
[0190] Clause 19: The method of clause 13, further comprising enabling a host processing system memory command block for the enhanced memory module prior to transferring the data from the first subset of memory to the second subset of memory.
[0191] Clause 20: The method of clause 19, further comprising disabling the host processing system memory command block for the enhanced memory module after transferring the data from the first subset of memory to the second subset of memory.
[0192] Clause 21: The method of any of clauses 12-20, wherein the data comprises: data for processing by a machine learning model; and machine learning model parameters.
[0193] Clause 22: The method of any of clauses 12-21, wherein the enhanced memory module comprises a dual in-line memory module (DIMM).
[0194] Clause 23: The method of any of clauses 12-22, further comprising de-allocating the first subset of memory and the second subset of memory; and allocating a third subset of memory of the one or more dual-mode memory elements as host processing system addressable memory and a fourth subset of memory of the one or more dual-mode memory elements as compute core addressable memory, wherein the third subset of memory is different from the first subset, and wherein the fourth subset of memory is different from the second subset.
[0195] Clause 24: A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a host processing system, cause the host processing system to perform a method for processing data with an enhanced memory module comprising a compute core, the method comprising: initializing the enhanced memory module by allocating a first subset of memory of the enhanced memory module as host processing system addressable memory and a second subset of memory of the enhanced memory module as compute core addressable memory; receiving, at the enhanced memory module, data from the host processing system; processing the data with the compute core on the enhanced memory module to generate processed data; and providing the processed data to the host processing system via the first subset of memory.
[0196] Clause 25: The non-transitory computer-readable medium of clause 24, wherein the method further comprises: storing the received data in the first subset of memory; transferring the received data from the first subset of memory to the second subset of memory prior to processing the data with the compute core; and transferring the processed data from the second subset of memory to the first subset of memory after processing the data with the compute core.
[0197] Clause 26: The non-transitory computer-readable medium of any of clauses 24-25, wherein the first subset of memory and the second subset of memory are associated with one or more dual-mode memory elements.
[0198] Clause 27: The non-transitory computer-readable medium of any of clauses 24-26, wherein the second subset of memory is further associated with one or more single-mode memory elements configured to be used by the compute core.
[0199] Clause 28: The non-transitory computer-readable medium of clause 25, wherein the method further comprises: setting a busy state on a data bus of the host processing system prior to transferring the data from the first subset of memory to the second subset of memory to indicate that a data transfer is in progress.
[0200] Clause 29: The non-transitory computer-readable medium of clause 29, wherein the method further comprises: setting an available state on the data bus of the host processing system after transferring the data from the first subset of memory to the second subset of memory to allow.
[0201] Clause 30: The non-transitory computer-readable medium of clause 25, wherein the method further comprises: enabling a host processing system memory command block for the enhanced memory module prior to transferring the data from the first subset of memory to the second subset of memory; and disabling the host processing system memory command block for the enhanced memory module after transferring the data from the first subset of memory to the second subset of memory.
[0202] Other Considerations
[0203] The previous description is provided to enable any person skilled in the art to practice the various embodiments described herein. The examples discussed herein are not limiting of the scope, applicability, or embodiments set forth in the claims. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments. For example, changes can be made in the function and arrangement of elements discussed without departing from the scope of the disclosure. Various examples can omit, substitute, or add various procedures or components as appropriate. For instance, the methods described can be performed in an order different from that described, and / or various steps can be added, omitted, or combined. Also, features described with respect to some examples can be combined in some other examples. For example, a device or method can be implemented using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover devices or methods which achieve the same process, using other structure, functionality, or structure and functionality. It is understood that any of the aspects of the disclosure disclosed herein can be embodied by one or more elements of a claim.
[0204] As used herein, the term “exemplary” means “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.
[0205] As used herein, the phrase referring to “at least one of’ a list of items means any combination of those items, including single members. For example, “at least one of a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination of items from a, b, and c (for example, a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c).
[0206] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” can include calculating, computing, processing, deriving, investigating, looking up (such as, for example, looking up in a table, a database or another data structure), ascertaining and the like. Also, “determining” can include receiving (such as, for example, receiving information), accessing (such as, for example, accessing data in a memory) and the like. Also, “determining” can include resolving, selecting, choosing, establishing and the like.
[0207] The methods disclosed herein comprise one or more steps or actions for achieving the methods. The method steps and / or actions can be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order is specified, the order or use of terms can be modified without departing from the scope of the claims. Additionally, the various operations can be performed by any suitable means capable of performing the corresponding functions. The means can include various hardware and / or software component(s) and / or module(s), including, but not limited to a circuit, an application specific integrated circuit (ASIC), or processor. Generally, where there are operations illustrated in figures, those operations can have corresponding counterpart means-plus-function components with similar numbering.
[0208] The following claims are not to be limited to the embodiments shown herein, but are to be accorded the full scope consistent with the language of the claims. In one claim, the singular forms “a,” “an,” and “the” are intended to mean “one or more” unless the context clearly indicates otherwise. The term “some” means one or more unless the context clearly indicates otherwise. None of the claims are intended to invoke 35 U.S.C. § 112(f) unless the exact phrase “means for” is recited in the claim. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether these disclosure is explicitly recited in the claims.
Claims
1. An enhanced memory module, comprising: Computing core; as well as One or more dual-mode memory elements, each of which can be dynamically allocated to the computing core for on-memory processing or allocated to the host processing system. The enhanced memory module is configured as follows: Receive at least a portion of the workload, wherein the portion of the workload is determined by the accelerator module based on the capacity of the enhanced memory module and the total number of enhanced memory modules used for the workload; Based on the received portion of the workload and dynamic configuration information from the host processing system, a first subset of the memory in the one or more dual-mode memory elements is allocated as host processing system addressable memory, and a second subset of the memory in the one or more dual-mode memory elements is allocated as compute core addressable memory. Receive data from the host processing system; The data is processed using the computing core to generate processed data; as well as The processed data is provided to the host processing system via the first subset of the memory.
2. The enhanced memory module according to claim 1, wherein the enhanced memory module is further configured to: The received data is stored in the first subset of the memory; Before processing the data using the computing core, the received data is transferred from the first subset of the memory to the second subset of the memory; as well as After processing the data using the computing core, the processed data is transferred from the second subset of the memory to the first subset of the memory.
3. The enhanced memory module of claim 1, wherein the second subset of the memory further includes one or more single-mode memories configured for use by the computing core.
4. The enhanced memory module according to claim 1, further comprising: The flash memory module includes firmware instructions for allocating a first subset of memory and a second subset of memory.
5. The enhanced memory module of claim 1, wherein the enhanced memory module is further configured to: set a busy state on the data bus of the host processing system to indicate that a data transfer is in progress before transferring the data from the first subset of the memory to the second subset of the memory.
6. The enhanced memory module of claim 5, wherein the enhanced memory module is further configured to: after transferring the data from the first subset of the memory to the second subset of the memory, set an available state on the data bus of the host processing system.
7. The enhanced memory module of claim 1, wherein the enhanced memory module is further configured to: enable a host processing system memory command block for the enhanced memory module before transferring the data from the first subset of the memory to the second subset of the memory.
8. The enhanced memory module of claim 7, wherein the enhanced memory module is further configured to: disable the host processing system memory command block for the enhanced memory module after transferring the data from the first subset of the memory to the second subset of the memory.
9. The enhanced memory module of claim 1, wherein the data includes: Data used by machine learning models; as well as Machine learning model parameters.
10. The enhanced memory module of claim 1, wherein the enhanced memory module comprises a dual in-line memory module (DIMM).
11. The enhanced memory module of claim 1, wherein the enhanced memory module is further configured to: Release the first subset of memory and the second subset of memory; and A third subset of the memory in the one or more dual-mode memory elements is allocated as host processing system addressable memory, and a fourth subset of the memory in the one or more dual-mode memory elements is allocated as computing core addressable memory. The third subset of the memory is different from the first subset, and The fourth subset of the memory is different from the second subset.
12. A method for processing data using an enhanced memory module, the enhanced memory module including a computing core and one or more dual-mode memory elements, each of the one or more dual-mode memory elements being dynamically allocated to the computing core for on-memory processing or allocated to a host processing system, the method comprising: Receive at least a portion of the workload, wherein the portion of the workload is determined by the accelerator module based on the capacity of the enhanced memory module and the total number of enhanced memory modules used for the workload; Based on the received portion of the workload and the dynamic configuration information of the host processing system, the enhanced memory module is initialized by allocating a first subset of the memory in the one or more dual-mode memory elements as host processing system addressable memory and a second subset of the memory in the one or more dual-mode memory elements as compute core addressable memory. At the enhanced memory module, data from the host processing system is received; The data is processed using the computing core on the enhanced memory module to generate processed data; as well as The processed data is provided to the host processing system via the first subset of the memory.
13. The method of claim 12, further comprising: The received data is stored in the first subset of the memory; Before processing the data using the computing core, the received data is transferred from the first subset of the memory to the second subset of the memory; as well as After processing the data using the computing core, the processed data is transferred from the second subset of the memory to the first subset of the memory.
14. The method of claim 12, wherein the first subset of the memory and the second subset of the memory are associated with one or more dual-mode memory elements.
15. The method of claim 12, wherein initializing the enhanced memory module comprises: Process firmware instructions stored in the flash memory module on the enhanced memory module.
16. The method of claim 12, wherein the second subset of the memory is further associated with one or more single-mode memory elements configured for use by the computing core.
17. The method of claim 13, further comprising: Before transferring the data from the first subset of memory to the second subset of memory, a busy state is set on the data bus of the host processing system to indicate that the data transfer is in progress.
18. The method of claim 17, further comprising: After the data is transferred from the first subset of memory to the second subset of memory, an availability state is set on the data bus of the host processing system.
19. The method of claim 13, further comprising: Before transferring the data from the first subset of the memory to the second subset of the memory, the host processing system memory command block for the enhanced memory module is enabled.
20. The method of claim 19, further comprising: After the data is transferred from the first subset of the memory to the second subset of the memory, the host processing system memory command block for the enhanced memory module is disabled.
21. The method of claim 12, wherein the data comprises: Data used for processing by machine learning models; and Machine learning model parameters.
22. The method of claim 12, wherein the enhanced memory module comprises a dual in-line memory module (DIMM).
23. The method of claim 12, further comprising: Release the first subset of memory and the second subset of memory; as well as A third subset of the memory of the enhanced memory module is allocated as host processing system addressable memory, and a fourth subset of the memory of the enhanced memory module is allocated as computing core addressable memory. The third subset of the memory is different from the first subset, and The fourth subset of the memory is different from the second subset.
24. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a host processing system, cause the host processing system to perform a method for processing data using an enhanced memory module, the enhanced memory module including a computing core and one or more dual-mode memory elements, each of the one or more dual-mode memory elements being dynamically allocated to the computing core for on-memory processing or allocated to the host processing system, the method comprising: Receive at least a portion of the workload, wherein the portion of the workload is determined by the accelerator module based on the capacity of the enhanced memory module and the total number of enhanced memory modules used for the workload; Based on the received portion of the workload and the dynamic configuration information received from the host processing system, the enhanced memory module is initialized by allocating a first subset of the memory in the one or more dual-mode memory elements as host processing system addressable memory and a second subset of the memory in the one or more dual-mode memory elements as compute core addressable memory. At the enhanced memory module, data from the host processing system is received; The data is processed using the computing core on the enhanced memory module to generate processed data; as well as The processed data is provided to the host processing system via the first subset of the memory.
25. The non-transitory computer-readable medium of claim 24, wherein the method further comprises: The received data is stored in the first subset of the memory; Before processing the data using the computing core, the received data is transferred from the first subset of the memory to the second subset of the memory; as well as After processing the data using the computing core, the processed data is transferred from the second subset of the memory to the first subset of the memory.
26. The non-transitory computer-readable medium of claim 24, wherein the first subset of the memory and the second subset of the memory are associated with one or more dual-mode memory elements.
27. The non-transitory computer-readable medium of claim 24, wherein the second subset of the memory is further associated with one or more single-mode memory elements configured for use by the computing core.
28. The non-transitory computer-readable medium of claim 25, wherein the method further comprises: Before transferring the data from the first subset of memory to the second subset of memory, a busy state is set on the data bus of the host processing system to indicate that the data transfer is in progress.
29. The non-transitory computer-readable medium of claim 28, wherein the method further comprises: After the data is transferred from the first subset of memory to the second subset of memory, an availability state is set on the data bus of the host processing system to allow it.
30. The non-transitory computer-readable medium of claim 25, wherein the method further comprises: Before transferring the data from the first subset of the memory to the second subset of the memory, enable the host processing system memory command block for the enhanced memory module; as well as After the data is transferred from the first subset of the memory to the second subset of the memory, the host processing system memory command block for the enhanced memory module is disabled.
Citation Information
Patent Citations
Apparatuses and methods for in-memory operations
US20190115063A1