Fine-grained data mover
By implementing the fine-grained data mover component on the memory controller, processing the data mover call and performing the data mover task, the problem of inefficiency of distributed memory systems in the prior art is solved, and more efficient memory access operations are achieved.
Patent Information
- Application Number
- CN202411246031.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2024-09-06
- Publication Date
- 2025-05-13
AI Technical Summary
Existing distributed memory systems are less efficient when handling memory access workloads, resulting in increased host processing overhead and excessive bandwidth usage of CXL components.
By implementing the Fine Grained Data Mover (FGDM) component on the memory controller, receiving data mover calls, creating and executing data mover tasks, including read and write phases, and executing multiple memory commands concurrently to improve efficiency.
It realizes the unloading of memory access workload from the host system, reduces host processing overhead, frees up CXL component bandwidth, and improves the overall efficiency of distributed memory systems.
Smart Images

Figure CN119987648A_ABST
Abstract
Description
[0001] Government Rights
[0002] This invention was made with U.S. Government support under Contract No. DE-NA0003525; U.S. Department of Energy Sandia National Laboratories subcontract 2168213. The U.S. Government has certain rights in this invention. Technical Field
[0003] Embodiments are directed to improving the efficiency of distributed memory systems. Background Art
[0004] Memory devices used in computers or other electronic devices can be classified as volatile and non-volatile memory. Volatile memory requires power to maintain its data and includes random access memory (RAM), dynamic random access memory (DRAM), or synchronous dynamic random access memory (SDRAM), etc. Non-volatile memory can save stored data when not powered, and includes flash memory, read-only memory (ROM), electrically erasable programmable ROM (EEPROM), static RAM (SRAM), erasable programmable ROM (EPROM), resistance variable memory, phase change memory, storage class memory, resistance random access memory (RRAM) and magnetoresistive random access memory (MRAM), etc. Persistent memory is an architectural attribute of the system, where the data stored in the media is available after a system reset or power cycle. In some examples, non-volatile memory media can be used to build a system with a persistent memory model.
[0005] The memory device may be coupled to a host (e.g., a host computing device) to store data, commands, and / or instructions for use by the host when the computer or electronic system is operating. For example, data, commands, and / or instructions may be transferred between the host and the memory device during operation of the computing or other electronic system.
[0006] Various protocols or standards may be applied to facilitate communication between a host and one or more other devices (such as a memory buffer, accelerator, or other input / output device). In an example, a protocol such as Compute Express Link (CXL) may be used to provide high bandwidth and low latency connectivity. Summary of the invention
[0007] One aspect of the present application relates to a method for offloading memory access workload from a host system in a distributed memory architecture, the method comprising: at a processor on a memory controller: receiving a data mover call, the data mover call comprising a request to copy a value, a request to aggregate a value, a request to scatter a value, or a request to set a value, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover call, the data mover tasks being created based on a specified mapping between the data mover call and the data mover tasks, each data mover task including a read phase and a write phase; executing each data mover task in the set by issuing a memory access command corresponding to each phase of each specific task of the set of data mover tasks to execute the phase of the specific one of the set of data mover tasks; concurrently executing multiple memory commands of the set of data mover tasks on multiple memory interface slices, the multiple memory commands targeting the memory locations specified in the data mover call; determining that all of the data mover tasks of the data mover call have been completed; and sending a response to the data mover call to the host.
[0008] Another aspect of the present application relates to a memory controller device for offloading memory access workload from a host system in a distributed memory architecture, the memory controller device comprising: a processor configured to perform operations, the operations comprising: receiving a data mover call, the data mover call comprising a request to copy a value, a request to aggregate a value, a request to scatter a value, or a request to set a value, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover call, the data mover tasks being based on the specified relationships between the data mover call and the data mover tasks; A host computer is provided for creating a mapping, each data mover task including a read phase and a write phase; executing each data mover task in the set of data mover tasks by issuing memory access commands corresponding to each phase of each specific task of the set of data mover tasks to execute the phase of the specific one of the set of data mover tasks; concurrently executing multiple memory commands of the set of data mover tasks on multiple memory interface slices, the multiple memory commands targeting the memory locations specified in the data mover call; determining that all the data mover tasks of the data mover call have been completed; and sending a response to the data mover call to the host.
[0009] Another aspect of the present application relates to a non-transitory machine-readable medium storing instructions for offloading memory access workload from a host system in a distributed memory architecture, the instructions, when executed by a processor of a memory controller, causing the memory controller to perform operations, the operations comprising: receiving a data mover call, the data mover call comprising a request to copy a value, a request to aggregate a value, a request to scatter a value, or a request to set a value, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover call, the data mover tasks being based on the data mover call and the data mover task; A host computer is provided to create a data mover call based on a specified mapping between tasks, each data mover task including a read phase and a write phase; execute each data mover task in the set of data mover tasks by issuing memory access commands corresponding to each phase of each specific task of the set of data mover tasks to execute the phase of the specific one of the set of data mover tasks; concurrently execute multiple memory commands of the set of data mover tasks on multiple memory interface slices, the multiple memory commands targeting the memory locations specified in the data mover call; determine that all of the data mover tasks of the data mover call have been completed; and send a response to the data mover call to the host. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. The same reference numerals with different letter suffixes may represent different instances of similar components. The drawings generally illustrate various embodiments discussed in this document by way of example and not limitation.
[0011] Figure 1A is a representation of a chiplet system mounted on a peripheral board according to some examples of the present disclosure.
[0012] Figure 1B is a block diagram of components in a tag chiplet system according to some examples of the present disclosure.
[0013] Figure 2 A distributed memory system according to some examples of the present disclosure is described.
[0014] Figure 3 A block diagram illustrating how data mover components interface with other components of a memory controller according to some examples of the present disclosure.
[0015] Figure 4 A block diagram illustrating a fine-grained data mover according to some examples of the present disclosure.
[0016] Figure 5 A block diagram illustrating a FGDM according to some examples of the present disclosure.
[0017] Fig. 6A The stages of a data mover task (DMT) initiated from a copy, scatter-span, or gather-span call according to some examples of the present disclosure are illustrated.
[0018] Figure 6B The stages of a data mover task resulting from scatter-address, scatter-index, gather-address, and gather-index calls according to some examples of the present disclosure are described.
[0019] Figure 7 A flow chart illustrating a method of processing calls in a fine-grained data mover according to some examples of the present disclosure.
[0020] Figure 8 is a block diagram illustrating an example of a machine upon which one or more embodiments may be implemented. DETAILED DESCRIPTION
[0021] Compute Express Link (CXL) is an open standard interconnect configured for high-bandwidth, low-latency connectivity between host devices and other devices such as accelerators, memory devices, and intelligent I / O devices. CXL is designed to facilitate high-performance computing workloads by supporting heterogeneous processing and memory systems. CXL implements coherency and memory semantics on top of PCI Express (PCIe)-based I / O semantics to optimize performance.
[0022] In some examples, CXL is used in applications such as artificial intelligence, machine learning, analytics, cloud infrastructure, edge computing devices, communication systems, and elsewhere. Data processing in such applications can use a variety of scalar, vector, matrix, and spatial architectures that can be deployed in CPUs, GPUs, FPGAs, smart NICs, and other accelerators that can be coupled using CXL links.
[0023] CXL supports dynamic multiplexing using a set of protocols including input / output (CXL.io, based on PCIe), cache (CXL.cache), and memory (CXL.memory) semantics. In an example, CXL can be used to maintain a unified, consistent memory space between a CPU (e.g., a host device or host processor) and any memory on an attached CXL device. This configuration allows the CPU and other devices to share resources and operate on the same memory region to improve performance, reduce data movement, and reduce software stack complexity. In an example, the CPU is primarily responsible for maintaining or managing consistency in a CXL environment. Therefore, CXL can be utilized to help reduce device cost and complexity, as well as the overhead traditionally associated with consistency across I / O links.
[0024] CXL runs on top of the PCIe PHY and provides full interoperability with PCIe. In an example, a CXL device starts link training at the PCIe Gen 1 data rate and negotiates CXL as its operating protocol if its link partner is capable of supporting CXL (e.g., using the alternate protocol negotiation mechanism defined in the PCIe 5.0 specification). As a result, devices and platforms can more easily adopt CXL by leveraging the PCIe infrastructure without having to design and validate PHYs, channels, channel extension devices, or other upper layers of PCIe.
[0025] In an example, CXL supports single-stage switching to achieve fan-out to multiple devices. This enables multiple devices in a platform to be migrated to CXL while maintaining CXL's backward compatibility and low latency characteristics. In an example, CXL can provide a standardized computing fabric that supports multiple logical devices (MLDs) and pooling of a single logical device, such as using a CXL switch connected to several host devices or nodes (e.g., root ports). This feature enables servers to pool resources, such as accelerators and / or memory that can be assigned based on workload. For example, CXL can help facilitate resource allocation or dedicating and release. In an example, CXL can help allocate and de-allocate memory to various host devices as needed. This flexibility helps designers avoid over-provisioning while ensuring optimal performance. The CXL protocol enables the construction of large, multi-host, fabric-attached memory systems. In addition, CXL memory systems can be built from multi-port, hot-pluggable devices and connected to hot-pluggable memory switches.
[0026] Some of the compute-intensive applications and operations mentioned herein may require or use large data sets. Memory devices storing such data sets may be configured for low latency and high bandwidth and persistence. One problem with load-store interconnect architectures includes ensuring persistence. CXL can help solve this problem using architectural flows of software and standard memory management interfaces, for example enabling persistent memory to move from a controller-based approach to direct memory management.
[0027] In some instances, such distributed memory systems may be constructed using chiplets. Chiplets are an emerging technology for integrating various processing functionalities. Typically, a chiplet system consists of discrete modules (each a "chiplet") that are integrated on an interposer and interconnected in many instances as needed through one or more established networks to provide the desired functionality to the system. The interposer and the included chiplets may be packaged together to facilitate interconnection with other components of the larger system. Each chiplet may include one or more individual integrated circuits or "chips" (ICs), which may be combined with discrete circuit components and are coupled together to a respective substrate to facilitate attachment to the interposer. Most or all of the chiplets in the system will be individually configured to communicate through one or more established networks.
[0028] Configuring chiplets as individual modules of a system is different from implementing the system on a single chip containing distinct device blocks (e.g., intellectual property (IP) blocks) on one substrate (e.g., a single die), such as a system on a chip (SoC) or multiple discrete packaged devices integrated on a printed circuit board (PCB). In general, chiplets provide better performance (e.g., lower power consumption, reduced latency, etc.) than discrete packaged devices, and chiplets provide greater production benefits than single die chips. These production benefits may include higher yields or reduced development costs and time.
[0029] A chiplet system may include, for example, one or more application (or processor) chiplets and one or more support chiplets. Here, the differences between application chiplets and support chiplets only relate to possible design cases for chiplet systems. Thus, for example, a synthetic vision chiplet system may include (by way of example only) an application chiplet that produces a synthetic vision output and a support chiplet, such as a memory controller chiplet, a sensor interface chiplet, or a communication chiplet. In a typical use case, a synthetic vision designer may design an application chiplet and obtain support chiplets from other parties. Thus, design expenditure (e.g., in terms of time or complexity) is reduced by avoiding the design and production of functionality embodied in support chiplets. Chipsets also support tight integration of IP blocks that may otherwise be difficult, such as IP blocks manufactured using different processing technologies or using different feature sizes (or utilizing different contact technologies or spacings). Thus, multiple ICs or IC assemblies with different physical, electrical, or communication characteristics can be assembled in a modular manner to provide an assembly that provides the desired functionality. A chiplet system can also facilitate adaptation to the needs of different larger systems into which the chiplet system will be incorporated. In an example, an integrated circuit or other assembly can optimize power, speed, or heating for a specific function—as might happen with a sensor—and can be more easily integrated with other devices than trying to do so on a single die. Additionally, by reducing the overall size of the die, the yield of a chiplet tends to be higher than the yield of a more complex single-die device.
[0030] Figure 1A and 1B An example of a chiplet system 110 according to an embodiment is described. Figure 1A1 is a representation of a chiplet system 110 mounted on a peripheral board 105, which may be connected to a larger computer system via, for example, a peripheral component interconnect express (PCIe). The chiplet system 110 includes a package substrate 115, an interposer 120, and four chiplets: an application chiplet 125, a host interface chiplet 135, a memory controller chiplet 140, and a memory device chiplet 150. Other systems may include many additional chiplets to provide additional functionality as will be apparent from the following discussion. The packaging of the chiplet system 110 is illustrated as having a lid or cover 165, but other packaging techniques and structures for chiplet systems may be used. Figure 1B is a block diagram that labels the components in a chiplet system for clarity.
[0031] The application chiplet 125 is illustrated as including a network on chip (NOC) 130 to support a chiplet network 155 for inter-chiplet communications. In an example embodiment, the NOC 130 may be included on the application chiplet 125. In an example, the NOC 130 may be defined in response to the selected supporting chiplets (e.g., chiplets 135, 140, and 150), enabling the designer to select the appropriate number or chiplet network connections or switches for the NOC 130. In an example, the NOC 130 may be located on a separate chiplet, or even within the interposer 120. In an example as discussed herein, the NOC 130 implements a chiplet protocol interface (CPI) network.
[0032] CPI is a packet-based network that supports virtual channels to achieve flexible and high-speed interactions between chiplets. CPI enables bridging from an intra-chiplet network to a chiplet network 155. For example, the Advanced Extensible Interface (AXI) is a widely used specification for designing intra-chip communications. However, the AXI specification covers a variety of physical design options, such as the number of physical channels, signal timing, power, etc. Within a single chip, these options are typically selected to meet design goals, such as power consumption, speed, etc. However, in order to achieve flexibility in chiplet systems, adapters such as CPI are used to interface between various AXI design options that can be implemented in various chiplets. CPI bridges the intra-chiplet network across the chiplet network 155 by enabling physical channel to virtual channel mapping and encapsulating time-based signaling with a packetized protocol.
[0033] CPI can use various different physical layers to transmit packets. The physical layer may include simple conductive connections, or may include drivers to increase voltage, or otherwise facilitate the transmission of signals over longer distances. An example of such a physical layer may include an advanced interface bus (AIB), which in various examples may be implemented in the intermediate layer 120. The AIB uses source synchronous data transmission with a forwarding clock to transmit and receive data. Packets are transmitted across the AIB at a single data rate (SDR) or double data rate (DDR) with respect to the transmitted clock. Various channel widths are supported by the AIB. When operating in SDR mode, the AIB channel width is a multiple of 20 bits (20, 40, 60, ...), and for DDR mode is a multiple of 40 bits: (40, 80, 120, ...). The AIB channel width includes transmission and reception signals. The channel can be configured to have a symmetrical number of transmission (TX) and reception (RX) input / output (I / O), or an asymmetrical number of transmitters and receivers (e.g., full transmitters or full receivers). A channel can be used as a master or slave AIB depending on which chiplet provides the master clock. The AIB I / O cell supports three clock modes: asynchronous (i.e., unclocked), SDR, and DDR. In various examples, the unclocked mode is used for the clock and some control signals. The SDR mode can use a dedicated SDR-only I / O cell or a dual-purpose SDR / DDR I / O cell.
[0034] In an example, a CPI packet protocol (e.g., point-to-point or routable) may use symmetrical receive and transmit I / O units within an AIB channel. The CPI streaming protocol allows for more flexible use of the AIB I / O units. In an example, an AIB channel for streaming mode may configure the I / O units as full TX, full RX, or half RX and half RX. The CPI packet protocol may use the AIB channel in SDR or DDR operating modes. In an example, the AIB channel is configured in increments of 80 I / O units (i.e., 40 TX and 40 RX) for SDR mode and 40 I / O units for DDR mode. The CPI streaming protocol may use the AIB channel in SDR or DDR operating modes. Here, in an example, the AIB channel increments by 40 I / O units for both SDR and DDR modes. In an example, each AIB channel is assigned a unique interface identifier. The identifier is used during CPI reset and initialization to determine paired AIB channels across adjacent chiplets. In an example, the interface identifier is a 20-bit value that includes a 7-bit chiplet identifier, a 7-bit column identifier, and a 6-bit link identifier. The AIB physical layer transmits the interface identifier using the AIB out-of-band shift register. The 20-bit interface identifier is transmitted in both directions across the AIB interface using bits 32 to 51 of the shift register.
[0035] AIB defines a group of stacked AIB channels as an AIB channel column. An AIB channel column has a certain number of AIB channels plus auxiliary channels. The auxiliary channels contain signals for AIB initialization. All AIB channels in a column (except auxiliary channels) have the same configuration (e.g., full TX, full RX, or half TX and half RX, and have the same number of data I / O signals). In an example, the AIB channels are numbered in a continuous increasing order starting from the AIB channel adjacent to the AUX channel. The AIB channel adjacent to the AUX is defined as AIB channel 0.
[0036] Typically, the CPI interface on individual chiplets may include serialization-deserialization (SERDES) hardware. SERDES interconnects are well suited for scenarios where high-speed signaling with low signal counts is desired. However, SERDES may result in additional power consumption and longer latency for multiplexing and demultiplexing, error detection or correction (e.g., using block-level cyclic redundancy checks (CRCs)), link-level retries, or forward error correction. However, when low latency or energy consumption is a primary concern for ultra-short reach chiplet-to-chiplet interconnects, a parallel interface with a clock rate that allows data transfer with minimal latency may be utilized. CPI includes elements that minimize both latency and energy consumption in these ultra-short reach chiplet interconnects.
[0037] CPI employs a credit-based technique for flow control. A receiver, such as application chiplet 125, provides credits representing available buffers to a sender, such as memory controller chiplet 140. In an example, for a given unit of time transmission, a CPI receiver includes a buffer for each virtual channel. Thus, if a CPI receiver supports 5 time messages and a single virtual channel, the receiver has 5 buffers arranged in 5 rows (e.g., 1 row per unit of time). If 4 virtual channels are supported, the receiver has 20 buffers arranged in 5 rows. Each buffer holds the payload of 1 CPI packet.
[0038] As the sender transmits to the receiver, the sender decrements the available credits based on the transmission. Once all of the receiver's credits are consumed, the sender stops sending packets to the receiver. This ensures that the receiver always has an available buffer to store the transmission. As the receiver processes received packets and frees up buffers, the receiver passes the available buffer space back to the sender. This credit return can then be used by the sender to allow the transmission of additional information.
[0039] Also illustrated is a chiplet mesh network 160 that uses direct chiplet-to-chiplet technology without the need for a NOC 130. The chiplet mesh network 160 may be implemented in CPI or another chiplet-to-chiplet protocol. The chiplet mesh network 160 typically enables a chiplet pipeline where one chiplet serves as an interface to the pipeline while the other chiplets in the pipeline simply interface themselves.
[0040] In addition, dedicated device interfaces such as one or more industry standard memory interfaces 145 (e.g., for example, synchronous memory interfaces such as DDR5, DDR6) may also be used to interconnect the chiplets. The connection of the chiplet system or individual chiplets to an external device (e.g., a larger system) may be through a desired interface (e.g., a PCIE interface). In an example, this external interface may be implemented by a host interface chiplet 135, which in the depicted example provides a PCIE interface external to the chiplet system 110. Such dedicated interfaces 145 are typically adopted when conventions or standards in the industry have converged on such interfaces. The illustrated example of a double data rate (DDR) interface 145 that connects the memory controller chiplet 140 to a dynamic random access memory (DRAM) memory device chiplet 150 happens to be this industry convention.
[0041] Among the various possible supporting chiplets, a memory controller chiplet 140 may be present in a chiplet system 110 due to the nearly ubiquitous use of memory devices for computer processing and the use of state-of-the-art technology for memory devices. Therefore, using a memory device chiplet 150 and a memory controller chiplet 140 produced by others enables a chiplet system designer to use a robust product produced by an advanced manufacturer. Typically, a memory controller chiplet 140 provides a memory device specific interface for reading, writing, or erasing data. Typically, a memory controller chiplet 140 may provide additional features, such as error detection, error correction, maintenance operations, or atomic operation execution. For some types of memory, maintenance operations tend to be specific to the memory device 150, such as garbage collection in NAND flash or storage class memory, temperature adjustment in NAND flash memory (e.g., cross temperature management). In an example, maintenance operations may include logical to physical (L2P) mapping or management to provide a level of indirection between the physical and logical representations of data. In other types of memory, such as DRAM, some memory operations, such as refresh, may be controlled at some times by a host processor of a memory controller and at other times by the DRAM memory devices or by logic associated with one or more DRAM devices, such as an interface chip (in an example, a buffer).
[0042] The memory device chiplet 150 may be a volatile memory device or a non-volatile memory, or any combination of volatile memory devices or non-volatile memories. Examples of volatile memory devices include, but are not limited to, random access memory (RAM), such as DRAM, synchronous DRAM (SDRAM), Graphic Double Data Rate 6 SDRAM (GDDR6 SDRAM), etc. Examples of non-volatile memory devices include, but are not limited to, NAND flash memory, storage class memory (e.g., phase change memory or memristor-based technology), ferroelectric RAM (FeRAM), etc. The illustrated example includes the memory device 150 as a chiplet, however, the memory device 150 may reside elsewhere, such as in a different package on the board 105. For many applications, multiple memory device chiplets may be provided. In an example, these memory device chiplets may each implement one or more memory technologies. In an example, a memory chiplet may include multiple stacked memory dies of different technologies, for example, one or more SRAM devices stacked or otherwise communicating with one or more DRAM devices. The memory controller chiplet 140 may also be used to coordinate operations between multiple memory chiplets in the chiplet system 110; for example, using one or more memory chiplets in one or more cache storage levels, and using one or more additional memory chiplets as main memory. The chiplet system 110 may also include multiple memory controller chiplets 140, such as may be used to provide memory control functionality for separate processors, sensors, networks, and the like. A chiplet architecture such as the chiplet system 110 provides the advantage of allowing adaptation to different memory storage technologies; and providing different memory interfaces by upgrading the chiplet configuration without requiring redesign of the rest of the system structure. The memory controller chiplet 140 may include processing hardware, working memory, and the like that enables the memory controller chiplet 140 to perform operations on data. For example, the memory controller chiplet 140 may include fine-grained data mover logic that implements many of the techniques described herein.
[0043] Many high performance computing applications may benefit from hardware architectures with more memory bandwidth, such as distributed memory architectures. Example applications include Page ranking algorithms for rating the importance of each web page on the Internet; sparse BLAS libraries such as NIST Sparse-BLAS; Stencil; and others. These applications benefit from building multi-host clusters with memory bandwidth greater than what the hosts can see.
[0044] Figure 2A distributed memory system 200 according to some examples of the present disclosure is illustrated. A memory fabric, such as a CXL fabric 212, is used to connect hosts 210-A, 210-B ... 210-P. The CXL fabric 212 is connected to a plurality of memory devices 214-A to 214-N. The memory devices may include a memory controller and a memory medium. In some examples, N>P, such that the number of memory devices 214 exceeds the hosts. In some examples, the distributed memory system 200 may be or include a chiplet system, such that one or more of the components shown may be chiplets. In some examples, some components may be chiplets and other components may be connected with other types of compute buses.
[0045] A host performing sparse operations (e.g., matrix operations) may issue a large number of commands that read and write small values of a single matrix operation. For example, the algorithm Stencil may execute thousands of memory commands, each command reading or writing less than 32, 64 bit floating point numbers. When the host issues these thousands of commands to various memory controllers through the CXL fabric 212, these commands may utilize fabric bandwidth. In addition, these commands utilize processing overhead on the host that needs to track and manage these commands.
[0046] In some examples, improvements to a memory controller on a distributed memory system are disclosed that includes a fine-grained data mover (FGDM) component that offloads management of memory commands that access multiple small values to the memory controller. By providing data mover calls that offload a portion of the work of accessing multiple small values to the FGDM, CXL fabric bandwidth can be freed up and host processing overhead can be reduced to achieve performance improvements. The host can send work requests in the form of data mover calls or commands to the FGDM without utilizing OS system calls. The FGDM is a virtually addressed data move engine whose architecture is designed to transfer data at high transfer rates. A data mover call is a command sent from a host or other processor to a memory controller that instructs a unit (FGDM) of the memory controller to perform a specified operation.
[0047] FGDM converts data mover calls into one or more transfer groups called data mover tasks and assigns these tasks to data move engines, such as AMBA-AXI4 slice data move engines. Each data mover task is responsible for transferring a small amount of data and AXI slices are designed to actively process multiple data mover tasks concurrently. Since multiple data mover tasks from multiple data mover calls can be actively processed at the same time, FGDM can maintain a high level of utilization even if each data mover call is small.
[0048] The Data Mover Task concept allows the hardware to be scheduled at a small granularity, which is useful when trying to achieve high performance with many small requests or to enforce per-tenant QoS policies when multiple tenants use a single Data Mover. In some instances, a Data Mover Task manages the movement of up to 256 bytes or up to 16 elements – whichever limit is reached first when the FGDM hardware breaks down a Data Mover call into Data Mover Tasks. A Data Mover Task has at least a read and write phase and may include an optional fetch phase for obtaining an address.
[0049] exist Figure 2 214-A). The memory controller (e.g., memory controller 1 214-A) may include a fine grained data mover (FGDM) component 240 on a chip network architecture. The memory controller (e.g., memory controller 1 214-A) may also include a host interface component 230, a CXL fabric interface component 232, an on-chip network interface component 234, a FAM control component 236, and a media control component 238. The various components may be implemented as hardware, hardware configured by software, or the like. The host interface component 230 may implement one or more protocols or interfaces with the host to receive memory commands including data mover calls. The FGDM component 240 may include one or more processors, memories, or the like to implement data mover calls as described herein. The CXL fabric interface component 232 may implement one or more protocols or interfaces to communicate through the CXL fabric 212. The on-chip network interface component 234 may implement one or more on-chip network interfaces as previously described. The media control component 238 may implement media control operations, such as implementing read and write scheduling, refresh control, and ECC. For example, the media control component 238 may be a DRAM controller. In some examples, the FAM control component 236 includes an address translation table and an access control table to translate addresses between various forms to route memory requests to the appropriate device and back.
[0050] Figure 3 Block diagram illustrating how a data mover component according to some examples of the present disclosure interfaces with other components of a memory controller. The on-chip network router 334 routes requests initiated by the fine-grained data mover component 340 to the correct address interleaved fabric plane 350-A, 350-B, 350-C, or 350-D. The host system (or other processor) issues data mover calls to the data mover using the command manager 360.
[0051] Figure 4A block diagram illustrating a fine-grained data mover component 400 according to some examples of the present disclosure. The FGDM component 400 may be an instance of the FGDM component 240, 340. An inbound interface 410 receives new call requests and configuration status register (CSR) read and write requests. An active call handler component 412 decomposes calls into data mover tasks; tracks completion of concurrent active calls; and formulates call returns for calls. Each data mover task has a read and write phase, and some data mover tasks (depending on the data mover call) may have an extract phase (e.g., extract an address list). Tasks may include:
[0052]
[0053]
[0054] The tasks can be decomposed based on the mapping between data mover calls and data mover tasks. For example, according to the following table:
[0055] Call DM Tasks copy AtoB,NumOps=1 Gather BtoA, NumOps = 1 to many dispersion AtoB, NumOps = 1 to multiple set up Set NumOps = 1 to multiple
[0056] In some instances, there may be different strategies for initiating a particular data mover task depending on the call length and element size being operated on. For example:
[0057]
[0058] Based on the launch strategy, you can adjust the allocation and launch priority. For example:
[0059]
[0060]
[0061] In some examples, the call handler logic that initiates the data mover task has a current slice indication. This current slice is incremented at the end of each call.
[0062]
[0063] The active data mover task handler component 414 schedules the reads and writes to be performed by each of the AXI slices 420-A; 420-B; 420-C; and 420-D. For example, based on the above strategy. Although four AXI slices are shown, other numbers of AXI slices may be utilized. The AXI slices 420-A; 420-B; 420-C; and 420-D decompose the data mover task into AXI read and write operations. Each AXI slice may focus on one stage of the data mover task before moving to different stages of different data mover tasks. In some examples, the AXI slice may generate read commands for one task while generating write commands for another task. Tasks in the read stage are scheduled by the task handler of the AXI slice to its read generation block. At the same time, there may be a task pool scheduled for write generation of the slice in the write stage. The slice conductor component 418 coordinates reads and writes between AXI slices to optimize memory usage. The block internal interconnect 416 may include one or more hardware or software structures to connect the inbound interface 410, the active call handler component 412, and the active data mover task handler component 414 to one or more AXI slices, such as AXI slices 420-A, 420-B, 420-C, and 420-D. In some examples, each AXI slice may support multiple outstanding read and write requests, where a read request may be assigned a read identifier.
[0064] As described above, the fine-grained data mover receives a call request and decomposes the call into data mover tasks. Each data mover task has several stages. These data mover tasks are then further decomposed into AXI read and AXI write requests. Although the AMBA-AXI fabric is used herein, a person of ordinary skill in the art will recognize that fabrics other than AMBA-AXI may be utilized. Figure 4 The dotted boxes are used to illustrate the working units of component operations. The call domain 430 includes the inbound interface 410, the block internal interconnect component 416, and the active call handler component 412. These components handle FGDM calls. The data mover task domain 435 includes the active call handler component 412, the active data mover task handler component 414, the block internal interconnect component 416, the AXI slice and slice conductor component 418. The AXI memory request domain 440 includes the AXI slice and slice conductor component 418.
[0065] As previously described, the call may be a basic data movement operation, such as: copy, scatter, gather, or set. The copy operation moves an integer number of bytes from one byte-aligned virtual address to another byte-aligned virtual address. In some instances, if the source buffer overlaps the destination buffer, the state of the main memory after the operation is completed may not be defined. The scatter operation moves data in a continuously addressed buffer to an integer number of destination element positions. In some instances, all elements in a scatter call may be restricted to the same size. A scatter call may use one of three different addressing modes - stride, address, or index. The gather operation moves data from an integer number of source element positions to a continuously addressed destination buffer. In some instances, all elements in a gather call may be restricted so that they are the same size. The gather operation uses the same three different addressing modes as the scatter call - stride, address, or index. The set operation writes a data pattern of a predefined size (e.g., 64B) to an integer number of destination element positions. In some instances, all elements in a gather call may be restricted to the same size. The gather operation uses the same three different addressing modes as the scatter call - stride, address, or index.
[0066] FGDM supports a general call / return interface that allows a requester to construct a call request and receive a return response when the FGDM operation is completed. The FGDM operation can be called on behalf of a user host process or by a host operating system. In some instances, due to the latency of translation lookaside buffer (TLB) misses and fault handling, FGDM can support multiple simultaneous transfers. To achieve this, a means of maintaining the state of multiple transfer contexts can be provided.
[0067] The scatter and gather commands support three modes that can define non-contiguous memory areas. The stride mode specifies the base virtual address, uniform stride size, element size, and number of elements in the call command. The number of addresses used is specified by the number of elements, and the addresses start at the base virtual address and increment by the amount specified by the stride size. The next mode (addressing mode) specifies a memory address with a virtual address list. A call utilizing the addressing mode contains an address list base virtual address, element size, and number of elements. The memory address is read from an address list in memory at the address list base virtual address. The offset mode utilizes an offset list, which is similar to an address list in that a memory location containing a storage address is called. In the offset list, the call specifies the base virtual address, the virtual address of the offset list, the element size, and the number of elements.
[0068] The data mover call may specify the addresses of the source and destination memory areas. A subset of the data mover functionality (which may only be available to the operating system) may also be used to enable block configuration status register (CSR) access operations. These operations will use the virtual address of one memory area (source or destination) and the physical address of the CSR. The physical address of the CSR may be specified directly or it may be generated by the memory management unit (MMU) during the normal translation process. One type of block CSR operation may be a scatter command, in which a value may be written to a group of CSRs in a striding access. Similarly, a CSR gather operation may use a striding access to read a group of CSRs and write the CSRs to a contiguous destination memory area.
[0069] Figure 5 A block diagram illustrating FGDM 500 according to some examples of the present disclosure. According to some examples, FGDM 500 can be an example of FGDM components 400, 340, and / or 240. FGDM 500 includes a data mover (DM) inbound interface component 510 that handles inbound calls received from one or more host systems through AXI interface component 508. DM inbound interface component 510 queues new calls into FIFO queue 511. The calls are then released to active DM call handler 512 in FIFO order.
[0070] The active call component 532 of the DM call handler 512 splits the DM call into data mover tasks and the DM task allocation distribution component 538 allocates the first phase (e.g., the extraction or read phase of each of these tasks) to the AXI slices 520-A; 520-B; 520-C; 520-D. The call responder component 534 can determine when the data mover call is completed and schedule AXI writes to return the results of the data mover call. When all write phases of the DM tasks of all calls are completed, the DM considers the call completed. This can be tracked via the DM task counter in the active DM task counter 536 of each of the active calls. In some examples, the calls can be completed in a different order than they were received.
[0071] The tasks of a particular call can be tracked and handled by the active DM task handler 514. The active DM task handler component 514 monitors all active data mover tasks and determines when the various phases of each of the active tasks should start and determines when those tasks are completed. For example, the active DM task handler component 514 can determine when a read should be issued for a read phase that begins with extracting an address or index list, determine when the write phase of each DM task begins, and detect when the DM task is completed. The DM task read issuance component 542 can be used to issue a read task, including reading an address list or other index stored in a memory by issuing a read to one or more AXI slices. The DM task write issuance component 540 can issue one or more write tasks to one or more AXI slices. Read and write tasks can utilize coordinated read and / or write phases. The active write counter component 544 and the active read counter component 546 allow the active DM task handler component 514 to track the progress of the task and report back to the active call handler component 532 when the task is completed.
[0072] As previously described, AXI slice components 520-A, 520-B, 520-C, and 520-D convert tasks into AXI read / write commands. In some instances, various slice-to-slice optimizations may be performed. For example, the system may utilize write combining, where a portion of the memory written in a previous DM task may be passed to a subsequent slice so that the write may be merged with a portion of the memory of the next DM task. This may reduce the number of AXI writes required for certain tasks. For example, a write to the last 1 to 15 bytes of a DM task may be passed to a subsequent slice so that the write may be merged with the first 1 to 15 bytes of the next DM task to reduce the number of AXI writes required. For large calls, the last slice in a data mover task group passes a portion of the write to the next data mover task group. To achieve more write combining, each slice may include one or more write data export buffers. Another optimization may be read data borrowing, where one DM task may obtain data from a read of another slice. For example, slice 0 borrows data from slice 1; slice 1 borrows data from slice 2; slice 2 borrows from slice 3. Slice 3 will borrow from slice 0, but it will borrow from the next data mover task group.
[0073] In some examples, the slice conductor component 518 coordinates reads and writes between AXI slices to optimize memory usage and enable write combining, read data borrowing, and address fetch data borrowing. The slice conductor component 518 detects when all slices are ready to start coordinated reads or coordinated writes and initiates these coordinated reads and writes.
[0074] As described above, FGDM breaks down calls into data mover tasks (DMTs), where each data mover task is responsible for reading and writing to memory. Fig. 6A The stages of a DMT initiated from a copy, scatter-span, or gather-span call according to some examples of the present disclosure are illustrated. First is the allocation stage 610, followed by the read stage 605 and the write stage 607. The read stage first sends a DM task to an AXI slice 612. At 614, the AXI slice uses the DMT to issue an AXI read. At 616, the slice waits for all AXI reads to complete. After the read stage 605 is completed, the write stage 607 begins. The write stage begins at 618 by issuing an AXI write and then waits for all AXI writes to complete at 620.
[0075] Figure 6B The stages of a DMT generated from scatter-address, scatter-index, gather-address and gather-index calls according to some examples of the present disclosure are illustrated. The allocation stage 630 is followed by an extraction stage 625, a read stage 627 and a write stage 629. The extraction stage 625 extracts an address or offset list (depending on the addressing mode) and begins at 632 by sending a DM task to an AXI slice. At 634, an AXI read is issued to read an address list, an offset list or the like from memory. At 636, the AXI slice waits for all AXI reads to complete. After the extraction stage 625 is completed, the read stage begins at 627. An AXI read is issued at 638, the system waits for all AXI reads to complete at 640 and once all reads have completed, the DM task moves to the write stage 629. The write stage begins by issuing a write command at 642. After issuing the write command, the system waits for all AXI writes to complete at 644.
[0076] Figure 7A flow chart illustrating a method 700 for processing calls in processing a fine-grained data mover according to some examples of the present disclosure. At operation 712, the FGDM receives a data mover call. The call may be received via an on-chip network interface. In some instances, the data mover call may be received from a host system. In some instances, the data mover call may include a copy command to move a specified number of bytes from one virtual address to another virtual address. In other instances, the data mover call may include a gather command to move data from several non-contiguous source locations to a continuously addressed memory. In yet other instances, the data mover call may include a scatter command to move data from a continuously addressed buffer to several non-contiguous memory locations. In yet additional instances, the data mover call may include a set command to write a data pattern to several destination elements. In some instances, the data mover call specifies a source and / or destination location. The data mover call may specify the source and / or destination location by providing an address to store the source and / or destination location.
[0077] At operation 714, the FGDM creates a set of data mover tasks based on the data mover call. For example, the FGDM may use a mapping table or algorithm and identify the set of data mover tasks by reading parameters (e.g., memory addresses) from the call itself. In some examples, each data mover task includes a read phase or a write phase. Depending on the format of the call, some data mover tasks may include an extraction phase that extracts address data from memory. The task read phase reads values from memory, and the write phase writes values to memory.
[0078] At operation 717, the FGDM executes one or more data mover tasks in the group. In some examples, operation 717 includes concurrently executing multiple data mover tasks from the same call. In some examples, one or more data mover tasks from multiple different calls may be executed concurrently. Each data mover task may be executed by first executing any extraction phase at operation 718. As described above, the extraction phase retrieves an address or offset (source and / or destination) from a memory. The extraction phase begins by sending a task to an AXI slice that issues an AXI read command. Then, the phase waits for all AXI reads to complete before completion. The next phase is the read phase 720, which begins by sending a task to an AXI slice that issues an AXI read command. Then, the phase waits for all AXI reads to complete before completion. Finally, at operation 722, the write phase begins after the extraction and read phases are completed. The write phase begins by sending a task to an AXI slice that issues an AXI write command. Then, the phase waits for all AXI write commands to complete before completion.
[0079] At operation 724, it is determined for a particular call whether all data mover tasks for the call are complete. For example, the system may set a counter value to the number of tasks created for the particular call and decrement the counter each time a task is completed. Once the counter is zero, the call is complete. If additional tasks are to be completed, they are executed at operation 717. Once all tasks are completed, then at operation 726, the FGDM sends a response to the data mover call to the host.
[0080] Figure 8 A block diagram of an example machine 800 on which any one or more of the techniques (e.g., methodology) discussed herein may be executed is illustrated. In alternative embodiments, the machine 800 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine 800 may operate as a server machine, a client machine, or both in a server-client network environment. In an example, the machine 800 may act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. The machine 800 may be in the form of a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smart phone, a network appliance, a network router, a switch or a bridge, or any machine capable of (sequentially or otherwise) executing instructions specifying actions taken by that machine. In addition, although only a single machine is described, the term "machine" should also be deemed to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to execute any one or more of the methodologies discussed herein (e.g., cloud computing, software as a service (SaaS), other computer cluster configurations). The machine 800 may be or be configured as a host system, a CXL fabric, a memory controller, an FGDM, a memory device, or the like. The machine 800 may be arranged as a distributed memory system. The FGDM may be configured or may include Figure 1A , 1B , 2, 3, 4, 5 components; implementation Fig. 6A and 6B stage, and Figure 7 method.
[0081] As described herein, an instance may include or may operate on one or more logical units, components, or mechanisms (hereinafter "components"). A component is a tangible entity (e.g., hardware) that is capable of performing a specified operation and may be configured or arranged in a certain manner. In an example, a circuit may be arranged in a specified manner (e.g., internally or relative to an external entity such as other circuits) as a component. In an example, all or part of one or more computer systems (e.g., independent client or server computer systems) or one or more hardware processors may be configured by firmware or software (e.g., instructions, application portions, or applications) as a component that operates to perform a specified operation. In an example, the software may reside on a machine-readable medium. In an example, the software, when executed by the underlying hardware of the component, causes the hardware to perform the specified operation of the component.
[0082] Thus, the term "component" is understood to encompass a tangible entity, which may be physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., temporarily) configured (e.g., programmed) to operate in a specified manner or to perform part or all of any operations described herein. Considering instances in which components are temporarily configured, each of the components need not be instantiated at any one time. For example, where a component includes a general-purpose hardware processor configured using software, the general-purpose hardware processor may be configured as corresponding different components at different times. The software may configure the hardware processor accordingly, for example, to constitute a particular module at one time and to constitute different components at different times.
[0083] The machine (e.g., computer system) 800 may include one or more hardware processors, such as processor 802. Processor 802 may be a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof. The machine 800 may include a main memory 804 and a static memory 806, some or all of which may communicate with each other via an interconnect (e.g., a bus) 808. An example of main memory 804 may include synchronous dynamic random access memory (SDRAM), such as double data rate memory, such as DDR4 or DDR5. Interconnect 808 may be one or more different types of interconnects, such that one or more components may be connected using a first type of interconnect and one or more components may be connected using a second type of interconnect. Example interconnects may include a memory bus, a peripheral component interconnect (PCI), a peripheral component interconnect express (PCIe) bus, a universal serial bus (USB), or the like.
[0084] The machine 800 may further include a display unit 810, an alphanumeric input device 812 (e.g., a keyboard), and a user interface (UI) navigation device 814 (e.g., a mouse). In an example, the display unit 810, the input device 812, and the UI navigation device 814 may be a touch screen display. The machine 800 may additionally include a storage device (e.g., a drive unit) 816, a signal generating device 818 (e.g., a speaker), a network interface device 820, and one or more sensors 821, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. The machine 800 may include an output controller 828, such as a serial (e.g., universal serial bus (USB), parallel or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection for communicating with or controlling one or more peripheral devices (e.g., printers, card readers, etc.).
[0085] The storage device 816 may include a machine-readable medium 822 on which is stored one or more sets of data structures or instructions 824 (e.g., software) embodying or utilized by any one or more of the techniques or functions described herein. The instructions 824 may also reside, completely or at least partially, within the main memory 804, within the static storage 806, or within the hardware processor 802 during execution thereof by the machine 800. In an example, one or any combination of the hardware processor 802, the main memory 804, the static storage 806, or the storage device 816 may constitute a machine-readable medium.
[0086] Although machine-readable medium 822 is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (eg, a centralized or distributed database, and / or associated caches and servers) configured to store one or more instructions 824.
[0087] The term "machine-readable medium" may include any medium capable of storing, encoding, or carrying instructions that are executed by the machine 800 and cause the machine 800 to perform any one or more of the techniques of the present disclosure, or capable of storing, encoding, or carrying data structures used by or associated with such instructions. Non-limiting machine-readable medium examples may include solid-state memory and optical and magnetic media. Specific examples of machine-readable media may include: non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable magnetic disks; magneto-optical disks; random access memory (RAM); solid-state drives (SSD); and CD-ROM and DVD-ROM disks. In some examples, machine-readable media may include non-transitory machine-readable media. In some examples, machine-readable media may include machine-readable media that is not a transient propagation signal.
[0088] The instructions 824 may be further transmitted or received via the network interface device 820 using a transmission medium over a communication network 826. The machine 800 may communicate with one or more other machines by wire or wireless communication using any of a number of transmission protocols, such as frame relay, Internet Protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc. Example communication networks may include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), a mobile telephone network (e.g., a cellular network), a plain old telephone (POTS) network, and a wireless data network (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards (known as ), IEEE 802.15.4 family of standards, 5G new radio (NR) family of standards, long term evolution (LTE) family of standards, universal mobile telecommunications system (UMTS) family of standards, peer-to-peer (P2P) networks, etc.). In an example, the network interface device 820 may include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas to connect to the communication network 826. In an example, the network interface device 820 may include multiple antennas to use at least one of single input multiple output (SIMO), multiple input multiple output (MIMO), or multiple input single output (MISO) technology for wireless communication. In some examples, the network interface device 820 may use multiple user MIMO technology for wireless communication.
[0089] Other notes and examples
[0090] Example 1 is a method for offloading memory access workload from a host system in a distributed memory architecture, the method comprising: at a processor on a memory controller: receiving a data mover call from a host, the data mover call comprising a copy command, a gather command, a scatter command, or a set command, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover call, the data mover tasks being created based on a specified mapping between the data mover call and the data mover tasks, each data mover task including a read phase and a write phase; executing each data mover task in the set by issuing a memory access command corresponding to each phase of each specific task of the set of data mover tasks to execute the phase of the specific one of the set of data mover tasks; concurrently executing multiple memory commands of the set of data mover tasks on multiple memory interface slices, the multiple memory commands targeting the memory locations specified in the data mover call; determining that all of the data mover tasks of the data mover call have been completed; and sending a response to the data mover call to the host.
[0091] In Example 2, the subject matter of Example 1 includes, wherein the data mover call includes a first memory location, a stride, an element size, and a number of elements, and wherein the set of data mover tasks accesses a plurality of memory locations indicated by the number of elements and starting from the first memory location and incrementing by the stride.
[0092] In example 3, the subject matter of examples 1-2 includes wherein the data mover call includes a location of an address list, and wherein one of the set of data mover tasks includes extracting the address list.
[0093] In Example 4, the subject matter of Examples 1-3 includes, wherein the data mover call comprises a first location and a second location storing an offset list, and wherein one of the set of data mover tasks comprises extracting the offset list.
[0094] In Example 5, the subject matter of Examples 1 to 4 includes wherein the data mover tasks include tasks of reading from a contiguous memory location and writing data to one to multiple contiguous memory locations, tasks of reading from one to multiple memory locations and writing data to one contiguous memory location, and tasks of writing to one to multiple memory locations.
[0095] In example 6, the subject matter of examples 1-5 includes, wherein commanding the memory interface slice simultaneously initiates the memory command.
[0096] In Example 7, the subject matter of Examples 1-6 includes executing, by the memory controller, a second data mover call from the host concurrently at the same time as executing the first data mover call.
[0097] Example 8 is a memory controller device for offloading memory access workload from a host system in a distributed memory architecture, the memory controller device comprising: a processor configured to perform operations, the operations comprising: receiving a data mover call from a host, the data mover call comprising a copy command, a gather command, a scatter command, or a set command, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover call, the data mover tasks being created based on a specified mapping between the data mover call and the data mover tasks, each data mover task comprising a read phase and a write phase; executing each data mover task in the set by issuing a memory access command corresponding to each phase of each specific task of the set of data mover tasks to execute the phase of the specific one of the set of data mover tasks; concurrently executing multiple memory commands of the set of data mover tasks on multiple memory interface slices, the multiple memory commands targeting the memory locations specified in the data mover call; determining that all of the data mover tasks of the data mover call have been completed; and sending a response to the data mover call to the host.
[0098] In Example 9, the subject matter of Example 8 includes, wherein the data mover call includes a first memory location, a stride, an element size, and a number of elements, and wherein the set of data mover tasks accesses a plurality of memory locations indicated by the number of elements and starting from the first memory location and incrementing by the stride.
[0099] In example 10, the subject matter of examples 8-9 includes wherein the data mover call includes a location of an address list, and wherein one of the set of data mover tasks includes extracting the address list.
[0100] In Example 11, the subject matter of Examples 8-10 includes, wherein the data mover call comprises a first location and a second location storing an offset list, and wherein one of the set of data mover tasks comprises extracting the offset list.
[0101] In Example 12, the subject matter of Examples 8 to 11 includes a task wherein the data mover tasks include a task of reading from a contiguous memory location and writing data to one to multiple contiguous memory locations, a task of reading from one to multiple memory locations and writing data to one contiguous memory location, and a task of writing to one to multiple memory locations.
[0102] In Example 13, the subject matter of Examples 8 to 12 includes, wherein the operation further includes classifying the data mover call into a selected category of one of a plurality of predetermined categories; and wherein creating the set of data mover tasks based on the data mover call includes utilizing the selected category to determine an initiation rate, a slice interleaving strategy, or an allocation strategy.
[0103] In example 14, the subject matter of examples 8-13 includes, wherein the data mover task includes an extraction phase.
[0104] Example 15 is a non-transitory machine-readable medium storing instructions for offloading memory access workload from a host system in a distributed memory architecture, the instructions, when executed by a processor of a memory controller, causing the memory controller to perform operations, the operations comprising: receiving a data mover call from a host, the data mover call comprising a copy command, a gather command, a scatter command, or a set command, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover call, the data mover tasks being based on the instructions between the data mover call and the data mover tasks; A host computer is provided for creating a data mover task based on a given mapping, each data mover task including a read phase and a write phase; executing each data mover task in the set of data mover tasks by issuing memory access commands corresponding to each phase of each specific task of the set of data mover tasks to execute the phase of the specific one of the set of data mover tasks; concurrently executing multiple memory commands of the set of data mover tasks on multiple memory interface slices, the multiple memory commands targeting the memory locations specified in the data mover call; determining that all the data mover tasks of the data mover call have been completed; and sending a response to the data mover call to the host.
[0105] In Example 16, the subject matter of Example 15 includes, wherein the data mover call includes a first memory location, a stride, an element size, and a number of elements, and wherein the set of data mover tasks accesses a plurality of memory locations indicated by the number of elements and starting from the first memory location and incrementing by the stride.
[0106] In Example 17, the subject matter of Examples 15-16 includes, wherein the data mover call includes a location of an address list, and wherein one of the set of data mover tasks includes extracting the address list.
[0107] In Example 18, the subject matter of Examples 15-17 includes, wherein the data mover call comprises a first location and a second location storing an offset list, and wherein one of the set of data mover tasks comprises extracting the offset list.
[0108] In Example 19, the subject matter of Examples 15 to 18 includes wherein the data mover tasks include tasks of reading from a contiguous memory location and writing data into one to multiple contiguous memory locations, tasks of reading from one to multiple memory locations and writing data into one contiguous memory location, and tasks of writing into one to multiple memory locations.
[0109] In Example 20, the subject matter of Examples 15 to 19 includes, wherein the operation further includes: classifying the data mover call into a selected category of one of a plurality of predetermined categories; and wherein creating the set of data mover tasks based on the data mover call includes utilizing the selected category to determine an initiation rate, a slice interleaving strategy, or an allocation strategy.
[0110] Example 21 is at least one machine-readable medium comprising instructions that, when executed by processing circuitry, cause the processing circuitry to perform operations implementing any of Examples 1-20.
[0111] Example 22 is an apparatus comprising means for implementing any of Examples 1-20.
[0112] Example 23 is a system for implementing any one of Examples 1-20.
[0113] Example 24 is a method for implementing any of Examples 1-20.
Claims
1. A method for offloading memory access workload from a host system in a distributed memory architecture, the method comprising: At the processor on the memory controller: receiving a data mover call, the data mover call comprising a request to copy a value, a request to gather a value, a request to scatter a value, or a request to set a value, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover calls, the data mover tasks being created based on a specified mapping between data mover calls and data mover tasks, each data mover task comprising a read phase and a write phase; executing each data mover task in the set by issuing memory access commands corresponding to each stage of each particular one of the set of data mover tasks to execute the stage of the particular one of the set of data mover tasks; concurrently executing a plurality of memory commands of the set of data mover tasks on a plurality of memory interface slices, the plurality of memory commands targeting memory locations specified in the data mover calls; Determining that all of the data mover tasks invoked by the data mover have been completed; and A response to the data mover call is sent to the host.
2. The method of claim 1 , wherein the data mover call comprises a first memory location, a stride, an element size, and a number of elements, and wherein the set of data mover tasks accesses a plurality of memory locations indicated by the number of elements and starting from the first memory location and incrementing by the stride.
3. The method of claim 1, wherein the data mover call includes the location of an address list, and wherein one of the set of data mover tasks includes extracting the address list.
4. The method of claim 1, wherein the data mover call comprises a first location and a second location storing an offset list, and wherein one of the set of data mover tasks comprises extracting the offset list.
5. The method of claim 1 , wherein the data mover tasks include tasks that read from one continuous memory location and write data into one to multiple continuous memory locations, tasks that read from one to multiple memory locations and write data into one continuous memory location, and tasks that write into one to multiple memory locations. The method of claim 1 , wherein commanding the memory interface slices simultaneously initiates the memory command.
7. The method according to claim 1, further comprising: A second data mover call from the host is executed by the memory controller concurrently at the same time as the first data mover call is executed.
8. A memory controller device for offloading memory access workload from a host system in a distributed memory architecture, the memory controller device comprising: A processor configured to perform operations comprising: receiving a data mover call, the data mover call comprising a request to copy a value, a request to gather a value, a request to scatter a value, or a request to set a value, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover calls, the data mover tasks being created based on a specified mapping between data mover calls and data mover tasks, each data mover task comprising a read phase and a write phase; executing each data mover task in the set by issuing memory access commands corresponding to each stage of each particular one of the set of data mover tasks to execute the stage of the particular one of the set of data mover tasks; concurrently executing a plurality of memory commands of the set of data mover tasks on a plurality of memory interface slices, the plurality of memory commands targeting memory locations specified in the data mover calls; determining that all of the data mover tasks invoked by the data mover have completed; and A response to the data mover call is sent to the host.
9. The memory controller device of claim 8, wherein the data mover call comprises a first memory location, a stride, an element size, and a number of elements, and wherein the set of data mover tasks accesses a plurality of memory locations indicated by the number of elements and starting from the first memory location and incrementing by the stride.
10. The memory controller device of claim 8, wherein the data mover call includes the location of an address list, and wherein one of the set of data mover tasks includes extracting the address list.
11. The memory controller device of claim 8, wherein the data mover call comprises a first location and a second location storing an offset list, and wherein one of the set of data mover tasks comprises extracting the offset list.
12. The memory controller device of claim 8 , wherein the data mover tasks include a task of reading from one continuous memory location and writing data into one to multiple continuous memory locations, a task of reading from one to multiple memory locations and writing data into one continuous memory location, and a task of writing into one to multiple memory locations.
13. The memory controller device of claim 8, wherein the operations further comprise classifying the data mover call into a selected category of one of a plurality of predetermined categories; and Wherein creating the set of data mover tasks based on the data mover call includes utilizing the selected category to determine an initiation rate, a slice interleaving strategy, or an allocation strategy.
14. The memory controller device of claim 8, wherein the data mover task includes a fetch phase.
15. A non-transitory machine-readable medium storing instructions for offloading memory access workload from a host system in a distributed memory architecture, the instructions, when executed by a processor of a memory controller, causing the memory controller to perform operations comprising: receiving a data mover call, the data mover call comprising a request to copy a value, a request to gather a value, a request to scatter a value, or a request to set a value, the data mover call specifying one or more source memory locations and one or more destination memory locations; creating a set of data mover tasks based on the data mover calls, the data mover tasks being created based on a specified mapping between data mover calls and data mover tasks, each data mover task comprising a read phase and a write phase; executing each data mover task in the set by issuing memory access commands corresponding to each stage of each particular one of the set of data mover tasks to execute the stage of the particular one of the set of data mover tasks; concurrently executing a plurality of memory commands of the set of data mover tasks on a plurality of memory interface slices, the plurality of memory commands targeting memory locations specified in the data mover calls; determining that all of the data mover tasks invoked by the data mover have been completed; and A response to the data mover call is sent to the host.
16. The non-transitory machine-readable medium of claim 15, wherein the data mover call comprises a first memory location, a stride, an element size, and a number of elements, and wherein the set of data mover tasks accesses a plurality of memory locations indicated by the number of elements and starting from the first memory location and incrementing by the stride.
17. The non-transitory machine-readable medium of claim 15, wherein the data mover call includes the location of an address list, and wherein one of the set of data mover tasks includes extracting the address list.
18. The non-transitory machine-readable medium of claim 15, wherein the data mover call comprises a first location and a second location storing an offset list, and wherein one of the set of data mover tasks comprises extracting the offset list.
19. The non-transitory machine-readable medium of claim 15, wherein the data mover tasks include tasks that read from one contiguous memory location and write data into one to multiple contiguous memory locations, tasks that read from one to multiple memory locations and write data into one contiguous memory location, and tasks that write into one to multiple memory locations.
20. The non-transitory machine-readable medium of claim 15, wherein the operations further comprise: classifying the data mover call into a selected category of one of a plurality of predetermined categories; and Wherein creating the set of data mover tasks based on the data mover call includes utilizing the selected category to determine an initiation rate, a slice interleaving strategy, or an allocation strategy.