Controller for memory expansion, memory expansion device, and data processing method therefor

The CXL.MEM protocol-based Memory-Mapped NDP architecture addresses the challenges of high overhead and cost in existing NDP structures by enabling low-overhead, universal NDP with enhanced flexibility and scalability, achieving efficient and cost-effective memory expansion.

WO2025095324A1PCT designated stage expired Publication Date: 2025-05-08POSTECH ACADEMY INDUSTRY FOUNDATION
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/013322
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-09-04
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing NDP structures face challenges such as high communication overhead, limited flexibility and scalability for various applications, and high costs associated with using existing GPU or CPU cores for memory-intensive workloads.

Method used

The proposed solution involves a controller and memory extension device that utilize the CXL.MEM protocol to enable low-overhead, universal NDP through the use of Memory-Mapped NDP (M2 NDP) architecture, which includes M2 FUNC for efficient communication and μTHR for cost-effective general-purpose NDP.

Benefits of technology

This approach reduces communication overhead, enhances flexibility and scalability for various applications, and achieves cost-effectiveness by minimizing resource waste and optimizing resource utilization in memory-bound workloads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024013322_08052025_PF_FP_ABST
    Figure KR2024013322_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a controller for memory expansion, which: receives a request transmitted from a host to a memory device; determines, on the basis of the address of the memory device included in the request, whether the request corresponds to a memory-related first request distinguished from a read request and a write request; if the request is the first request, identifies the type of the first request on the basis of the address of the memory device included in the request; and, if the request is the first request, executes the identified first request on the basis of parameters included in the request.
Need to check novelty before this filing date? Find Prior Art

Description

Controller for memory expansion, memory expansion device, and data processing method therefor

[0001] The present invention relates to a controller for memory expansion, a memory expansion device, and a data processing method therefor, and in particular, to a technology for general-purpose proximity data processing with low overhead for a CXL memory expansion device.

[0002] The material described in this section merely provides background information for the present embodiment and does not constitute prior art.

[0003] Compute Express Link (CXL) is emerging as a widely adopted new interconnect standard for high-performance communication between processors, accelerators, and memory expansion units in a system.

[0004] CXL memory expansion devices can cost-effectively expand overall system memory for a variety of workloads that require large amounts of memory. While the additional latency and limited bandwidth of the CXL interface can significantly impact latency-sensitive and bandwidth-intensive workloads, the NDP of the CXL memory expansion device offers significant opportunities to effectively address this issue. Compute-bound workloads or workloads with small working sets that fit within on-chip caches can run more efficiently on the host. Memory-bound workloads tend to perform few calculations on each data access, so they are inherently low in arithmetic intensity (e.g., FLOPs / byte).

[0005] For bandwidth-intensive workloads exhibiting high memory traffic loads, GPUs are currently the most widely used devices for acceleration. A common pattern for executing these workloads on GPUs is to assign each thread a specific piece of data from a large (multidimensional) array using the Single-Program Multiple-Data (SPMD) model. At kernel startup, each thread computes the address of the allocated data using a linear equation of the thread block ID, block size, and thread ID. However, address generation can account for a significant portion of the dynamic instruction count in various GPU workloads.

[0006] The CXL standard shares the physical layer of the PCIe standard and protocol and defines three protocols. CXL.io is functionally identical to PCIe and is used for device discovery, enumeration, and management. CXL.cache allows CXL devices to access host memory via a cache coherency protocol. CXL.mem enables memory expansion via CXL. Specifically, CXL.mem is a memory-semantic protocol that allows the processor to access data in a CXL memory expansion device by simply executing load / store instructions, while providing lower latency compared to CXL.io.

[0007] The CXL specification also defines three device types based on the supported protocols. All types must natively support CXL.io for device management. The first type of device is a memory-less accelerator (e.g., smart NIC) that uses CXL.cache. The second type of device is a cache-coherent accelerator (e.g., GPUs and FPGAs) with memory that uses CXL.cache and CXL.mem. The third type of device is a memory expansion device that supports CXL.mem.

[0008] When using CXL.mem, the host can access the CXL memory expansion device using the Host Physical Address (HPA). The memory of the CXL memory expansion device is managed by the host processor and may be referred to as Host-managed Device Memory (HDM). A third type of CXL memory expansion device may use the HDM-H (host-only coherent) or HDM-DB (device coherent using back-invalidation) coherence model. The HDM-H is for passive memory expansion devices that do not manipulate memory exposed to the host. In contrast, the HDM-DB assumes that the CXL memory expansion device has a Device Coherence Agent (DCOH) with a snoop filter that can track the host's HDM caching and perform back-invalidation (BI) when necessary. Therefore, the above HDM-DB is suitable for a CXL memory expansion device with NDP functionality, and the present invention assumes an HDM-DB model. A CXL memory expansion device can perform BI for the host cache, but if BI occurs frequently during NDP, performance may deteriorate.

[0009] CXL 3.0 also supports direct peer-to-peer access, which allows a CXL device to directly access the HDM of another CXL device through a CXL switch. Routing by the switch is performed using the HDM decoder registers. This feature can be useful for NDP of multiple CXL memory expansion devices. For virtual-to-physical address translation of a CXL device, the host can obtain translation information using the Address Translation Service (ATS) defined in PCIe. However, this can incur latency of several μs due to protocol overhead and the host's page table walk. To reduce overhead, the CXL device may have an Address Translation Cache (ATC) that stores recently used translation information. When the page table is updated, the host can invalidate the ATC of the CXL device to prevent incorrect translations. In other words, frequent use of the address translation service via the CXL.io protocol can incur high performance overhead.

[0010] Meanwhile, Samsung Electronics Co., Ltd.'s CXL-PNM utilizes a computational unit in CXL memory to perform computations that accelerate the execution of generative language models. It offloads matrix multiplication operations via the CXL.io protocol.

[0011] CXL-ANNS accelerates approximate nearest neighbor search operations for proximity data processing on CXL-based memory devices.

[0012] When performing compute offloading using existing protocols like CXL.io or PCIe, multiple software and hardware steps are required, resulting in significant communication overhead with the host processor. This overhead can be significant, especially for fine-grained offloading, making it crucial to reduce it.

[0013] Existing NDP architectures are limited in their ability to support general-purpose operations. They are designed to accelerate specific applications, limiting flexibility and scalability for diverse applications.

[0014] Using existing GPU cores and CPU cores for NDP can be costly and power-intensive, and there are issues with running memory-intensive workloads cost-effectively.

[0015] On the other hand, the method of adding a new packet format by extending the existing standard has the problem that it cannot be used in host processors that only support the existing standard.

[0016] The present invention aims to provide a controller for memory expansion capable of reducing communication overhead, a memory expansion device, and a data processing method therefor.

[0017] In the present invention, M is provided to enable a universal NDP in a CXL memory expansion device. 2 I would like to provide NDP.

[0018] The present invention seeks to provide an architecture that does not require host processor hardware modification based on an unmodified CXL.mem protocol.

[0019] In the present invention, M 2 func and M 2 We would like to provide a memory expansion device that provides μthr functionality.

[0020] In the present invention, M 2 We aim to provide a controller and memory expansion unit that enables the NDP kernel to effectively avoid cold DRAM-TLB misses that cause long delays in ATS, with proposed OS support for DRAM-TLB pre-loading using func.

[0021] The present invention aims to provide a controller and memory expansion device capable of executing various general-purpose operations while taking cost efficiency and programmability into consideration.

[0022] The present invention seeks to provide a controller and memory expansion device that enables μthreads to initiate data access without duplicate address calculations performed across threads of a GPU warp by using one of the input data arrays of a data-parallel workload as a μthread pool area.

[0023] In the present invention, M can increase the utilization of computational units more than GPU through fine-grained resource management. 2 We want to provide μthr functionality.

[0024] According to one aspect of the present invention, a controller for memory expansion is provided, wherein the controller receives a request transmitted from a host to a memory device, determines whether the request corresponds to a first memory-related request that is distinguished from a read request and a write request based on an address of the memory device included in the request, and, if the request is the first request, identifies the type of the first request based on the address of the memory device included in the request, and, if the request is the first request, executes the identified first request based on an argument included in the request.

[0025] At this time, if it is determined that an additional read request has been received for the address of the first request after the request has been determined to be the first request, the processing status for the first request stored in the address of the first request of the memory device can be transmitted to the host.

[0026] At this time, the controller includes a packet filter, and an address area for distinguishing the first request is allocated and stored in the packet filter, and the packet filter can determine that the request is the first request if an address included in the request is included in the address area.

[0027] At this time, an address area for distinguishing the first request may be allocated to the packet filter, and a predetermined protocol may be used when receiving the request.

[0028] At this time, when assigning an address area to distinguish the first request to the packet filter, a predetermined first protocol from the host may be used, and when receiving the request from the host, a predetermined second protocol may be used.

[0029] At this time, the controller may include an NDP controller that executes the first request; and a plurality of NDP units controlled by the NDP controller. When the first request is a kernel launch request, the NDP units (generators) generate a plurality of microthreads that are mapped to different addresses in a first area (base&bound) of virtual memory, and each of the microthreads may execute an operation on the mapped memory address.

[0030] At this time, all microthreads running within any NDP unit among the plurality of NDP units can share data with each other through one or more on-chip scratchpad memories among the on-chip scratchpad memories included in each of the plurality of NDP units.

[0031] At this time, the controller may further include a DRAM-TLB for converting a virtual address of the virtual memory into a physical address of the memory device. If the request is a DRAM-TLB preloading request among the types of the first request, the DRAM-TLB may be used by preloading the DRAM-TLB.

[0032] According to one aspect of the present invention, a memory expansion device may include a memory device; and a controller disposed between a host and the memory device, the controller including a plurality of NDP units. When the controller receives a kernel launch request, the NDP units (generators) generate a plurality of micro-threads mapped to different addresses in a first area (base&bound) of a virtual memory, each of the micro-threads executes an operation on the mapped memory address, and the result of the execution may be stored in a result storage address defined in the kernel of the memory device.

[0033] At this time, all microthreads running within any NDP unit among the plurality of NDP units can share data with each other through one or more on-chip scratchpad memories among the on-chip scratchpad memories included in each of the plurality of NDP units.

[0034] At this time, the controller receives a request transmitted from the host to the memory device, and determines whether the request corresponds to a first memory-related request that is distinguished from a read request and a write request based on an address of the memory device included in the request, and if the request is the first request, the controller identifies the type of the first request based on the address of the memory device included in the request, and if the request is the first request, executes the identified first request based on an argument included in the request.

[0035] At this time, the NDP unit further includes a DRAM-TLB for converting a virtual address of the virtual memory into a physical address of the memory device, and when the request is a DRAM-TLB preloading request among the types of the first request, the DRAM-TLB can be used by preloading the DRAM-TLB.

[0036] According to one aspect of the present invention, a data processing method for a memory expansion device may include the steps of: receiving, by a controller, a request transmitted from a host to a memory device; determining, based on an address of the memory device included in the request, whether the request corresponds to a first memory-related request distinguished from a read request and a write request; identifying, if the request is the first request, a type of the first request based on the address of the memory device included in the request; and executing, if the request is the first request, the identified first request based on an argument included in the request.

[0037] At this time, after the request is determined to be the first request, the controller may further include a step of determining that an additional read request has been received for the address of the first request; and a step of the controller transmitting a processing status for the first request stored in the address of the first request of the memory device to the host.

[0038] At this time, the controller may further include a step of allocating and storing an address area for distinguishing the first request within the controller, and the determining step may be a step of determining the request as the first request if the address included in the request is included in the address area.

[0039] At this time, an address area for distinguishing the first request may be allocated to the packet filter, and a predetermined protocol may be used when receiving the request.

[0040] At this time, the step of allocating and storing an address area for distinguishing the first request may be a step of using a predetermined first protocol from the host, and the step of receiving the request from the host may be a step of using a predetermined second protocol.

[0041] At this time, if the first request is a kernel launch request, the controller may further include a step of controlling a plurality of NDP units included in the controller to create a plurality of micro-threads mapped to different addresses in a first area (base&bound) of virtual memory; and a step of causing the controller to cause each of the micro-threads to perform an operation on the mapped memory address.

[0042] At this time, all microthreads running within any NDP unit among the plurality of NDP units can share data with each other through one or more on-chip scratchpad memories among the on-chip scratchpad memories included in each of the plurality of NDP units.

[0043] At this time, if the request is a DRAM-TLB preloading request among the types of the first request, the controller further includes a step of preloading the DRAM-TLB, and the DRAM-TLB can be used to convert a virtual address of the virtual memory into a physical address of the memory device.

[0044] According to the present invention, a controller for memory expansion capable of reducing communication overhead, a memory expansion device, and a data processing method therefor can be provided.

[0045] According to the present invention, to enable universal NDP in a CXL memory expansion device, M 2 NDP can be provided.

[0046] According to the present invention, an architecture can be provided that does not require host processor hardware modification based on an unmodified CXL.mem protocol.

[0047] According to the present invention, M 2 func and M 2 A memory expansion device providing μthr functionality may be provided. M 2 func supports low-overhead NDP offloading and management on the host processor via CXL.mem, overcoming the high overhead of CXL.io for fine-grained NDP offloading while maintaining standards compatibility. M 2μthr implements a lightweight FGMT using RISC-V with vector extensions, reducing the overhead of redundant address computation compared to dedicated SIMT GPUs. Its wide on-chip scratchpad memory range allows for cost-effective NDP, reducing DRAM or SRAM traffic. Furthermore, initializing shared memory in each thread block requires additional intra-block synchronization, but the wide on-chip scratchpad memory range reduces the frequency of synchronization, improving performance. Furthermore, fine-grained μthread creation prevents resource waste due to thread block-granular resource allocation.

[0048] According to the present invention, M 2 With the proposed OS support for DRAM-TLB preloading using func, the NDP kernel can effectively avoid cold DRAM-TLB misses that cause long delays in ATS.

[0049] The present invention provides a controller and memory expansion device capable of executing various general-purpose operations while considering cost efficiency and programmability. Accordingly, the present invention provides a controller and memory expansion device with enhanced data processing capabilities. This cost-effective structure allows for improved processing times and reduced energy and hardware costs in large-scale cloud systems and large-scale computational processing systems.

[0050] According to the present invention, a controller and a memory expansion device can be provided that enable a μthread to start accessing data without duplicate address calculations performed across threads of a GPU warp by using one of the input data arrays of a data-parallel workload as a μthread pool area.

[0051] According to the present invention, M 2The μthr feature can be used to achieve higher compute unit utilization than the GPU through fine-grained resource management.

[0052] FIG. 1 is a drawing for explaining the configuration of a memory expansion device according to one embodiment of the present invention.

[0053] Figure 2 is an M according to one embodiment of the present invention. 2 This is a diagram to explain the func.

[0054] Figure 3 is an M according to one embodiment of the present invention. 2 This table shows predefined NDP management functions for various offsets based on the func area.

[0055] Figure 4 is a CXL.io or M according to one embodiment of the present invention. 2 A timeline showing NDP offloading using func.

[0056] Figure 5 illustrates a CPU, GPU, and M according to one embodiment of the present invention. 2 This table shows an architectural comparison between NDPs.

[0057] FIG. 6 is a graph showing the ratio of active threads running on SM or NDP units over time for the main kernel of the pagerank benchmark according to one embodiment of the present invention.

[0058] FIG. 7 illustrates the microarchitecture of an NDP unit according to one embodiment of the present invention.

[0059] FIG. 8 is a drawing for explaining the structure of an NDP kernel according to one embodiment of the present invention.

[0060] Figure 9 is an M according to one embodiment of the present invention. 2 This is a diagram to explain the system when using func.

[0061] Fig. 10 is an M according to another embodiment of the present invention. 2This is a diagram to explain the system when using func.

[0062] Figure 11 is an M according to one embodiment of the present invention. 2 This is a flowchart to explain the data processing method for a memory expansion device when using func.

[0063] Figure 12 is an M according to one embodiment of the present invention. 2 M without using func 2 This is a diagram to explain the system when only μthr is used.

[0064] Figure 13 is an M according to one embodiment of the present invention. 2 M without using func 2 This is a drawing to explain the case where only μthr is used.

[0065] Figure 14 is an M according to one embodiment of the present invention. 2 This is a flowchart explaining the data processing method for a memory expansion device when using μthr.

[0066] FIG. 15 is a conceptual diagram illustrating an example of a generalized controller, memory expansion device, or computing system capable of performing at least a portion of the processes of FIGS. 1 to 14.

[0067] In addition to the above purpose, other objects and features of the present invention will become apparent through the description of embodiments with reference to the attached drawings.

[0068] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.

[0069] Terms such as first, second, A, and B may be used to describe various components, but the components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component. The term "and / or" includes any combination of multiple related listed items or any one of multiple related listed items.

[0070] In the embodiments of the present application, “at least one of A and B” may mean “at least one of A or B” or “at least one of combinations of one or more of A and B.” Furthermore, in the embodiments of the present application, “at least one of A and B” may mean “at least one of A or B” or “at least one of combinations of one or more of A and B.”

[0071] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0072] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0073] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0074] Hereinafter, with reference to the attached drawings, preferred embodiments of the present invention will be described in more detail. In order to facilitate an overall understanding in describing the present invention, identical reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted.

[0075] A primary use case for CXL is memory expansion via the memory-semantic CXL.mem protocol, which enables low-latency remote memory access via load / store instructions. CXL.mem's latency is reportedly significantly lower than PCIe and comparable to cross-socket NUMA latency, allowing a host to cost-effectively increase its memory capacity beyond its limited DIMM slots. This capability can be particularly useful for workloads with large memory footprints, including large language models, big data, and graph analytics. CXL is already supported by mainstream CPU products, and several CXL memory prototypes with capacities up to 512 GB and a bandwidth (BW) of PCIe 5.0 have been announced.

[0076] However, the CXL link between the host and the device can be a bottleneck for bandwidth-intensive applications because its link BW is lower than the internal BW within the CXL memory expansion device. Furthermore, memory access via CXL.mem incurs additional latency through the protocol stack, whereas this latency overhead can be avoided within the CXL memory expansion device.

[0077] Therefore, compared to directly accessing the data in the CXL memory expansion device for computation on the host processor, performing NDP (Near-Data Processing) within the CXL memory expansion device can provide significant speedups for memory-bound workloads with low arithmetic intensity. However, previous approaches implement application-specific NDP hardware logic in the CXL memory expansion device. Since the main motivation for the CXL memory expansion device is to cost-effectively increase system memory capacity, introducing different ASICs for different NDP targets may not be a practical approach due to the high overall area and NRE costs. FPGAs can be tuned to target workloads, but they suffer from programmability issues. On the other hand, using general-purpose CPU or GPU cores for NDP in the CXL memory expansion device can significantly increase the cost of the CXL memory expansion device, making it unsuitable for cost-effective memory expansion with NDP functionality.

[0078] The CXL memory expansion unit employed in recent systems offers significant opportunities for performance enhancement by performing Near Data Processing (NDP) on the CXL memory expansion unit. However, because CXL memory expansion units are proposed as a cost-effective alternative to using larger or more CPUs for greater memory capacity, the industrial viability of these NDP architectures depends on supporting general-purpose computing while reducing overhead to avoid the high costs associated with producing multiple ASICs.

[0079] To reduce overhead, the present invention uses M 2 It can provide a low-overhead general-purpose NDP architecture for CXL memory, called NDP (Memory-Mapped NDP).

[0080] Because the present invention targets workloads with large memory footprints that do not fit within the host cache, it assumes that HDM data is not cached on the host when NDP is initiated. If necessary, the host can quickly flush HDM data from the cache using the current CPU's hardware support.

[0081] For memory mapping to CXL memory expansion units, we assume fine-grained 256-byte interleaving using hashing across memory channels to balance load within the CXL memory expansion units. If the system has multiple CXL memory expansion units, we assume that each page (4KB or 2MB) is mapped to a single CXL memory expansion unit, as is the case in current NUMA or multi-GPU systems.

[0082] The present invention targets memory-bound (bandwidth-bound or latency-bound) workloads with large memory spaces that do not fit into the on-chip caches of the host processor.

[0083] The present invention relates to computer architecture, memory systems, and large-scale data processing.

[0084] FIG. 1 is a drawing for explaining the configuration of a memory expansion device according to one embodiment of the present invention.

[0085] The above memory expansion device (CXL Memory Expander with M 2 The NDP) may include a CXL controller (1000) and a memory device (e.g., a plurality of DRAMs) (100) connected to the CXL controller (1000). In the embodiment of FIG. 1, the memory expansion device is illustrated as including a DRAM, but an SSD or the like may be connected to the CXL controller instead of the DRAM.

[0086] The host (CPU) (200) can access the memory expansion device via the CXL.io and CXL.mem protocols.

[0087] The host (200) can be connected to any memory expansion device among multiple memory expansion devices via a CXL switch (300) (link).

[0088] The CXL controller (1000) may include a packet filter (1100), an NDP controller (1200), multiple NDP units (1300), a cross bar (1400), multiple caches, and multiple memory controllers (1600).

[0089] Each of the multiple NDP units (1300) can be connected to any of the caches (1500) via a crossbar (1400). When connected to any of the caches, the DRAM (100) can be accessed through a memory controller (1600) matching the cache.

[0090] While the NDP of the above memory expansion device can significantly improve system performance for various applications, conventional technologies propose application-specific logic, resulting in high NRE and hardware cost when supporting diverse applications. While general-purpose NDPs could potentially overcome these limitations, utilizing existing CPU or GPU cores for NDPs can still result in high costs, as certain technologies are not necessarily optimized for memory-bound workloads. In other words, these specific technologies are designed for compute- and memory-bound workloads.

[0091] In order to achieve high performance / cost efficiency of general purpose NDP, the present invention uses the above-described CXL-M 2 M of CXL memory expansion device called NDP 2We propose a Memory-Mapped Near-Data Process (NDP). In the present invention, a CXL memory expansion device is a memory expansion device, a CXL memory, or a CXL-M. 2 It can be referred to as NDP. The above M 2 NDP has two mechanisms: 1) low-overhead NDP management based on unmodified CXL.mem and 2) M for offloading. 2 func(Memory-Mapped function) and 2) M for cost-effective NDP microarchitecture 2 It can be configured with μthr (Memory-Mapped μthreading).

[0092] Specifically, the above M 2 func is a CXL.mem-compatible low-overhead communication mechanism / method (i.e., low-cost communication mechanism) between the host processor (200) of the CXL memory expansion device and the NDP controller (1200). The M 2 func example (hereinafter, M 2 According to the func), the packet filter (1100) located at the input port of the CXL memory expansion device can be used to check whether a memory request corresponds to a pre-allocated memory range, and the NDP management function connected to the corresponding address can be executed. Since the NDP management function can be executed simply through memory access in the host (200), the delay time of NDP offloading can be minimized compared to the existing CXL.io / PCIe-based offloading method without modifying the CXL.mem standard.

[0093] Above M 2μthr is a multithreading method that minimizes address calculation overhead by creating a number of lightweight microthreads (μthreads) that allocate only as much of the architectural registers as necessary to realize a cost-effective general-purpose NDP, executing them concurrently, and distributing operations through mapping each to a specific memory address. In other words, the above M 2 The μthr embodiment can enable the design of low-cost NDP devices for general-purpose computing by introducing lightweight microthreads (μthreads) that support concurrent execution of NDP kernels while minimizing resource waste.

[0094] Above M 2 According to the μthr embodiment, the DRAM access latency can be hidden by enabling highly parallel execution of the kernel with minimal resource waste through lightweight microthreads that allocate only as much of the architectural registers as needed. 2 According to the μthr embodiment, address calculations can be reduced and the RISC-V vector extension ISA can be utilized through microthreads directly linked to specific memory locations. The above M 2 μthr Example (hereinafter, M 2 According to μthr, all microthreads running on the same NDP unit can share data in scratchpad memory, reducing DRAM traffic compared to the GPU. At this time, M 2 The μthr mapping unit is identical to the DRAM access unit, which avoids bottlenecks while maintaining high vector ALU unit utilization. However, in other embodiments, if the DRAM access unit is, for example, 32B, the M2μthr mapping unit can be 4B or 8B, which are smaller than the DRAM access unit, or 64B, which is larger than the DRAM access unit. Even in these cases, performance can be advantageous.

[0095] Above M2 NDP can be implemented on the controller chip of a CXL memory expansion device that also supports generic CXL.mem transactions.

[0096] Figure 2 is an M according to one embodiment of the present invention. 2 This is a diagram to explain the func.

[0097] Figure 3 is an M according to one embodiment of the present invention. 2 This table shows predefined NDP management functions for various offsets based on the func area.

[0098] The fields in the table in Figure 3 are Offset, Description, Privileged, M 2 func may contain arguments and a return value.

[0099] Hereinafter, the description will be given with reference to FIGS. 1 to 3.

[0100] The present invention can utilize the unmodified CXL.mem protocol. Conventionally, to utilize NDP for both fine-grained computation offloading and rough-grained offloading, the host and CXL-M 2 Communication latency between NDPs must be minimized. While the CXL.mem protocol offers low latency, the standard only defines packet types for general CXL memory access and cannot be directly used for other communications. Extending CXL.mem to support custom packet types would break compatibility across host processors and hinder widespread adoption of the NDP architecture. In contrast, CXL.io can be used for ad hoc communications, but it incurs higher latency in the protocol stack and requires context switches to the OS for privileged IO device communications, further increasing latency.

[0101] So, using unmodified CXL.mem, CXL-M on the host 2The M of the present invention is to enable low overhead and flexible communication with NDP. 2 func can be used, i.e. M 2 It is to reserve some physical memory space of the memory expansion device for host communication, called func area. There are two purposes of packets: general memory access (read request / write request), or M 2 In order to distinguish between a call to a func and a packet filter (1100), a packet filter (1100) is introduced to the input port of the memory expansion device to inspect all packets entering the memory expansion device and determine whether the packet is a general memory access or an M based on the packet address. 2 Determines whether it should be interpreted as a func call.

[0102] Below, the method of distinguishing between the two purposes through the packet filter is described in detail.

[0103] Figure 2 shows the host code, the host processor, and the CXL-M. 2 NDP is shown.

[0104] The above host code is a general vector instruction code that executes a vector store instruction to store the v1 vector register value at the address stored in a scalar register (e.g., x7).

[0105] In the above host processor box, the x7 register value points to address 0x10040. The v1 register stores the content to be written, for example, information required for kernel execution.

[0106] CXL-M as a general vector write instruction via CXL.mem write packet (Addr:[0x10040], Data:[0, 1, 0xA000, 0xA1FF, ...]) 2NDP, that is, when data is transmitted to the memory expansion device, the packet filter (1100) of the memory expansion device examines the address area of ​​the CXL.mem write packet (e.g., the address value stored in x7), and if the address of the write packet is included in the address area pre-registered in the packet filter (1100), a pre-registered function can be executed.

[0107] Specifically, as shown in the table of the packet filter (1100) of Fig. 2, an address area for distinguishing the first request may be allocated and stored in the packet filter (1100). The field of the table is M 2 It may contain a func area field and an ASID field. M 2 The func area field can be divided into areas such as 0x10000-0x1FFFF, 0x20000-0x2FFFF, etc. An ASID (Address Space Identifier) ​​can be matched and stored in the address area of ​​each record.

[0108] For example, in the case of the above write packet, the x7 address is 0x10040, so the above M 2 It can be seen that the first area (i.e., the first record) of the func area is included. Therefore, the above write packet is not a write packet that simply accesses the memory itself (i.e., not a general memory access), but a memory-related request (i.e., M) to execute pre-registered functions. 2 You can see that it is a func call.

[0109] M in Fig. 2 2In the table representing the func area, the first record defines 0x10000-0x1FFFF. Referring to FIGS. 2 and 3 together, for example, functions can be pre-registered such that 0x10000 matches offset 0<<5, 0x10020 matches offset 1<<5, and 0x10040 matches offset 2<<5. In the embodiment of FIG. 2, the address of the packet points to 0x10040, so it corresponds to the NDP kernel launch corresponding to offset 2<<5, and the arguments can include Synchronicity(Sync / async), NDPKernelID, μthreadPoolRegion(base, bound), KernelArgSize(bytes), and KernelArguments. The return value at this time can be a kernel instance ID or -1 indicating an error. The return value can be stored in the address of the packet, for example, the address 0x10040 of the memory device. Alternatively, in another embodiment, the return value may be stored in the address of the packet, for example, in the address of a memory buffer within the controller (1000).

[0110] For example, the kernel launch in Fig. 2 may be for a VectorAdd NDP kernel that computes C=A+B. At this time, vectors A, B, and C may be placed in OxA000, 0xB000, and OxC000, respectively. At this time, OxA000, 0xB000, and OxC000 may all point to address areas of virtual memory. Each μthread (μthr) can compute a 32B (8x4B) partial vector output. For the sake of brevity, other data path components are not shown in Fig. 2. In the case of Fig. 2, an actual implementation may use μthreads that use SIMD (single instruction, multiple data) operations and a larger intersection placement interval. For example, 32B units of μthreads may be placed in larger 64B units between NDP units. For example, uthr0 and uthr1 in FIG. 2 may be placed in the NDP unit (1310), and uthr2 and uthr3 may be placed in the NPU unit (1320).

[0111] For example, in another embodiment, M of the packet filter of FIG. 2 2 It can be assumed that an address area containing 0x00FF0000 (or 0x00FF0020) is pre-registered in the func area. At this time, the address of the write packet, 0x00FF0000 (or 0x00FF0020), can match offset 0 (or offset 1<<5) in Fig. 3. Therefore, NDP kernel registration (NDP kernel deregistration) corresponding to offset 0<<5 (or offset 1<<5) can be executed.

[0112] As shown in Fig. 3, M 2The func call can perform various functions including registering (offset 0<<5) the NDP kernel, unregistering (offset 1<<5), or executing (offset 2<<5), and the execution of these functions can be performed by the NDP controller (1200). In Fig. 3, only functions 0 to 5 are shown, but more diverse functions can be pre-registered and used. That is, M of the CXL.mem packet 2 You can call different functions using addresses with different offsets based on the func area.

[0113] If the request received from the host (200) is a general memory access request, the packet filter (1100) can execute the request by being connected to the cross bar (1400). For example, if the packet filter (1100) receives the general memory access request to the DRAM, if the address of the request is a read request for data being processed while the current kernel is being executed, an incomplete result of the processing may be returned to the host (200), and if the address of the request is an address for data unrelated to the kernel execution, simple access to the data may be possible. That is, according to the present invention, DRAM access is possible even during NDP processing such as kernel execution.

[0114] Meanwhile, the NDP controller (1200) can be implemented with a low-cost core similar to the microcontroller of the GPU. M 2 For func initialization, each user-level process of the host (200) must cache M that cannot be cached in the above memory expansion device. 2 func allocates a region. The CXL memory driver can use CXL.io to insert the region's address range into the packet filter. Once initialized, CXL.io is no longer needed for NDP, and CXL.mem is used for general read / write and NDP-related communications (M 2func call) or through the above CXL memory driver. 2 You can also use func to insert the address range of a region into the packet filter. In this case, CXL.io may not be used. In this example, M 2 When using func, it is described as using CXL.mem, but it can also be based on other protocols similar to CXL, such as CCIX.

[0115] As described above, M 2 For a func call, the write request format is used to include arguments in the data write portion of the request. To send the request, the host process (200) executes a store command using a register that holds the arguments. Vector / SIMD registers can be used to send multiple arguments up to the size of the vector register. M 2 Since the func area is not cacheable, writes bypass the host cache. However, the response to a write request cannot include the return value data of the NDP controller (1200) using CXL.mem (i.e., the return value for each function in FIG. 3). Therefore, a subsequent read request to the same address is used to access the return value of the latest function call by the current process. Since the return value is accessed through a normal memory access, the NDP controller (1200) can store the return value of the function at that memory address and process the read request as a normal access. For the correct order, the host code must have M 2 There must be a fence command between the func request and the read request.

[0116] As described in the NDP kernel structure below, different NDP kernels may require different amounts of register and scratchpad memory resources. Therefore, they must be specified as metadata when registering the kernel. In addition, the kernel argument size must be specified so that the arguments can be properly extracted from the kernel launch packet. The metadata of registered NDP kernels is also stored in the M of the current host process. 2 It is stored in the func area, and in addition to the offsets used in Figure 3, it is pre-registered and starts at a predetermined location. Therefore, the host can easily access the kernel metadata from memory when needed. M 2 Because func's memory area is allocated for each process, it is protected from other processes by the host's virtual memory system.

[0117] Figure 4 is a CXL.io or M according to one embodiment of the present invention. 2 A timeline showing NDP offloading using func.

[0118] The timeline on the left is when using CXL.io, and the timeline on the right is when using M 2 This is when using func.

[0119] Here, we assume a synchronous launch and an NDP kernel execution time of 10 μs. We also assume a 3 μs latency for round-trip CXL.io / PCIe and kernel overhead, and a 180 ns load-use latency for host CXL.mem accesses.

[0120] Compared to NDP kernel launch using CXL.io, which causes a context switch to the kernel and longer CXL.io protocol latency, as shown in the timeline on the left, M 2 Kernel launch using func is performed in user space and has lower latency overall, i.e. M compared to when using the CXL.io protocol.2 When using func, we can achieve ~42% speedup for a 10μs kernel. Although a separate CXL.mem read request (CXL.mem rd req) and a Barrier (barrier between CXL.mem wr rsq and CXL.mem rd req) are required, they overlap with kernel execution. That is, although the requirement of a barrier (or fence) may seem to incur a large overhead, this overhead is overlapped (obscured) with the NDP kernel execution time as shown in Figure 4, so it does not affect the end-to-end NDP execution time from the user's perspective. If the NDP kernel is extremely short, the barrier may not be completely obscured, but CXL.mem will still have a lower overall latency than CXL.io. In this case, the Barrier can correspond to the fence instruction described above.

[0121] In the embodiment of FIG. 4, the host (200) received a read response (CXL.mem rd rsp) from the CXL memory expansion device (specifically, the CXL controller (1000)) after the NDP kernel was terminated, so the synchronization argument of the kernel launch may be 'sync' with reference to FIG. 3. For example, referring to FIGS. 2 and 3 together, the synchronization argument may have a value of '0' if it is a synchronous state, and a value of '1' if it is an asynchronous state.

[0122] The present invention focuses on supporting NDP offloading using CXL.mem to minimize overhead as described above, while not excluding the use of CXL.io to provide NDP management functions (e.g., M 2(used when registering address ranges in the func area to packet filters). For a rough NDP kernel, the CXL.io overhead can be well amortized over long kernel runtimes.

[0123] Hereinafter, the NDP kernel launch will be described with reference to FIGS. 1 to 4.

[0124] Thanks to the low-overhead CXL.mem protocol, M 2 func can launch the NDP kernel with minimal overhead. The NDP kernel launch uses a store instruction that generates a write request containing kernel launch arguments as shown in Fig. 2, and M corresponding to offset 2<<5 in Fig. 3. 2 It can be performed by calling func. At this time, in Fig. 3, M 2 The func argument and the NDP kernel argument are shown together, and the above M 2 The func argument is an argument to the kernel launch function, which determines how to launch the kernel, and the NDP kernel argument is different in that it is an argument used directly in the NDP kernel code. Large data for the kernel (e.g., an array) can also be stored in a separate memory location in the CXL memory expansion device, and a pointer to it can be passed as an argument. Each NDP kernel instance is M 2 The μthread mechanism can be used to associate a virtual memory area for input or output data arrays, called μthread pool area, provided in the kernel launch call for NDP kernel launch. Referring to Fig. 4, M 2 After launching the kernel with a write packet using func, the NDP controller (1200) can always immediately send back an acknowledge packet (CXL.mem wr rsp).

[0125] Referring to FIG. 4, the host (200) has the same M 2func offset 2<<5 may have a memory fence (Barrier) and a load command (CXL.mem rd req) to get the return value for the kernel launch function. If the CXL memory expansion device received a write request from the host (200) for kernel launch, this time it receives a read request from the host (200). However, the read response (CXL.mem rd rsp) including the return value may be transmitted to the host (200) differently depending on the synchronization argument (see FIG. 3) provided at the time of kernel launch. At this time, the return value may be returned immediately in the case of an asynchronous launch, but may be returned after the kernel is terminated in the case of a synchronous launch. The asynchronous launch may be useful when the host (200) needs to perform other calculations that overlap with the NDP kernel. Then, the host (200) may later perform M for offset 3<<5 of FIG. 3. 2 Completion can be confirmed through the func request (NDP kernel status poll).

[0126] At this time, there may be three cases in which a return value for the kernel is returned. The first is when a read packet is received, which is an asynchronous kernel launch and returns immediately. The second is when a read packet is received and returns after the kernel launch is completed because it is a synchronous launch. The last case is when a read packet is received after a write packet for the offset 3<<5 (NDP kernel status poll) is received, and the return time for the last case can also be determined depending on synchronous / asynchronous. In the last case, since the return value cannot be sent to the host for a write packet request, the CXL controller can send the return value to the host when it receives a read packet request. Referring to FIG. 3, in the case of the offset 3<<5 (NDP kernel status poll), the CXL controller can receive a request for a write packet including an NDPKernelInstanceID as an argument from the host. The CXL controller can determine which kernel's status to poll through the argument when multiple instances of NDP kernels are running. The CXL controller can store the state of the determined kernel at the address of the write packet, and when a read packet is received thereafter, return the state of the determined kernel stored at the address to the host.

[0127] If the available resources of the NDP unit (1300) are insufficient for kernel launch due to another currently running kernel, the kernel launch request may be buffered and served after the previous kernel is completed. Referring to offset 2<<5 in Fig. 3, when the buffer is full, M 2 Since the func request kernel launch cannot be executed, the CXL controller (1000) may return -1 as a return value to the host (200) to indicate an error.

[0128] Figure 5 illustrates a CPU, GPU, and M according to one embodiment of the present invention. 2 This table shows an architectural comparison between NDPs.

[0129] For example, the GPU in FIG. 5 may be an NVIDIA GPU.

[0130] At this time, the above M 2 NDP may refer to the memory expansion device described above in Fig. 1.

[0131] In the comparison table in Figure 5, each row is for, in order, Thread creation granularity, Flynn's taxonomy, Per-thread registers, Thread creation, Thread scheduling, Out-of-order exec., Scratchpad memory scope, and Thread Identification.

[0132] To maximize memory bandwidth utilization for NDP kernels running on CXL memory expansion devices, a large number of memory accesses must be performed concurrently to hide memory latency. While out-of-order cores can perform multiple memory accesses concurrently, their high control logic overhead makes them unsuitable for cost-effective NDP.

[0133] Referring to the thread creation granularity of the first row, CPU threads are created in units of fine-grained threads, while GPU threads are created in units of coarse-grained thread blocks, and the present invention (M 2 In the case of NDP, it can be seen that each micro-thread is generated in detail.

[0134] Looking at the Flynn classification in the second row, we can see that the CPU and the present invention utilize a combination of SISD and SIMD, while the GPU utilizes only SIMD (SIMT).

[0135] Fine-Grained Multithreading (FGMT) can efficiently provide high concurrent processing, especially when there are many threads, such as GPUs. However, the SIMT-only execution of GPUs may be inefficient because threads perform redundant calculations within a warp, as it does not provide scalar operations (e.g., loop variable management, address calculation). In this case, PTX, SSAS, etc. can be used as the ISA of the GPU. Meanwhile, x86, Arm, RISC-V, etc. can be used as the ISA of the CPU. In order to solve the above-mentioned problem, the present invention adopts a RISC-V ISA (i.e., RV64GV) with vector extensions to efficiently support both scalar and SIMD operations, and the present invention's M based on concurrent FGMT 2 It can be modified to support μthr without limitation. However, in other embodiments, Arm, etc. may be used as an ISA in addition to RISC-V. In the present invention, M 2 μthr can provide a very flexible execution environment with few limitations. Therefore, an extended RISC-V ISA, i.e., RV64IMAFDV, may be adopted in other embodiments of the present invention.

[0136] For CPUs, the OS is responsible for thread creation and management, but with a large number of threads, especially short-lived ones, the overhead can be significant. For example, creating a thread on a modern OS incurs a μs-scale delay per thread.

[0137] Looking at the third row in the comparison table in Fig. 5, since CPU threads require the entire register set defined by the ISA, the register file cost can increase linearly with the number of HW threads. On the other hand, memory-bound workloads tend to require fewer registers than compute-bound workloads due to their low arithmetic intensity. Therefore, in the present invention, the register file cost can be reduced by using threads managed by GPU-style HW without a conventional OS for the CPU and providing the number of registers for each thread as specified by the SW (i.e., compiler) or the user when the NDP kernel is registered, as in offset 0<<5 in Fig. 3. For example, if 5 integer registers and 3 vector registers are specified, the kernel can only use the x0-x4 registers and the v0-v2 registers. In the present invention, this type of thread can be referred to as a microthread because of its low resource usage.

[0138] Referring to the fourth row, for CPUs, thread creation is performed by the OS. Conversely, μthread creation can be performed in hardware as quickly as on a GPU. In the present invention, on-chip scratchpad memory is introduced to facilitate efficient communication between μthreads in the NDP unit (1300).

[0139] Although there are similarities, μthreads of the present invention differ from GPU threads in several ways, beyond the differences in ISA. That is, while GPUs use a dedicated SIMT GPU ISA, the present invention uses SISD and SIMD based on a RISC-V ISA with vector extensions for μthreads.

[0140] First, GPU threads are identified by multi-dimensional thread blocks and thread indices, whereas μthreads can be identified by addresses mapped to the μthread pool area. As shown in Figure 2, by using one of the input data arrays of a data-parallel workload as the μthread pool area, a μthread can initiate data access without redundant address calculations performed across threads in a GPU warp, which can account for ~30% of the number of dynamic instructions according to previous studies. Since the mapped address is provided as a base address and offset pair, the offset can also be used to access different input / output data with different base addresses. In another embodiment, the μthreads can be identified by addresses mapped to an unallocated dummy memory address area, not the μthread pool area. In this case, since it would be an abnormal memory access for the created μthread to actually access the dummy memory address area, the access is not performed, and the address mapped to the dummy memory address area (i.e., the offset of the x2 register) can be used as the μthread ID to calculate another memory address to actually access. At this time, a kernel for this can also be implemented.

[0141] Second, as mentioned above, GPU threads are created in coarse-grained thread blocks, whereas μthreads are created in fine-grained units of individual threads. Coarse-grained thread creation can lead to resource fragmentation and underutilization due to inter-warp divergence. That is, unused resources from completed warps in a thread block will remain unused until the entire thread block completes and the resources are freed for the next thread block.

[0142] FIG. 6 is a graph showing the ratio of active threads running on SM or NDP units over time for the main kernel of the pagerank benchmark according to one embodiment of the present invention.

[0143] At this time, the above SM may mean a streaming multiprocessor of a GPU core.

[0144] The horizontal axis of the graph represents time (time(ⅹ1000cycles)), and the vertical axis represents the ratio of active threads (contexts).

[0145] The graphs represent the ratio of active threads for GPU SM (TB size: 32), GPU SM (TB size: 64), and GPU SM (TB size: 128), respectively, and the graph of the NDP unit of the present invention is M 2 μthr represents the ratio of μthreads.

[0146] At this time, the maximum number of thread blocks per SM limits the active warp ratio to 32 for the thread block (TB) size.

[0147] For example, Figure 6 shows that the active warp ratio in the GPU SM used for NDP varies over time between 0.5 and 1.0 depending on the thread block size (TB size). In contrast, according to the present invention, the resources of a completed μthread are immediately released so that they can be used by the next μthread, thereby improving resource utilization and performance / cost. While reducing the thread block size on the GPU may sometimes improve resource utilization, it may make it more difficult to use CUDA shared memory effectively because different thread blocks cannot share data through shared memory. Consequently, global memory traffic may increase. By eliminating the thread block hierarchy, the need to optimize the thread block dimension, which can have a significant impact on performance, is also eliminated.

[0148] Third, the range of the on-chip scratchpad memory of the NDP unit (1300) is larger in μthread than in CUDA. CUDA shared memory is not shared between thread blocks even when they run on the same SM. In other words, even if they are within the same core, they cannot reference each other's shared memory values ​​if they correspond to different thread blocks. Therefore, for example, when performing a histogram operation, when the execution of a thread block is finished, the data is stored in DRAM, and another thread block reads the data stored in the DRAM, counts it, and then stores it in DRAM again, which results in a lot of memory traffic. On the other hand, according to the present invention, the contents of the scratchpad memory are shared across the entire same NDP unit, the scratchpad memory is initialized once before all μthreads are executed, and after the execution of all μthreads is finished, the data stored in the scratchpad memory can be stored in DRAM through the finalizer of FIG. 8. That is, all μthreads running on the same NDP unit can share data through the on-chip scratchpad memory, further reducing off-chip memory traffic.

[0149] The size of the data associated with a μthread may or may not be the same as the memory access granularity of the DRAM interface (e.g., 64 B for DDR5, 32 B for LPDDR5). To load balance across the NDP units (1300), μthreads are mapped to the NDP units (1300) in an interleaved manner using the memory access granularity. μthreads execute concurrently in a massively synchronous parallel model without ordering guarantees, similar to CUDA threads in GPU kernels. Therefore, the NDP kernel must be written accordingly.

[0150] Meanwhile, looking at thread scheduling, the CPU can use ST / SMT / FGMT / CGMT methods, and the GPU and the M of the present invention can use ST / SMT / FGMT / CGMT methods. 2 NDP uses the FGMT method.

[0151] FIG. 7 illustrates the microarchitecture of an NDP unit according to one embodiment of the present invention.

[0152] The NDP unit (1300) is designed to be low cost while supporting general purpose calculations.

[0153] The NDP unit (1300) may include a plurality of NDP sub-cores, a microthread generator, a scratchpad memory (scratchpad memory, spadmem), an L1 instruction cache (L1 I $), an L1 data cache (L1 data $), an L1 instruction TLB (L1 I TLB), and an L1 data TLB (L1 D TLB). In this case, the scratchpad memory may also be referred to as an 'on-chip scratchpad memory'.

[0154] The above NDP sub-core may include an L0 instruction cache (L0 I $), a decoder / register renaming unit, multiple micro-thread slots, a dispatch (4-way), a scalar ALU, a scalar SFU, a scalar LSU, a vector ALU, a vector SFU, a vector LSU, and a register file.

[0155] The above L0 instruction cache may be a cache that primarily checks instructions accessed by each NDP sub-core. If the requested instruction is not in the L0 instruction cache, the L0 cache controller may forward a request for the instruction to the L1 instruction cache.

[0156] Referring to FIGS. 2 and 7 together, when the NDP kernel starts, the NDP controller (1200) instructs the μthread creator to create a μthread of the size of the data to be executed by allocating μthread slot and register file resources to one of the sub-cores of the NDP unit (1300). The sub-core design is used to reduce the complexity of the dispatch unit.

[0157] At this time, the micro thread slot (μthread slot) may include a PC (program counter), an Opcode and register ID (RegIDs) of the currently decoded instruction, a State, and a CSR (configure and status register) of RISC-V.

[0158] The above register ID may include the first register ID required by the currently decoded instruction and a base ID specified for renaming registers to be used in the corresponding μthread. That is, the base ID may be a base (start) ID for integer / floating point / vector registers.

[0159] At this time, the base ID, which is the register ID used by each μthread, all points to the same address, but in reality, it must be stored in a separate register. Therefore, the register ID in the actual hardware may be different. Therefore, a process to align the base ID with the hardware ID is required, and this process can be performed in the register renaming unit.

[0160] The above State can store the execution-related state of μthread (e.g., ready, stall, etc.).

[0161] The above base ID, i.e., base register IDs, are determined when each μthread is created and the registers required by the kernel are allocated. Logical registers are renamed to physical registers for access simply by adding a logical ID to the base ID. In addition, the first two non-zero scalar registers (i.e., x1 and x2) are initialized to the address and offset of the μthread pool associated with the μthread. After a microthread slot is allocated to a μthread, the PC of the microthread slot is initialized to the kernel code location and an instruction is fetched through the instruction cache for execution. For example, the instruction cache may be an L0 instruction cache, and the L0 instruction cache may be accessed primarily to fetch an instruction. If a miss occurs in the L0 instruction cache, the L1 instruction cache may be accessed.

[0162] Additionally, the on-chip scratchpad memory can be used by each NDP unit for data sharing among all μthreads running on the same NDP unit. The scratchpad memory may be a separate memory space that is not address mapped to the DRAM (100). Data to be shared among multiple threads may be stored in advance by the user in the scratchpad memory.

[0163] For example, referring to FIGS. 2 and 7 together, all μthreads running in 'NDP unit 0' can share data via the on-chip scratchpad memory included in 'NDP unit 0'. All μthreads running in 'NDP unit 1' can share data via the on-chip scratchpad memory included in 'NDP unit 1'. For example, μthreads of 'NDP unit 0' can also share data with μthreads included in other NDP units (e.g., 'NDP unit 1') via the on-chip scratchpad memory included in the other NDP units. That is, μthreads of any NDP unit can access the scratchpad memories of other NDP units.

[0164] At this time, the scratchpad memory can be allocated according to the amount specified by the NDP kernel for each NDP unit.

[0165] At this time, kernel arguments are also placed in the scratchpad memory after the microthread slot is allocated. The scratchpad memory is mapped to an unused area of ​​the virtual memory layout and can be accessed using normal load / store operations. At this time, when μthread accesses the scratchpad memory, the address can be added to the base address of the allocated scratchpad memory.

[0166] A Load / Store Unit (LSU) for scratchpad memory equipped with atomic operation capability is also provided within the NDP unit (1300) to manipulate shared data (e.g., conversion operation by multiple μthreads). That is, the scratchpad memory can be provided for safe operation of shared data by multiple μthreads within the NDP unit. To avoid coherence issues, global memory atomicity is performed in the memory-side L2 cache (cache (1500) of FIG. 1). At this time, address translation is performed using on-chip TLBs, DRAM-TLB, and ATS. The NDP unit (1300) can access all memory locations in the CXL memory expansion device of the system through on-chip and off-chip interconnects.

[0167] Without the scratchpad memory described above, when multiple μthreads need to perform operations on the same value, they all access, for example, DRAM (100). Since accessing DRAM (100) takes longer than accessing SRAM, using scratchpad memory can reduce the access time.

[0168] While μthread instructions are executed sequentially, one at a time, different μthreads independently issue instructions via the FGMT. This avoids the need for complex dependency checking between instructions or data transfer logic, thereby avoiding high control logic overhead. Using a sufficient number of microthread slots (e.g., 64 per NDP unit) to execute memory-bound kernels allows for significant utilization of the memory bandwidth within the CXL memory expansion unit. When a μthread completes execution, another μthread from the μthread pool is created in an idle slot.

[0169] As described above, the NDP unit (1300) may include various scalar and vector function units. For example, for efficiency, the width of the vector unit may match the DRAM access granularity (e.g., 32B for LPDDR5) to avoid computational bottlenecks. However, in other embodiments, the width of the vector unit may not match the DRAM access granularity.

[0170] Hereinafter, the cache hierarchy will be described with reference to FIG. 1 and FIG. 7 together.

[0171] To avoid the complexity of cache coherency, the cache hierarchy of the GPU is adopted. The L1 data cache (L1 D $) of the NDP unit (1300) is read-only and can use a write-evict policy or a write-through policy among write policies. For example, the L0 instruction cache (L0 I $), the L1 instruction cache (L1 I $), and the L1 data cache can be configured as a virtual cache and SRAM. The L1 data cache and the L1 instruction cache are physically separate from the DRAM, and fetch and store values ​​in the DRAM so that they can be accessed within the NDP unit (1300) more quickly.

[0172] The capacity of the L1 data cache of the above NDP unit (1300) can be configured to be partitioned and used as a general L1 data cache and scratchpad memory. The L2 cache (1500) is arranged in front of the memory controller (1600) as shown in FIG. 1 to prevent cache coherency issues and also supports global memory atomic operations on DRAM data. The NDP unit (1300) uses a small instruction cache because data-parallel, memory-bound workloads have relatively smaller instruction footprints than compute-bound workloads. To prevent access to old code, the instruction caches are flushed when the NDP kernel is deregistered via offset 1<<5 in FIG. 3. However, this occurs rarely and has a minimal impact on performance.

[0173] As described above, the above-described effects of the present invention can be obtained by newly proposing NDP sub-cores, a micro-thread generator, a scratchpad memory, and a plurality of micro-thread slots of the NDP sub-core in the micro-architecture of the NDP unit of the present invention.

[0174] FIG. 8 is a drawing for explaining the structure of an NDP kernel according to one embodiment of the present invention.

[0175] Figure 8 illustrates an example of an NDP kernel for large-scale data reduction operations. It can be assumed that the scratchpad memory is mapped to 0x10000000, and the final result is stored at the location designated by 0x10000008 in the scratchpad memory. The AMOADD instruction can perform atomic memory operations.

[0176] An NDP kernel may contain an Initializer, a Kernel Body, and a Finalizer to support various use cases.

[0177] The above initializer may be executed only once, before any μthreads for the kernel body are created. That is, the initializer may be executed only once when the NDP kernel is launched to initialize the scratchpad memory (if necessary) (i.e., to store shared data to be used) and to perform any necessary pre-computation before the main computation. For the initializer, one μthread is created for each μthread slot using a unique ID in the x2 (or offset) register. In other embodiments, one μthread slot may be included per NDP unit, and only one μthread may be created per NDP unit. In yet other embodiments, even if the NDP unit includes multiple μthread slots, the initializer may only create one μthread. When the initializer is called, the NDP kernel arguments may be available at the starting address of the scratchpad memory allocated to the kernel for each NDP unit (1300). The execution of the initializer may be performed in various ways other than the above-described manner.

[0178] Once the above initializer completes, the μthread constructor can start creating μthreads for the μthread pool area given at kernel startup time to execute the kernel body. At this time, there can be multiple kernel bodies, such that when the execution of the kernel body for all μthreads is completed, all μthreads are recreated for execution of the next kernel body. The kernel body can be executed once by each μthread created to perform the main computation.

[0179] After execution of all kernel bodies (i.e., when execution of all μthreads is completed), the finalizer is executed in the same manner as the initializer, or in a predetermined other manner, but may be executed to perform post-processing and, if necessary, store kernel-level output in memory (e.g., DRAM).

[0180] Meanwhile, the kernel can directly or indirectly access all memory locations in the HDM (including memory locations in peer CXL memory). This allows pointer tracking for irregular workloads (e.g., graph analysis). Due to the lack of protocol support, direct access to host-side memory from the NDP kernel using CXL.mem is not possible, but M 2 NDP can adopt page fault handling support on GPUs with PCIe and host drivers / runtimes.

[0181] Hereinafter, virtual memory and DRAM-TLB preloading will be described with reference to FIGS. 2, 3, and 7.

[0182] CXL memory expansion device (i.e. M 2 The NDP architecture) can efficiently support a virtual memory system. Since the CXL.mem protocol uses host physical addresses for memory access, no address translation is required in the memory expansion device for general CXL.mem requests. However, a virtual address is provided for the μthread pool area, and the NDP kernel code also uses the virtual address for memory access within the kernel. For example, as shown in Fig. 2, the entire memory area can be divided into small units (e.g., 0x20) and assigned to each μthread, which is a lightweight thread, to configure a μthread pool. Each μthread can be evenly distributed to the NDP unit (1300) and execute NDP operations for the memory area allocated to it.

[0183] For example, when using virtual memory, the NDP sub-core of the NDP unit (1300) accesses and references the L1 cache, and when a miss occurs in the L1 cache, the L2 cache can be referenced. For example, the L1 cache can be implemented as a virtual cache using a virtual memory address, VIPT (Virtual Indexed, Physically Tagged), etc. For example, the L2 cache can be a physical cache using a physical address. That is, when a miss occurs in the L1 cache, the L2 cache can be accessed using a physical address. At this time, address translation through the TLB is required to access the L2 cache.

[0184] Therefore, the NDP unit (1300) of the present invention can use the TLB of FIG. 7 for address translation. However, the on-chip TLB may not be sufficient for the NDP kernel that processes a large amount of data in CXL memory expansion devices. Although CXL.io supports ATS, frequent use may result in performance degradation due to round-trip communication with the host and page table walks possible in the host. Therefore, in order to cost-effectively improve the TLB reach of the NDP unit (1300) in the present invention, DRAM-TLB (L1 I TLB and L1 D TLB) is adopted to minimize the miss penalty of the on-chip TLBs.

[0185] According to the present invention, as a process for accessing an L2 cache, when a miss occurs in an L1 instruction cache, the L1 instruction TLB (L1 I TLB) is first accessed, and if a miss occurs in the L1 instruction TLB, the DRAM-TLB can be accessed. If a miss occurs in the DRAM-TLB, information can be requested from the host using ATS. Similarly, when a miss occurs in an L1 data cache, the L1 data TLB (L1 D TLB) is accessed, and if a miss occurs in the L1 data TLB, the DRAM-TLB is accessed, and if a miss occurs in the DRAM-TLB, information about address translation must be requested from the host using the above-described ATS.

[0186] Although one embodiment of the present invention has been described as using only the L1 (I / D) TLB, other embodiments may utilize TLBs composed of multiple layers. For example, an L1 (I / D) TLB, an L2 (I / D) TLB, an L3 (I / D) TLB, etc. may be included, and the present invention may be utilized in such a manner that when a miss occurs in the L1 TLB, the L2 TLB is accessed, and when a miss occurs in the L2 TLB, the L3 TLB is accessed. For example, the L3 TLB is the last TLB, and when a miss occurs in the L3 TLB, the DRAM-TLB can be accessed. When a miss occurs in the DRAM-TLB, information can be requested to the host using the ATS.

[0187] Each DRAM-TLB entry can store an ASID, tag, physical page number, and other attributes (e.g., permission bits) using 16 bytes. If the system has multiple CXL memories, the DRAM-TLB entry can be placed in the same CXL memory to which the page is mapped to localize DRAM-TLB access. The physical location (L) of a DRAM-TLB entry for a virtual page number (VPN) can be simply obtained using the following mathematical expression (1).

[0188]

[0189] Here, S may be the TLB entry size (in bytes), R may be the size of the area allocated to the DRAM-TLB, and Base may be the physical base address of the DRAM-TLB. To support multiple page sizes, the present invention utilizes a POM-TLB approach that partitions DRAM-TLBs for various page sizes and uses a highly accurate page size predictor to determine which partition to access first.

[0190] DRAM-TLB can be implemented with low overhead. Even for the smallest 4KB page size, the overhead of storing DRAM-TLB entries is only 16B / 4KB=0.4%, and for 2MB pages, the overhead can be negligible. If the DRAM-TLB area is sufficiently sized so that the TLB range is similar to the memory capacity of the CXL memory expansion device, DRAM-TLB misses in hashed location calculations are virtually eliminated after the DRAM-TLB is warmed up.

[0191] Additionally, to avoid an initial burst of cold DRAM-TLB misses, the DRAM-TLB can be preloaded when the application's data is first loaded into the memory expansion device (e.g., when the embedding table of a recommended model is loaded into the CXL memory expansion device). As shown in Fig. 3, for efficient DRAM-TLB preloading, the OS receives a virtual address range and M 2 A system call can be provided to preload DRAM-TLB entries using func. If such OS support is not provided, the user can also perform preloading by implementing an NDP kernel that touches all pages in a given virtual address range, generating a DRAM-TLB miss. Later, when a DRAM-TLB miss occurs, the ATS can be used to obtain translation information and populate the DRAM-TLB entry. The host can also provide M for TLB shutdown. 2 Both on-chip TLBs and DRAM TLBs can be invalidated using func.

[0192] CXL-M 2 The on-chip TLB and DRAM TLB of the NDP may also maintain address translation information for address regions mapped to other CXL memory expansion devices, if these devices are present. If the mapping of a page changes, all CXL-M 2 TLB shutdown must be performed for NDPs, but this is unlikely to occur for in-memory data as assumed in the present invention (e.g., no swap to disk).

[0193] Meanwhile, referring to Fig. 1, it can be assumed that multiple memory expansion devices are used. At this time, NDP kernels use direct P2P access between CXL devices through the CXL switch (300) to access other CXL-M 2Large data sets can be processed by accessing data from NDPs. However, CXL interface bandwidth can become a bottleneck for frequent P2P access, so data partitioning across multiple CXL memory expansion devices must be carefully performed. Since different workloads exhibit different memory access patterns, data partitioning schemes are typically specialized for the target workloads. For optimal performance, current multi-GPU systems also require user-level software that partitions data across GPUs and executes separate kernels. Therefore, data is placed on CXL memory expansion devices by software, and the NDP kernel is allocated to each CXL-M for multi-device expansion. 2 It can be similarly assumed that they are launched from NDP. However, since NDP units can directly access other CXL memory expansion units for GPU-like reads and atomic operations, data localization does not need to be perfect. While CXL 3.0 allows for fine-grained address interleaving across CXL memory expansion units, the present invention can assume page-level data placement by the user across CXL memory expansion units to increase data access locality.

[0194] Figure 9 is an M according to one embodiment of the present invention. 2 This is a diagram to explain the system when using func.

[0195] Fig. 10 is an M according to another embodiment of the present invention. 2 This is a diagram to explain the system when using func.

[0196] Figure 11 is an M according to one embodiment of the present invention. 2 This is a flowchart to explain the data processing method for a memory expansion device when using func.

[0197] Below, with reference to FIGS. 9 to 11, M2 Explains the case where only func is used.

[0198] Referring to FIG. 9, a controller (1000) may be placed between a host (200) and a memory device (100). The memory device (100) may be DRAM or flash memory (e.g., SSD). In another embodiment, referring to FIG. 10, a memory element (e.g., a memory buffer (SRAM)) within the controller (1000) may function as a memory device. In another embodiment of the present invention, the controller may be implemented to be included within the memory device.

[0199] For convenience of explanation, it has been described that one controller (1000) controls and manages the operations of the embodiments of the present invention. However, in an alternative embodiment of the present invention, some of the operations of the embodiments of the present invention (particularly FIGS. 9 to 11) may be distributed and executed by a plurality of control logics distributed and arranged at the interface between the host (200) and the memory device (100).

[0200] For example, some functions may be divided and executed by the control logic close to the host (200), and the remaining functions may be divided and executed by the control logic close to the memory device (100), and an interface controller that controls and manages the interface between the host (200) and the memory device (100) may also divide and execute some functions.

[0201] M 2 The configuration for func may be a configuration required for communication between the host (200) and the memory expansion device. That is, M 2 The controller (1000), which is a configuration for func, is a configuration for communication between a host (200) and an IO device, and can be utilized in an IO device that is based on PCIe, such as an NDP, a general GPU, a network card (NIC), an SSD, etc., and supports communication with a CXL memory.

[0202] M by controller (1000) 2 The execution process of func may include steps (S410) to (S430).

[0203] In step (S410), a request transmitted from a host (200) to a memory device (100) is received, and based on an address of the memory device (100) included in the request, it can be determined whether the request corresponds to a first memory-related request that is distinguished from a read request and a write request. At this time, referring to FIG. 2 together, the request (CXL.mem write packet) may include the remaining data excluding the base and bound (base&bound) indicating the uthread pool area among the address (Addr) and data (Data).

[0204] At this time, the first memory-related request may refer to a request for an SSD read / write function, a network card packet transmission / reception function, etc. At this time, an offset may be matched for each function. For example, the function may be defined in the manner illustrated in FIG. 3.

[0205] At this time, the address of the memory device may be, for example, an address mapped to DRAM, an address mapped to a memory buffer within the controller (1000), or an address mapped to flash memory.

[0206] In step (S420), if the request is the first request, the type of the first request can be identified based on the address of the memory device (100) included in the request.

[0207] In step (S430), if the request is the first request, the identified first request can be executed based on the argument included in the request (e.g., the argument of v1 in FIG. 2).

[0208] At this time, after the request is determined to be the first request, if the controller (1000) determines that an additional read request has been received for the address of the first request, the processing status for the first request stored in the address of the first request of the memory device (100) can be transmitted to the host (200). The read request may be to confirm whether the first request (i.e., function call) has been performed normally. At this time, the processing status may mean the result data of executing the kernel, or the function being terminated / in progress, or the function being successful / error, as shown in FIG. 3.

[0209] At this time, the controller (1000) may include a packet filter (1100).

[0210] At this time, the controller (1000) may allocate and store an address area for distinguishing the first request within the controller (1000). That is, an address area for distinguishing the first request may be allocated and stored in the packet filter (1100). At this time, step (S410) may be a step of determining the request as the first request if the address included in the request is included in the address area. At this time, step (S410) may be specifically executed in the packet filter (1100) of the controller (1000). For example, referring to FIG. 2, when allocating the address area of ​​the packet filter (1100), addresses 0x10000 to 0x1FFFF of the first record can be allocated to the address area mapped to DRAM, addresses 0x20000 to 0x2FFFF of the second record can be allocated to the address area mapped to the memory buffer inside the controller (1000), and addresses 0x30000 to 0x3FFFF of the third record can be allocated to the address area mapped to the flash memory. The above description is one example and address areas can be allocated by mapping addresses to various types of memory in various ways.

[0211] At this time, the present invention is characterized in that the request (i.e., data of the packet) includes an argument for a function.

[0212] At this time, a fence command may be included between the first request and the additional read request for the address of the first request.

[0213] At this time, an address area for distinguishing the first request may be allocated to the packet filter, and a predetermined protocol may be used when receiving the request. The predetermined protocol may be the CXL.mem protocol or another protocol similar to CXL, such as CCIX.

[0214] At this time, when assigning an address area to distinguish the first request to the packet filter (1100), a predetermined first protocol (e.g., CXL.io protocol, M) from the host (200) is used. 2 The CXL.mem protocol for using func, or another protocol similar to CXL) may be used, and the request from the host (200) may use a second protocol (e.g., the CXL.mem protocol, or another protocol similar to CXL).

[0215] Meanwhile, conventional technologies primarily require entry into kernel mode to receive and execute requests. In contrast, the present invention allows for the identification and execution of the type of first request by receiving command information for an address area for identifying the first request without entering kernel mode.

[0216] According to the present invention, the above-described M 2 The effect is that overhead can be reduced by using the configuration for func.

[0217] Previously M 2 If func is not used, function call is not possible with only the storage command, but according to the present invention, M 2Because func is used, the host can receive function arguments for a function type with just a simple save command.

[0218] According to the present invention, the host hardware (200) is M 2 For example, various functions such as kernel registration, kernel launch, DRAM-TLB entry preload, etc. of func may not be distinguished and transmitted to the controller (1000). If the host hardware (200) distinguishes the various functions, the host hardware may have to be changed, which is a major limitation. Therefore, the present invention is to distinguish M at the software level stage of a user (e.g., a developer). 2 It is characterized by using func and does not require any changes to the host hardware.

[0219] According to the present invention, not only the functions for the NDP shown in FIG. 3, but also communication with the host through the existing PCIe (e.g., communication for reading and writing files to SSD) is performed by the M of the present invention. 2 can be replaced with configurations for func. That is, the present invention can provide greater effectiveness when fine-grained communication tasks are required.

[0220] In another embodiment, referring to FIGS. 1 and 9 together, a CXL switch (300) (e.g., a semiconductor chip) positioned between a host (200) and CXL memories including a conventional general CXL controller and connecting the host (200) and a specific CXL memory may include a packet filter (1100), an NDP controller (1200), NDP units (1300), and a cross bar (1400). That is, M 2 func and M 2uthr can be implemented together through a CXL switch (300). At this time, a small SRAM can be used for the address area of ​​the packet filter (1100). At this time, a plurality of CXL memories can be selectively connected by a cross bar (1400). A conventional CXL controller included in the CXL memory connected to the CXL switch (300) can include, for example, caches (1500) and memory controllers (1600). In one embodiment of the present invention, the CXL switch can be used for a workload (e.g., an ML model) that does not require simultaneous shared data manipulation between the host and the NDP to avoid consistency issues with the host.

[0221] Figure 12 is an M according to one embodiment of the present invention. 2 M without using func 2 This is a diagram to explain the system when only μthr is used.

[0222] Figure 13 is an M according to one embodiment of the present invention. 2 M without using func 2 This is a drawing to explain the case where only μthr is used.

[0223] Figure 14 is an M according to one embodiment of the present invention. 2 This is a flowchart explaining the data processing method for a memory expansion device when using μthr.

[0224] Below, with reference to FIGS. 12 to 14, M 2 Describes the case where only μthr is used.

[0225] The memory expansion device may include a memory device (100) and a controller (1000).

[0226] The controller (1000) may include an NDP controller (1200) and a plurality of NDP units (1300). In the embodiment of FIG. 13, it may be assumed that the CXL controller (1000) does not include a packet filter (1100).

[0227] In one embodiment of FIGS. 12 to 14, the controller (1000) is described as being applied to NDP, but in other embodiments, it can be applied to GPU, DPU, SmartNIC, etc. and utilized as various accelerators.

[0228] A memory expansion device (specifically, a CXL controller (1000)) can receive an NDP function request from a host (200) via, for example, a PCIe or CXL.io protocol.

[0229] For example, the request may be a request to execute an NDP function for the memory area 0xA000~0xA1FF. The request may include synchronous / asynchronous information, a kernel ID, a μthread pool area (base&bound), argument size, NDP kernel arguments, etc. In this case, the addresses of the μthread pool area and the NDP kernel arguments may be addresses included in the virtual memory area.

[0230] In step (S510), when the controller (1000) receives an operation request (kernel launch request), the controller (1000) can control a plurality of NDP units (1300) included in the controller (1000) to create a plurality of micro threads mapped to different addresses in the first area (base&bound) of the virtual memory.

[0231] In step (S520), the controller (1000) can cause each of the above micro threads to perform an operation on the mapped memory address.

[0232] In step (S530), the controller (1000) may store the executed result in a result storage address defined in the kernel of the memory device. At this time, the result storage address defined in the kernel may be an address of the memory device (100). At this time, the controller (1000) further includes a memory controller, and the executed result may be stored in the memory device (100) via the memory controller.

[0233] For example, the NDP controller (1200) of the controller (1000) can divide the 0xA000~0xA1FF memory area into small units (e.g., 0x20) and assign each unit to a lightweight thread, μthread, to configure a μthread pool. Each μthread is evenly distributed to the NDP units (1300) and can execute NDP operations for the virtual memory area allocated to it. For example, μthread0, μthread4, etc. may be distributed to the NDP unit (1310), μthread1, μthread5, etc. may be distributed to the NDP unit (1320), and μthread3, μthread7, etc. may be distributed to the NDP unit (1340). For example, the virtual memory address of μthread0 may be 0xA000, and the virtual memory address of μthread1 may be 0xA020.

[0234] For example, the execution result of the kernel operation of μthread0 may be stored in the virtual memory address of μthread0, for example, 0xA000.

[0235] At this time, all microthreads running within any NDP unit among the plurality of NDP units (1300) can share data with each other through at least one on-chip scratchpad memory among the on-chip scratchpad memories included in each of the plurality of NDP units. At this time, when the kernel launch is completed, the operation result can be temporarily stored in the on-chip scratchpad memory, and then the on-chip scratchpad memory can be erased.

[0236] At this time, the controller (1000) may further include a DRAM-TLB for converting the virtual address of the virtual memory into a physical address of the memory device (100). For example, the virtual address of the virtual memory may mean the address of the first area of ​​the virtual memory to which μthread is mapped, may mean the address of another area other than the first area of ​​the virtual memory used by μthread, or may mean the address of the NDP kernel arguments of another area of ​​the virtual memory.

[0237] The CXL.io protocol can be used for this, but M through the CXL.mem protocol 2 func can be used, M 2 The func is as described with reference to FIG. 11. That is, the controller (1000) can receive a request transmitted from the host (200) to the memory device (100), and determine whether the request corresponds to a first memory-related request that is distinguished from a read request and a write request based on the address of the memory device (100) included in the request. If the request is the first request, the controller (1000) can identify the type of the first request based on the address of the memory device (100) included in the request, and if the request is the first request, can execute the identified first request based on the argument included in the request. At this time, if the request is a DRAM-TLB preloading request among the types of the first request, the DRAM-TLB can be used by preloading the DRAM-TLB.

[0238] In the embodiments of Figs. 12 to 14, M 2 M without using func 2 I explained the case of using μthr. If M 2 func M2 When used with μthr, the first request may be, for example, a request for a kernel launch corresponding to offset 2<<5 in Fig. 3. For example, M may be used to launch another NDP kernel while μthread is running. 2 A write packet can be generated for the purpose of using func. At this time, the packet can be checked by the packet filter and the kernel launch function can be executed.

[0239] In another embodiment of the present invention, M 2 When using only func and M 2 func M 2 When used with μthr, the (CXL) controller (1000) can send requests (packets) to itself and receive the sent packets. In this respect, the controller (1000) can also perform the functions of the host (200).

[0240] The above CXL.io utilizes the physical design and protocol stack of PCIe, so it is expected to have similar performance characteristics. Computation offloading via PCIe involves several software and hardware steps that can incur significant overhead in terms of latency and host processor utilization, especially for fine-grained offloads.

[0241] For example, launching a GPU kernel requires a user-mode CUDA / OpenCL runtime and a kernel-mode GPU device driver. The runtime records the kernel launch GPU command in a user buffer, and the driver pushes a packet pointing to the GPU command into a ring buffer in kernel space. The host then updates the head pointer of the ring buffer to notify the GPU of the new command, which incurs additional latency over PCIe. Overall, kernel execution can have a latency of several μs (e.g., ~4.5 μs). Polling or interrupts can be used to check for kernel completion, both of which consume additional host processor cycles. Polling can incur an overhead of 2-3 μs, while interrupts have a similar or higher overhead, depending on the bottom-half mechanism used. Therefore, the total latency for kernel launch and completion check can be significantly longer than 5 μs or 10,000 cycles on a 2 GHz CPU. While this latency may be acceptable for coarse-grained NDP offloading, it may be excessively high for fine-grained NDP kernels that are sensitive to latency. Therefore, to address the aforementioned issues, the present invention provides a low-overhead offloading mechanism based on CXL.mem that is effective for both fine-grained and coarse-grained offloading.

[0242] While CXL.io can have high overhead for frequent, granular communications, CXL.mem messages can be transmitted with low latency and CPU usage. The current CXL.mem protocol defines several unused bits in the packet format. Therefore, one could consider using these bits to encode information needed to implement specialized functions not defined in the standard (e.g., NDP management).

[0243] However, enabling this customized communication requires modifying the host processor hardware to support special usage of reserved bits. Therefore, it cannot be utilized on commodity processors that only support standard protocols. Furthermore, transmitting special packets requires the introduction of special instructions into the host's ISA, as with previous technologies. Appropriately extending the standard protocol or the host's ISA hinders widespread adoption.

[0244] Therefore, the present invention provides an NDP architecture based on the unmodified CXL.mem protocol, as described above, for optimal compatibility with various host processors. It can be assumed that data to be processed by NDP is stored in a CXL memory expansion device. For data residing in the host's local memory, the host has high BW access rights to the data, and NDP is not required.

[0245] Meanwhile, the present invention provides a new M for general-purpose NDP of CXL memory with low overhead. 2 It can provide NDP (Memory-Mapped NDP) architecture. As described above, M 2 NDP includes M2func for low overhead communication between the host (200) and the CXL memory expansion device and M for cost-effective NDP kernel execution. 2 There are two main mechanisms that can provide μthr.

[0246] CXL.io can be used for various communication between the host and CXL memory (e.g., NDP offload commands), but its protocol stack incurs higher overhead than CXL.mem and requires expensive kernel-level operations on the host, potentially wasting CPU cycles. Meanwhile, CXL.mem offers low latency and can be used without kernel intervention, but only supports basic memory read / write transactions.

[0247] The present invention provides an M that selectively repurposes read and write packets defined in CXL.mem for efficient host-device communication beyond memory transactions. 2 func can be provided. The high overhead of CXL.io can be avoided by encapsulating NDP management commands (e.g., function calls) of CXL.mem requests into predetermined addresses. M 2 The main activating element of the func is a packet filter (1100) placed at the input port of the CXL memory expansion device. It can check whether the memory address of the incoming request matches the memory range pre-allocated exclusively for each host process. Then, various NDP management functions can be triggered for the matching request depending on the address. Therefore, NDP management function calls (e.g., kernel registration, execution, and status polling) can be performed simply by executing memory accesses on the host (200). As a result, the high overhead of the CXL.io protocol stack and kernel operations can be avoided for low-overhead NDP offloading, especially for fine-grained NDP. In addition, there is no need to modify the CXL.mem standard for the best compatibility with the host (CPU). In order to use CXL.io, there is an overhead going through the OS, which is M 2 func can be directly available in userspace. Also, M 2 func is M 2 It is highly versatile as it can be used for various functions (e.g., storage and network access) on CXL devices in addition to the functions basically supported for NDP.

[0248] The present invention also provides M for intuitive abstraction and cost-effective NDP. 2can provide μthreads. Memory-bound workloads tend to use fewer registers than compute-bound workloads. Therefore, the present invention can provide μthreads, which are lightweight threads that contain a subset of architectural registers as execution units. By reducing register usage, the NDP unit can execute many μthreads concurrently, hiding DRAM access latency without excessive physical register file costs. In addition, memory-bound data parallel workloads can be implemented such that each thread is associated with specific data to be processed. In traditional programming environments such as CUDA, the association between a thread and a memory location is expressed indirectly through code (e.g., calculating the index of an array element for a thread using CUDA's thread block ID, block dimension, and thread ID). In contrast, M 2 When using μthr, each μthread is created with a direct association to a specific memory location. In other words, μthread is memory-mapped. As a result, the initial address calculation code can be eliminated within the kernel. The way it is scheduled and executed on the hardware is similar to CUDA's warp, but CUDA warp only supports SIMD (vector) mode, where all threads within it are executed by default. On the other hand, μthread supports vector instructions in scalar threads by default, so both scalar and vector operations can be used as needed without wasting resources.

[0249] The architecture of the NDP unit (1300) of the present invention is based on a RISC-V ISA with vector extensions that support scalar operations to fully utilize the DRAM BW within the CXL memory expansion unit at a low cost by utilizing SIMD function units while avoiding redundant address calculations in SIMT-only GPUs. Many memory-mapped μthreads are executed with fine-grained multithreading (FGMT) to hide memory access latency. In addition, DRAM traffic can be reduced by sharing data in on-chip scratchpad memory that provides a wider range than the GPU's shared memory. Unlike thread block creation in a GPU, which can waste resources due to inter-warp divergence, μthreads are created individually.

[0250] Additionally, to provide flexibility for various workloads, M 2 The NDP architecture supports virtual memory. However, address translation can cause significant overhead to the NDP because the host (200) must access page table entries through the Address Translation Service (ATS) via CXL.io. To avoid overhead, the CXL memory expansion device uses a DRAM-TLB. In the present invention, since a cold miss of the DRAM-TLB can be serious, M 2 DRAM-TLB preloading by the operating system using func can also be provided. According to the present invention, a significant speedup of up to 171 times and energy reduction of 81.3% can be achieved for various workloads compared to a host processor with a passive CXL memory expansion device. 2 It can provide the efficiency of NDP.

[0251] M 2 func and M 2 By combining μthr and utilizing OS support, the M of the present invention2 The NDP architecture enables low-overhead, general-purpose NDP in CXL memory expansion devices. The design's effectiveness is demonstrated across a variety of workloads, including in-memory online analytical processing (OLAP), deep-learning recommendation models (DLRM), graph workloads, and critical kernels in large-scale language models (LLMs).

[0252] In another embodiment, M 2 NDP can achieve up to 171x speedups across a variety of workloads compared to baseline systems using passive CXL memory expansion devices, while reducing energy consumption by up to 94.2%.

[0253] The controller and / or memory expansion device of the present invention can be applied to generation AI, data centers, cloud computing, big data and analysis, in-memory online analytical processing (OLAP), recommendation systems, and graph analysis processing, such as ChatGPT.

[0254] FIG. 15 is a conceptual diagram illustrating an example of a generalized controller, memory expansion device, or computing system capable of performing at least a portion of the processes of FIGS. 1 to 14.

[0255] At least a part of the process of the data processing method for a memory expansion device according to one embodiment of the present invention can be executed by the computing system (2000) of FIG. 15.

[0256] Referring to FIG. 15, a computing system (2000) according to one embodiment of the present invention may be configured to include a processor (2100), a memory (2200), a communication interface (2300), a storage device (2400), an input interface (2500), an output interface (2600), and a bus (2700).

[0257] A computing system (2000) according to one embodiment of the present invention may include at least one processor (2100) and a memory (2200) that stores instructions that instruct the at least one processor (2100) to perform at least one step. At least some steps of a method according to one embodiment of the present invention may be performed by the at least one processor (2100) loading and executing instructions from the memory (2200).

[0258] The processor (2100) may mean a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor on which methods according to embodiments of the present invention are performed.

[0259] Each of the memory (2200) and the storage device (2400) may be configured with at least one of a volatile storage medium and a non-volatile storage medium. For example, the memory (2200) may be configured with at least one of a read-only memory (ROM) and a random access memory (RAM).

[0260] Additionally, the computing system (2000) may include a communication interface (2300) that performs communication via a wireless network.

[0261] Additionally, the computing system (2000) may further include a storage device (2400), an input interface (2500), an output interface (2600), etc.

[0262] Additionally, each component included in the computing system (2000) can communicate with each other by being connected by a bus (2700).

[0263] Examples of the computing system (2000) of the present invention may include a desktop computer, a laptop computer, a notebook, a smart phone, a tablet PC, a mobile phone, a smart watch, a smart glass, an e-book reader, a portable multimedia player (PMP), a portable game console, a navigation device, a digital camera, a digital multimedia broadcasting (DMB) player, a digital audio recorder, a digital audio player, a digital video recorder, a digital video player, a PDA (Personal Digital Assistant), etc.

[0264] The operations of the method according to an embodiment of the present invention can be implemented as a computer-readable program or code on a computer-readable recording medium. A computer-readable recording medium includes any type of recording device that stores information readable by a computer system. Furthermore, a computer-readable recording medium can be distributed across network-connected computer systems, allowing the computer-readable program or code to be stored and executed in a distributed manner.

[0265] Additionally, the computer-readable recording medium may include hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, flash memory, etc. The program instructions may include not only machine language codes produced by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.

[0266] While some aspects of the present invention have been described in the context of a device, they may also represent a description of a corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method may also be described as a corresponding block or item or a feature of a corresponding device. Some or all of the method steps may be performed by (or using) a hardware device, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, at least one or more of the most important method steps may be performed by such a device.

[0267] In embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In embodiments, the field-programmable gate array may operate in conjunction with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by some hardware device.

[0268] Although the present invention has been described above with reference to preferred embodiments thereof, it will be understood by those skilled in the art that various modifications and changes may be made to the present invention without departing from the spirit and scope of the present invention as set forth in the claims below.

Claims

1. Receive a request transmitted from a host to a memory device, and determine whether the request corresponds to a first memory-related request that is distinguished from a read request and a write request based on an address of the memory device included in the request; If the request is the first request, the type of the first request is identified based on the address of the memory device included in the request, If the above request is the first request, executing the identified first request based on the arguments included in the request; Controller for memory expansion.

2. A controller for memory expansion, wherein, in the first paragraph, if it is determined that an additional read request has been received for the address of the first request after the request is determined to be the first request, the controller transmits the processing status for the first request stored in the address of the first request of the memory device to the host.

3. In the first paragraph, a packet filter is included, The above packet filter has an address area allocated and stored to distinguish the first request, The packet filter determines that the request is the first request if the address included in the request is included in the address area. Controller for memory expansion.

4. A controller for memory expansion in the third paragraph, wherein an address area for distinguishing the first request is allocated to the packet filter, and a predetermined protocol is used when the request is received.

5. A controller for memory expansion, wherein, in the third paragraph, when assigning an address area for distinguishing the first request to the packet filter, a predetermined first protocol from the host is used, and when receiving the request from the host, a predetermined second protocol is used.

6. In paragraph 1, NDP controller executing the first request; and A plurality of NDP units controlled by the above NDP controller; Includes, If the above first request is a kernel launch request, The above NDP units create multiple microthreads that are mapped to different addresses in the first area of ​​virtual memory, Each of the above microthreads performs operations on mapped memory addresses, Controller for memory expansion.

7. A controller for memory expansion in the 6th paragraph, wherein all microthreads running within any NDP unit among the plurality of NDP units share data with each other through one or more on-chip scratchpad memories among the on-chip scratchpad memories included in each of the plurality of NDP units.

8. In paragraph 6, Further comprising a DRAM-TLB for converting a virtual address of the virtual memory into a physical address of the memory device, If the above request is a DRAM-TLB preloading request among the types of the first request, the DRAM-TLB is used by preloading the DRAM-TLB. Controller for memory expansion.

9. Memory device; and A controller disposed between a host and the memory device, the controller including a plurality of NDP units; Includes, When the above controller receives a kernel launch request, The above NDP units create multiple microthreads that are mapped to different addresses in the first area of ​​virtual memory, Each of the above microthreads performs an operation on the mapped memory address, The result of the above execution is stored in the result storage address defined in the kernel of the memory device. Memory expansion device.

10. A memory expansion device in accordance with claim 9, wherein all microthreads running within any NDP unit among the plurality of NDP units share data with each other through one or more on-chip scratchpad memories among the on-chip scratchpad memories included in each of the plurality of NDP units.

11. In paragraph 9, The controller receives a request transmitted from the host to the memory device, and determines whether the request corresponds to a first memory-related request that is distinguished from a read request and a write request based on an address of the memory device included in the request, The controller identifies the type of the first request based on the address of the memory device included in the request, if the request is the first request, If the above request is the first request, executing the identified first request based on the arguments included in the request; Memory expansion device.

12. In the 11th paragraph, the NDP unit further includes a DRAM-TLB for converting a virtual address of the virtual memory into a physical address of the memory device, If the above request is a DRAM-TLB preloading request among the types of the first request, the DRAM-TLB is used by preloading the DRAM-TLB. Memory expansion device.

13. A step in which the controller receives a request transmitted from the host to the memory device, and determines whether the request corresponds to a first memory-related request that is distinguished from a read request and a write request based on an address of the memory device included in the request; If the request is the first request, a step of identifying the type of the first request based on the address of the memory device included in the request; and If the request is the first request, executing the identified first request based on the arguments included in the request; including, Data processing method for a memory expansion device.

14. In paragraph 13, after the request is determined to be the first request, A step in which the controller determines that an additional read request has been received for the address of the first request; and A step in which the controller transmits a processing status for the first request stored in the address of the first request of the memory device to the host; including more, A method of processing data using a memory expansion device.

15. In paragraph 13, The controller further includes a step of allocating and storing an address area for distinguishing the first request within the controller, The above-determining step is a step of determining the request as the first request if the address included in the request is included in the address area. A method of processing data using a memory expansion device.

16. A data processing method using a memory expansion device in paragraph 14, wherein an address area for distinguishing the first request is allocated to the packet filter, and a predetermined protocol is used when the request is received.

17. A data processing method using a memory expansion device, wherein in paragraph 15, the step of allocating and storing an address area for distinguishing the first request is a step of using a predetermined first protocol from the host, and the step of receiving the request from the host is a step of using a predetermined second protocol.

18. In paragraph 13, If the above first request is a kernel launch request, A step in which the controller controls a plurality of NDP units included in the controller to create a plurality of micro-threads mapped to different addresses of a first area of ​​virtual memory; and A step in which the controller causes each of the micro threads to perform an operation on a mapped memory address; including more, A method of processing data using a memory expansion device.

19. A data processing method using a memory expansion device in claim 18, wherein all microthreads running within any NDP unit among the plurality of NDP units share data with each other through one or more on-chip scratchpad memories among the on-chip scratchpad memories included in each of the plurality of NDP units.

20. In paragraph 18, If the above request is a DRAM-TLB preloading request among the types of the first request, the controller further includes a step of preloading the DRAM-TLB, The DRAM-TLB is used to convert the virtual address of the virtual memory into the physical address of the memory device. A method of processing data using a memory expansion device.

Citation Information

Patent Citations

  • Kernel stack distribution method and device

    CN109857677A

  • Memory control method

    KR1020150092676A

  • Memory unit for emulated shared memory architectures

    KR1020160010580A

  • System and method for diagnosing and evaluating core competencies for each university major based on mathematical learning results

    KR1020230109591A

  • Food thermal cooling and storage system

    KR102662992B1