Data processing unit

The DPU addresses memory inefficiencies in data centers by providing hardware-accelerated remote memory access and management, enhancing efficiency and reducing costs through flexible memory utilization and expansion.

WO2026057153A1PCT designated stage Publication Date: 2026-03-19HUAWEI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Modern data centers face high memory costs and inefficiencies due to architectural constraints of server CPUs, limited memory upgrades, and stranded memory resources, leading to suboptimal resource utilization and increased costs.

Method used

A Data Processing Unit (DPU) that provides hardware acceleration for remote memory access, enabling byte-addressable memory access through parallel processing and dynamic translation of virtual to physical addresses, caching, and buffering, with network processors handling memory access requests and managing memory management tasks independently of the external CPU.

Benefits of technology

Enhances resource efficiency, reduces costs, and improves data-center performance by optimizing memory utilization and reducing stranded memory, allowing flexible and scalable memory expansion without hardware modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024075304_19032026_PF_FP_ABST
    Figure EP2024075304_19032026_PF_FP_ABST
Patent Text Reader

Abstract

A Data Processing Unit, DPU, a corresponding Central Processing Unit, CPU, and a corresponding data server are disclosed. The DPU is configured to provide hardware acceleration for memory accesses to remote memory by, under the control of an internal CPU of the DPU: presenting to an external CPU a virtually addressable memory; receiving from the external CPU memory access requests for the virtually addressable memory; translating virtual addresses of received memory access requests to corresponding physical addresses of remote memory; and performing memory accesses from the physical addresses of the remote memory.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DATA PROCESSING UNIT

[0002] TECHNICAL FIELD

[0003] This disclosure relates to a Data Processing Unit, DPU, a corresponding Central Processing Unit, CPU, and a corresponding data server.

[0004] BACKGROUND

[0005] Modem data-center applications are increasingly characterized by their extensive memory requirements. This trend is prominently observed in various domains such as artificial intelligence, in-memory key-value stores, data analytics workloads, and graph analytics workloads. The memory demands of these applications have highlighted several critical challenges and inefficiencies within current data-center infrastructures.

[0006] One of the foremost issues is the high cost of host memory, which constitutes approximately 40% of the total cost of ownership for data centers. This significant expense is exacerbated by the architectural constraints of server CPUs, which have a limited number of pins, thereby restricting the number of dual in-line memory modules that can be supported. Traditional server architectures couple CPUs tightly with memory technology, typically preventing independent upgrades of memory and processors. This rigidity can lead to suboptimal resource utilization and increased costs.

[0007] One phenomenon observed in modem data centers is the presence of substantial levels of ‘ stranded memory ’ . Stranded memory refers to memory resources that remain underutilized due to fragmented resource allocation across data-center servers. This fragmentation often results from over-provisioning memory to accommodate peak utilization scenarios, leaving considerable portions of memory unused during typical operations. Additionally, some memory pages may become “cold” and infrequently accessed, further contributing to inefficiency.

[0008] To address these issues, there is a growing interest in remote, tiered, disaggregated, far, and swap memory approaches. These methodologies aim to decouple memory from individual servers, allowing for more flexible and efficient memory utilization across the data center. By enabling byte-addressable memory accesses to other devices within the data center, servers can leverage shared memory for more efficient communication and processing.

[0009] An object of the present disclosure is to provide a novel and / or improved processing architecture.

[0010] SUMMARY

[0011] In accordance with the present disclosure, there is provided a Data Processing Unit, DPU, a corresponding Central Processing Unit, CPU, and a corresponding data server as claimed in the accompanying claims.

[0012] More specifically, a DPU is provided which is configured to provide hardware acceleration for memory accesses to remote memory by, under the control of an internal Central Processing Unit, CPU, of the DPU: presenting to an external CPU a virtually addressable memory; receiving from the external CPU memory access requests for the virtually addressable memory; translating virtual addresses of received memory access requests to corresponding physical addresses of remote memory; and performing memory accesses from the physical addresses of the remote memory. Dynamic random-access memory may be provided for caching and / or buffering data of memory accesses. The DPU may be further configured to cache data may be speculatively cached in anticipation of future memory access requests for that data.

[0013] The DPU may further comprise a plurality of network processors configured to perform memory accesses in parallel and under the control of the internal CPU. For example, at least one network processor may be configured to perform a memory access by establishing a Remote Direct Memory Access, RDMA, or other communication and data offloading connection with a remote memory. At least one network processor and the internal CPU may be configured to handle an error in performing a memory access without requiring an action of the external CPU.

[0014] The DPU may be further configured to determine if a memory access request from the external CPU will or is likely to experience delay in fulfilling the memory accesses. Based on the determining, perform an external CPU related action may include at least one of: notifying the external CPU of the delay; notifying the external CPU of temporary buffering in the DPU to accommodate the delay; notifying the external CPU of remote memory error or unavailability ; interrupting the external CPU, and signalling an exception to the external CPU.

[0015] The DPU may be further configured to, under the control of the internal CPU, maintain a translation table which maps the virtually addressable memory presented to the external CPU with corresponding physical addresses of remote memory of the storage server (SPAs) or local cached data inside the DPU.

[0016] Where this is the case, the DPU may be further configured to, under the control of the internal CPU, cache address translations in an address translation cache; and search the address translation cache upon receiving a memory access request.

[0017] The DPU may comprise a Base Address Register, BAR, to receive memory access requests from the external CPU.

[0018] The remote memory may include available on-board memory in hardware residing on a remote host, with the DPU being further configured to, under the control of the internal CPU, perform memory accesses from such hardware.

[0019] The DPU may be further configured to update the presentation to the external CPU of the virtually addressable memory upon a change in the provision of remote memory.

[0020] BRIEF DESCRIPTION OF DRAWINGS

[0021] Figures lAto 1C illustrate the architecture of a DPU according to the present disclosure, as used in a server system.

[0022] DETAILED DESCRIPTION OF DRAWINGS

[0023] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.

[0024] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate. The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and “the” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof.

[0025] Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein.

[0026] Referring to figure 1A, an architecture of a server system is illustrated including a compute server comprising a CPU (CPUe with the suffix ‘e’ denoting that the CPU is external to the DPU). CPUe is connected to memory (DRAM) and, via a serial PCI Express / Compute Express Link (PCIe / CXL), to a hard drive (SSD). CXL is an open standard for high-speed, high capacity CPU-to-device and CPU-to-memory connections, and is designed for high performance data center computers. Also connected to CPUe via the PCIe / CXL is a DPU according to the present disclosure. Although a PCIe / CXL interconnect is show, others high-speed interconnects may be used.

[0027] The compute server is connected to a storage server (which in practice may be a farm of collocated and / or remote general and / or storage servers). The storage server shown comprises a network I / O adapter which supports protocols such as RDMA including by handling RDMA metadata as well as network linkage. As with the computer server, there is a CPU connected to an SSD via a PCIe / CXL interconnect. The CPU is connected to memory (DRAM) and also onboard memory of a GPU.

[0028] Fig. 1 A also illustrates the DPU of the compute server is greater detail comprising a CPU (CPU; with the suffix ‘i1denoting that the CPU is internal to the DPU), memory (DRAM) and a PCIe / CXL interface. DRAM allows the DPU to partition it into designated regions for page-caching and buffering of memory access data, and free space for general operational needs. The DPU is further provided with a plurality of network processing units (NPUs) configured to perform memory accesses in parallel, under the control of CPU; of the DPU. Each NPU can leverage its internal hardware capabilities to perform address translations.

[0029] On a high level, the DPU functions to provide CPUe of the compute server with hardware acceleration for memory accesses to remote memory of the storage server under the control of CPU; of the DPU. As illustrated in fig. IB, this is done by the DPU presenting to the CPUe of the compute server with a virtually addressable memory (DPU VA) on the CPUe address space. CPUe of the compute server issues and the DPU receives memory access requests for the virtually addressed memory, just as if the CPU was performing a regular memory access to its DRAM. The virtual addresses of received memory access requests are translated in to corresponding physical addresses, either of the DPU itself in cases where the requested data resides and is cached in the DPU DRAM, or the address of remote memory directly, i.e. byte addressable. Finally, the NPUs of the DPU perform memory accesses from the physical addresses of the remote memory on behalf of the compute server.

[0030] For the NPUs of the DPU to translates virtual addresses of received memory access requests to corresponding physical addresses of either the local DPU DRAM or remote memory, the NPUs access the translation table which are managed by the DPU. The translation table maps the virtually addressable memory presented to the external CPU with corresponding physical addresses of DPU DRAM or remote memory. In cases where the data does not reside on the DPU DRAM, the NPUs will fetch the remote data either by utilizing the CPU of the DPU, or they may use established protocols such as RDMA or other data offloading connection protocols, without requiring constant action of the CPUe of the computer server. For example, an RDMA connection may be established, a WQE created and submit, CQE polled for and then an application running on CPUe of the computer server notified that the data is available.

[0031] Under the control of CPU; of the DPU, prior address translations can be cached in an address translation cache (e.g. in DRAM), and the address translation cache searched upon receiving a memory access request prior to performing a full lookup of the address table. If a cached translation is found, this may obviate the need for such a full lookup.

[0032] Implementing memory management with the DPU architecture of the present disclosure as detailed above can significantly enhance resource efficiency, reduce costs, and improve the overall performance of data-center applications (or indeed other memory intensive applications to which the disclosure is applied).

[0033] Stranded Memory and Memory Expansion

[0034] The present disclosure can enable stranded memory to be more readily utilised via swapping mechanisms without unduly burdening the CPUe of the compute server. For example, swapping mechanisms involve moving inactive or less frequently accessed memory pages to alternative storage locations, such as disks or remote memory devices, to free up local memory for active processes. Conventionally, such swapping would be performed by an operating system running on the CPUe of the computer server. This has the disadvantage that, as a typically software-only approach to memory management, data cannot be accessed directly in its remote location by the CPU, and must be ‘swapped- in’ to the DRAM of a compute server before it can be utilized. Moreover, modem remote memory technologies and protocols like RDMA can be leveraged to enable efficient remote swapping with the NPUs of the DPU.

[0035] Furthermore, byte-addressable memory technologies enable seamless memory expansion. Unlike traditional memory that is tied to specific servers, byte-addressable memory can be accessed remotely, using protocols like CXL. These technologies allow servers to access additional memory resources on-demand from a centralized memory pool or other machine’s memory, effectively expanding their memory capacity without physical upgrades. Conventionally, this would require intensive support from the CPUe of the computer server to be able to access non-local memory transparently and / or in a cacheable manner. Further, CPUs that incorporate such CXL technology do not offload the memory management operations such as hotness detection, migration, and memory tiering, which burdens the CPUs with additional overheads.

[0036] By leveraging byte-addressable technologies, data centers can optimize memory utilization, reduce costs associated with overprovisioning, and improve the performance and scalability of applications. These advancements contribute to more efficient and cost-effective data center management. For example, the DPU can function as memory expansion for the server, accessible through interfaces like Direct Access (DAX). This capability allows for scalable memory solutions without the need for extensive hardware modifications. The DPU (with or without initialization help from the compute server) can expose the memory ranges of the registered devices to the DPU’s BAR memory presentation.

[0037] Address Translation And Mapping To Data Center Resources

[0038] The DPU exposes a memory region within the processor's address space, making it byte-addressable and optionally cacheable by the processor. This allows the CPU to directly access remote memory and other device memory and resources without the overhead of traditional API calls or complex software support, with the support of the DPU and this disclosure.

[0039] As mentioned, under the control of CPU; of the DPU, prior address translations can be cached in an address translation cache, and the address translation cache searched upon receiving a memory access request prior to performing a full lookup of the address table. Such a cache may be implemented by an NPU having integral Static Random -Access Memory (SRAM), akin to a translation lookaside buffer (TLB). Hardware acceleration may also be used in address translation including by using Content Addressable Memory (CAM) or Ternary CAM (TCAM).

[0040] For fast-path data serving where the translation is available, and the data can be retrieved in a timely manner (e.g., it is cached on the DPU’s DRAM, or is directly accessible via the NPU from the address table), the NPU will respond to the request with the required information (i.e., data responses to loads)

[0041] For slow-path data serving where the NPU cannot provide an address translation, or the data is not available to be accessed in a timely manner, the NPU may trigger a handling routine performed by the DPU’s CPUi. CPU; will then handle the request in a similar fashion as traditional operating systems handle page -faults, but has the advantage of performing it in a streamlined fashion, such as using high-performance software or leveraging HW offloading engines incorporated into the DPU such as NPUs. It determines whether the data is readily available but the translation in the NPU has not been updated yet with its corresponding DA (minor page- fault). In this case it will update the translation structures and serve the memory request in case the memory access is valid. If the memory access is invalid (i.e., accesses to a unmapped region of the DPUs address space due to no backing-device being mapped to the accessed region), it will return a fault (i.e., poison / interrupt).

[0042] If a memory request is valid but the data is not readily available (i.e., remote memory not cached in the DPU), the data can be fetched it from its backing device (assuming registered with the DPU). Also, in cases the registered device with the DPU was configured as “cacheable”: an optional ‘poison’ or interrupt-like mechanism can be returned to the compute server to notify it that its request cannot be served immediately, relieving the CPU to perform other tasks while it waits for the data to be ready; CPU; of the DPU can initialize a data transfer request of the corresponding location into its ‘page cache’ partition in its local DRAM, in some granularity. CPU; can then update the translation table. In case of writes, the processor may cache it separately in its DRAM and perform the required operations asynchronously

[0043] If the memory access is invalid (i.e., accesses to an unmapped region of the DPU’s address space due to no backing-device being mapped to the accessed region), it will return a fault (i.e., poison / interrupt).

[0044] The DPU’s processor can manage the page-cache and buffer partitions. This would include applying an eviction / prefetch policy based on heuristics and access data for which the policies could be configurable (e.g. Al learned or explicitly defined by a user). Temporary buffers for uncached medias may serve as pinned ‘bounce buffers’ .

[0045] Translation of the DPU virtual address within to corresponding Storage, GPU, or other device’s memory can use separate data structures. Fine-grained based translation can be used for data that is cached in the DPU DRAM. Coarse-grained based translation can be used for uncached data to locate the origin of the requested data in an efficient manner, with an optional corresponding fine-grained data structure for access permissions.

[0046] To manage permissions (whether the compute server can access all of the mapped ranged or parts of it), additional, fine -grained structures can be defined access permissions - such as a bitmap per range for access control.

[0047] Since the compute server will utilize its own translation tables to access memory from its processes, and the exposed DPU memory window / PCIe BAR can be large, there is no need for fine-grained translations and mapping in the DPU from the DPU VA (which corresponds to the CPU’s view of the DPUs addresses) to the remote resource’s addresses. These can be performed on the compute server side. The remote resource addresses is comprised of (from the most significant bit to the least significant bit): a descriptor defining the backing device (occupying the most significant bits of the address of the address), the offset within the device. This structure will resemble a MTT (memory translation table) / range translation table, since the device mappings to the exposed linear address space will be relatively large. Caching, Buffering

[0048] The DPU exposes a memory region to the CPUe of the computer servers processor’s address space that is byte-addressable, and (optionally) cacheable by the CPU; of the DPU. The exposed memory regions will be inherently mapped to specific data center resources, including but not limited to: remote memory or local / remote accelerator memory etc. Multiple resources can be mapped to the DPU’s exposed memory region(s)

[0049] Upon receipt of a memory access, the DPU will either serve the request from its internal, managed page-cache, or will fetch it from the appropriate source to the page cache / temporary buffer, and return the cache-line requested / serve the write.

[0050] The DPU may implement custom prefetching / caching algorithms fit for the use of the current CPUe application and / or destination data source. The DPU can serve as memory expansion to the server, or can be explicitly utilized via memory accesses to it through interfaces (not limited to) such as DAX. The DPU does not have to implement any cache coherence protocols since it can use existing technologies such as PCIe, CXL etc.

[0051] The DPU can generate an exception / interrupt as a response to a memory operation via synchronous mechanisms such as returning a poison. This is done so long-latency accesses will not stall the CPU for an extended period of time, allowing the OS to schedule another process in the meantime, or future CPU hardware support to schedule another task. It will then notify the OS when the requested data is available

[0052] The present DPU architecture may provide flexible granularity of data transfers for caching and data serving purposes (e.g. free from granularity constraints of typically 4KiB to be transferred). It may also provide direct access other remote memory types such as onboard GPU memory which might otherwise be inaccessible. Indeed, referring to fig. 1C, the presentation to CPUe of the computer server of a virtually addressable memory (DPU VA) is shown corresponding to remote memory composed of the storage server’s DRAM (STORAGE DRAM) and accessible onboard GPU memory. In this scenario, the CPUe is able to access in a byte-addressable manner the memory of the GPU and the remote memory of the storage server, using regular load and store instructions, without performing API calls or implementing transport-specific functions.

[0053] The remote memory may include available on-board memory in hardware residing on a remote host or heterogenous computational devices which include memory, the DPU being further configured to, under the control of the internal CPU, perform memory accesses from such hardware. The remote memory in itself does not have to be byte-addressable, but the DPU may present it in a byte-addressable manner to the CPUe. The DPU may perform data transfers to the remote memory at any granularity that is required by the remote resource.

[0054] The DPU may be further configured to update the presentation to the external CPU of the virtually addressable memory upon a change in the provision of remote memory. This allows the resources which are available to the CPUe to be dynamically reconfigured, attached, and detached, offering additional resource provisioning flexibility.

[0055] The DPU may be further configured to speculatively cache data in anticipation of future memory access requests for that data. Traditional or advanced, Al based prefetching may be included and accelerated by the DPU. The DPU will also incorporate configurable cache replacement algorithms.

[0056] Hardware Acceleration

[0057] Parallel processing in the DPU can independently manage issues such as data retrievals and page-fault handling which would otherwise impact the performance of CPUe of the compute server. I.e. CPUe would normally have to handle page-faults itself which would involve page table manipulations, data transfer initialization etc., amounting to a substantial processing overhead. Also, NPUs can be programmable using micro-code which typical computes faster than general purposes OS applications running on CPUe.

[0058] NPUs and CPU; of the DPU may communicate with the CPUe of the compute server. For example, CPU; may be configured to determine if there is or there is likely to be a delay in fulfilling a memory access. Upon doing so, the DPU may perform an action for CPUe of the computer server including at least one of: notifying the external CPU of the delay; notifying the external CPU of temporary buffering in the DPU to accommodate the delay; notifying the external CPU of remote memory error or unavailability; interrupting the external CPU; and signalling an exception to the external CPU.

[0059] To prevent long-latency accesses from stalling CPUe of the computer server, CPU; of the DPU can generate exceptions or interrupts, signalling the OS to perform other tasks while waiting for data retrieval. This allows for better utilization of CPU resources and improved system responsiveness.

[0060] The present disclosure would typically be used in data center environments where efficient and flexible memory and resource management are critical, and DPUs leveraged to enable byte -addressable memory accesses to remote and stranded memory (especially in a fragmented cloud data center), as well as other devices' resources, such as GPUs.

[0061] It is also likely that the present disclosure would find application in high-performance computing where efficient memory access and management are crucial for performance. I.e. the present disclosure reducing latency and overhead burden associated with traditional memory access methods . In this regard, Al and machine learning models often require large amounts of memory and quick access to data stored on accelerators like GPUs. The disclosure facilitates direct, byte-addressable access to these resources.

[0062] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive.

Claims

CLAIMS1. A Data Processing Unit, DPU, configured to provide hardware acceleration for memory accesses to remote memory by, under the control of an internal Central Processing Unit, CPU, of the DPU: presenting to an external CPU a virtually addressable memory; receiving from the external CPU memory access requests for the virtually addressable memory; translating virtual addresses of received memory access requests to corresponding physical addresses of remote memory; and performing memory accesses from the physical addresses of the remote memory.

2. A DPU according to claim 1 comprising dynamic random-access memory for caching and / or buffering data of memory accesses.

3. A DPU according to claim 1 or claim 2 further configured to speculatively cache data in anticipation of future memory access requests for that data.

4. A DPU according to any preceding claims further comprising a plurality of network processors configured to perform memory accesses in parallel and under the control of the internal CPU.

5. A DPU according to claim 4 wherein at least one network processor is configured to perform a memory access by: establishing a Remote Direct Memory Access, RDMA, or other communication and data offloading connection with a remote memory.

6. A DPU according to claim 4 or claim 5 wherein at least one network processor and the internal CPU are configured to handle an error in performing a memory access without requiring an action of the external CPU.

7. A DPU according to any preceding claim further configured to: determine if a memory access request from the external CPU will or is likely to experience delay in fulfilling the memory accesses.

8. A DPU according to claim 7 further configured to, based on the determining, perform an external CPU related action including at least one of: notifying the external CPU of the delay; notifying the external CPU of temporary buffering in the DPU to accommodate the delay; notifying the external CPU of remote memory error or unavailability; interrupting the external CPU; and signalling an exception to the external CPU.

9. A DPU according to any preceding claim further configured to, under the control of the internal CPU: maintain a translation table which maps the virtually addressable memory presented to the external CPU with corresponding physical addresses of remote memory or local cached data inside the DPU.

10. A DPU according to claim 9 further configured to, under the control of the internal CPU: cache address translations in an address translation cache, and search the address translation cache upon receiving a memory access request.

811. A DPU according to any preceding claim, comprising a Base Address Register, BAR, to receive memory access requests from the external CPU.

12. A DPU according to any preceding claim, wherein the remote memory includes available on-board memory in hardware residing on a remote host, the DPU being further configured to, under the control of the internal CPU, perform memory accesses from such hardware.

13. A DPU according to any preceding claim further configured to update the presentation to the external CPU of the virtually addressable memory upon a change in the provision of remote memory.

14. A CPU configured to issue virtually addressed, memory access requests to a DPU according to any of the preceding claims.

15. A data server comprising: a CPU according to claim 14: and a DPU according to any of claims 1 to 13.

Citation Information

Patent Citations

  • Apparatus, method and computer program product for efficient software-defined network accelerated processing using storage devices which are local relative to a host

    US11940935B2

  • Buffer allocation

    US20240211392A1