Network interface device-based memory access

By identifying memory node characteristics and network information in advanced network interface devices and merging data access requests, the problem of increased write amplification factor in virtualization environments is solved, thereby improving memory access efficiency and saving resources.

CN121879896APending Publication Date: 2026-04-17INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTEL CORP
Filing Date
2025-09-12
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In a highly virtualized environment, server resources are heavily consumed by hypervisors, container engines, networking, and storage functions, leading to an increase in the write amplification factor (WAF) and impacting system performance.

Method used

By introducing advanced network interface devices (such as IPUs and DPUs), dedicated programmable cores accelerate infrastructure functions, and network interface devices identify the characteristics of memory nodes and network information, data access requests are merged to reduce the write amplification factor.

Benefits of technology

It effectively reduces the write amplification factor, improves memory access efficiency, reduces server resource consumption, and enhances system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121879896A_ABST
    Figure CN121879896A_ABST
Patent Text Reader

Abstract

The invention relates to network interface device-based memory access. An apparatus is disclosed comprising a network interface device comprising: a processor to implement network interface device functionality, and communication protocol engine circuitry, where the network interface device is to: receive a request to write data to a memory node, the memory node communicatively coupled to the network interface device; identifying network information corresponding to the request, where the network information includes at least one of a Quality of Service (QoS), a Physical Function (PF), a Virtual Function (VF), a Name Space Identifier (NSID), a Flow ID, a Service Level Target (SLO), or a Process Address Space ID (PASID); identifying a characteristic of the memory node, where the characteristic includes at least a page size of the memory node; and merging the data with other data on the memory node based on the network information and the characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to memory access for reducing write amplification factor and provides proof of such access for network interface-based devices. Background Technology

[0002] In highly virtualized environments, a significant amount of server resources are consumed by processing tasks outside of user applications. These tasks can include hypervisors, container engines, networking and storage functions, security, and substantial network traffic. To handle these diverse processing tasks, advanced network interface devices (NICs) with enhanced accelerators and network connectivity have been introduced. These NICs are known as Infrastructure Processing Units (IPUs), Data Processing Units (DPUs), programmable network devices, and so on. NICs can use dedicated programmable cores deployed within the device to accelerate and manage infrastructure functions. NICs can provide infrastructure offloading and an additional layer of security by acting as a host control point for running infrastructure applications. By using NICs, the overhead associated with running infrastructure tasks can be offloaded from server devices. Summary of the Invention

[0003] According to embodiments of this disclosure, an apparatus is provided, comprising: a network interface device, including: one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine circuits, wherein the network interface device is configured to: receive a request to write data to a memory node, the memory node being communicatively coupled to the network interface device; identify network information corresponding to the request, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); identify characteristics of the memory node, wherein the characteristics include at least the page size of the memory node; and, based on the network information and the characteristics, merge the data with other data pointing to the memory node.

[0004] According to embodiments of this disclosure, a method is provided, comprising: receiving a request by a network interface device to write data to a memory node, wherein the memory node is communicatively coupled to the network interface device, the network interface device including: one or more processors for implementing network interface device functions, and one or more communication protocol engine circuits; identifying network information corresponding to the request by the network interface device, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); identifying characteristics of the memory node by the network interface device, wherein the characteristics include at least the page size of the memory node; and merging the data with other data on the memory node by the network interface device based on the network information and the characteristics.

[0005] According to embodiments of this disclosure, a system is provided for facilitating memory access based on a network interface device to reduce WAF and provide proof, the system comprising: a processing unit cluster; and a network interface device communicatively coupled to the processing unit cluster, wherein the network interface device includes: one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine intellectual property (IP), and wherein the network interface device is configured to: receive a request to write data to a memory node, the memory node being communicatively coupled to the network interface device; identify network information corresponding to the request, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); identify characteristics of the memory node, wherein the characteristics include at least the page size of the memory node; and, based on the network information and the characteristics, merge the data with other data pointing to the memory node.

[0006] According to embodiments of the present disclosure, at least one machine-readable medium is provided, comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to perform the method as described above.

[0007] According to embodiments of this disclosure, an apparatus is provided for facilitating memory access based on a network interface device to reduce WAF and provide proof, including means for performing the method described above. Attached Figure Description

[0008] The embodiments described herein are illustrated in the accompanying drawings by way of example and not limitation. Similar reference numerals in the drawings indicate similar elements, and in the drawings:

[0009] Figure 1 This is a block diagram illustrating a computer system configured to implement one or more aspects of the embodiments described herein;

[0010] Figure 2 It is a block diagram of a system including selected components of a data center;

[0011] Figure 3 This is a block diagram of a data center as described in one or more examples of this specification.

[0012] Figure 4 The diagram illustrates a forwarding element that includes a control plane and a programmable data plane;

[0013] Figure 5 An example network interface device is described;

[0014] Figure 6 It is a block diagram illustrating the programmable network interface and data processing unit;

[0015] Figure 7 This is a block diagram illustrating an IP core development system;

[0016] Figure 8 The diagram illustrates a block diagram of an example computing environment based on the implementation described in this paper, which is used to provide memory access based on advanced network interface devices to reduce the write amplification factor (WAF) and provides proof.

[0017] Figure 9 This is a block diagram of an example Advanced Network Interface Device (ANID) implemented according to the method described in this paper, which is used to provide ANID-based memory access to reduce WAF and provide proof.

[0018] Figure 10 It is a block diagram depicting a computing environment supported by a partitioned namespace (ZNS) implemented by ANID, according to the implementation described in this paper;

[0019] Figure 11 This is a flowchart illustrating an embodiment of a method by which ANID provides flexible data placement (FDP) to reduce WAF;

[0020] Figure 12 This is a flowchart illustrating an embodiment of a method for reducing WAF by supporting ZNS with ANID; and

[0021] Figure 13This is a flowchart illustrating an embodiment of a method for providing data erasure proof by ANID. Detailed Implementation

[0022] In the following description, numerous specific details are set forth to provide a more thorough understanding. However, it will be apparent to those skilled in the art that the embodiments described herein may be implemented without one or more of these specific details. In other instances, well-known features have not been described to avoid obscuring the details of these embodiments.

[0023] Figure 1 This is a block diagram illustrating a computing system 100 configured to implement one or more aspects of the embodiments described herein. The computing system 100 includes a processing subsystem 101 having one or more processors 102 and a system memory 104, which communicate via interconnect paths that may include a memory hub 105. The memory hub 105 may be a separate component within a chipset assembly or may be integrated within one or more processors 102. The memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107 that enables the computing system 100 to receive input from one or more input devices 108. Furthermore, the I / O hub 107 enables a display controller to provide output to one or more display devices 110A, which may be included within one or more processors 102. In one embodiment, one or more display devices 110A coupled to I / O hub 107 may include local, internal, or embedded display devices.

[0024] For example, processing subsystem 101 includes one or more parallel processors 112 coupled to memory hub 105 via communication links 113, such as buses or fabrics. Communication links 113 can be any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or vendor-specific communication interfaces or architectures. The one or more parallel processors 112 can form a computationally centralized parallel or vector processing system that may include a large number of processing cores and / or processing clusters, such as many integrated core (MIC) processors. For example, the one or more parallel processors 112 form a graphics processing subsystem that can output pixels to one or more display devices 110A coupled via I / O hub 107. The one or more parallel processors 112 may also include a display controller and display interface (not shown) for direct connection to one or more display devices 110B.

[0025] Within the I / O subsystem 111, system storage unit 114 can be connected to I / O hub 107 to provide storage for computing system 100. I / O switch 116 can be used to provide an interface mechanism for connecting I / O hub 107 to other components (e.g., network adapter 118 and / or wireless network adapter 119 that can be integrated into the platform, and various other devices that can be added via one or more add-on devices 120). Add-on devices 120 may also include, for example, one or more external graphics processing units, graphics cards, and / or computing accelerators. Network adapter 118 can be an Ethernet adapter or another wired network adapter. Wireless network adapter 119 can include one or more of the following: Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more radio devices.

[0026] The computing system 100 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 107. Figure 1 The communication paths between the various interconnected components can be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect) based protocols (e.g., PCI Fast), or any (one or more) other bus or point-to-point communication interface and / or protocol, such as NVLink high-speed interconnect, compute fast link, etc. TM (Compute Express Link TM CXL TM(e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Ultra Ethernet Transport (UET), Intel Quick Path Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) Interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerator Cache Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G, and their variants, or wired or wireless interconnect protocols known in the art. In some examples, protocols such as architecture-based non-volatile memory express (NVMe) (NVMe over Fabrics, NVMe-oF) or NVMe can be used to copy or store data to virtualized storage nodes. In one embodiment, time-aware communication protocols are supported, including time-aware RDMA, time-aware NVME, and time-aware NVME-oF, where precise time and data consumption rates are used to control data transmission.

[0027] One or more parallel processors 112 may include circuitry optimized for graphics and video processing, such as video output circuitry, and constitute a graphics processing unit (GPU). Alternatively or additionally, one or more parallel processors 112 may include circuitry optimized for general-purpose processing while retaining the underlying computing architecture described in more detail herein. Components of the computing system 100 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more parallel processors 112, memory hub 105, processor(s) 102, and I / O hub 107 may be integrated into a system-on-a-chip (SoC) integrated circuit. Alternatively, components of the computing system 100 may be integrated into a single package to form a system-in-package (SIP) configuration. In one embodiment, at least a portion of the components of the computing system 100 may be integrated into a multi-chip module (MCM) that can interconnect with other MCMs to form a modular computing system.

[0028] In some configurations, in addition to processor(s) 102 and parallel processor(s) 112, computing system 100 includes one or more accelerator devices 130 coupled to memory hub 105. Accelerator devices 130 are configured to perform domain-specific acceleration of workloads to handle computationally intensive or high-throughput tasks. Accelerator devices 130 can alleviate the burden on processor(s) 102 and / or parallel processor(s) 112 of computing system 100. Accelerator devices 130 may include, but are not limited to, intelligent network interface cards, data processing units, cryptographic accelerators, storage accelerators, artificial intelligence (AI) accelerators, neural processing units (NPUs), and / or video transcoding accelerators.

[0029] It will be understood that the computing system 100 shown herein is illustrative only and can be varied and modified. The connection topology (including the number and arrangement of bridges, the number of processors(one or more) 102, and the number of parallel processors(one or more) 112) can be modified as needed. For example, system memory 104 can be directly connected to processors(one or more) 102 instead of via bridges, while other devices communicate with system memory 104 via memory hub 105 and processors(one or more) 102. In other alternative topologies, parallel processors(one or more) 112 are connected to I / O hub 107 or directly to one of the processors(one or more) 102 instead of memory hub 105. In other embodiments, I / O hub 107 and memory hub 105 can be integrated into a single chip. Two or more sets of processors 102 can also be attached via multiple sockets, which can couple to two or more instances of parallel processors(one or more) 112.

[0030] Some specific components shown in this document are optional and may not be included in all implementations of the computing system 100. For example, any number of add-on cards or peripherals may be supported, or some components may be eliminated. Furthermore, some architectures may differ in their implementations. Figure 1 The similar components shown use different terminology.

[0031] Figure 2 This is a block diagram of system 200, which includes selected components of a data center. The components of the data center shown may reside, for example, in a cloud service provider (CSP) or another data center. As a non-limiting example, this other data center could be a traditional enterprise data center, an enterprise "private cloud," or a "public cloud," providing services such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), or Software as a Service (SaaS). System 200 includes a number of workload clusters, including but not limited to workload cluster 218A and workload cluster 218B. Workload clusters can be clusters of individual servers, blade servers, rack servers, or any other suitable server topology.

[0032] System 200 may include workload clusters 218A-218B. Each workload cluster 218A-218B may include rack 248, which houses multiple servers (e.g., server 246). The racks 248 and servers of workload clusters 218A-218B may conform to rack unit (“U”) standards, where one rack unit conforms to a 19-inch wide rack frame, and a full-size industry-standard rack accommodates 42 device units (42U). A device unit (1U) (e.g., a 1U server) may be 1.75 inches high and approximately 36 inches deep. In various configurations, computing resources such as processors, memory, storage devices, accelerators, and switches may be installed in multiple rack units within rack 248.

[0033] Each server 246 can host a standalone operating system configured to provide server functionality, or the server can be virtualized. Virtualized servers can be controlled by a Virtual Machine Manager (VMM), hypervisor, and / or orchestrator, and can host one or more virtual machines, virtual servers, or virtual devices. Workload clusters 218A-218B can be co-located in a single data center or can be situated in different geographic data centers. Depending on contractual agreements, some servers may be dedicated to specific enterprise customers or tenants, while others may be shared.

[0034] Various devices in a data center can be interconnected via a switching architecture 270, which may include one or more high-speed routing and / or switching devices. The switching architecture 270 can provide north-south traffic 202 (e.g., traffic to and from a wide area network (WAN) such as the Internet) and east-west traffic 204 (e.g., traffic across data centers). Historically, north-south traffic 202 accounted for the majority of network traffic, but as network services have become more complex and distributed, the volume of east-west traffic 204 has increased. In many data centers, east-west traffic 204 now accounts for the majority of traffic. Furthermore, the volume of traffic may increase further as the capacity of each server 246 increases. For example, each server 246 may provide multiple processor sockets, each socket accommodating a processor with four to eight cores, and sufficient memory for these cores. Therefore, each server can host multiple VMs, and each VM can be a source of traffic generation.

[0035] To accommodate the high traffic volumes in data centers, a high-performance switching architecture 270 can be provided. The illustrated switching architecture 270 is an example of a flat network where each server 246 can be directly connected to a top-of-rack switch (ToR switches 220A-220B) (e.g., a "star" configuration). A first ToR switch 220A can be connected to a first workload cluster 218A, and a second ToR switch 220B can be connected to a second workload cluster 218B. Each ToR switch 220A-220B can be coupled to a core switch 260. This two-layer flat network architecture is shown as an illustrative example, and other architectures can be used, such as a three-layer star or leaf-ridge topology (also known as a "fat tree" topology) based on a "Clos" architecture, a hub-and-spoke topology, a mesh topology, a ring topology, or a three-dimensional mesh topology (as non-limiting examples).

[0036] The switching architecture 270 can be provided by any suitable interconnect using any suitable interconnect protocol. For example, each server 246 may include some type of architecture interface (FI), network interface card (NIC), or other host interface. The host interface itself may be coupled to one or more processors via an interconnect or bus (e.g., PCI, PCIe, or the like), and in some cases, the interconnect bus may be considered part of the switching architecture 270. The switching architecture may also use a PCIe physical interconnect to implement more advanced protocols, such as Compute Fast Link (CXL).

[0037] Interconnect technologies can be provided by a single interconnect or a hybrid interconnect, where PCIe provides on-chip communication, 1Gb or 10Gb copper Ethernet provides a relatively short connection to the ToR switches 220A-220B, and fiber optic cable provides a relatively long connection to the core switch 260. Interconnect technologies include (as a non-limiting example) Super Path Interconnect (UPI), Fibre Channel, Ethernet, Fibre Channel over Ethernet (FCoE), InfiniBand, PCIe, NVLink, or fiber optics, etc. Some of these technologies will be better suited for certain deployments or functions than others, and choosing the appropriate architecture for the immediate application is a skill of the average technician.

[0038] In one embodiment, the switching elements of architecture 270 are configured to implement switching technologies to improve network performance in high-usage scenarios. Example advanced switching technologies include, but are not limited to, adaptive routing, adaptive fault recovery, and adaptive and / or telemetry-based congestion control.

[0039] Adaptive routing enables ToR 220A-220B switches and / or core switch 260 to select the output port to which traffic is switched based on the load on the selected port (assuming unconstrained port selection is enabled). Adaptive routing tables can configure the forwarding tables of switches in architecture 270 to select among multiple ports between switches when multiple connections exist between a given set of switches in an adaptive routing group. Adaptive fault recovery (e.g., self-healing) allows for the automatic selection of a backup port when the port selected in the forwarding table is faulty or inactive, enabling rapid recovery in the event of port failures between switches. Notifications can be sent to neighboring switches when adaptive routing or adaptive fault recovery becomes active in a given switch. Adaptive congestion control configures a switch to send notifications to neighboring switches when port congestion on that switch exceeds a configured threshold, allowing these neighboring switches to adaptively switch to an uncongested port on that switch or a switch associated with an alternative route to the destination.

[0040] Telemetry-based congestion control uses real-time monitoring of telemetry from network devices (e.g., switches within architecture 270) to detect when congestion will begin to impact the performance of architecture 270 and proactively adjusts the switching tables within the network devices to prevent or mitigate impending congestion. ToR220A-220B switches and / or core switches 260 can implement built-in telemetry-based congestion control algorithms, or provide an API through which programmable telemetry-based congestion control algorithms can be implemented. Continuous feedback loops can be implemented, where the telemetry-based congestion control system continuously monitors the network and adjusts traffic flow in real time based on ongoing telemetry data. The telemetry-based congestion control system can learn and adapt, meaning it can adapt to changing network conditions and improve its congestion control strategy based on historical data and trends.

[0041] However, it should be noted that while this document presents a high-end architecture as an example, more generally, the switching architecture 270 may include any suitable interconnect or bus for a particular application, including traditional interconnects for implementing local area networks (LANs), synchronous fiber optic networks (SONETs), asynchronous transfer mode (ATM) networks, wireless networks (e.g., Wi-Fi and Bluetooth), 5G wireless, DSL interconnects, MOCA, or the like. It is also explicitly anticipated that new network technologies will emerge in the future to complement or replace some of the technologies listed herein, and any such future network topologies and technologies may become or form part of the switching architecture 270.

[0042] Figure 3This is a block diagram of a portion of a data center 300 according to one or more examples of this specification. The portion of data center 300 shown is not intended to include all components of a data center. The shown portion may be repeated multiple times within data center 300, and / or data center 300 may include portions other than those shown, depending on the capacity and functionality that data center 300 is intended to provide. In various embodiments, data center 300 may include... Figure 2 The system consists of components of a 200-data center, or may be different data centers.

[0043] Data center 300 comprises multiple logical elements that form multiple nodes, each of which can be provided by a physical server, a group of servers, or other hardware. Each server may also host one or more virtual machines depending on its application. Architecture 370 is provided to interconnect the various aspects of data center 300. Architecture 370 can be provided by any suitable interconnect technology, including but not limited to InfiniBand, Ethernet, PCIe, or CXL. Architecture 370 of data center 300 can be... Figure 2 A certain version of the architecture 270 of system 200, and / or including Figure 2 The system 200 has an architecture 270 of components. The data center 300 has an architecture 370 that can interconnect data center components, including: server nodes 304, 306, 308, 310, accelerator 330, gateways 340A-340B to other architectures, architectures or interconnect technologies, and orchestrator 360.

[0044] Server nodes 304, 306, 308, and 310 in data center 300 may include, but are not limited to, storage server node 304, heterogeneous computing server node 306, CPU server node 308, and storage server node 310. Heterogeneous computing server node 306 and CPU server node 308 can perform independent operations for different tenants, or they can collaboratively perform operations for a single tenant. Heterogeneous computing server node 306 and CPU server node 308 can also host virtual machines, which provide virtual server functionality to tenants in the data center.

[0045] Each of server nodes 304, 306, 308, and 310 can be connected to architecture 370 via architecture interface 372. The specific type of architecture interface 372 used depends at least in part on the technology or protocol used to implement architecture 370. For example, if architecture 370 is an Ethernet architecture, each architecture interface 372 can be an Ethernet network interface controller. If architecture 370 is a PCIe-based architecture, the architecture interface can be a PCIe-based interconnect. If architecture 370 is an InfiniBand architecture, the architecture interface 372 for heterogeneous computing server node 306 and CPU server node 308 can be a host channel adapter (HCA), while the architecture interface 372 for memory server node 304 and storage server node 310 can be a destination channel adapter (TCA). Various architecture interfaces can be implemented as intellectual property (IP) blocks, which can be inserted as modular units into integrated circuits, just as other circuits within data center 300 can do.

[0046] The heterogeneous computing server node 306 includes multiple CPU sockets, each CPU socket can accommodate a CPU 319, and each CPU 319 can (but is not limited to) include multiple cores. Xeon TM Processor. The CPU 319 could also be, for example, a multi-core, data center-grade processor. CPU, for example Grace TM CPU. The heterogeneous computing server node 306 includes a memory device 318 to store data for runtime execution and a storage device 316 to implement persistent storage of data in a non-volatile memory device. The heterogeneous computing server node 306 can perform heterogeneous processing via GPUs (e.g., GPU 317), which can be used, for example, to perform high-performance computing (HPC), media server, cloud gaming server, and / or machine learning computing operations. In one configuration, the GPUs can interconnect with each other via interconnect technologies such as PCIe, CXL, or NVLink, and interconnect with the CPU of the heterogeneous computing server node 306.

[0047] CPU server node 308 includes multiple CPUs (e.g., CPU 319), memory (e.g., memory device 318), and storage devices (storage device 316) for executing application and other program code that provides server functionality (e.g., a web server or other types of functionality remotely accessible to clients of CPU server node 308). CPU server node 308 can also execute program code that provides services or microservices that implement complex enterprise functionalities. Architecture 370 will be equipped with sufficient throughput to allow CPU server node 308 to be accessed by a large number of clients simultaneously, while also reserving sufficient throughput for heterogeneous computing server node 306 and enabling both heterogeneous computing server node 306 and CPU server node 308 to utilize memory server node 304 and storage server node 310. Furthermore, in one configuration, CPU server node 308 may primarily rely on distributed services provided by memory server node 304 and storage server node 310, because the memory and storage devices of CPU server node 308 may be insufficient for all operations that CPU server node 308 intends to perform. Conversely, large, high-speed, or dedicated storage pools can be dynamically configured across multiple nodes, allowing each node to access a large pool of resources that are not idle when a particular node is not using them. This type of distributed architecture is possible and potentially advantageous due to the high speed and low latency offered by modern data center architectures, as there is no reason to over-provision resources for each server node.

[0048] Memory server node 304 may include memory node 305 with memory technology suitable for storing data used during the execution of program code by heterogeneous computing server node 306 and CPU server node 308. Memory node 305 may include volatile memory modules (e.g., DRAM modules) and / or non-volatile memory technologies that can operate at speeds similar to DRAM, enabling these modules to have sufficient throughput and latency performance metrics to serve as a system memory layer during runtime. Memory server node 304 may be linked to heterogeneous computing server node 306 and / or CPU server node 308 via a technology such as CXL.mem, allowing host-to-device memory access. In this configuration, CPU 319 of heterogeneous computing server node 306 and CPU server node 308 may be linked to memory server node 304 and access memory node 305 of memory server node 304 in a manner similar to, for example, how CPU 319 of heterogeneous computing server node 306 can access the device memory of the GPU within heterogeneous computing server node 306. For example, memory server node 304 can provide remote direct memory access (RDMA) to memory node 305, where, for example, CPU server node 308 can use DMA operations to access memory resources on memory server node 304 via architecture 370 in a manner similar to how a CPU accesses its own onboard memory.

[0049] Heterogeneous computing server node 306 and CPU server node 308 can use memory server node 304 to extend runtime memory available during memory-intensive activities, such as training machine learning models. A tiered memory system can be enabled, where model data can be swapped into memory device 318 of heterogeneous computing server node 306 with higher performance and / or lower latency than local storage (e.g., storage device 316), and swapped out from memory device 318 of heterogeneous computing server node 306 to memory of memory server node 304. During workload execution setup, the entire working dataset can be loaded into one or more memory nodes 305 of memory server node 304, and loaded into memory device 318 of heterogeneous computing server node 306 during heterogeneous workload execution.

[0050] Storage server node 310 provides storage functionality to heterogeneous computing server node 306, CPU server node 308, and potentially storage server node 304. Storage server node 310 can provide networked bundle of disks (NBOD), programmable flash (PFM), redundant array of independent disks (RAID), redundant array of independent nodes (RAIN), network attached storage (NAS), or other non-volatile memory solutions. In one configuration, storage server node 310 can be coupled to heterogeneous computing server node 306, CPU server node 308, and / or storage server node 304 (e.g., NVMe-oF), thereby allowing the NVMe protocol to be implemented on architecture 370. In this configuration, the architecture interface 372 of these servers can be a smart interface that includes hardware for accelerating NVMe-oF operation.

[0051] Accelerator 330 within data center 300 can provide various acceleration functions, including hardware or coprocessor acceleration for functions such as packet processing, encryption, decryption, compression, decompression, network security, or other acceleration functions within the data center. In some examples, accelerator 330 may include a deep learning accelerator, such as a neural processing unit (NPU), which can offload matrix multiplication operations or other neural network operations from heterogeneous computing server node 306 or CPU server node 308. In some configurations, accelerator 330 may reside in a dedicated accelerator server or be distributed across various server nodes within data center 300. For example, an NPU may be directly attached to one or more CPU cores within heterogeneous computing server node 306 or CPU server node 308. In some configurations, accelerator 330 may include or be included within intelligent network controllers, infrastructure processing units(s), or data processing units that combine network controller functionality with accelerator, processor, or coprocessor functionality.

[0052] In one configuration, data center 300 may include gateways 340A-340B connecting architecture 370 to other architectures, architectures, or interconnect technologies. For example, if architecture 370 is an InfiniBand architecture, gateways 340A-340B may be gateways to an Ethernet architecture. If architecture 370 is an Ethernet architecture, gateways 340A-340B may include routers for routing data to other parts of data center 300 or a larger network (e.g., the Internet). For example, a first gateway 340A may connect to different networks or subnets within data center 300, while a second gateway 340B may be a router to the Internet.

[0053] Orchestrator 360 manages the provisioning, configuration, and operation of network resources within Data Center 300. Orchestrator 360 may include hardware or software running on a dedicated orchestration server. Orchestrator 360 may also be embodied in software, such as software running on CPU server node 308, which configures the software-defined networking (SDN) capabilities of components within Data Center 300. In various configurations, Orchestrator 360 can automate the provisioning and configuration of components within Data Center 300 by performing network resource allocation and template-based deployment. Template-based deployment is a method of provisioning and managing IT resources using predefined templates, which may be based on standard templates used by governments, service providers, financial institutions, standards, or customers. The template may also specify Service Level Agreements (SLAs) or Service Level Obligations (SLOs). Orchestrator 360 may also perform functions, including but not limited to load balancing and traffic engineering, network segmentation, security automation, real-time telemetry monitoring, and adaptive switching management (including telemetry-based adaptive switching). In some configurations, Orchestrator 360 can also provide multi-tenancy and virtualization support by enabling virtual network management (including creating and deleting virtual LANs (VLANs) and virtual private networks (VPNs)) and tenant isolation in multi-tenant data centers.

[0054] Figure 4 The illustration shows a forwarding element 400 comprising a control plane and a programmable data plane. Forwarding element 400 can be configured to forward data messages within a network based on a user-provided program. In some embodiments, the program includes instructions for forwarding data messages and performing other processes (e.g., firewall, denial-of-service attack protection, and load balancing operations). Forwarding element 400 can be any type of forwarding element, including but not limited to switches, switching chips, routers, or bridges. Forwarding element 400 can forward data messages associated with various technologies (e.g., but not limited to Ethernet, Super Ethernet, InfiniBand, or NVLink).

[0055] In various network configurations, forwarding elements are deployed as non-edge forwarding elements within the network to forward data messages from source devices to destination devices. In other network configurations, forwarding element 400 is deployed as an edge forwarding element at the network edge to connect computing devices (e.g., standalone computers or hosts) that serve as the source and destination of data messages. As a non-edge forwarding element, forwarding element 400 forwards data messages between forwarding elements within the network, for example, through an intermediate network architecture. As an edge forwarding element, forwarding element 400 forwards data messages between edge computing devices, to other edge forwarding elements, and / or to non-edge forwarding elements.

[0056] The forwarding element 400 includes circuitry for implementing a data plane 402, which performs forwarding operations to forward data messages received by the forwarding element 400 to other devices. The forwarding element 400 also includes circuitry for implementing a control plane 404, which configures the data plane circuitry. Furthermore, the forwarding element 400 includes physical ports 406 that receive and send data messages from and to devices outside the forwarding element 400. The data plane 402 includes ports 408 that receive and process data messages from the physical ports 406. The processed data messages are then forwarded to another port on the data plane 402, which is connected to another physical port of the forwarding element 400. In addition to being associated with the physical ports of the forwarding element 400, some ports 408 on the data plane 402 may also be associated with other modules of the data plane 402.

[0057] The data plane is implemented by programmable packet processor circuitry that provides several programmable message processing stages. These stages can be configured to perform data plane forwarding operations of forwarding element 400 to process data messages and forward them to their destinations. These message processing stages perform these forwarding operations by processing the data tuples (e.g., message headers) associated with the data messages received by data plane 402 to determine how to forward the messages. Each message processing stage includes a Matching Action Unit (MAU) that attempts to match the message's data tuples (e.g., header vectors) with a table record specifying the action to be performed on the data tuples. In some embodiments, the table record is populated by control plane 404 and is unknown when configuring the data plane to execute programs provided by network users. The programmable message processing circuitry is grouped into multiple message processing pipelines. These message processing pipelines can be ingress or egress pipelines, located before or after the traffic management stage of the forwarding element, which directs messages from the ingress pipeline to the egress pipeline.

[0058] The hardware details of data plane 402 depend on the communication protocol implemented via forwarding element 400. Ethernet switches use application-specific integrated circuits (ASICs) designed to handle Ethernet frames and the TCP / IP protocol stack. These ASICs are optimized for various traffic types, including unicast, multicast, and broadcast. Ethernet switch ASICs are typically designed to balance cost, power consumption, and performance, although high-end Ethernet switches may support more advanced features such as deep packet inspection and advanced QoS (Quality of Service). InfiniBand switches use dedicated ASICs designed for ultra-low latency and high throughput. These ASICs support features optimized for processing the InfiniBand protocol, such as RDMA, and provide support for other features that utilize precise timing and high-speed data processing. However, high-end Ethernet switches may support RoCE (RDMA over Converged Ethernet), which offers similar advantages to InfiniBand but with higher latency compared to native InfiniBand RDMA.

[0059] The forwarding element 400 can also be configured as an NVLink switch (e.g., an NVSwitch), which interconnects multiple graphics processors via the NVLink connectivity protocol. When configured as an NVLink switch, the forwarding element 400 can provide increased GPU-to-GPU bandwidth to GPU servers (relative to GPU servers interconnected via InfiniBand). NVLink switches can reduce network traffic hotspots that may occur when interconnected GPU-equipped servers perform operations such as distributed neural network training.

[0060] Generally, when data plane 402 and a program executing on data plane 402 (e.g., a program written in P4 language) cooperate to perform message or packet forwarding operations on incoming data, control plane 404 determines how messages or packets should be forwarded. The behavior of the program executing on data plane 402 is partially determined by control plane 404, which populates a matching action table with specific forwarding rules. The forwarding rules used by the program executing on data plane 402 are independent of the data plane program itself. In one configuration, the control plane may be coupled to management port 410, which allows an administrator to configure forwarding element 400. The data connection established via management port 410 is separate from the data connections used for ingress and egress data ports. In one configuration, management port 410 may be connected to management plane 405, which facilitates administrative access to the device, supports analysis of device status and health, and supports device reconfiguration. Management plane 405 may be part of control plane 404 or communicate directly with control plane 404. In one implementation, the administrator does not have direct access to the components of control plane 404. Instead, information is collected by management plane 405, and changes to control plane 404 are executed by management plane 405.

[0061] Figure 5 An example network interface device 500 is depicted. In one configuration, the network interface device 500 may include a transceiver 502, a transmit queue 507, a receive queue 508, a memory 510, a bus interface 512, and a DMA engine 552. The network interface device 500 may also include a system-in-package (SiP) 550, which includes a processor 505 for implementing intelligent network interface device functions, and an accelerator 506 for various acceleration functions, such as NVMe-oF or RDMA. In some implementations, the network interface device 500 may include a SoC in place of and / or supplement the SiP 550. The specific configuration of the network interface device 500 depends on the protocol implemented via the network interface device 500.

[0062] In various configurations, network interface device 500 can be configured to interface with networks including, but not limited to, InfiniBand, Ethernet, or NVLink. For example, transceiver 502 can receive and transmit packets according to InfiniBand, Ethernet, or NVLink protocols, although other protocols may also be used. Transceiver 502 can receive packets from and transmit packets to the network via the network medium. Transceiver 502 may include PHY circuitry 514 and Media Access Control (MAC) circuitry 516. PHY circuitry 514 may include encoding and decoding circuitry to encode and decode data packets according to applicable physical layer specifications or standards. MAC circuitry 516 can be configured to assemble data packets to be transmitted into packets that include destination and source addresses, network control information, and error detection hashes.

[0063] SiP 550 may include processors, which may be any combination of the following: CPU processor, graphics processing unit (GPU), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or other programmable hardware devices that allow programming of network interface device 500. For example, an intelligent network interface may use processor 505 to provide packet processing capabilities in the network interface. The operational configuration of processor 505 (including a programmable data plane processor) may be programmed using the following: a protocol-independent packet processor (P4), C, Python, Broadcom Network Programming Language (NPL), x86, or ARM-compatible executable binaries, or other executable binaries.

[0064] Packet distributor 524 can use time slot allocation to provide distribution of received packets for processing by multiple CPUs or cores. Interrupt merging circuitry 522 can perform interrupt regulation, wherein interrupt merging circuitry 522 waits for multiple packets to arrive or waits for timeouts to expire before generating an interrupt to the host system to process the received(one or more) packets(s). Network interface device 500 can perform receive segment merging (RSC), wherein the individual parts of an incoming packet are combined into a segment of the packet. Network interface device 500 can then provide the merged packet to the application. DMA engine 552 can copy packet headers, packet payloads, and / or descriptors directly from host memory to the network interface and vice versa, instead of copying packets to an intermediate buffer at the host and then copying them from the intermediate buffer to the destination buffer using another copy operation. Memory 510 can be any type of volatile or non-volatile storage device and can store any queues or instructions used to program network interface device 500. Transmit queue 507 can include data to be transmitted by the network interface or a reference to that data. Receive queue 508 can include data received by the network interface from the network or a reference to that data. Descriptor queue 520 may include descriptors referencing data or packets in transmit queue 507 or receive queue 508. Bus interface 512 can provide an interface with a host device. For example, bus interface 512 may be PCI Fast compatible, although other interconnect standards may also be used.

[0065] Figure 6 This is a block diagram illustrating a programmable network interface 600 and a data processing unit. The programmable network interface 600 is a programmable network engine that can be used to accelerate network-based computing tasks in a distributed environment. The programmable network interface 600 can be coupled to a host system via a host interface 670. The programmable network interface 600 can be used to accelerate network or storage operations of the host system's CPU or GPU. The host system can be, for example, a node in a distributed learning system used to perform distributed training, such as... Figure 6 As shown. The host system can also be a data center node within a data center.

[0066] In one embodiment, a programmable network interface 600 can accelerate access to a remote storage device containing model data. For example, the programmable network interface 600 can be configured to present the remote storage device as a local storage device on the host system. The programmable network interface 600 can also accelerate RDMA operations performed between the GPU of the host system and the GPU of the remote system. In one embodiment, the programmable network interface 600 can implement storage functions, such as, but not limited to, NVMe-oF. The programmable network interface 600 can also accelerate encryption, data integrity, compression, and other operations on remote storage on behalf of the host system, thereby allowing the remote storage to approach the latency of storage devices directly attached to the host system.

[0067] The programmable network interface 600 can also perform resource allocation and management on behalf of the host system. Storage security operations can be offloaded to the programmable network interface 600 and performed in conjunction with the allocation and management of remote storage resources. Network-based operations for managing access to remote storage devices can be performed by the programmable network interface 600 instead of by the host system's processor.

[0068] In one embodiment, network and / or data security operations can be offloaded from the host system to the programmable network interface 600. Data center security policies for the data center node can be handled by the programmable network interface 600 instead of the host system's processor. For example, the programmable network interface 600 can detect and mitigate attempted network-based attacks (e.g., DDoS) against the host system, preventing attacks from compromising the host system's availability.

[0069] The programmable network interface 600 may include a system-on-a-chip (SoC 620) that executes an operating system via multiple processor cores 622. The processor cores 622 may include general-purpose processor (e.g., CPU) cores. In one embodiment, the processor cores 622 may also include one or more GPU cores. The SoC 620 may execute instructions stored in a memory device 640. The memory device 650 may store local operating system data. The memory device 650 and memory device 640 may also be used to cache remote data for the host system. Network ports 660A-660B enable connectivity to a network or architecture and facilitate network access to the SoC 620 and the host system via the host interface 670. In one configuration, a first network port 660A may connect to a first forwarding element, while a second network port 660B may connect to a second forwarding element. Alternatively, both network ports 660A-660B may connect to a single forwarding element using a link aggregation protocol (LAG). The programmable network interface 600 may also include an I / O interface 675, such as a USB interface. I / O interface 675 can be used to couple external devices to programmable network interface 600 or as a debug interface. Programmable network interface 600 also includes management interface 630, which enables software on the host device to manage and configure programmable network interface 600 and / or SoC 620. In one embodiment, programmable network interface 600 may also include one or more accelerators or GPUs 645 to accept offloading of parallel computing tasks from SoC 620, the host system, or remote systems coupled via network ports 660A-660B. For example, programmable network interface 600 may be configured with a graphics processor and participate in general-purpose or graphics computing operations in a data center environment.

[0070] One or more aspects can be implemented by representative code stored on a machine-readable medium, which represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions representing various logics within a processor. When read by a machine, these instructions can cause the machine to fabricate logic to perform the techniques described herein. This representation, referred to as an "IP core," is a reusable logic unit for an integrated circuit, which can be stored on a tangible machine-readable medium as a hardware model describing the structure of the integrated circuit. This hardware model can be provided to various customers or manufacturing facilities, which load the hardware model onto fabrication machines that manufacture integrated circuits. Integrated circuits can be fabricated such that the circuit performs the operations described in association with any of the embodiments described herein.

[0071] Figure 7This diagram illustrates a block diagram of an IP core development system 700. This IP core development system 700 can be used to fabricate integrated circuits to perform the operations of the architectures and data center components described herein. The IP core development system 700 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to construct entire integrated circuits (e.g., SOC integrated circuits). Design facility 730 can generate software simulations 710 of the IP core designs in a high-level programming language (e.g., C / C++). Software simulation 710 can be used to design, test, and verify the behavior of the IP cores using simulation model 712. Simulation model 712 can include functional, behavioral, and / or timing simulations. Register transfer level designs (RTL designs 715) can then be created or synthesized from simulation model 712. RTL design 715 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers, including the associated logic performed using the modeled digital signals. In addition to RTL design 715, lower-level designs at the logic or transistor levels can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.

[0072] The RTL design 715 or its equivalent can also be synthesized by a design facility into a hardware model 720, which may take the form of a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. The IP core design can be stored and delivered to a fabrication facility 765 using non-volatile memory 740 (e.g., hard disk, flash memory, or any non-volatile storage medium). The fabrication facility 765 can be a third-party fabrication facility. Alternatively, the IP core design can be transmitted via a wired connection 750 or a wireless connection 760 (e.g., via the Internet). The fabrication facility 765 can then fabricate an integrated circuit at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations according to at least one embodiment described herein.

[0073] Memory access based on advanced network interface devices for reducing write amplification factor and providing proof.

[0074] In highly virtualized environments, a significant amount of server resources are consumed by processing tasks outside of user applications. These tasks can include hypervisors, container engines, networking and storage functions, security, and substantial network traffic. To handle these diverse processing tasks, advanced network interface devices with enhanced accelerators and network connectivity have been introduced. These advanced network interface devices are known as Infrastructure Processing Units (IPUs), Data Processing Units (DPUs), programmable network devices, and so on.

[0075] Advanced network interface devices (ANICs) can accelerate and manage infrastructure functions using dedicated programmable cores deployed within the device. ANICs can provide infrastructure offloading and an additional layer of security by acting as a host control point for running infrastructure applications. By using ANICs, the overhead associated with running infrastructure tasks can be offloaded from server devices.

[0076] In the implementation described in this article, the Advanced Network Interface Device (ANID) may be referred to as an Advanced Network Interface Device (ANID), a network interface device, a programmable network interface device, an IPU, or a DPU. For the purposes of this discussion, the Advanced Network Interface Device will be referred to as ANID.

[0077] One infrastructure task that ANID can handle is memory access. Memory access can be a costly part of any computing system. Memory commands can include storage commands, which can be generated by a requesting device (which may include circuitry for generating the storage commands) and sent to a target device (which may include memory and circuitry, such as a controller for decoding and processing the storage commands). The data returned by the storage device can be consumed (e.g., processed) by a consuming device, which may be the same as the requesting device or a different device.

[0078] Storage commands can be transmitted to and / or executed by any suitable memory node, including addressable memory. For example, such a memory node can include memory devices and / or storage devices, such as storage drives (e.g., SSDs, hard disk drives, etc. with flash memory); storage devices; host memory (e.g., storing data of applications running on the XPU; in some cases, host memory may be DRAM or other volatile memory); caches (e.g., L1 cache, L2 cache, last-level cache, other dedicated caches, etc.); first-in-first-out (FIFO) structures within caches; scratch memory within caches; memory cards; universal serial bus (USB) drives; dual in-line memory modules (DIMMs), such as non-volatile DIMMs (NVDIMMs); ​​storage devices integrated into devices such as smartphones, cameras, or media players; or other suitable storage devices. In various implementations, storage devices can be used in any suitable configuration, such as memory pools, two-level memory (2LM), multi-level memory, compute fast link (CXL) connections, multi-tenancy, and scalable I / O virtualization (e.g., scalable IOV) environments.

[0079] The memory of a storage device may include non-volatile memory and / or volatile memory. Non-volatile memory is a storage medium that does not utilize electricity to maintain the state of the data stored on it; therefore, even if power is interrupted to the device housing the memory, non-volatile memory can maintain a defined state. Non-limiting examples of non-volatile memory may include any one or a combination of the following: 3D crosspoint memory, phase change memory (e.g., memory using chalcogenide glass phase change materials in the memory cells), ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, polymer memory (e.g., ferroelectric polymer memory), ferroelectric transistor random access memory (Fe-TRAM), ovonic memory, antiferroelectric memory, nanowire memory, electrically erasable programmable read-only memory (EEPROM), memristor, single-stage or multi-stage phase change memory. PCM memory, spin Hall effect magnetic RAM (SHE-MRAM), spin-transfer torque magnetic RAM (STTRAM), resistive memory, magnetoresistive random access memory (MRAM) incorporating memristor technology, resistive memory including metal oxide base, oxygen vacancy base and conductive bridge random access memory (CB-RAM), spintronic junction-based devices, magnetic tunnel junction (MTJ-based devices, DW (domain wall) and SOT (spin-orbit transfer) based devices, thyristor-based memory devices, or any combination of the above, or other memory.

[0080] Volatile memory is a storage medium that uses electricity to maintain the state of the data stored in it (therefore, volatile memory is a memory whose state (and the data stored therein) is indeterminate if power is interrupted to the device housing it). Dynamically volatile memory requires refreshing the data stored in the device to maintain its state. An example of dynamically volatile memory includes DRAM (Dynamic Random Access Memory) or certain variants, such as Synchronous DRAM (SD RAM). The memory subsystem described herein is compatible with a variety of memory technologies, such as DDR3 (Double Data Rate version 3), DDR4 (DDR version 4), DDR4E (DDR version 4, extended), LPDDR3 (Low Power DDR version 3), LPDDR4 (Low Power Double Data Rate (LPDDR) version 4), WIO2 (Wide I / O 2), HBM (High Bandwidth Memory DRAM), DDR5 (DDR version 5), HBM2 (HBM version 2), or other memory technologies or combinations thereof, as well as technologies based on derivatives or extensions of such specifications.

[0081] In some memory protocols, memory removal commands (e.g., release commands, erase commands, delete commands, data consistency for write refresh, etc.) can provide the ability to quickly remove large areas of memory (e.g., in storage drives such as SSDs). In some cases, for SSDs, such as flash memory, a phenomenon known as write amplification can occur. Write amplification refers to the amount of actual information physically written to the storage medium being a multiple of the expected amount of logical information written (called the write amplification factor (WAF)). Because SSDs and flash memory should be erased before they can be rewritten, and erase operations are much coarser in granularity than write operations, the processes used to perform these operations result in multiple moves (or rewrites) of user data and metadata. This multiplication effect (or WAF) increases the number of writes used over the lifetime of the SSD, which shortens the time it can operate reliably. The increased writes also consume bandwidth in the SSD and / or flash memory, which degrades the write performance of the SSD and / or flash memory.

[0082] In some current methods for mitigating WAF (Web Application Firewall), SSD controllers can implement Flexible Data Placement (FDP) and / or Partition Namespaces (ZNS). FDP allows the host to provide hints about where to place data in memory via virtual handles or pointers. ZNS is used with NVMe... TM The protocol's SSD command set exposes a partitioned block storage interface between the host and the SSD, allowing the SSD to precisely align data to its media. ZNS divides the SSD into logically independent and individually addressable storage spaces, with each ZNS having its own I / O queue. However, traditional FDP and ZNS methods are not network-aware solutions and do not consider network-related information in their data placement methods.

[0083] Another technical challenge encountered when integrating with SSDs involves proving the erasure of SSD memory. Typically, logical addresses are used when addressing SSDs and flash memory externally. When a command to erase SSD memory (e.g., flash memory) is received, the proof of erasure is the erasure of physical addresses. Traditionally, to improve performance, some SSDs and flash memory confirm the memory has been erased by declaring that logical addresses are no longer assigned to physical addresses. The physical address is sent to a garbage collection unit, which erases the data at the physical address over time for reuse by other data. However, proving to the end user (application, customer, etc.) that data has been erased requires confirmation that the data at the physical address has been erased.

[0084] Therefore, the implementation of this disclosure addresses the aforementioned technical problems by providing memory access based on advanced network interface devices to mitigate WAF and provide proof. In the implementation described herein, ANID is used to support FDP and ZNS in decision-making processes within FDP and ZNS methods using network-based information. For example, ANID can use flow, precise time, and other network-related / network-known information available to ANID (e.g., Quality of Service (QoS), Physical Function (PF) identifier, Virtual Function (VF) identifier, Namespace ID (NSID), Flow ID, Service Level Objective (SLO), Process Space Address ID (PASID), etc.) to notify the SSD (e.g., via hints) to allow data to be placed in the SSD to improve WAF. Furthermore, ANID can implement ZNS by utilizing the application's PASID to perform region lookup on the SSD's ZNS. Finally, the implementation described herein enables ANID to perform rapid proof of SSD erasure by notifying ANID using hints between the SSD and ANID when both logical and physical memory locations in the SSD have been erased. Then, ANID can prove that the data is no longer located on the SSD.

[0085] The technical advantages of the implementation of this disclosure include increased memory (e.g., SSD (e.g., flash memory)) lifetime due to improved WAF, improved data placement due to optimized data placement, thereby reducing length write and erase cycles and / or improving latency, and enhanced security due to proof of data erasure at physical addresses.

[0086] The following is for reference. Figures 9 to 13 More details are provided regarding the implementation of memory access based on advanced network interface devices for mitigating WAF and providing proof.

[0087] Figure 8 This diagram illustrates a block diagram of an example computing environment 800 according to an implementation described herein, which provides memory access based on advanced network interface devices to mitigate WAF and provide proofs. In one implementation, computing environment 800 illustrates an example computing environment that can use storage commands. Computing environment 800 may include various clusters (e.g., 840A-840C) of processing units 845A-845C (e.g., GPUs, TensorFlow processors, other types of accelerators, etc.). Cluster 840 may also include one or more ANIDs 850 to facilitate communication between processing units 845 and network 830. Network 830 may be further coupled to various storage devices 820A-820C and orchestrator 810.

[0088] Figure 8Elements having the same or similar names as elements in any other figure herein describe the same elements in the other figures, can operate or function in a similar manner, may include the same components, and may be linked to other entities as, but not limited to, those described elsewhere herein. Therefore, the discussion of any features of the graphics processor herein also discloses corresponding combinations with the example computing environment 800, but is not limited to.

[0089] In various embodiments, components of the computing environment 800 (including requesting devices, target devices, and / or consuming devices) can be coupled together via one or more networks (e.g., networks) that include any number of intermediate network nodes, such as routers, switches, or other computing devices. The network, requesting devices, and / or target devices can be part of any suitable network topology (e.g., data center network, wide area network, local area network, edge network, or enterprise network).

[0090] Storage commands can be transmitted from the requesting device to the target device via any suitable communication protocol (or multiple protocols), and / or data read in response to a storage command can be transmitted from the target device to the consumer device via any suitable communication protocol (or multiple protocols), such as Peripheral Component Interconnect (PCI), PCIe, CXL, Universal Serial Bus (USB), Serial Attached SCSI (SAS), Serial ATA (SATA), InfiniBand, Fibre Channel (FC), IEEE 802.3, IEEE 802.11, Super Ethernet, or other current or future signaling protocols. Storage commands can include, but are not limited to, commands for writing data, reading data, and / or erasing data. In certain embodiments, storage commands conform to logical device interface specifications (also referred to herein as network communication protocols), such as High Speed ​​Non-Volatile Memory (NVMe) or Advanced Host Controller Interface (AHCI).

[0091] A computing platform (e.g., computing environment 800) may include one or more requesting devices, consuming devices, and / or target devices. Such a device may include one or more processing units (e.g., processing unit 845) for generating storage commands, decoding and processing storage commands, and / or consuming (e.g., processing) data requested by storage commands. As used herein, the terms “processor unit,” “processing unit,” “processor,” or “processing element” may refer to any device or part of a device that performs the following operation: processing electronic data from registers and / or memory to convert that electronic data into other electronic data that can be stored in registers and / or memory. The processing unit may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), general-purpose GPUs (GPGPUs), accelerated processing units (APUs), field-programmable gate arrays (FPGAs), neural network processing units (NPUs), edge processing units (EPUs), vector processing units, software-defined processing units, video processing units, data processing units (DPUs), memory processing units, storage processing units, accelerators (e.g., graphics accelerators, compression accelerators, artificial intelligence accelerators, network accelerators), controller cryptographic processors (dedicated processors that execute cryptographic algorithms within hardware), server processors, I / O controllers, NICs (e.g., SmartNICs), infrastructure processing units (IPUs), microcode engines, memory controllers (e.g., cache controllers, host memory controllers, DRAM controllers, SSD controllers, hard disk drive (HDD) controllers, non-volatile memory controllers, etc.), or any other suitable type of processor unit. Therefore, a processor unit may be referred to as an XPU. In some implementations, the computing environment 800 may include components for implementing networks such as Ethernet, Super Ethernet, CXL, networks using proprietary network protocols, or other suitable networks.

[0092] In some embodiments, computing environment 800 may be a data center or other similar environment, wherein any combination of components may be housed together in a rack or shared within a data center enclosure. In various embodiments, computing environment 800 may represent a telecommunications environment, wherein any combination of components may be encapsulated together in a roadside / street facility or enterprise wiring closet.

[0093] In some embodiments, orchestrator 810 may act as a requesting device and send storage commands, as described herein, to storage devices 820A-820C, which act as target devices. Some of these commands may read data, which is then provided to processing unit 845, which acts as a consuming device. In some embodiments, processing unit 845 or ANID 850 may act as a requesting device. Therefore, processing unit 845 can be either a requesting device or a consuming device. In one implementation, ANID 850 may be associated with the storage devices described herein. Figure 5 The network interface device is the same as 500, or the same as the one described in this article. Figure 6 The programmable network interface 600 is the same as the data processing unit and may be referred to as an IPU or DPU, for example. As previously discussed, the ANID 850 in this implementation is configured to support FDP and ZNS to reduce WAF, and to support fast proof to confirm that physical addresses in memory have been erased.

[0094] Figure 9 This is a block diagram of an example ANID 900 based on the implementation described in this article. The ANID 900 is used to provide ANID-based memory access to mitigate WAF and provide proof. In one implementation, the ANID 900 can be associated with... Figure 8 The ANID 850 described herein is identical. In some implementations, the ANID 900 may be the same as described herein. Figure 5 Network interface device 500 and / or the network interface device described herein Figure 6 The programmable network interface 600 is the same as the data processing unit, and in some examples it may be referred to as IPU or DPU.

[0095] In one configuration, the ANID 900 may include a host interface 910, a communication protocol engine IP 920, and a network interface 970. The ANID 900 may also include a SoC 960, which includes a processor 962 for implementing intelligent network interface device functions, and an accelerator 964 for various acceleration functions, such as NVMe-oF or RDMA. The specific configuration of the ANID 900 depends on the protocol implemented via the ANID 900.

[0096] In various configurations, the ANID 900 can be configured to interface with networks including, but not limited to, InfiniBand, Ethernet, or NVLink. The SoC 960 may include processors, which can be any combination of: a CPU processor, a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other programmable hardware devices that allow programming of the ANID 900. For example, an intelligent network interface can use processor 962 to provide packet processing capabilities within the network interface. The operational configuration of processor 962 (including a programmable data plane processor) can be programmed using: a protocol-independent packet processor (P4), C, Python, the Broadcom Network Programming Language (NPL), x86, or ARM-compatible executable binaries, or other executable binaries. In the implementations described herein, the ANID 900 may provide one or more of the following: encryption services, compression-related services, storage security services, or access control, etc.

[0097] The communication protocol engine IP 920 can copy packet headers, packet payloads, and / or descriptors directly from host memory to the network interface and vice versa, instead of copying packets to an intermediate buffer at the host and then using another copy operation to copy them from the intermediate buffer to the destination buffer. In some implementations, multiple communication protocol engine IPs 920 can be implemented in the ANID 900. For example, an NVME (including NVMe-oF) communication protocol engine IP and / or an RDMA communication protocol engine IP can be implemented, although other communication protocols can also be used. The host interface 910 can provide an interface to host devices. For example, the host interface 910 can be PCIe compliant, although other interconnect standards can also be used.

[0098] As previously discussed, ANID 900 can provide ANID-based memory access to mitigate WAF and provide proof. In some implementations, ANID 900 can provide an application programming interface (API) through which ANID-based memory access can be implemented to mitigate WAF and provide proof. In some implementations, the API can query whether the computing system supports ANID-based memory access capabilities (to mitigate WAF and provide proof). For example, the API can query whether ANID 900 provides FDP or ZNS functionality as described herein. In some implementations, the API can enable or disable this functionality. For example, the API can be configured to enable and / or disable FDP or ZNS functionality in ANID 900.

[0099] As shown in the figure, ANID 900 includes an FDP circuit 930, a ZNS circuit 940, and a proof circuit 950. According to the implementation described herein, the FDP circuit 930 can provide ANID-based FDP storage hints to improve the WAF, the ZNS circuit 940 can provide ANID-based ZNS, and the proof circuit 950 can provide rapid proof of data erasure operations.

[0100] Regarding the FDP circuit 930, the ANID 900 can utilize the FDP circuit 930 to provide a data placement process for writing data to memory (e.g., SSD memory, including flash memory). In one implementation, the FDP circuit 930 can work in conjunction with the host interface 910 and a memory controller (not shown in the figure) (e.g., an SSD controller) hosted on the ANID 900 to merge (combine, join, adjacent storage, etc.) data to be written to memory with other data based on network-related information known to the ANID 900 due to its functionality and its placement in the computing environment. The ANID 900 is aware of network-related information such as QoS, PF, VF, NSID, stream ID, precise timing, SLO, PASID, etc. Furthermore, since the ANID 900 operates as a network interface device, it is aware of the memory / drive characteristics (e.g., memory page size, etc.) of the communication-coupled memory nodes (memory devices or storage devices, etc.).

[0101] Using network-related information and memory node characteristic data, FDP can merge data to be written to a memory node with other data on the memory node that conforms to network-related information. In one implementation, merging data may include providing the memory node with hints regarding data placement. In some implementations, merging data takes into account memory drive characteristics (e.g., flash page size) to inform where the data should be placed on the memory drive. In some implementations, FDP circuitry 930 can merge data targeting the same memory region (e.g., bank, sector, address, etc.).

[0102] In some implementations, the FDP circuit 930 can provide a namespace for the data merging process. The FDP circuit 930 can also support standard flash device functionality as part of the FDP process implemented at ANID 900.

[0103] In some implementations, the FDP circuit 930 utilizes memory node characteristics to create media alignment as part of the FDP process. For example, the FDP circuit 930 can align the data being merged according to the page size. Furthermore, the FDP circuit 930 can enable erase operations at any media boundary.

[0104] In some implementations, when the memory node is flash memory, the FDP circuitry 930 can provide RAID or XOR scrambling for the memory node (e.g., SSD or flash memory). Furthermore, the FDP circuitry 930 can perform flash garbage collection on all SSDs communicatively coupled to the ANID 900.

[0105] Regarding the ZNS circuit 940, ANID can utilize the ZNS circuit 940 to implement the ZNS described in this paper. As previously discussed, the ZNS uses NVMe. TM The protocol's SSD storage command set exposes a partitioned block storage interface between the host and the SSD, allowing the SSD to precisely align data to its media. ZNS divides the SSD into logically independent and individually addressable storage spaces, with each ZNS having its own I / O queue.

[0106] ANID 900 performs logical block address (LBA)-based placement in SSD memory (e.g., flash memory) or other memory nodes external to flash memory. ZNS circuitry 940 can support ZNS by leveraging identifiers (e.g., the process address space ID (PASID) of an application) to perform region lookups on the SSD's ZNS. A PASID is a feature that allows a single endpoint device to be shared among multiple processes while providing each process with a complete virtual address space (e.g., 64-bit). In the implementation described herein, each process has its own PASID, and each process can use its PASID to distinguish itself from other processes when sending commands (e.g., PCIe commands, such as write commands) to the device (e.g., ANID 900).

[0107] In one implementation, the ZNS circuit 940 can implement a horizontal scaling memory algorithm that looks up a specific region of a lookup table based on the PASID. Data is placed in a memory node based on which application (process) is accessing the data. The PASID of that application / process is used for the lookup to map the data to a specific region of the memory node. In the implementation described herein, the lookup process can be performed in hardware as a series of programmable lookups that can be programmed in software to be performed in the hardware and / or software paths of the ANID 900 and / or the ZNS circuit 940.

[0108] In one implementation, the ZNS circuit 940 can utilize hints that the target can send back to the initiator. For example, the initiator SW can use these hints to change how and to which applications (processes) a specific partition namespace is allocated. The target SW's memory can send these hints back to the initiator SW, which can then re-adjust and reprogram the tables for these lookup settings. Subsequently, the application (process) can be mapped to a region in the ZNS using its PASID. In one implementation, the hints may include, but are not limited to, indications of the region to use, that a specific region is slowing down and a different region should be used, WAF information associated with a specific region that could allow mappings to be changed to another region, and so on.

[0109] Figure 10 This is a block diagram depicting a computing environment 1000 supported by ZNS implemented by ANID, according to the implementation described herein. The computing environment 1000 may include memory nodes, such as SSD 1030, which includes multiple flash memory devices 1035A-1035C. The SSD 1030 can be associated with... Figure 8 The described storage devices are the same as 820A-820C. The SSD 1030 can be communicatively coupled to the ANID 1020. The ANID 1020 can be with... Figure 8 ANID 850 and / or Figure 9 It is the same as ANID 900. ANID 1020 may include ZNS circuitry 1022 and SSD controller 1024.

[0110] In one example implementation, multiple processes (applications) 1010A-1010C can access SSD 1030 via memory requests sent through ANID 1020. ZNS circuitry 1022 can work in conjunction with SSD controller 1024 to enable ZNS support for SSD 1030. As shown in computing environment 1000, process 1 1010A is associated with partition namespace 1 1035A, process 2 1010B is associated with partition namespace 2 1035B, and process 3 1010C is associated with partition namespace 3 1035C. ZNS circuitry 1022 can perform PASID lookup 1040 to identify partition namespaces 1035A-1035C associated with processes 1010A-1010C that request data from / to SSD 1030.

[0111] Regarding the proof circuit 950, the ANID 900 can utilize the proof circuit 950 to perform rapid proof of data erasure on a memory node. In some implementations, this proof can be completed quickly (e.g., every time the memory node is accessed). In one example use case, this functionality provides a security check that allows customers in a cloud service provider (CSP) data center to know that their data has been erased by having the ANID 900 prove the erasure.

[0112] In one implementation, the proof circuit 950 can coordinate the passing of hints between the memory node (e.g., an SSD) and the proof circuit 950. These hints can be configured to allow the proof circuit 950 to be notified when both the logical and physical memory locations within the memory node have been erased. The proof circuit 950 can then prove that the data is no longer located on that memory node.

[0113] As mentioned earlier, proving data erasure via the proof circuit 950 improves security by proving the erasure of data at the physical address. However, this is only a minimal result. In some cases, the proven portion of a flash entry is erased (related to the proof), but another portion is rewritten to a different location in the flash memory. For example, if there is a 4KB sector and the implementation attempts to prove 2KB has been erased, then the entire 4KB will be erased, while the remaining 2KB will be written back to a new physical flash memory location.

[0114] Figure 11 This is a flowchart illustrating an embodiment of method 1100, which uses ANID to provide FDP to reduce WAF. Method 1100 can be executed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (e.g., instructions running on a processing device), or a combination thereof. For simplicity and clarity, the processes of method 1100 are illustrated as a linear sequence; however, it is contemplated that any number of processes may be executed in parallel, asynchronously, or in different orders. Furthermore, for simplicity and ease of understanding, details regarding... Figures 1 to 10 The described components and processes. In one implementation, ANID (e.g., Figure 8 ANID 850 and / or Figure 9 The ANID 900 can execute method 1100.

[0115] Method 1100 begins with processing block 1110, where the ANID can receive a request to write data to an SSD memory communicatively coupled to the ANID. Then, at block 1120, the ANID can identify network information corresponding to the request. In one implementation, the network information includes one or more of the following: QoS, PF, VF, NSID, SLO, PASID, or flow ID.

[0116] Subsequently, at block 1130, ANID can identify the characteristics of the SSD storage. In one implementation, these characteristics include at least the page size of the SSD storage. Finally, at block 1140, ANID can merge the data with other data on the SSD storage based on network information and characteristics.

[0117] Figure 12 This is a flowchart illustrating an embodiment of method 1200, which uses ANID to support ZNS to reduce WAF. Method 1200 can be executed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (e.g., instructions running on a processing device), or a combination thereof. For simplicity and clarity, the processes of method 1200 are illustrated as a linear sequence; however, it is contemplated that any number of processes may be executed in parallel, asynchronously, or in different orders. Furthermore, for simplicity and ease of understanding, details regarding... Figures 1 to 11 The described components and processes. In one implementation, ANID (e.g., Figure 8 ANID 850 and / or Figure 9 The ANID 900 can execute method 1200.

[0118] Method 1200 processes block 1210, where the ANID can receive a request to write data to an SSD memory communicatively coupled to the ANID. Then, at block 1220, the ANID can identify the process address space ID (PASID) corresponding to the request.

[0119] Subsequently, at block 1230, ANID can identify the region in the SSD memory corresponding to PASID via a lookup process. Finally, at block 1240, ANID can write data to that region of the SSD memory.

[0120] Figure 13This is a flowchart illustrating an embodiment of method 1300, which provides data erasure proof by ANID. Method 1300 can be executed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (e.g., instructions running on a processing device), or a combination thereof. For simplicity and clarity, the process of method 1300 is illustrated as a linear sequence; however, it is contemplated that any number of processes may be executed in parallel, asynchronously, or in different orders. Furthermore, for simplicity and ease of understanding, details regarding... Figures 1 to 12 The described components and processes. In one implementation, ANID (e.g., Figure 8 ANID 850 and / or Figure 9 The ANID 900 can execute method 1300.

[0121] Method 1300 begins with processing block 1310, where the ANID can receive a request to erase data in an SSD memory communicatively coupled to the ANID. Then, at block 1320, the ANID can send a request to the SSD memory to erase the data. At block 1330, the ANID can use at least one hint between the ANID and the SSD memory to confirm that the physical address of the data has been erased.

[0122] Subsequently, at block 1340, ANID can receive confirmation that the physical address of the data has been erased in response to at least one prompt. Finally, at block 1350, ANID can respond to this confirmation to prove that the data has been erased.

[0123] The following examples relate to further embodiments. Example 1 is an apparatus for facilitating memory access based on a network interface device to reduce WAF and provide proof. The apparatus of Example 1 includes a network interface device comprising one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine intellectual property (IP), wherein the network interface device is configured to: receive a request to write data to a memory node, the memory node being communicatively coupled to the network interface device; identify network information corresponding to the request, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); identify characteristics of the memory node, wherein the characteristics include at least the page size of the memory node; and, based on the network information and the characteristics, merge the data with other data pointing to the memory node.

[0124] In Example 2, the subject matter of Example 1 may optionally include: wherein the network interface device for merging the data further includes: the network interface device for providing write prompts to the memory node for media placement. In Example 3, the subject matter of any one of Examples 1-2 may optionally include: wherein the network interface device is used to provide a namespace used when merging the data. In Example 4, the subject matter of any one of Examples 1-3 may optionally include: wherein the memory node includes at least a solid-state drive (SSD) memory, the SSD memory including flash memory, and wherein the network interface device is used to support standard flash memory device functionality of the flash memory.

[0125] In Example 5, the subject matter of any one of Examples 1-4 may optionally include: wherein the network interface device is further configured to: merge the data for at least one of the same bank, sector, or address of the flash memory. In Example 6, the subject matter of any one of Examples 1-5 may optionally include: wherein the network interface device is further configured to: provide at least one of RAID or XOR scrambling for the memory node. In Example 7, the subject matter of any one of Examples 1-6 may optionally include: wherein the network interface device is further configured to create media alignment for the data, and wherein the network interface device is further configured to enable erasure at the media boundary of the memory node.

[0126] In Example 8, the subject matter of any one of Examples 1-7 may optionally include: wherein the network interface device is further configured to provide at least one of the following: encryption service, compression-related service, storage security service, or access control. In Example 9, the subject matter of any one of Examples 1-8 may optionally include: wherein the network interface device is further configured to perform garbage collection on the memory node communicatively coupled to the network interface device. In Example 10, the subject matter of any one of Examples 1-9 may optionally include: wherein an application programming interface (API) is provided for at least one of the following operations: querying whether the Flexible Data Placement (FDP) function of the network interface device for enabling data merging is implemented on the network interface device, or enabling / disabling the FDP function on the network interface device.

[0127] Example 11 is a method for facilitating memory access based on a network interface device to reduce WAF and provide proof. The method of Example 10 may include: receiving a request by a network interface device to write data to a memory node, wherein the memory node is communicatively coupled to the network interface device, the network interface device including one or more processors for implementing network interface device functions, and one or more communication protocol engine intellectual property (IP); identifying network information corresponding to the request by the network interface device, wherein the network information includes at least one of: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); identifying characteristics of the memory node by the network interface device, wherein the characteristics include at least the page size of the memory node; and merging the data with other data on the memory node by the network interface device based on the network information and the characteristics.

[0128] In Example 12, the subject matter of Example 11 may optionally include: wherein, enabling the data merging further includes: the network interface device providing the memory node with write prompts for media placement. In Example 13, the subject matter of Examples 11-12 may optionally include: wherein the memory node includes at least a solid-state drive (SSD) memory with flash memory, and wherein the network interface device is configured to support standard flash memory device functionality for the flash memory.

[0129] In Example 14, the subject matter of Examples 11-13 may optionally include: providing at least one of RAID or XOR scrambling for the flash memory. In Example 15, the subject matter of Examples 11-14 may optionally include: creating media alignment for the data. In Example 16, the subject matter of Examples 11-15 may optionally include: enabling erasure at the media boundaries of the memory node.

[0130] Example 17 is a non-transitory computer-readable storage medium for facilitating memory access based on a network interface device to reduce WAF and provide proof. The non-transitory computer-readable storage medium of Example 17 stores instructions thereon that, when executed by one or more processors, cause the one or more processors to perform operations including: receiving a request by a network interface device to write data to a memory node, wherein the memory node is communicatively coupled to the network interface device, the network interface device including the one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine intellectual property (IP); identifying network information corresponding to the request by the network interface device, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); identifying characteristics of the memory node by the network interface device, wherein the characteristics include at least the page size of the memory node; and merging the data with other data on the memory node by the network interface device based on the network information and the characteristics.

[0131] In Example 18, the subject matter of Example 17 may optionally include: wherein, causing the data merging further includes: the network interface device providing the memory node with write prompts for media placement. In Example 19, the subject matter of Examples 17-18 may optionally include: wherein the memory node includes at least a solid-state drive (SSD) memory, the SSD memory including flash memory, and wherein the network interface device is configured to support standard flash memory device functionality of the flash memory. In Example 20, the subject matter of Examples 17-19 may optionally include: wherein the operation further includes providing at least one of RAID or XOR scrambling to the flash memory.

[0132] Example 21 is a system for facilitating memory access based on a network interface device to reduce WAF and provide proof. The system of Example 21 may optionally include: a cluster of processing units; and a network interface device communicatively coupled to the cluster of processing units, wherein the network interface device includes: one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine intellectual property (IP), and wherein the network interface device is configured to: receive a request to write data to a memory node, the memory node being communicatively coupled to the network interface device; identify network information corresponding to the request, wherein the network information includes at least one of: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); identify characteristics of the memory node, wherein the characteristics include at least the page size of the memory node; and, based on the network information and the characteristics, merge the data with other data pointing to the memory node.

[0133] In Example 22, the subject matter of Example 21 may optionally include: wherein the network interface device for merging the data further includes: the network interface device for providing write prompts to the memory node for media placement. In Example 23, the subject matter of any one of Examples 21-22 may optionally include: wherein the network interface device is used to provide a namespace used when merging the data. In Example 24, the subject matter of any one of Examples 21-23 may optionally include: wherein the memory node includes at least a solid-state drive (SSD) memory, the SSD memory including flash memory, and wherein the network interface device is used to support standard flash memory device functionality of the flash memory.

[0134] In Example 25, the subject matter of any one of Examples 21-24 may optionally include: wherein the network interface device is further configured to: merge the data for at least one of the same bank, sector, or address of the flash memory. In Example 26, the subject matter of any one of Examples 21-25 may optionally include: wherein the network interface device is further configured to: provide at least one of RAID or XOR scrambling for the memory node. In Example 27, the subject matter of any one of Examples 21-26 may optionally include: wherein the network interface device is further configured to create media alignment for the data, and wherein the network interface device is further configured to enable erasure at the media boundary of the memory node.

[0135] In Example 28, the subject matter of any one of Examples 21-27 may optionally include: wherein the network interface device is further configured to provide at least one of the following: encryption service, compression-related service, storage security service, or access control. In Example 29, the subject matter of any one of Examples 21-28 may optionally include: wherein the network interface device is further configured to perform garbage collection on the memory node communicatively coupled to the network interface device. In Example 30, the subject matter of any one of Examples 21-29 may optionally include: wherein an application programming interface (API) is provided for at least one of the following operations: querying whether the Flexible Data Placement (FDP) function of the network interface device for enabling data merging is implemented on the network interface device, or enabling / disabling the FDP function on the network interface device.

[0136] Example 31 is an apparatus for facilitating memory access based on a network interface device to reduce WAF and provide proof, comprising: means for receiving a request to write data to a memory node using the network interface device, wherein the memory node is communicatively coupled to the network interface device, the network interface device including one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine intellectual property (IP); means for identifying network information corresponding to the request using the network interface device, wherein the network information includes at least one of: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); means for identifying characteristics of the memory node using the network interface device, wherein the characteristics include at least the page size of the memory node; and means for merging the data with other data on the memory node using the network interface device based on the network information and the characteristics. In Example 32, the subject matter of Example 31 may optionally include: means further configured to perform the method of any one of Examples 12 to 16.

[0137] Example 33 is at least one machine-readable medium comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to perform the method according to any one of Examples 11 to 16. Example 34 is an apparatus for facilitating memory access of a network interface device to mitigate a WAF and provide proof, configured to perform the method of any one of Examples 11 to 16. Example 35 is an apparatus for facilitating memory access of a network interface device to mitigate a WAF and provide proof, including means for performing the method of any one of Examples 11 to 16. Details in these examples may be used anywhere in one or more embodiments.

[0138] The above description and accompanying drawings should be considered exemplary, not restrictive. Those skilled in the art will understand that various modifications and changes can be made to the embodiments described herein without departing from the broader spirit and scope of the features set forth in the appended claims.

Claims

1. An apparatus comprising: A network interface device includes: one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine circuits, wherein the network interface device is used for: Receive a request to write data to a memory node, the memory node being communicatively coupled to the network interface device; Identify network information corresponding to the request, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); Identify characteristics of the memory node, wherein the characteristics include at least the page size of the memory node; and Based on the network information and the characteristics, the data is merged with other data pointing to the memory node.

2. The apparatus of claim 1, wherein, The network interface device for enabling the data merging further includes: the network interface device for providing write prompts to the memory node regarding media placement.

3. The apparatus as described in any one of claims 1 to 2, wherein, The network interface device is used to provide a namespace for use when merging the data.

4. The apparatus according to any one of claims 1 to 3, wherein, The memory node includes at least a solid-state drive (SSD) memory, which includes flash memory, and wherein the network interface device is configured to support standard flash memory device functions of the flash memory.

5. The apparatus according to any one of claims 1 to 4, wherein, The network interface device is also used to merge the data for at least one of the same storage bank, sector, or address of the flash memory.

6. The apparatus according to any one of claims 1 to 5, wherein, The network interface device is also used to provide at least one of RAID or XOR scrambling for the memory node.

7. The apparatus according to any one of claims 1 to 6, wherein, The network interface device is also configured to create media alignment for the data, and wherein the network interface device is also configured to enable erasure at the media boundary of the memory node.

8. The apparatus according to any one of claims 1 to 7, wherein, The network interface device is also used to provide at least one of the following: encryption services, compression-related services, storage security services, or access control.

9. The apparatus according to any one of claims 1 to 8, wherein, The network interface device is also used to perform garbage collection on the memory node communicatively coupled to the network interface device.

10. The apparatus according to any one of claims 1 to 9, wherein, An application programming interface (API) is provided for at least one of the following operations: querying whether the Flexible Data Placement (FDP) function of the network interface device for enabling the data to be merged is implemented on the network interface device, or enabling / disabling the FDP function on the network interface device.

11. A method comprising: A network interface device receives a request to write data to a memory node, wherein the memory node is communicatively coupled to the network interface device, and the network interface device includes: one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine circuits. The network interface device identifies network information corresponding to the request, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID). The network interface device identifies characteristics of the memory node, wherein the characteristics include at least the page size of the memory node; and The network interface device, based on the network information and the characteristics, merges the data with other data on the memory node.

12. The method of claim 11, wherein, The data merging also includes the network interface device providing the memory node with write prompts for media placement.

13. The method according to any one of claims 11 to 12, wherein, The memory node includes at least a solid-state drive (SSD) memory with flash memory, and the network interface device is configured to support standard flash memory device functions of the flash memory.

14. The method of any one of claims 11 to 13, further comprising: Provide at least one of RAID or XOR scrambling to the memory node.

15. The method of any one of claims 11 to 14, further comprising: Create a media alignment for the data.

16. The method of any one of claims 11 to 15, further comprising: Erase is enabled at the media boundary of the memory node.

17. A system for facilitating memory access based on a network interface device to reduce WAF and provide proof, the system comprising: Processing unit cluster; as well as A network interface device communicatively coupled to the processing unit cluster, wherein the network interface device includes: one or more processors for implementing the functions of the network interface device, and one or more communication protocol engine intellectual property (IP), and wherein the network interface device is used for: Receive a request to write data to a memory node, the memory node being communicatively coupled to the network interface device; Identify network information corresponding to the request, wherein the network information includes at least one of the following: Quality of Service (QoS), Physical Function (PF), Virtual Function (VF), Namespace Identifier (NSID), Flow ID, Service Level Objective (SLO), or Process Address Space ID (PASID); Identify characteristics of the memory node, wherein the characteristics include at least the page size of the memory node; and Based on the network information and the characteristics, the data is merged with other data pointing to the memory node.

18. The system of claim 17, wherein, The network interface device for enabling the data merging further includes: the network interface device for providing write prompts to the memory node regarding media placement.

19. The system as claimed in any one of claims 17 to 18, wherein, The network interface device is used to provide a namespace for use when merging the data.

20. The system as claimed in any one of claims 17 to 19, wherein, The memory node includes at least a solid-state drive (SSD) memory, which includes flash memory, and wherein the network interface device is configured to support standard flash memory device functions of the flash memory.

21. The system as claimed in any one of claims 17 to 20, wherein, The network interface device is also used to provide at least one of RAID or XOR scrambling for the memory node.

22. The system as claimed in any one of claims 17 to 21, wherein, The network interface device is further configured to create media alignment for the data, and wherein the network interface device is further configured to enable erasure at the media boundary of the memory node, and wherein the network interface device is further configured to perform garbage collection on the memory node communicatively coupled to the network interface device.

23. The system as claimed in any one of claims 17-22, wherein, An application programming interface (API) is provided for at least one of the following operations: querying whether the Flexible Data Placement (FDP) function of the network interface device for enabling the data to be merged is implemented on the network interface device, or enabling / disabling the FDP function on the network interface device.

24. At least one machine-readable medium comprising a plurality of instructions, the instructions being responsive to execution on a computing device to cause the computing device to perform the method as claimed in any one of claims 11 to 16.

25. An apparatus for facilitating memory access based on a network interface device to reduce WAF and provide proof, comprising means for performing the method as claimed in any one of claims 11 to 16.