RDMA Network Interface Controller Multi-Destination Memory Writes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face limitations in efficiently transferring data to multiple memory domains in clustered systems with compute accelerators, such as GPUs, due to unreliable multicast and the need for multiple network transfers, which interrupt host processor operations and incur latency.
Innovation Solution
A protocol and technology that enables network interface controllers (NICs) with RDMA offload capabilities to perform a single data transfer to multiple memory domains, including shared and independent memory domains, directly into GPU memory without interrupting the host CPU, reducing network operations and latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If multiple network transfers are performed to transfer data to multiple memory domains, then data distribution is achieved, but network operations increase and latency increases
Solution Approach 1:
The patent combines multiple separate network transfer operations into a single network transfer operation. The NIC receives one packet containing data for multiple memory domains and uses RDMA to simultaneously write the data to multiple destination memory regions across different memory domains, eliminating the need for multiple separate network transfers and reducing both network operations and latency.
Solution Approach 2:
The patent enables a single network interface controller to perform multiple functions by writing the same data packet to multiple different memory domains (shared memory and independent memory domains) through a single network operation. This multi-functional approach allows one transfer operation to achieve what previously required multiple specialized transfers.
2Quantity of substance
If multiple network transfers are performed to transfer data to multiple memory domains, then data distribution is achieved, but host processor operations are interrupted
Solution Approach 1:
The patent extracts the data copying function from the host processor and transfers it to the NIC's RDMA engine. The host CPU only needs to initiate a single network receive operation, while the RDMA hardware automatically handles the complex task of copying data to multiple memory domains without CPU intervention, thus eliminating interruptions to host processor operations.
Solution Approach 2:
The system enables self-service by allowing the NIC and RDMA engine to autonomously perform data distribution to multiple memory domains without requiring host processor intervention. The single network receive operation automatically triggers RDMA writes that distribute data to multiple destinations independently of the host CPU, maintaining processor operation integrity.
3Loss of time
If multicast is used to transfer data to multiple destinations, then network operations are reduced, but reliability is insufficient for RDMA operations
Solution Approach 1:
Instead of using unreliable multicast to send one packet to multiple destinations, the patent inverts the approach by using reliable unicast RDMA writes from a single destination (the network packet) to multiple sources (multiple memory domains). This inversion maintains reliability by using RDMA's guaranteed delivery semantics while still achieving efficient multi-destination data distribution.
Data Source
AI summary
Examples described herein relate to configuring a target network interface to recognize packets that are to be written directly from the network interface to multiple memory destinations. A packet can include an identifier that a portion of the packet is to be written to multiple memory devices at specific addresses. The packet is validated to determine if the target network interface is permitted to directly copy the portion of the packet to memory of the target. The target network interface can perform a direct copy to multiple memory locations of a portion of the packet.


