Global Memory Address Map for PCIe Interconnect Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional interconnect networks, such as PCIe networks, face latency and CPU interaction issues during internode data transfers between endpoint processors, like GPUs, due to the need for network fabric and bounce buffers, which hinder efficient data transfer and resource utilization.
Innovation Solution
The implementation of a global memory address map that dynamically partitions computing resources and enables Direct Memory Access (DMA) operations without buffers or IOMMU transactions, allowing direct inter-node communications by mapping non-overlapping memory address spaces across devices in the network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If conventional PCIe interconnect networks use network fabric and bounce buffers for internode data transfers, then data transfer between endpoint processors is enabled, but latency and CPU interaction overhead increase
Solution Approach 1:
The patent extracts and removes the bounce buffer intermediate storage requirement from the data transfer path. By implementing a global memory address map that provides direct address translation between endpoint processors across nodes, the system eliminates the need for temporary buffer storage and CPU-mediated data copying, enabling direct memory access between remote endpoint processors.
Solution Approach 2:
The patent introduces a global memory address map as an intermediary translation mechanism that directly maps memory addresses across different nodes without requiring CPU involvement. This address map serves as a mediator that translates source endpoint processor addresses to destination endpoint processor addresses, enabling direct transfers while removing the traditional CPU bounce buffer intermediary.
2Productivity
If bounce buffers are used for internode data transfers, then data transfer is enabled, but CPU interaction is required which reduces resource utilization
Solution Approach 1:
The patent implements self-service by enabling endpoint processors to perform Direct Memory Access (DMA) operations autonomously without CPU intervention. The global memory address map provides the necessary address translation information, allowing endpoint processors to directly access remote memory resources and complete data transfers independently, thereby freeing CPU resources for computational tasks.
3Adaptability or versatility
If conventional memory address mapping is used, then device-specific address spaces are maintained, but efficient inter-node communication is hindered
Solution Approach 1:
The patent implements universality by creating a global memory address map that provides a unified address space view across multiple nodes and multiple endpoint processors. This single address map structure serves multiple functions: it maintains device-specific address spaces, enables cross-node address translation, and supports direct memory access operations, thereby unifying the memory access interface while preserving individual device addressing requirements.
Data Source
AI summary
A device in an interconnect network is provided. The device comprises an end point processor comprising end point memory and an interconnect network link in communication with an interconnect network switch. The device is configured to issue, by the end point processor, a request to send data from the end point memory to other end point memory of another end point processor of another device in the interconnect network and provide, to the interconnect network switch, the request using memory addresses from a global memory address map which comprises a first global memory address range for the end point processor and a second global memory address range for the other end point processor.


