Global Memory Address Map for PCIe Interconnect Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional interconnect networks, such as PCIe networks, face latency and CPU interaction issues during internode data transfers between endpoint processors, like GPUs, due to the need for network fabric and bounce buffers, which hinder efficient data transfer and resource utilization.

Innovation Solution

The implementation of a global memory address map that dynamically partitions computing resources and enables Direct Memory Access (DMA) operations without buffers or IOMMU transactions, allowing direct inter-node communications by mapping non-overlapping memory address spaces across devices in the network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If conventional PCIe interconnect networks use network fabric and bounce buffers for internode data transfers, then data transfer between endpoint processors is enabled, but latency and CPU interaction overhead increase

Engineering Contradiction:
Improvedata transfer speedVSAvoidtransfer latency
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent extracts and removes the bounce buffer intermediate storage requirement from the data transfer path. By implementing a global memory address map that provides direct address translation between endpoint processors across nodes, the system eliminates the need for temporary buffer storage and CPU-mediated data copying, enabling direct memory access between remote endpoint processors.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a global memory address map as an intermediary translation mechanism that directly maps memory addresses across different nodes without requiring CPU involvement. This address map serves as a mediator that translates source endpoint processor addresses to destination endpoint processor addresses, enabling direct transfers while removing the traditional CPU bounce buffer intermediary.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If bounce buffers are used for internode data transfers, then data transfer is enabled, but CPU interaction is required which reduces resource utilization

Engineering Contradiction:
Improveresource utilizationVSAvoidCPU interaction requirement
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent implements self-service by enabling endpoint processors to perform Direct Memory Access (DMA) operations autonomously without CPU intervention. The global memory address map provides the necessary address translation information, allowing endpoint processors to directly access remote memory resources and complete data transfers independently, thereby freeing CPU resources for computational tasks.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If conventional memory address mapping is used, then device-specific address spaces are maintained, but efficient inter-node communication is hindered

Engineering Contradiction:
Improveaddress space flexibilityVSAvoidinter-node communication efficiency
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The patent implements universality by creating a global memory address map that provides a unified address space view across multiple nodes and multiple endpoint processors. This single address map structure serves multiple functions: it maintains device-specific address spaces, enables cross-node address translation, and supports direct memory access operations, thereby unifying the memory access interface while preserving individual device addressing requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12001370B2Multi-node memory address space for PCIe devices
Publication Date: 2024.06.04 ADVANCED MICRO DEVICES INC
  • US12001370B2 patent drawing
  • US12001370B2 patent drawing
  • US12001370B2 patent drawing

AI summary

A device in an interconnect network is provided. The device comprises an end point processor comprising end point memory and an interconnect network link in communication with an interconnect network switch. The device is configured to issue, by the end point processor, a request to send data from the end point memory to other end point memory of another end point processor of another device in the interconnect network and provide, to the interconnect network switch, the request using memory addresses from a global memory address map which comprises a first global memory address range for the end point processor and a second global memory address range for the other end point processor.