IOMMU Pre-Translation for DMA Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing systems face significant latency issues due to the overhead of address translations during direct memory access (DMA) operations, particularly in virtualized environments where nested translations are required, leading to performance degradation and increased CPU cycles.

Innovation Solution

The implementation of a parallel pipeline architecture that pre-translates memory addresses before DMA requests are made, using an Input-Output Memory Management Unit (IOMMU) to cache translations, thereby reducing the latency associated with address translation and improving overall system performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If address translation is performed during DMA operations in virtualized environments, then memory access security and virtualization support are improved, but latency increases and performance decreases

Engineering Contradiction:
Improvememory access securityVSAvoidaddress translation latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs address translation before DMA operations by having the IOMMU pre-walk page tables and populate translation caches with virtual-to-physical address mappings. This preliminary action ensures that when DMA requests occur, the translations are already available, eliminating latency during actual memory access operations while maintaining virtualization security

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If nested address translations are performed for virtualized DMA operations, then virtualization support and memory isolation are improved, but processing time and CPU overhead increase

Engineering Contradiction:
Improvevirtualization supportVSAvoidtranslation processing time
Core Design Contradiction:
Adaptability or versatilityVSDuration of action of moving object

Solution Approach 1:

The system performs nested page table walks in advance during device driver initialization or memory mapping operations, populating the IOMMU translation cache with both guest virtual-to-physical and host virtual-to-physical translations. This eliminates the need for real-time nested translations during DMA operations, reducing processing time while maintaining full virtualization support

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The IOMMU acts as an intermediary that caches address translations between the guest virtual address space and host physical address space. By maintaining translation caches and performing pre-walking of nested page tables, the IOMMU mediates the complex nested translation process, isolating the overhead from the DMA operation itself and enabling fast direct memory access while preserving memory isolation guarantees

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If page table walks are performed during DMA memory access, then correct physical address mapping is achieved, but access speed and throughput decrease

Engineering Contradiction:
Improveaddress mapping accuracyVSAvoidmemory access speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The IOMMU performs page table walks and validates address mappings in advance, storing verified virtual-to-physical translations in its translation cache. When DMA operations occur, the system uses these pre-validated translations directly, maintaining address mapping accuracy while achieving high-speed memory access without real-time page table walks

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240020241A1Apparatus and method for address pre-translation to enhance direct memory access by hardware subsystems
Publication Date: 2024.01.18 INTEL CORP
  • US20240020241A1 patent drawing
  • US20240020241A1 patent drawing
  • US20240020241A1 patent drawing

AI summary

Apparatus and method for performing address pre-translation to enhance direct memory access by hardware subsystems is described herein. An apparatus embodiment includes a processor to execute an enqueue instruction to submit, to a hardware subsystem, a job descriptor describing a job to be performed. The job descriptor includes virtual addresses of memory locations in which data required to perform the job are stored. An input-output memory management unit (IOMMU) is to obtain the address translations for the virtual addresses responsive to a pre-translation request from the processor. The address translations is obtained by the IOMMU prior to receiving a memory access request from the hardware subsystem. The IOMMU is to retrieve the data from the memory location using the address translations and to provide the retrieved data to the hardware subsystem to fulfill the request.