Multi-Frame DMA Support for Accelerator Memory Latency Bottlenecks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional processing accelerators, such as vector processing units, face inefficiencies due to system latencies in memory access and constrained bit widths, leading to bottlenecks in data processing and increased energy consumption, particularly in applications like vehicle automation.

Innovation Solution

Implementing a pixel processing engine with a 2D array of processing elements and direct memory access (DMA) systems that allow independent data processing and batch multiple DMA transfers, reducing memory access latencies and enabling scalable bit width adjustments without increasing memory size.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional accelerators use common memory source for data processing, then memory access is simplified, but system latencies increase and processing efficiency decreases

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidmemory access latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments the memory system into multiple independent memory sources (first memory and second memory) that can be accessed simultaneously by different processing elements. This segmentation allows parallel data retrieval operations, eliminating the bottleneck of a single common memory source and reducing overall system latency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension to memory access by implementing a two-dimensional array of processing elements with dual memory interfaces. Each processing element can access both first and second memories independently, creating a multi-dimensional memory access topology that increases throughput and reduces latency compared to conventional single-dimension access.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If conventional accelerators use fixed bit width configuration, then hardware design is simplified, but processing versatility and scalability are limited

Engineering Contradiction:
Improvebit width scalabilityVSAvoidaccelerator configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic bit width configuration where processing elements can be programmatically configured to operate at different bit widths (e.g., 8-bit, 16-bit, 32-bit) based on the specific processing task. This dynamic reconfigurability allows the same hardware to adapt to various precision requirements without physical redesign, enhancing versatility while maintaining manageable complexity through software control.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal processing element design that can handle multiple data types and bit widths using the same hardware infrastructure. By incorporating configurable arithmetic logic units and data paths that support variable precision operations, the system achieves multi-functionality, allowing a single accelerator to serve diverse applications from low-precision image processing to high-precision scientific computing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If conventional systems process each DMA transfer independently, then transfer configuration is simplified, but computational overhead increases and efficiency decreases

Engineering Contradiction:
ImproveDMA transfer efficiencyVSAvoidframe format configuration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges multiple independent DMA transfer operations into a single batched transfer operation by introducing a frame format structure that describes sequences of transfers. Instead of configuring and executing each DMA transfer separately, the system processes multiple transfers atomically using shared configuration data, reducing computational overhead and improving efficiency while managing complexity through structured data formats.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements preliminary action by pre-configuring frame formats that describe sequences of DMA transfers before actual data transfer occurs. The frame format structure allows advance specification of transfer parameters, patterns, and sequences, enabling the DMA system to execute multiple transfers with reduced real-time configuration overhead and improved throughput.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260037462A1Processing data using accelerators with multi-frame support
Publication Date: 2026.02.05 NVIDIA CORP
  • US20260037462A1 patent drawing
  • US20260037462A1 patent drawing
  • US20260037462A1 patent drawing

AI summary

In various examples, systems and methods are disclosed that relate to processing data using accelerators in a system on a chip. For example, a direct memory access (DMA) system can be programmed to perform one or more DMA transfers between source memory and destination memory in a sequence. The DMA system can signal to an accelerator that the DMA transfers are complete and that the data is available in the destination memory. In some examples, the DMA system can be configured to perform the one or more data transfers in accordance with frame formats associated with one of a plurality of frame types.