Heterogeneous xPU and FPGA Workload Scheduling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data centers processing AI and HPC workloads face inefficiencies due to operations that do not align with the core architecture of xPUs like GPUs and TPUs, leading to reduced value and performance.

Innovation Solution

Integrating FPGAs with xPUs via a high-bandwidth interconnect, allowing FPGAs to implement domain-specific accelerators for operations not aligned with xPU architectures, thereby maximizing silicon area for core xPU functions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If domain-specific acceleration operations are implemented directly on xPUs, then processing speed for specific operations improves, but silicon area available for core xPU functions decreases and operational efficiency deteriorates

Engineering Contradiction:
Improveprocessing speed for domain-specific operationsVSAvoidadaptability of xPU architecture
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system segments processing operations into two distinct categories: core xPU functions executed on xPUs and domain-specific acceleration operations executed on separate FPGAs. This segmentation allows each component to be optimized for its specific function without compromising the other, resolving the contradiction between processing speed for specific operations and adaptability of the xPU architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

FPGAs serve as intermediary devices between the xPUs and domain-specific operations. The FPGAs are configured to implement domain-specific acceleration operations and communicate with xPUs through a high-bandwidth interconnect, allowing xPUs to maintain their core architecture while still achieving accelerated processing for domain-specific operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If xPUs are used for all operations, then system simplicity is maintained, but processing efficiency for operations not aligned with xPU architecture deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidprocessing efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The system achieves multi-functionality by combining xPUs for core operations with FPGAs that can be reconfigured for different domain-specific operations. The high-bandwidth interconnect enables seamless communication between components, allowing the system to handle diverse workloads efficiently without sacrificing too much complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The FPGA components provide dynamic adaptability, allowing the system to reconfigure acceleration operations based on the specific workload requirements. This dynamic capability enables the system to optimize processing efficiency for different operations while maintaining a relatively simple overall architecture.

Inventive Principle:
Principle #15Dynamics

3Productivity

If FPGAs are integrated with xPUs, then processing efficiency for domain-specific operations improves, but system complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system merges xPUs and FPGAs into a unified heterogeneous processing platform, allowing both components to work together on the same system. This merging enables efficient processing of both core xPU functions and domain-specific operations while managing complexity through integrated design and high-bandwidth interconnect.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250117358A1Heterogenous Acceleration of Workloads using xPUs and FPGAs
Publication Date: 2025.04.10 ALTERA CORP
  • US20250117358A1 patent drawing
  • US20250117358A1 patent drawing
  • US20250117358A1 patent drawing

AI summary

The embodiment disclosed herein include a system for a data center that includes processing units (xPUs) and programmable logic devices. The xPUs may implement a design in hardware to perform specialized operations suited for the design. The programmable logic devices may implement different designs based on operations to be performed. For example, a schedule may be generated to process a workload received by the system. The schedule may include multiple phases that may be mapped to either the xPUs or the programmable logic devices based on an efficiency of performing operations of the phases using the xPU or the programmable logic device. In this way, the system may operate at maximum efficiency when processing the workload.