Embedded FPGA ISA for Low-Latency Coherent Offload Computing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing compute models for offload computing in integrated circuits face limitations due to latency, memory coherency, and flexibility issues, particularly in interconnects like PCIE/Ethernet and UPI/IAL/CCIX-based accelerators, which have high latency, lack memory coherency, or limited flexibility.

Innovation Solution

A flexible instruction set architecture (ISA) is implemented with an embedded programmable fabric (FPGA) to enhance processor functionality, providing flexibility, lower latency, and cache coherency, and enabling custom instructions for workload acceleration through the Advanced Matrix Extension (AMX) architecture.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If PCIE/Ethernet-based accelerators are used for offload computing, then device flexibility is improved, but latency increases significantly (100 μs)

Engineering Contradiction:
Improvedevice flexibilityVSAvoidlatency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent merges the CPU and accelerator into a single integrated processor package, eliminating the need for external interconnects like PCIE/Ethernet. The accelerator is directly coupled to the CPU cores through an on-die interconnect, combining the benefits of flexibility (through programmable accelerator) and low latency (through direct integration) that were previously traded off against each other.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If UPI/IAL/CCIX-based accelerators are used for offload computing, then latency is reduced (1 μs) and memory coherency is provided, but device flexibility is limited through fine-grained memory sharing

Engineering Contradiction:
Improvememory coherencyVSAvoiddevice flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal memory space that can be accessed by both CPU cores and accelerator logic without requiring fine-grained sharing configurations. The unified memory interface allows the same memory resources to serve multiple functions and workloads dynamically, providing both low latency and full flexibility simultaneously through a multi-functional memory architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If accelerators are integrated into core software before utilization, then memory coherency is achieved, but flexibility is reduced and integration complexity increases

Engineering Contradiction:
Improvememory coherencyVSAvoidintegration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where the processor automatically manages memory coherency and accelerator integration through hardware-supported virtualization and memory management. The system self-configures the accelerator resources and memory access paths without requiring manual software integration, reducing integration complexity while maintaining coherency through automated hardware-firmware cooperation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250315079A1Flexible Instruction Set Architecture Supporting Varying Frequencies
Publication Date: 2025.10.09 ALTERA CORP
  • US20250315079A1 patent drawing
  • US20250315079A1 patent drawing
  • US20250315079A1 patent drawing

AI summary

A semiconductor device may include a programmable fabric and a processor. The processor may utilize one or more extension architectures. At least one of these extension architectures may be used to integrate and/or embed the programmable fabric into the processor as part of the processor. Systems and methods for transitioning data between the programmable fabric and the processor associated with different clock domains is described.