Flexible ISA With Integrated Programmable Fabric for Coherent Offload
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing compute models for integrated circuits, such as CPUs and microprocessors, face limitations due to high latency, lack of memory coherency, and inflexibility in offload computing implementations, particularly in interconnects like PCIE/Ethernet and UPI/IAL/CCIX-based accelerators.
Innovation Solution
Incorporating a programmable fabric, such as an FPGA, into the processor's architecture to create a flexible instruction set architecture (ISA) that enables efficient offloading of computations, provides cache coherency, and reduces latency through clock transition circuitry for coherent communication across different clock domains.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If PCIE/Ethernet-based accelerators are used for offload computing, then device flexibility is improved, but communication latency increases significantly (100 μs)
Solution Approach 1:
The patent merges the accelerator device with the CPU by integrating it into the same interconnect fabric and memory space. This allows the accelerator to communicate with the CPU and access memory through the same high-speed pathways, eliminating the need for separate PCIE/Ethernet interfaces and reducing communication latency while maintaining flexibility through programmable logic.
Solution Approach 2:
The patent introduces a gateway or interface layer that mediates between the accelerator and the CPU/memory system. This intermediary enables the accelerator to participate in the CPU's memory-coherent address space, allowing low-latency access to memory and registers through the existing high-speed interconnect rather than through external PCIE/Ethernet interfaces.
2Reliability
If UPI/IAL/CCIX-based accelerators are used for offload computing, then communication latency is reduced (1 μs) and memory coherency is provided, but device flexibility is limited due to fine-grained memory sharing requirements
Solution Approach 1:
The patent creates a universal interface that allows the accelerator to function in multiple modes - it can access memory through the coherent interconnect like UPI/IAL/CCIX accelerators, but it can also be configured with different memory access patterns and communication protocols. The programmable logic can be adapted to support various workloads and communication requirements without being constrained to fine-grained memory sharing only.
Solution Approach 2:
The patent introduces dynamic configurability to the accelerator, allowing it to change its operational characteristics based on workload requirements. The programmable logic can be reconfigured to optimize for different access patterns, communication protocols, and memory sharing granularities, providing flexibility while maintaining low latency and memory coherency when needed.
3Reliability
If UPI/IAL/CCIX-based accelerators are integrated into core software before utilization, then memory coherency is achieved, but ease of operation is reduced due to integration complexity
Solution Approach 1:
The patent enables the accelerator to participate in the CPU's memory-coherent address space through automatic address translation and memory management. The system's existing memory management infrastructure handles coherency and address mapping for the accelerator without requiring manual integration into core software, reducing operational complexity while maintaining memory coherency.
4Productivity
If a programmable fabric is incorporated into the processor architecture, then adaptability and computation efficiency are improved, but device complexity increases
Solution Approach 1:
The patent segments the programmable fabric into modular functional units that can be independently configured and managed. This segmentation allows the complex functionality to be broken down into manageable blocks that can be programmed and controlled through standardized interfaces, reducing the operational complexity despite the increased adaptability and computation efficiency.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The flexible ISA enhances computation efficiency by allowing the programmable fabric to handle specific compute types effectively, reduces latency, and provides flexibility for future workload acceleration with custom instructions, while maintaining memory coherency and lowering communication latency.
Implementation Method 1
a first fractional PLL communicatively coupled to the CPU and the PLD wherein the first fractional PLL performs clock synchronization operations comprising: receive an indication of the CPU clock frequency; receive an indication of the maximum PLL clock frequency; determine the first PLL clock frequency below the maximum PLL clock frequency, wherein the CPU clock frequency is an integer multiple value of the first PLL clock frequency; provide an indication of the first PLL clock frequency to the first PLL to set the first PLL clock frequency; and provide a clock ratio between the CPU clock frequency and the first PLL clock frequency to the CPU to enable the CPU to communicate with the PLD
Data Source
AI summary
A semiconductor device may include a programmable fabric and a processor. The processor may utilize one or more extension architectures. At least one of these extension architectures may be used to integrate and/or embed the programmable fabric into the processor as part of the processor. Systems and methods for transitioning data between the programmable fabric and the processor associated with different clock domains is described.


