FPGA Streaming Links for Custom Instruction Acceleration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current FPGA-based SoCs face challenges in communication among multiple hardware and software processors due to high overhead from bus arbitration, limiting the complexity of custom instructions that can be performed within a single clock cycle.

Innovation Solution

Implementing a method that uses configurable uni-directional serial links formed from FIFO streaming communication networks to enable fast linked processing modules, allowing multiple custom operations to be performed within a single clock cycle through custom instructions, and providing a streaming communication channel between processors and a System on Chip communication bus.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If bus arbitration is used for communication among multiple processors in FPGA-based SoC, then communication can be implemented, but communication overhead increases and processing speed decreases

Engineering Contradiction:
Improvecommunication capabilityVSAvoidprocessing speed
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent segments the communication system by eliminating the shared bus arbitration mechanism and replacing it with dedicated point-to-point communication links between processors. Each processor has direct dedicated links to communicate with specific other processors, eliminating the need for bus arbitration and reducing communication overhead while maintaining communication capability.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If custom instructions are added to extend processor instruction sets, then application-specific operations can be performed, but processor complexity increases and clock speed is limited

Engineering Contradiction:
Improvecustom operation capabilityVSAvoidprocessor complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges custom instruction functionality with dedicated hardware accelerator modules that are integrated into the FPGA architecture. Instead of adding complex instructions to the processor core, the custom operations are implemented as separate hardware modules that can be triggered by simpler instructions, thereby maintaining processor simplicity while enabling complex custom operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces hardware accelerator modules as intermediaries between the processor and the complex custom operations. These accelerator modules handle the computationally intensive custom instructions, allowing the processor to remain simple while still performing complex operations through the intermediary hardware components.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If multiple custom operations are performed within a single clock cycle, then processing speed improves, but the complexity of implementing such operations increases

Engineering Contradiction:
Improveprocessing speedVSAvoidimplementation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments complex custom operations into separate hardware accelerator modules, each handling specific custom operations. This segmentation allows multiple operations to be performed in parallel across different modules within a single clock cycle, improving processing speed while managing implementation complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs custom operations in parallel across multiple dedicated links and hardware modules simultaneously, transitioning from sequential operation execution to parallel execution. This dimensional change from time-sequential to space-parallel processing enables multiple custom operations to complete within a single clock cycle, significantly improving processing speed.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS7676661B1Method and system for function acceleration using custom instructions
Publication Date: 2010.03.09 XILINX INC
  • US7676661B1 patent drawing
  • US7676661B1 patent drawing
  • US7676661B1 patent drawing

AI summary

A fast linked multiprocessor network including a plurality of processing modules implemented on a field programmable gate array and a plurality of configurable uni-directional links coupled among at least two of the plurality processing modules provide a streaming communication channel between at least two of the plurality of processing modules. Such configuration provides a function accelerator that can feed at least one processor with data values using one custom instruction to put data values on at least one uni-directional serial link and that can extract data values from at least one processor using one custom instruction to get data values from the at least one uni-directional serial link.