Bufferless Deflection Router for FPGA Network-on-Chip

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current FPGA network-on-chip (NOC) designs are inefficient in terms of resource usage and latency, struggling to achieve full-bandwidth data transmission across numerous client cores and high-bandwidth interfaces, with prior solutions being complex, resource-intensive, and unsuitable for large-scale implementations.

Innovation Solution

The Hoplite router and NOC system implement a 64-bit wide 4x4 directional torus deflection router using 1230 6-LUTs with a latency of 2-3 ns, featuring a directional torus topology, bufferless deflection routing, and optimized technology mapping, enabling efficient interconnection of hundreds of client cores over high-bandwidth links with reduced resource consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If prior-art buffered Virtual Channel routers are used, then reliability and throughput are improved, but device complexity and area consumption increase significantly

Engineering Contradiction:
ImprovethroughputVSAvoidrouter complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the buffer component from the router architecture, transitioning from buffered Virtual Channel routers to bufferless deflecting routers. This extraction eliminates the complexity associated with buffer management, virtual channel arbitration, and flow control mechanisms while maintaining network throughput through deflection routing that redirects traffic around congested nodes.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of using buffers to absorb traffic variations and ensure throughput (conventional approach), the patent inverts the approach by using deflection routing to actively redirect traffic away from congested areas. The router actively manages traffic flow by dynamically changing routing paths rather than passively absorbing variations with buffers.

Inventive Principle:
Principle #13The other way round (Inversion)

2Productivity

If buffered Virtual Channel routers with multiple virtual channels are implemented, then throughput is improved, but latency increases

Engineering Contradiction:
ImprovethroughputVSAvoidrouting latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent removes the buffer stage that inherently introduces latency in buffered routers. By eliminating buffers, messages are forwarded immediately through the network without waiting for buffer availability or arbitration delays, significantly reducing routing latency while maintaining throughput through deflection mechanisms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The bufferless deflecting router rushes messages through the network by eliminating the buffering stage entirely. Messages are forwarded immediately without being held in buffers, skipping the latency-introducing stages present in conventional buffered routers while deflection routing ensures messages still reach their destinations even under congestion.

Inventive Principle:
Principle #21Skipping (Rushing through)

3Area of stationary object

If conventional torus router designs are used, then area efficiency is improved, but multicast capability and flexibility are limited

Engineering Contradiction:
Improverouter areaVSAvoidmulticast support
Core Design Contradiction:
Area of stationary objectVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal router design that handles both point-to-point and multicast traffic using the same bufferless deflecting architecture. The router achieves multi-functionality by using dimension-order routing that naturally supports both unicast and multicast operations without requiring separate specialized hardware or complex control logic, maintaining area efficiency while enhancing versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The router employs dynamic routing decisions that can adapt to different traffic types (unicast or multicast) on the fly. The dimension-order routing algorithm dynamically determines the routing dimension based on destination coordinates, enabling flexible support for various communication patterns without sacrificing area efficiency or requiring static specialized structures.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3298740B1Directional two-dimensional router and interconnection network for field programmable gate arrays
Publication Date: 2023.04.12 GRAY RESEARCH LLC
  • EP3298740B1 patent drawingFigure 1
  • EP3298740B1 patent drawingFigure 2A~2B
  • EP3298740B1 patent drawingFigure 3

AI summary

A configurable directional 2D router for Networks on Chips (NOCs) is disclosed. The router, which may be bufferless, is designed for implementation in programmable logic in FPGAs, and achieves theoretical lower bounds on FPGA resource consumption for various applications. The router employs an FPGA router switch design that consumes only one 6-LUT or 8-input ALM logic cell per router per bit of router link width. A NOC comprising a plurality of routers may be configured as a directional 2D torus, or in diverse ways, network sizes and topologies, data widths, routing functions, performance-energy tradeoffs, and other options. System on chip designs may employ a plurality of NOCs with different configuration parameters to customize the system to the application or workload characteristics. A great diversity of NOC client cores, for communication amongst various external interfaces and devices, and on-chip interfaces and resources, may be coupled to a router in order to efficiently communicate with other NOC client cores. The router and NOC enable feasible FPGA implementation of large integrated systems on chips, interconnecting hundreds of client cores over high bandwidth links, including compute and accelerator cores, industry standard IP cores, DRAM/HBM/HMC channels, PCI Express channels, and 10G/25G/40G/100G/400G networks.