Self-Synchronous Bus Interconnect for Programmable Processing Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current semiconductor devices face challenges in achieving high performance, low power consumption, and programmability, with ASICs being slow to market, DSPs being inefficient in data transfer, and FPGAs being expensive and power-hungry, while GALS circuits offer flexibility but lack programmability and efficient signal processing.

Innovation Solution

A semiconductor device with a M×N matrix of synchronously operating processing clusters connected by self-synchronous local and global buses, allowing for programmability using HDL or software, and featuring Arithmetic Logic Units, enabling a flexible and efficient programmable fabric that combines the advantages of ASICs and processors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If ASICs are used for high performance signal processing, then speed and power efficiency are improved, but time to market and adaptability deteriorate

Engineering Contradiction:
Improvesignal processing speedVSAvoidprogrammability
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The device is segmented into multiple processing clusters arranged in an M×N matrix, where each cluster can be independently configured. This segmentation allows the system to achieve ASIC-like performance for specific functions while maintaining reconfigurability through the array of clusters, resolving the contradiction between speed and adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The processing clusters are designed to be dynamically reconfigurable using HDL or software, allowing the system to adapt its functionality after manufacturing. This dynamic reconfiguration capability enables the device to maintain high performance for specific applications while also providing versatility for different signal processing tasks.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If FPGAs are used for programmability and reconfiguration, then adaptability is improved, but power consumption and cost worsen

Engineering Contradiction:
Improvereconfiguration capabilityVSAvoidpower consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

Each processing cluster is designed with local reconfiguration capability rather than requiring full-system reconfiguration like traditional FPGAs. This local quality approach allows individual clusters to be reconfigured independently, reducing the overall power consumption and resource overhead while maintaining adaptability where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses replicated processing clusters throughout the M×N matrix, where each cluster contains essential processing functionality. This copying approach allows the system to achieve FPGA-like reconfigurability through multiple identical units that can be independently programmed, rather than using a single large reconfigurable fabric that consumes excessive power.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If DSPs are used for programmability, then adaptability is improved, but processing efficiency worsens due to data transfer overhead

Engineering Contradiction:
ImproveprogrammabilityVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent transitions from the traditional von Neumann architecture (used in DSPs) to a data-flow architecture organized in an M×N matrix of processing clusters. This dimensional change in architectural organization allows data to flow directly between clusters through local interconnects, eliminating the bottleneck of centralized data transfer and improving processing efficiency while maintaining programmability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The local interconnect fabric acting between processing clusters serves as an intermediary that enables direct peer-to-peer data exchange. This intermediary structure replaces the centralized memory-bus architecture of DSPs, allowing processing elements to communicate efficiently without the overhead of global data transfer, thus improving productivity while maintaining adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Stability of the object's composition

If global clock synchronization is used in highly integrated circuits, then coordination between functional blocks is improved, but clock skew and power consumption worsen

Engineering Contradiction:
ImprovesynchronizationVSAvoidclock distribution power
Core Design Contradiction:
Stability of the object's compositionVSUse of energy by stationary object

Solution Approach 1:

The system is segmented into multiple independent clock domains, with each processing cluster operating from its own local clock. This segmentation eliminates the need for a single global clock distribution network, reducing clock skew issues and the power consumption associated with distributing a global clock signal across the entire chip.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Asynchronous wrapper circuits serve as intermediaries between different clock domains, enabling communication between processing clusters that operate on different local clocks. These wrappers translate between synchronous interfaces and asynchronous data flow, allowing the system to maintain synchronization and coordination without requiring a global clock signal.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7765382B2Propagating reconfiguration command over asynchronous self-synchronous global and inter-cluster local buses coupling wrappers of clusters of processing module matrix
Publication Date: 2010.07.27 HARRIS CORP
  • US7765382B2 patent drawing
  • US7765382B2 patent drawing
  • US7765382B2 patent drawing

AI summary

A semiconductor device includes a plurality of processing clusters that operate synchronously internally and arranged in a M×N matrix. Each processing cluster is formed as a plurality of processing elements and clocked buses that interconnect the processing elements within each processing cluster. A self-synchronous cluster wrapper is operative with the processing elements such that each processing cluster forms a programmable module. Self-synchronous global and local buses interconnect the processing clusters for communicating externally. An input/output circuit interconnects the global and local buses.