Self-Synchronous Bus Interconnect for Programmable Processing Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current semiconductor devices face challenges in achieving high performance, low power consumption, and programmability, with ASICs being slow to market, DSPs being inefficient in data transfer, and FPGAs being expensive and power-hungry, while GALS circuits offer flexibility but lack programmability and efficient signal processing.
Innovation Solution
A semiconductor device with a M×N matrix of synchronously operating processing clusters connected by self-synchronous local and global buses, allowing for programmability using HDL or software, and featuring Arithmetic Logic Units, enabling a flexible and efficient programmable fabric that combines the advantages of ASICs and processors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If ASICs are used for high performance signal processing, then speed and power efficiency are improved, but time to market and adaptability deteriorate
Solution Approach 1:
The device is segmented into multiple processing clusters arranged in an M×N matrix, where each cluster can be independently configured. This segmentation allows the system to achieve ASIC-like performance for specific functions while maintaining reconfigurability through the array of clusters, resolving the contradiction between speed and adaptability.
Solution Approach 2:
The processing clusters are designed to be dynamically reconfigurable using HDL or software, allowing the system to adapt its functionality after manufacturing. This dynamic reconfiguration capability enables the device to maintain high performance for specific applications while also providing versatility for different signal processing tasks.
2Adaptability or versatility
If FPGAs are used for programmability and reconfiguration, then adaptability is improved, but power consumption and cost worsen
Solution Approach 1:
Each processing cluster is designed with local reconfiguration capability rather than requiring full-system reconfiguration like traditional FPGAs. This local quality approach allows individual clusters to be reconfigured independently, reducing the overall power consumption and resource overhead while maintaining adaptability where needed.
Solution Approach 2:
The patent uses replicated processing clusters throughout the M×N matrix, where each cluster contains essential processing functionality. This copying approach allows the system to achieve FPGA-like reconfigurability through multiple identical units that can be independently programmed, rather than using a single large reconfigurable fabric that consumes excessive power.
3Adaptability or versatility
If DSPs are used for programmability, then adaptability is improved, but processing efficiency worsens due to data transfer overhead
Solution Approach 1:
The patent transitions from the traditional von Neumann architecture (used in DSPs) to a data-flow architecture organized in an M×N matrix of processing clusters. This dimensional change in architectural organization allows data to flow directly between clusters through local interconnects, eliminating the bottleneck of centralized data transfer and improving processing efficiency while maintaining programmability.
Solution Approach 2:
The local interconnect fabric acting between processing clusters serves as an intermediary that enables direct peer-to-peer data exchange. This intermediary structure replaces the centralized memory-bus architecture of DSPs, allowing processing elements to communicate efficiently without the overhead of global data transfer, thus improving productivity while maintaining adaptability.
4Stability of the object's composition
If global clock synchronization is used in highly integrated circuits, then coordination between functional blocks is improved, but clock skew and power consumption worsen
Solution Approach 1:
The system is segmented into multiple independent clock domains, with each processing cluster operating from its own local clock. This segmentation eliminates the need for a single global clock distribution network, reducing clock skew issues and the power consumption associated with distributing a global clock signal across the entire chip.
Solution Approach 2:
Asynchronous wrapper circuits serve as intermediaries between different clock domains, enabling communication between processing clusters that operate on different local clocks. These wrappers translate between synchronous interfaces and asynchronous data flow, allowing the system to maintain synchronization and coordination without requiring a global clock signal.
Data Source
AI summary
A semiconductor device includes a plurality of processing clusters that operate synchronously internally and arranged in a M×N matrix. Each processing cluster is formed as a plurality of processing elements and clocked buses that interconnect the processing elements within each processing cluster. A self-synchronous cluster wrapper is operative with the processing elements such that each processing cluster forms a programmable module. Self-synchronous global and local buses interconnect the processing clusters for communicating externally. An input/output circuit interconnects the global and local buses.


