Clustered Emulation Processor Interconnect Architecture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As the number of emulation processors in design verification systems increases, the required chip area and power consumption for interconnect structures also rise significantly, leading to inefficiencies and increased costs, particularly evident in processor-based systems where the number of multiplexers and their width must increase to handle the additional processors, resulting in a severe area and power burden.
Innovation Solution
The solution involves clustering processors into groups, reducing the width of multiplexers while maintaining the same number, with each multiplexer inputting only one Node Bit Out (NBO) from each cluster, eliminating redundant selections and allowing processors within a cluster to access shared data storage arrays, thereby reducing the complexity and resources needed for interconnects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the number of emulation processors is increased to handle complex logic designs, then the verification capability is improved, but the chip area and power consumption for interconnect structures increase significantly
Solution Approach 1:
The system divides the large number of emulation processors into multiple clusters, where each cluster contains a subset of processors. This segmentation allows the interconnect structure to be organized hierarchically, with cluster-level multiplexers handling local connections and a smaller number of inter-cluster multiplexers handling global connections, thereby reducing the total chip area required for interconnect structures.
Solution Approach 2:
The patent introduces a hierarchical dimension to the interconnect architecture by adding cluster-level abstraction. Instead of a flat N:1 multiplexer structure, the system creates a two-level hierarchy: first organizing processors into M clusters, then creating inter-cluster connections. This dimensional change transforms the interconnect problem from O(N) to O(sqrt(N)) or better, reducing chip area.
2Productivity
If the number of emulation processors is increased to handle complex logic designs, then the verification capability is improved, but the power consumption for interconnect structures increases significantly
Solution Approach 1:
By segmenting processors into clusters and implementing cluster-level multiplexing, the patent reduces the total number of multiplexer inputs and outputs that need to be driven across the chip. This segmentation decreases switching activity and signal routing complexity, thereby reducing power consumption in the interconnect structures.
Solution Approach 2:
The patent merges multiple processor outputs within a cluster into a single cluster output through cluster-level multiplexers. This merging reduces the number of long-distance signal routes required, decreasing capacitive loading and dynamic power consumption associated with driving signals across the entire chip.
3Adaptability or versatility
If the width of multiplexers is increased to handle more processors, then the interconnect capability is improved, but the chip area and complexity increase severely
Solution Approach 1:
The patent segments the large N:1 multiplexer into multiple smaller multiplexers organized in a hierarchy. Each multiplexer handles a manageable subset of signals, reducing individual multiplexer complexity while maintaining overall interconnect capability through the hierarchical structure. This segmentation makes the system more adaptable and easier to implement.
Data Source
AI summary
The present system and methods are directed to the interconnection of clusters of emulation processors comprising emulation processors in a software-driven hardware design verification system. The processors each output one NBO output signal. The clusters are interconnected by partitioning a common NBO bus into a number of smaller NBO busses, each carrying unique NBO signals but together carrying every NBO. Each of the smaller NBO busses are passed into a series of multiplexers, each dedicated to a particular processor. The multiplexers select a signal for output back to the emulation clusters. The multiplexers that handle these smaller NBO busses are narrower than was previously required, thus reducing the amount of power, interconnect, and area required by the multiplexer array and dedicated interconnect.


