Cell-Based Multi-Processor Interconnect Architecture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multi-processor computer systems face performance limitations due to shared bus architectures and crossbar switch connectivity constraints, which result in latency and bandwidth issues, particularly in systems with a large number of CPUs and memory resources.
Innovation Solution
A cell-based architecture with multiple cells, I/O backplanes, and crossbar networks, including global and local crossbar networks, that provide high-speed data links and direct cell-to-cell connections to minimize latency and maximize bandwidth through efficient routing of messages across multiple crossbar hops.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a shared bus architecture is used to interconnect CPUs and memory resources, then device complexity is reduced, but communication bandwidth and system performance deteriorate due to bus contention and latency
Solution Approach 1:
The system segments the interconnection architecture into multiple dedicated crossbar switches, each handling specific communication paths. This segmentation eliminates the single shared bus bottleneck by providing multiple parallel communication channels, thereby improving system performance while maintaining manageable complexity through modular organization
Solution Approach 2:
Dedicated crossbar switches act as intermediary devices between CPUs and memory resources, providing direct switching paths that eliminate bus contention. These intermediaries enable simultaneous communications without conflict, resolving the trade-off between simplicity and performance
2Productivity
If multiple crossbar circuits are used to increase connectivity for larger numbers of CPUs and memory resources, then bandwidth is improved, but latency increases due to multiple crossbar hops
Solution Approach 1:
The system performs preliminary routing decisions at each crossbar switch to optimize path selection. By pre-determining the most efficient routes through the multiple crossbar circuits, the architecture minimizes the number of hops required for data transfer, thereby reducing latency while maintaining high bandwidth capacity
Solution Approach 2:
The architecture introduces dimensional organization by grouping crossbar switches into hierarchical levels or domains. This dimensional structure allows communications to be routed through optimized paths that minimize hops, effectively adding a routing dimension that reduces latency while preserving bandwidth
3Productivity
If more crossbar switches are added to provide sufficient bandwidth for CPU-memory interconnect, then bandwidth is improved, but device complexity and cost increase
Solution Approach 1:
The interconnect architecture is segmented into multiple smaller crossbar switches rather than using a single large crossbar. This segmentation provides sufficient total bandwidth through parallel paths while keeping each individual crossbar circuit manageable in complexity, avoiding the need for an excessively large complex switch
Solution Approach 2:
Each crossbar switch in the architecture is designed to be multi-functional, handling various types of communications (CPU-to-memory, CPU-to-CPU, memory-to-memory) through standardized interfaces. This universality allows the system to achieve high bandwidth with fewer specialized components, reducing overall complexity
Data Source
AI summary
In an embodiment, a multi-processor computer system includes multiple cells, where a cell may include one or more processors and memory resources. The system may further include a global crossbar network and multiple cell-to-global-crossbar connectors, to connect the multiple cells with the global crossbar network. In an embodiment, the system further includes at least one cell-to-cell connector, to directly connect at least one pair of the multiple cells. In another embodiment, the system further includes one or more local crossbar networks, multiple cell-to-local-crossbar connectors, and local input/output backplanes connected to the one or more local crossbar networks.


