Double-Wing Multiprocessor Architecture Bandwidth Balance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing close-coupling shared storage architectures face challenges in maintaining a balance between processor bandwidth and network bandwidth, leading to increased average delay and limited scalability when expanding the number of processors, as current methods either mismatch bandwidths or increase interconnection hops.
Innovation Solution
A double-wing expandable multiprocessor architecture is implemented, where each processor module is formed by coupling and cross-jointing processors, with each processor directly connected to a node controller through one link, and node controllers connected to processors and interconnect networks in a balanced manner, ensuring non-blocking communication and network transmission, and utilizing multiple cross switch route chips to maintain relative balance between processor and network bandwidths.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the number of processors is doubled using traditional connection methods, then processor scale is improved, but bandwidth mismatch between processor and network increases
Solution Approach 1:
The system divides processors into groups of 2, with each group sharing a common node controller. This segmentation allows the node controller to aggregate bandwidth from multiple processors while maintaining balanced connection ratios to the interconnection network, resolving the bandwidth mismatch problem when scaling processor count.
Solution Approach 2:
Multiple processor connections are merged through a single node controller before reaching the interconnection network. By combining multiple processor links at the node controller level, the system achieves bandwidth aggregation that maintains proportional balance between processor capacity and network capacity even as processor count increases.
2Productivity
If more links are added to maintain bandwidth balance, then device complexity increases
Solution Approach 1:
The node controller serves multiple functions: it manages connections to multiple processors, aggregates their bandwidth, and interfaces with the interconnection network. This multi-functionality allows a single component to handle bandwidth balancing for multiple processors without requiring separate control logic for each processor, reducing overall system complexity.
3Speed
If node controller is placed close to processor, then processor access speed is improved, but network bandwidth utilization decreases
Solution Approach 1:
Multiple processor connections are merged at the node controller before reaching the interconnection network. This merging allows the system to maintain fast local processor access while simultaneously optimizing network bandwidth utilization through aggregated traffic flow, resolving the contradiction between local speed and network efficiency.
Data Source
AI summary
A close-coupling shared storage architecture of double-wing expandable multiprocessor is provided in the close-coupling shared storage architecture with p processors scale, the close-coupling shared storage architecture of double-wing expandable multiprocessor comprises: j processor modules PMs; wherein, each processor module is formed by coupling and cross-jointing i processors Cs, and each processor is directly connected with a node controller NC through only one link; each processor module PM comprises 2 pairing node controllers NCs, and each node controller NC is connected with the processors through m links and is connected with an interconnect network through n links; the interconnect network comprises two groups, and each group comprises k cross switch route chips NRs, each of which has q ports. By adopting the connection method above, the close-coupling shared storage architecture of double-wing expandable multiprocessor is formed. On the premise that the processor scale is kept expandable, the balance between the processor bandwidth and the network bandwidth is achieved, and the lower average delay of the interconnect network is kept simultaneously.


