Reconfigurable Compute Pods with Optical Circuit Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Rigid interconnect networks in supercomputers limit scalability, availability, and performance, especially when processing nodes fail, and do not optimize for varying computational workloads in size and complexity.
Innovation Solution
Utilizing optical networks to dynamically configure clusters of compute nodes, allowing for flexible arrangement and substitution of faulty nodes, and enabling various shapes of workload clusters with appropriate types and numbers of compute nodes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If rigid interconnect networks are used to connect processing nodes, then structural stability and simplicity are improved, but adaptability and scalability deteriorate when workloads vary in size and complexity
Solution Approach 1:
The patent implements dynamic reconfigurability by allowing the interconnect network topology to change based on workload requirements. Processing nodes can be dynamically added, removed, or repositioned within the network, transforming the rigid static structure into a flexible dynamic system that adapts to varying computational demands while maintaining operational stability
Solution Approach 2:
The system segments the supercomputer into independent processing nodes that can be individually configured and repositioned. This segmentation allows the network to be reorganized into different topologies (mesh, torus, hypercube, etc.) based on specific workload requirements, enabling adaptability without compromising the overall structural integrity of the system
2Ease of manufacture
If rigid interconnect networks with specific arrangements are used, then ease of manufacture and initial setup are improved, but availability and performance deteriorate when processing nodes fail
Solution Approach 1:
The system employs dynamic topology reconfiguration that automatically adjusts the network layout when node failures occur. Remaining healthy nodes can be dynamically repositioned and reconnected to maintain optimal communication paths, ensuring continued system availability and performance without requiring manual intervention or fixed predetermined arrangements
Solution Approach 2:
The patent changes the topological parameters of the interconnect network dynamically based on system state. When nodes fail, the network parameters (connection patterns, routing paths, dimensional configurations) are modified to accommodate the reduced node count and maintain optimal performance, transforming the system from a static to a adaptive configuration
3Adaptability or versatility
If dynamic reconfiguration of compute nodes is enabled, then adaptability and performance are improved, but device complexity and routing configuration increase
Solution Approach 1:
The system implements feedback mechanisms where the interconnect network continuously monitors its own state and automatically adjusts routing configurations based on performance metrics and topology requirements. This self-regulating feedback loop manages the complexity of dynamic reconfiguration by using real-time information to optimize node arrangements and communication paths without requiring external control for every change
Solution Approach 2:
The patent enables the interconnect network to self-configure and self-optimize its topology and routing paths based on workload characteristics and system state. The network performs automatic topology discovery, path optimization, and resource allocation, reducing the need for manual configuration management and external control while handling the complexity of dynamic reconfiguration internally
4Speed
If optical networks are used instead of traditional interconnects, then speed and latency are improved, but device complexity and infrastructure requirements increase
Solution Approach 1:
The patent replaces traditional electrical/copper-based interconnect mechanisms with optical communication systems. This substitution uses photons instead of electrons for data transmission, providing significantly higher bandwidth and lower latency while reducing signal degradation and heat generation, despite the increased complexity of optical infrastructure
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Methods, systems, and apparatus, including an apparatus for generating clusters of building blocks of compute nodes using an optical network. In one aspect, a method includes receiving request data specifying requested compute nodes for a computing workload. The request data specifies a target n-dimensional arrangement of the compute nodes. A selection is made, from a superpod that includes a set of building blocks that each include an m-dimensional arrangement of compute nodes, a subset of the building blocks that, when combined, match the target arrangement specified by the request data. The set of building blocks are connected to an optical network that includes one or more optical circuit switches. A workload cluster of compute nodes that includes the subset of the building blocks is generated. The generating includes configuring, for each dimension of the workload cluster, respective routing data for the one or more optical circuit switches.