DWDM HPC Network Using MEMS Optical Circuit Switches
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity and hardware intensity of training AI systems, particularly with large language models, result in significant costs due to the large number of optical components, switches, and cables required for configuring graphics processing units (GPUs) in these networks.
Innovation Solution
An ultra-scalable high-performance computing (HPC) network based on dense wavelength-division multiplexing (DWDM) is introduced, which includes an interconnection of GPUs, multiplexer/demultiplexer devices, amplifiers, wavelength selective switches, and optical circuit switches. This configuration allows for the formation of pathways to connect GPUs in various network topologies for AI system training and inferencing, using fewer optical components and switches while providing greater bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional optical networking components are used to configure GPUs for AI training, then the system can be built, but the number of optical components, switches, and cables becomes extremely large, resulting in high equipment and labor costs
Solution Approach 1:
The patent combines multiple optical channels onto a single fiber using wavelength-division multiplexing (WDM), where each wavelength carries independent data streams. This merging of multiple communication channels into one physical medium dramatically reduces the number of optical components, switches, and cables needed while maintaining or increasing total bandwidth capacity
Solution Approach 2:
The photonic circuit switch fabric serves multiple functions simultaneously: it routes optical signals between GPU devices, performs wavelength multiplexing/demultiplexing, and enables dynamic reconfiguration of network topologies. This multi-functionality eliminates the need for separate dedicated components for each function, reducing overall system complexity
2Productivity
If more optical components and switches are added to increase bandwidth, then the bandwidth increases, but the equipment costs and labor costs increase significantly
Solution Approach 1:
By merging multiple wavelength channels onto single fiber links and using compact integrated photonic circuits, the system achieves high bandwidth without proportionally increasing the number of discrete optical components. This consolidation directly reduces equipment costs and assembly labor
Solution Approach 2:
The patent replaces traditional mechanical optical switching systems with photonic circuit switching that uses optical signals to control optical routing. This substitution eliminates mechanical moving parts, reducing component complexity, equipment costs, and labor requirements for assembly and maintenance
3Productivity
If a large number of GPUs are interconnected using traditional optical networking, then the computing power increases, but the number of cables and switches becomes unmanageably large
Solution Approach 1:
The system merges multiple data streams at different wavelengths onto single fiber cables connecting GPU devices. This allows a large number of GPUs to be interconnected with far fewer physical cables than traditional single-wavelength systems, making the quantity of cabling manageable even as computing power scales
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The HPC network achieves over an order of magnitude greater bandwidth compared to conventional systems with significantly fewer optical components, switches, and cables, while maintaining high scalability and reducing equipment and labor costs.
Implementation Method 1
Each GPU device includes one or more optoelectronic devices that are configured to convert and output data signals from the one or more GPUs into optical output signals and configured to convert optical input signals into data signals to be input into the one or more GPUs
Implementation Method 2
Each OCS includes a second number of I/O ports and a third number of MEMS mirrors, each MEMS mirror being configured to selectively route optical signals between one I/O port and another I/O port among the second number of I/O ports
Implementation Method 3
a plurality of WSSs, each WSS including a first plurality of WSS mux/demux devices and a second plurality of WSS mux/demux devices... Each WSS is configured to selectively route optical signals between one of the first plurality of WSS mux/demux devices and one of the second plurality of WSS mux/demux devices
Data Source
AI summary
Systems and methods are provided for implementing an ultra-scalable high-performance computing (“HPC”) network using dense wavelength-division multiplexing (“DWDM”). The HPC system includes an interconnection of GPU devices, multiplexer/demultiplexer (“mux/demux”) devices, amplifiers, wavelength selective switches (“WSSs”), and optical circuit switches (“OCSs”). Each OCS includes a plurality of micro-electromechanical systems (“MEMS”) mirrors and a plurality of input/output (“I/O”) ports each communicatively coupled to one WSS mux/demux device one WSS. Each WSS mux/demux device is either communicatively coupled to one of the I/O ports of an OCS or one of a plurality of GPU mux/demux devices via an amplifier. Each GPU mux/demux device is communicatively coupled to a number of GPU devices, each including another number GPUs and one or more optoelectronic devices. Selectively controlling the MEMS mirrors of the OCSs and the WSS mux/demux devices of the WSSs allows connecting the GPUs in a network topology for computing a series of computations.


