Multi-Node Network Plane Switching for Bandwidth-Power Balance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern computing systems face performance degradation due to network contention among multiple applications competing for resources, leading to increased latencies and reduced bandwidth, which can be exacerbated by power consumption issues in network components.
Innovation Solution
A multi-node computing system with multiple network planes is implemented, allowing dedicated network planes to be deactivated when utilization thresholds are met, reducing power consumption without affecting connectivity, and utilizing a fabric manager to manage traffic routing across active planes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple network planes are kept active to handle traffic from multiple applications, then network bandwidth and connectivity are improved, but power consumption increases
Solution Approach 1:
The system dynamically adjusts the number of active network planes based on real-time traffic utilization metrics. When utilization falls below a threshold, network planes are deactivated to save power; when utilization exceeds the threshold, planes are reactivated to maintain performance. This dynamic adaptation resolves the contradiction between maintaining high bandwidth availability and reducing power consumption.
Solution Approach 2:
The system changes the operational state parameter of network planes from always-active to conditionally-active based on utilization thresholds. By monitoring traffic patterns and adjusting the active/inactive state of network planes accordingly, the system optimizes the balance between network capacity and energy consumption.
2Use of energy by moving object
If network utilization is monitored and planes are deactivated to save power, then power consumption is reduced, but network latency may increase due to traffic rerouting
Solution Approach 1:
The system performs preliminary actions by gradually deactivating network planes based on utilization thresholds before complete overload occurs. The fabric manager proactively reroutes traffic through remaining active planes and monitors their utilization, preventing sudden latency spikes by maintaining headroom in active planes before deactivation.
Solution Approach 2:
The system implements continuous feedback loops where the fabric manager monitors utilization metrics of active network planes and dynamically adjusts traffic routing decisions. When utilization approaches thresholds, the system receives feedback and redistributes traffic to prevent overload, thereby maintaining low latency while still achieving power savings through selective plane deactivation.
3Use of energy by moving object
If network planes are deactivated based on utilization thresholds, then power consumption is optimized, but system complexity increases due to management overhead
Solution Approach 1:
The fabric manager implements self-service mechanisms by automatically monitoring utilization metrics, making deactivation/reactivation decisions, and rerouting traffic without manual intervention. The system serves itself by autonomously optimizing network plane utilization based on predefined thresholds, reducing the need for complex external management while achieving power optimization.
4Productivity
If traffic is continuously monitored to determine plane deactivation, then network performance is optimized, but processing overhead increases
Solution Approach 1:
The system applies partial monitoring by focusing utilization measurement only on active network planes rather than all planes in the system. This selective monitoring approach maintains optimal performance tracking for planes that are currently in use while reducing the processing overhead associated with monitoring inactive planes, thereby balancing performance optimization with energy efficiency.
Data Source
AI summary
A multi-node computing system. In some embodiments, the system includes: a first compute board, a second compute board, a plurality of compute elements, a plurality of memories, a first network plane connecting the first compute board and the second compute board, and a second network plane connecting the first compute board and the second compute board. The plurality of memories may store instructions that, when executed by the plurality of compute elements, cause the plurality of compute elements to: determine that a criterion for deactivating the first network plane is met, and deactivate the first network plane.


