Dynamic Compute Node Composition for Data Center Workload Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data centers face operational complexity and inefficiency due to diverse server types and networks, leading to marooned resources and high operational costs, as they struggle to dynamically manage and reconfigure servers to meet varying workload demands.
Innovation Solution
The technique involves composing and recomposing compute nodes from pools of general purpose diskless and NIC-less physical servers, allowing for dynamic reconfiguration and redeployment while preserving state, enabling efficient utilization of resources across workload silos and improving scale-out responsiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data centers use diverse server types (with local storage, without local storage, with GPUs, with FPGAs, etc.) to meet different workload requirements, then workload-specific performance is improved, but operational complexity and management cost increase significantly
Solution Approach 1:
The system segments servers into modular components: processing circuitry (compute nodes) and peripheral devices (storage, networking, accelerators). This allows independent management and reconfiguration of components, reducing operational complexity while maintaining diverse server configurations for different workloads.
Solution Approach 2:
The composition engine provides a universal management platform that can provision and manage diverse server types through a single interface. The system universally applies the same composition and reconfiguration processes across different server configurations, simplifying operations despite hardware diversity.
2Reliability
If data centers provision dedicated servers for different workload types (virtualized clusters for general applications, bare metal for containers/AI/ML), then workload performance is optimized, but resource utilization efficiency decreases due to inflexible infrastructure silos
Solution Approach 1:
The system dynamically reconfigures servers by detaching and reattaching peripheral devices to different processing circuitry based on real-time workload demands. This dynamic composition allows resources to be reassigned from underutilized servers to workload-intensive tasks, improving both performance and utilization efficiency.
Solution Approach 2:
The system changes the configuration parameters of servers by modifying which peripheral devices are attached to which processing circuitry. This allows the same physical hardware to be reconfigured for different workload types (e.g., switching from AI training to database operations), enhancing resource utilization while maintaining workload optimization.
3Stability of the object's composition
If data centers maintain hardware configurations fixed after procurement and deployment, then hardware stability is ensured, but adaptability to changing workload demands decreases, leading to marooned resources
Solution Approach 1:
The system extracts peripheral devices from their current processing circuitry attachments, allowing them to be removed and reattached to different compute nodes. This extraction capability enables hardware reconfiguration without physical hardware changes, maintaining stability while improving adaptability to workload demands.
Solution Approach 2:
The composition engine acts as an intermediary layer between physical hardware and workload requirements. It manages the detachment and reattachment of peripheral devices, mediating between the stable physical hardware layer and the dynamic workload demands, enabling reconfiguration without affecting hardware integrity.
4Reliability
If data centers procure and deploy additional spare physical servers to support workload peaks, then service availability is improved, but capital expenditure and operational costs increase
Solution Approach 1:
The system recovers and reallocates peripheral devices from idle or underutilized servers to support workload peaks. Instead of discarding (leaving idle) spare capacity, the system actively recovers resources from non-critical workloads and redistributes them where needed, maintaining service availability with fewer total servers.
Data Source
AI summary
This disclosure describes techniques that include provisioning compute nodes within a data center out of available pools of hardware. In one example, this disclosure describes a method that includes monitoring, by a computing system, a first workload executing on a first compute node, wherein the first compute node includes processing circuitry and first node secondary storage; monitoring, by the computing system, a second workload executing on a second cluster of compute nodes; expanding, by the computing system, the second cluster of compute nodes to include a second compute node that includes second node secondary storage; redeploying the processing circuitry included within the first compute node to the second compute node; and enabling, by the computing system, the second workload to continue executing on the second cluster of compute nodes including the second compute node.


