Accelerator Coherency Logic for Multi-Sled Workload Composition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In systems with multiple processors, resource utilization is inefficient due to mismatched allocation of workloads to available processors, leading to underutilization and wastage, especially in cloud data centers where customers pay for predefined quality of service metrics.
Innovation Solution
Implementing a data center architecture with robotically accessible sleds featuring disaggregated resources, high-bandwidth, low-latency optical networking, and dynamic resource pooling, enabling efficient allocation and utilization of processors across multiple compute sleds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If workloads are executed on a fixed number of processors within a single compute device, then device complexity is reduced and ease of operation is improved, but resource utilization deteriorates leading to processor underutilization or overutilization
Solution Approach 1:
The system segments the compute device into multiple independent compute sleds, each containing processors that can be independently allocated to workloads. This segmentation enables dynamic composition of compute resources, allowing the system to optimize resource utilization while maintaining operational simplicity through modular architecture.
Solution Approach 2:
The system implements dynamic processor allocation where the number of processors assigned to a workload is not fixed but can be adjusted based on actual workload demands. The orchestrator server dynamically composes managed nodes by selecting processors from available compute sleds, enabling the system to adapt resource allocation in real-time to optimize both ease of operation and resource utilization efficiency.
2Productivity
If more processors are allocated to a workload, then productivity and quality of service are improved, but resource wastage increases when the workload underutilizes the allocated processors
Solution Approach 1:
The system employs dynamic resource allocation where processor allocation to workloads is adjusted based on actual performance needs and workload characteristics. The orchestrator server monitors workload performance and re-composes managed nodes to optimize the balance between productivity and resource utilization, preventing both underutilization and over-provisioning of processors.
Solution Approach 2:
The system implements feedback mechanisms where the orchestrator server monitors workload performance metrics and resource utilization patterns. Based on this feedback, the system dynamically adjusts processor allocation by re-composing managed nodes, ensuring that productivity is maintained while minimizing resource wastage through data-driven optimization decisions.
3Adaptability or versatility
If processors are physically distributed across multiple compute sleds, then resource allocation flexibility is improved, but device complexity and inter-processor communication complexity increase
Solution Approach 1:
The orchestrator server acts as an intermediary that manages the complexity of distributed processor coordination. It handles processor selection, allocation, and coherence maintenance across multiple compute sleds, shielding workload execution from the underlying distributed system complexity while maintaining high resource allocation flexibility.
Solution Approach 2:
The system implements virtualization where physical processors across multiple compute sleds are copied or replicated as virtual processor resources in the orchestrator's resource pool. This abstraction layer allows flexible resource allocation while hiding the physical distribution complexity, enabling processors to be allocated and managed as if they were local resources.
4Speed
If processors are located on the same compute device, then inter-processor communication speed is improved, but adaptability to varying workload demands deteriorates
Solution Approach 1:
The system transitions from a single-dimension local communication model to a multi-dimensional architecture where processors can communicate both locally within compute sleds and remotely across the network. The orchestrator intelligently selects processors based on workload requirements, optimizing for communication speed when processors are co-located while enabling flexible scaling by accessing processors across multiple compute sleds when needed.
Data Source
AI summary
Technologies for composing a managed node with multiple processors on multiple compute sleds to cooperatively execute a workload include a memory, one or more processors connected to the memory, and an accelerator. The accelerator further includes a coherence logic unit that is configured to receive a node configuration request to execute a workload. The node configuration request identifies the compute sled and a second compute sled to be included in a managed node. The coherence logic unit is further configured to modify a portion of local working data associated with the workload on the compute sled in the memory with the one or more processors of the compute sled, determine coherence data indicative of the modification made by the one or more processors of the compute sled to the local working data in the memory, and send the coherence data to the second compute sled of the managed node.


