Accelerator Coherency Logic for Multi-Sled Workload Composition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In systems with multiple processors, resource utilization is inefficient due to mismatched allocation of workloads to available processors, leading to underutilization and wastage, especially in cloud data centers where customers pay for predefined quality of service metrics.

Innovation Solution

Implementing a data center architecture with robotically accessible sleds featuring disaggregated resources, high-bandwidth, low-latency optical networking, and dynamic resource pooling, enabling efficient allocation and utilization of processors across multiple compute sleds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If workloads are executed on a fixed number of processors within a single compute device, then device complexity is reduced and ease of operation is improved, but resource utilization deteriorates leading to processor underutilization or overutilization

Engineering Contradiction:
Improveworkload execution simplicityVSAvoidresource utilization efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system segments the compute device into multiple independent compute sleds, each containing processors that can be independently allocated to workloads. This segmentation enables dynamic composition of compute resources, allowing the system to optimize resource utilization while maintaining operational simplicity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements dynamic processor allocation where the number of processors assigned to a workload is not fixed but can be adjusted based on actual workload demands. The orchestrator server dynamically composes managed nodes by selecting processors from available compute sleds, enabling the system to adapt resource allocation in real-time to optimize both ease of operation and resource utilization efficiency.

Inventive Principle:
Principle #15Dynamics

2Productivity

If more processors are allocated to a workload, then productivity and quality of service are improved, but resource wastage increases when the workload underutilizes the allocated processors

Engineering Contradiction:
Improveworkload execution performanceVSAvoidprocessor resource wastage
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system employs dynamic resource allocation where processor allocation to workloads is adjusted based on actual performance needs and workload characteristics. The orchestrator server monitors workload performance and re-composes managed nodes to optimize the balance between productivity and resource utilization, preventing both underutilization and over-provisioning of processors.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms where the orchestrator server monitors workload performance metrics and resource utilization patterns. Based on this feedback, the system dynamically adjusts processor allocation by re-composing managed nodes, ensuring that productivity is maintained while minimizing resource wastage through data-driven optimization decisions.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If processors are physically distributed across multiple compute sleds, then resource allocation flexibility is improved, but device complexity and inter-processor communication complexity increase

Engineering Contradiction:
Improveresource allocation flexibilityVSAvoiddistributed system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The orchestrator server acts as an intermediary that manages the complexity of distributed processor coordination. It handles processor selection, allocation, and coherence maintenance across multiple compute sleds, shielding workload execution from the underlying distributed system complexity while maintaining high resource allocation flexibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements virtualization where physical processors across multiple compute sleds are copied or replicated as virtual processor resources in the orchestrator's resource pool. This abstraction layer allows flexible resource allocation while hiding the physical distribution complexity, enabling processors to be allocated and managed as if they were local resources.

Inventive Principle:
Principle #26Copying

4Speed

If processors are located on the same compute device, then inter-processor communication speed is improved, but adaptability to varying workload demands deteriorates

Engineering Contradiction:
Improveinter-processor communication speedVSAvoidworkload scaling flexibility
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system transitions from a single-dimension local communication model to a multi-dimensional architecture where processors can communicate both locally within compute sleds and remotely across the network. The orchestrator intelligently selects processors based on workload requirements, optimizing for communication speed when processors are co-located while enabling flexible scaling by accessing processors across multiple compute sleds when needed.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12619465B2Maintaining workload data coherency between local and remote accelerators
Publication Date: 2026.05.05 INTEL CORP
  • US12619465B2 patent drawing
  • US12619465B2 patent drawing
  • US12619465B2 patent drawing

AI summary

Technologies for composing a managed node with multiple processors on multiple compute sleds to cooperatively execute a workload include a memory, one or more processors connected to the memory, and an accelerator. The accelerator further includes a coherence logic unit that is configured to receive a node configuration request to execute a workload. The node configuration request identifies the compute sled and a second compute sled to be included in a managed node. The coherence logic unit is further configured to modify a portion of local working data associated with the workload on the compute sled in the memory with the one or more processors of the compute sled, determine coherence data indicative of the modification made by the one or more processors of the compute sled to the local working data in the memory, and send the coherence data to the second compute sled of the managed node.