Disaggregated Resource Management via Distributed Sled Intelligence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In data centers with disaggregated resources, the central compute device faces a significant burden due to the large amount of data it needs to process, leading to potential thermal issues and quality of service (QoS) problems when reducing data communication to alleviate this load.

Innovation Solution

The implementation of a system where resources such as compute, accelerator, and storage sleds are managed to form 'managed nodes' that can dynamically allocate and deallocate resources based on QoS targets, using advanced resource management and thermal management logic to optimize resource utilization and thermal conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the central compute device receives detailed operating condition data from each resource to maintain QoS targets, then QoS reliability is improved, but the data communication burden on the central compute device and network paths increases

Engineering Contradiction:
ImproveQoS target satisfactionVSAvoiddata communication volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent divides the centralized resource management function into distributed management units located at each resource component (compute devices, accelerator devices, storage devices). Each management unit independently monitors its local resource's operating conditions and makes local decisions, eliminating the need for all data to be centralized at one compute device while maintaining QoS through distributed intelligence

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces management units as intermediary components between the resources and the central orchestration system. These management units aggregate and process local operating condition data before transmitting to the central system, reducing the volume of raw data that needs to be communicated while ensuring QoS requirements are met

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If data communication to the central compute device is reduced to alleviate its burden, then device complexity and energy consumption are improved, but thermal conditions on resources may fall outside desired ranges

Engineering Contradiction:
Improvecentral compute device burdenVSAvoidresource thermal conditions
Core Design Contradiction:
Device complexityVSTemperature

Solution Approach 1:

The patent enables each resource's management unit to autonomously monitor its own thermal conditions and adjust its operation accordingly. Each management unit independently makes thermal management decisions based on local sensor data, eliminating the need for continuous centralized control while maintaining thermal conditions within desired ranges

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements local feedback loops where each management unit continuously monitors its resource's thermal conditions and adjusts operations in real-time. This distributed feedback mechanism ensures thermal conditions remain within desired ranges without requiring all data to be communicated to the central compute device

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11115497B2Technologies for providing advanced resource management in a disaggregated environment
Publication Date: 2021.09.07 INTEL CORP
  • US11115497B2 patent drawing
  • US11115497B2 patent drawing
  • US11115497B2 patent drawing

AI summary

Technologies for providing advanced resource management in a disaggregated environment include a compute device. The compute device includes circuitry to obtain a workload to be executed by a set of resources in a disaggregated system, query a sled in the disaggregated system to identify an estimated time to complete execution of a portion of the workload to be accelerated using a kernel, and assign, in response to a determination that the estimated time to complete execution of the portion of the workload satisfies a target quality of service associated with the workload, the portion of the workload to the sled for acceleration.