Disaggregated Server Scheduling for Reliable Resource Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data centers face challenges in resource allocation efficiency and reliability due to hardware disaggregation, which can lead to increased failure patterns and network dependency, limiting the scale of disaggregation and affecting service continuity.
Innovation Solution
A server with a DDC architecture and a task scheduler module that utilizes a DDC hardware monitor to detect hardware information and allocate resources based on mixed-integer linear programming (MILP) methods, addressing both static and dynamic scenarios to maximize reliability and resource availability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If hardware disaggregation is implemented to improve resource efficiency, then resource utilization improves, but system reliability deteriorates due to increased failure patterns and network dependency
Solution Approach 1:
The patent segments the data center into independent processing modules that can be individually monitored and managed. Each module is treated as a separate unit with its own resource pool, allowing selective allocation and isolation of failures. This segmentation enables the system to maintain resource efficiency through fine-grained allocation while improving reliability by containing failures within specific modules rather than affecting the entire system.
Solution Approach 2:
The patent implements preliminary monitoring of hardware information including topology, loading status, and failure/repair states before making resource allocation decisions. The task scheduler module proactively analyzes this pre-collected hardware information to anticipate potential failures and make informed allocation decisions, rather than reacting to failures after they occur. This preliminary action allows the system to maintain both high resource efficiency and improved reliability through preventive resource management.
2Quantity of substance
If the scale of disaggregation is increased to improve resource pooling, then resource availability improves, but network dependency increases leading to more failure patterns
Solution Approach 1:
The patent applies local quality by implementing differentiated monitoring and allocation strategies for different regions of the disaggregated system. The hardware monitor tracks specific local conditions (topology, loading, failure states) of individual processing modules, and the task scheduler applies localized allocation decisions based on these specific conditions rather than uniform system-wide policies. This allows the system to scale resource pooling while managing failure patterns through localized responses to local conditions.
Solution Approach 2:
The patent introduces the task scheduler module as an intermediary between the hardware monitor and the processing modules. This intermediary analyzes hardware information including failure states and makes intelligent allocation decisions, acting as a buffer that prevents direct propagation of failure patterns across the network. The intermediary enables larger resource pooling by mediating between the increased network dependency and the need for reliable service delivery.
3Adaptability or versatility
If dynamic resource allocation is implemented to improve adaptability, then resource flexibility improves, but allocation complexity increases
Solution Approach 1:
The patent implements a feedback mechanism where the hardware monitor continuously collects hardware information (topology, loading, failure/repair states) and feeds this information back to the task scheduler module. The scheduler uses this feedback to dynamically adjust resource allocation decisions in real-time, achieving high adaptability and resource flexibility. The structured feedback loop manages allocation complexity by providing the scheduler with organized, relevant information rather than requiring complex ad-hoc analysis of raw system states.
Data Source
AI summary
A server and a resource scheduling method for use in a server. The server comprises a plurality of processing modules each having predetermined resources for processing tasks handled by the server, wherein the plurality of processing modules are interconnected by communication links forming a network of processing modules having a Disaggregated Data Centers (DCC) architecture; a DCC hardware monitor arranged to detect hardware information associated with the network of processing modules during an operation of the server, and a task scheduler module arranged to analysis a resource allocation request associated with each respective task and the hardware information, and to facilitate processing of the task by one or more of the processing modules selected based on the analysis.


