Cross-Node Accelerator Resource Scheduling for Storage Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing storage systems face inefficiencies due to insufficient accelerator resources, leading to processing delays, especially in active-inactive HA systems and scalable multi-node systems, where accelerator resources are underutilized or overloaded, and current job scheduling solutions fail to address this issue from a global system perspective.
Innovation Solution
A method is introduced to determine local processing delays and select remote accelerator resources based on combined processing delays and round-trip times across nodes, allowing for the efficient distribution of jobs across nodes to balance processing pressure and improve resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If accelerator resources are allocated locally to each node, then processing speed is improved, but resource utilization efficiency deteriorates due to underutilization or overload
Solution Approach 1:
The patent introduces a new dimension of resource allocation by allowing accelerator resources to be scheduled across multiple nodes rather than being confined to local nodes. The scheduling decision considers both local processing speed and global resource utilization by evaluating remote nodes' accelerator resources, effectively moving from a single-node dimension to a multi-node distributed dimension.
Solution Approach 2:
The patent changes the scheduling parameters from purely local metrics to a combination of local processing delay, remote processing delay, and round-trip time. This parameter transformation enables the system to make informed decisions about whether to allocate resources locally or remotely, balancing speed and utilization efficiency.
2Ease of operation
If jobs are scheduled based on local accelerator resources only, then local processing is optimized, but system-wide resource allocation becomes inefficient
Solution Approach 1:
The scheduling system is enhanced to perform multiple functions: it evaluates local accelerator resources for immediate processing, assesses remote accelerator resources for potential offloading, and makes unified scheduling decisions that optimize both local operations and global resource utilization. This multi-functional approach resolves the contradiction between local optimization and system-wide efficiency.
Solution Approach 2:
The system incorporates feedback mechanisms by monitoring processing delays and round-trip times in real-time. This feedback allows the scheduler to dynamically adjust resource allocation decisions, learning from actual performance data to improve both local processing optimization and system-wide allocation efficiency over time.
3Productivity
If remote accelerator resources are utilized, then resource utilization efficiency is improved, but network transmission delay increases
Solution Approach 1:
The patent transforms the decision-making parameters to include quantitative metrics of processing delay and round-trip time. By changing from qualitative scheduling to quantitative parameter-based scheduling, the system can objectively determine when the trade-off between resource utilization efficiency and network transmission delay is favorable, making informed decisions that optimize overall system performance.
Solution Approach 2:
The system applies partial offloading by selectively utilizing remote accelerator resources only when necessary and beneficial, rather than consistently distributing all jobs remotely. This partial action approach allows the system to maintain local processing for most jobs while leveraging remote resources for specific workloads where resource utilization efficiency outweighs the network transmission penalty.
Data Source
AI summary
Embodiments of the present disclosure provide a resource utilization method, an electronic device, and a computer program product. A resource utilization method comprises: at a first node of a storage system, determining whether a local processing delay of a first accelerator resource of the first node exceeds a first threshold delay or not; if it is determined that the local processing delay exceeds the first threshold delay, determining at least one remote processing delay respectively corresponding to at least one second node of the storage system, wherein each remote processing delay comprises a processing delay of a second accelerator resource of a corresponding second node and a round-trip time between the first node and the corresponding second node; and at least based on the at least one remote processing delay, selecting a second accelerator resource, from the second accelerator resources of the at least one second node, to execute a target job of the first node. In this way, the calling of the accelerator resources across nodes may be implemented, thereby not only improving the processing efficiency of the jobs but also increasing the overall utilization rate of system resources.


