Dynamic Hardware Accelerator Reassignment for Workload Prioritization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional workload management systems fail to effectively and efficiently utilize available hardware accelerator resources across multiple systems, leading to suboptimal performance and resource underutilization.
Innovation Solution
A dynamic resource management method and system that initially assigns hardware accelerators to information processing systems and monitors jobs to reassigned resources from one system to another based on performance goals and priorities, ensuring that critical jobs receive the necessary resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If hardware accelerators are assigned to specific host systems, then each system has dedicated computing resources, but resource utilization across multiple systems is suboptimal
Solution Approach 1:
The patent implements dynamic reassignment of hardware accelerators from idle or low-utilization host systems to systems with high-priority jobs. The workload management system continuously monitors accelerator utilization and dynamically reallocates resources based on current workload demands, transforming static assigned resources into dynamic shared resources that adapt to changing conditions.
Solution Approach 2:
The patent creates a universal accelerator pool that can serve multiple host systems based on demand. Instead of dedicating accelerators to specific hosts, the system allows any accelerator in the pool to be assigned to any host that needs it, making the resources multi-functional and adaptable to different workload requirements across the system.
2Device complexity
If accelerators are statically assigned to systems, then resource allocation is simple, but performance goals cannot be met when workloads change
Solution Approach 1:
The system transitions from static accelerator assignment to dynamic reassignment based on real-time workload monitoring. When performance goals are not met, the workload management system automatically identifies and reallocates accelerators from systems with excess capacity to systems needing additional resources, enabling the system to adapt to changing performance requirements.
Solution Approach 2:
The patent implements a feedback mechanism where the workload management system monitors whether jobs are meeting their performance goals and uses this information to trigger accelerator reassignment. The system continuously checks performance metrics and adjusts resource allocation accordingly, creating a closed-loop control system that responds to performance feedback.
3Quantity of substance
If accelerator resources are distributed across multiple systems, then each system has access to computing power, but available resources are not effectively utilized when some systems have idle accelerators
Solution Approach 1:
The patent merges accelerator resources from multiple host systems into a shared pool that can be dynamically allocated. By combining previously isolated accelerator resources into a unified pool, the system eliminates idle capacity at individual systems and ensures that accelerators are always assigned to productive workloads somewhere in the system.
Solution Approach 2:
The system recovers idle accelerator resources from systems that do not currently need them and reallocates these resources to systems with active workloads. This continuous recovery and reallocation process ensures that accelerator resources are constantly productive and minimizes idle time across the system.
Data Source
AI summary
A method, information processing system, and computer readable storage medium are provided for dynamically managing accelerator resources. A first set of hardware accelerator resources is initially assigned to a first information processing system, and a second set of hardware accelerator resources is initially assigned to a second information processing system. Jobs running on the first and second information processing systems are monitored. When one of the jobs fails to satisfy a goal, at least one hardware accelerator resource in the second set of hardware accelerator resources from the second information processing system are dynamically reassigned to the first information processing system.


