Workload Manager Resource Reallocation for System Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In high-availability computer systems, workload migration to another partition after a failure often results in inadequate resources, leading to over-provisioning and increased costs due to the need for spare resources to maintain system availability.
Innovation Solution
A global workload manager method that characterizes virtual machines as clustered or unclustered, reallocates resources from unclustered virtual machines to ensure sufficient resources are available for migrated clustered virtual machines, allowing for efficient resource reallocation and minimizing the need for spare resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If workload is migrated to another partition after failure, then system availability is maintained, but resources become inadequate and over-provisioning is required
Solution Approach 1:
The patent segments workloads into clustered and non-clustered categories, allowing differential resource allocation strategies. Clustered workloads receive priority resource allocation from the shared pool during failure scenarios, while non-clustered workloads can have their resources reallocated or terminated, eliminating the need for over-provisioning spare resources across all workloads.
Solution Approach 2:
The patent implements dynamic resource allocation where the shared resource pool can be flexibly allocated between clustered and non-clustered workloads based on failure conditions. The system dynamically adjusts resource distribution in real-time, transitioning from static over-provisioning to adaptive resource management that maintains availability only when needed.
2Reliability
If each partition has sufficient resources to run extra workload, then failure recovery is enabled, but resource cost increases due to over-provisioning
Solution Approach 1:
The patent merges multiple partitions into a clustered group that shares a common resource pool. Instead of each partition maintaining independent spare resources, the cluster consolidates resources that are dynamically allocated to any member partition during failure scenarios. This combining effect eliminates redundant spare resources while maintaining recovery capability.
Solution Approach 2:
The shared resource pool serves multiple functions: it supports normal operation of clustered workloads, provides failover resources when partitions fail, and can be dynamically reallocated based on system conditions. This multi-functional resource pool replaces the need for dedicated spare resources at each partition.
3Productivity
If resources are reallocated from non-clustered to clustered virtual machines, then resource efficiency improves, but non-clustered workload availability may be affected
Solution Approach 1:
The patent applies different quality levels of service to different workload types. Clustered workloads receive high-priority resource allocation with guaranteed availability, while non-clustered workloads receive best-effort service where resources can be reclaimed during failures. This local differentiation optimizes resource efficiency without requiring uniform high availability across all workloads.
Data Source
AI summary
The present invention provides for check-pointing an non-clustered workload to make room for a clustered workload that was running on a computer system that has suffered a hardware failure.


