Global Workload Manager for Reliability-Based Resource Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current workload management systems face challenges in dynamically allocating computer resources to ensure high performance and reliability, particularly when dealing with increasing error rates and hardware failures, as they struggle to effectively migrate workloads without overburdening destination partitions and maintaining uninterrupted operation.
Innovation Solution
A computer system with partitions and a global workload manager that monitors resource utilization, power consumption, and error rates, using error-correcting channels and memories to track non-fatal errors, plans resource allocation based on management policies, and migrates workloads or reallocates hardware to maintain reliability and performance by prioritizing critical workloads and quarantining unreliable components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If workloads are migrated to maintain high performance and reliability, then system reliability is improved, but system complexity increases due to the need for monitoring, decision-making, and migration management
Solution Approach 1:
A global workload manager is introduced as an intermediary component that centralizes the monitoring, decision-making, and migration management functions. This manager collects reliability indications from multiple partitions, applies management policies to determine optimal workload placements, and executes migration decisions. By consolidating these complex functions into a dedicated intermediary system, the patent manages the increased complexity through specialization and centralization rather than distributing it across the entire system.
Solution Approach 2:
The system implements continuous feedback loops where the global workload manager monitors reliability indications (such as error rates) from partitions, evaluates current workload placements against management policies, and dynamically adjusts workload allocations. This feedback mechanism enables the system to respond automatically to changing conditions, maintaining reliability while managing complexity through automated closed-loop control rather than manual intervention.
2Reliability
If critical workloads are prioritized and migrated away from unreliable partitions, then reliability is improved, but productivity decreases due to migration overhead and resource reallocation
Solution Approach 1:
The system proactively monitors reliability indications and identifies partitions showing signs of degradation before critical failures occur. By detecting early warning signs (such as increasing error rates) and preemptively migrating critical workloads away from at-risk partitions, the system prevents failures rather than reacting to them. This preliminary action reduces the need for emergency migrations and maintains productivity by avoiding service interruptions.
Solution Approach 2:
The workload management system dynamically adjusts workload allocations based on real-time reliability conditions rather than following static assignments. The global workload manager continuously evaluates reliability indications and management policies to determine optimal placements, allowing the system to adapt flexibly to changing conditions. This dynamic approach enables the system to maintain productivity by optimizing workload placements for both reliability and performance rather than following rigid allocation rules.
3Reliability
If error rates are monitored and used to guide resource allocation, then reliability is improved, but measurement precision requirements increase for detecting and responding to errors
Solution Approach 1:
The system implements differentiated monitoring and response strategies based on workload criticality levels. Critical workloads receive enhanced monitoring with lower error rate thresholds and more frequent checks, while non-critical workloads use standard monitoring. Management policies allow the system to apply different measurement precision requirements to different workloads and partitions, allocating measurement resources locally where they are most needed rather than uniformly across the entire system.
Solution Approach 2:
The global workload manager dynamically adjusts error rate thresholds and monitoring parameters based on workload criticality and current system conditions. For critical workloads, the system uses lower thresholds that trigger migrations at lower error rates, while non-critical workloads tolerate higher error rates before triggering migrations. This parameter adaptation allows the system to maintain high reliability for critical functions while reducing measurement precision requirements for less important workloads, balancing overall system needs.
Data Source
AI summary
A computer system has plural partitions for running respective workloads. Reliability-indicating events are monitored and the resulting data is used by a workload manager in allocating computer resources to workloads.

