Workload Manager Resource Reallocation for System Failure Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In high-availability computer systems, workload migration to another partition after a failure often results in inadequate resources, leading to over-provisioning and increased costs due to the need for spare resources to maintain system availability.

Innovation Solution

A global workload manager method that characterizes virtual machines as clustered or unclustered, reallocates resources from unclustered virtual machines to ensure sufficient resources are available for migrated clustered virtual machines, allowing for efficient resource reallocation and minimizing the need for spare resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If workload is migrated to another partition after failure, then system availability is maintained, but resources become inadequate and over-provisioning is required

Engineering Contradiction:
Improvesystem availabilityVSAvoidspare resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments workloads into clustered and non-clustered categories, allowing differential resource allocation strategies. Clustered workloads receive priority resource allocation from the shared pool during failure scenarios, while non-clustered workloads can have their resources reallocated or terminated, eliminating the need for over-provisioning spare resources across all workloads.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic resource allocation where the shared resource pool can be flexibly allocated between clustered and non-clustered workloads based on failure conditions. The system dynamically adjusts resource distribution in real-time, transitioning from static over-provisioning to adaptive resource management that maintains availability only when needed.

Inventive Principle:
Principle #15Dynamics

2Reliability

If each partition has sufficient resources to run extra workload, then failure recovery is enabled, but resource cost increases due to over-provisioning

Engineering Contradiction:
Improvefailure recovery capabilityVSAvoidspare resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent merges multiple partitions into a clustered group that shares a common resource pool. Instead of each partition maintaining independent spare resources, the cluster consolidates resources that are dynamically allocated to any member partition during failure scenarios. This combining effect eliminates redundant spare resources while maintaining recovery capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared resource pool serves multiple functions: it supports normal operation of clustered workloads, provides failover resources when partitions fail, and can be dynamically reallocated based on system conditions. This multi-functional resource pool replaces the need for dedicated spare resources at each partition.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If resources are reallocated from non-clustered to clustered virtual machines, then resource efficiency improves, but non-clustered workload availability may be affected

Engineering Contradiction:
Improveresource efficiencyVSAvoidnon-clustered workload availability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies different quality levels of service to different workload types. Clustered workloads receive high-priority resource allocation with guaranteed availability, while non-clustered workloads receive best-effort service where resources can be reclaimed during failures. This local differentiation optimizes resource efficiency without requiring uniform high availability across all workloads.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS7653833B1Terminating a non-clustered workload in response to a failure of a system with a clustered workload
Publication Date: 2010.01.26 HEWLETT PACKARD ENTERPRISE DEV LP
  • US7653833B1 patent drawing
  • US7653833B1 patent drawing
  • US7653833B1 patent drawing

AI summary

The present invention provides for check-pointing an non-clustered workload to make room for a clustered workload that was running on a computer system that has suffered a hardware failure.