Master Workload Management Process Failure Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing workload management systems fail to maintain resource utilization efficiency when the global workload management process encounters a system failure, leading to reduced resource sharing between virtual partitions until administrative intervention occurs.

Innovation Solution

A cluster organization of computing partitions with a selected master workload management process that reallocates resources between nodes, using a 'heartbeat' signal to monitor the master's operation and electing a replacement upon failure, ensuring continuous resource allocation and reorganization without requiring a system reset.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a master workload management process is used to reallocate resources between virtual partitions, then resource utilization efficiency is improved, but system reliability deteriorates when the master process fails

Engineering Contradiction:
Improveresource utilization efficiencyVSAvoidsystem reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements preliminary action by having non-master workload management processes monitor the master process through heartbeat signals before failure occurs. When the master process becomes inoperable, the monitoring processes are already positioned and authorized to elect a replacement master, eliminating the need for manual intervention and ensuring continuous resource allocation between virtual partitions.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If manual administrative intervention is required to reset the system after master process failure, then system complexity is reduced, but loss of time increases

Engineering Contradiction:
Improvesystem complexityVSAvoiddowntime
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the workload management system to automatically detect master process failure through heartbeat monitoring and autonomously elect a replacement master process. This self-healing mechanism eliminates the need for manual administrative intervention to reset the system, thereby reducing downtime while maintaining manageable system complexity through automated protocols.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS7979862B2System and method for replacing an inoperable master workload management process
Publication Date: 2011.07.12 HEWLETT PACKARD ENTERPRISE DEV LP
  • US7979862B2 patent drawing
  • US7979862B2 patent drawing
  • US7979862B2 patent drawing

AI summary

In one embodiment, a method comprises executing respective workload management processes within a plurality of computing compartments to allocate at least processor resources to applications executed within the plurality of computing compartments, selecting a master workload management process to reallocate processor resources between the plurality of computing compartments in response to requests from the workload management processes to receive additional resources, monitoring operations of the master workload management process by the other workload management processes, detecting, by the other workload management processes, when the master workload management process becomes inoperable, and selecting a replacement master workload management process by the other workload management processes in response to the detecting.