Master Workload Management Process Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing workload management systems fail to maintain resource utilization efficiency when the global workload management process encounters a system failure, leading to reduced resource sharing between virtual partitions until administrative intervention occurs.
Innovation Solution
A cluster organization of computing partitions with a selected master workload management process that reallocates resources between nodes, using a 'heartbeat' signal to monitor the master's operation and electing a replacement upon failure, ensuring continuous resource allocation and reorganization without requiring a system reset.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a master workload management process is used to reallocate resources between virtual partitions, then resource utilization efficiency is improved, but system reliability deteriorates when the master process fails
Solution Approach 1:
The patent implements preliminary action by having non-master workload management processes monitor the master process through heartbeat signals before failure occurs. When the master process becomes inoperable, the monitoring processes are already positioned and authorized to elect a replacement master, eliminating the need for manual intervention and ensuring continuous resource allocation between virtual partitions.
2Device complexity
If manual administrative intervention is required to reset the system after master process failure, then system complexity is reduced, but loss of time increases
Solution Approach 1:
The patent implements self-service by enabling the workload management system to automatically detect master process failure through heartbeat monitoring and autonomously elect a replacement master process. This self-healing mechanism eliminates the need for manual administrative intervention to reset the system, thereby reducing downtime while maintaining manageable system complexity through automated protocols.
Data Source
AI summary
In one embodiment, a method comprises executing respective workload management processes within a plurality of computing compartments to allocate at least processor resources to applications executed within the plurality of computing compartments, selecting a master workload management process to reallocate processor resources between the plurality of computing compartments in response to requests from the workload management processes to receive additional resources, monitoring operations of the master workload management process by the other workload management processes, detecting, by the other workload management processes, when the master workload management process becomes inoperable, and selecting a replacement master workload management process by the other workload management processes in response to the detecting.


