High Availability Cluster Probability of Breach Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing high availability clusters face challenges when additional resources are needed, particularly when poisoning problems occur due to software errors, as they struggle to detect and respond to failures effectively, leading to potential service level agreement breaches.

Innovation Solution

A method and apparatus that calculate a probability of breach based on event messages from the cluster, updating a data model and recommending reprovisioning to maintain high availability by filtering and translating event messages into probability data using a service class factor and number of active/failed servers, thereby avoiding poisoning issues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing H/A clusters use simple failover mechanisms with limited redundancy, then device complexity is reduced and ease of operation is improved, but reliability deteriorates when poisoning problems occur or when all standby resources fail

Engineering Contradiction:
Improvehigh availabilityVSAvoidcluster management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where the cluster monitors its own health status, calculates probability of breach values based on event messages, and automatically communicates with the provisioning manager to request additional resources before service level agreements are breached. This closed-loop feedback system enables proactive resource management without requiring complex manual intervention.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary actions by calculating probability of breach values and requesting reprovisioning resources before actual failures occur. The cluster proactively identifies potential poisoning problems and resource exhaustion scenarios, and triggers resource allocation in advance to prevent service disruptions, rather than reacting after failures happen.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If the cluster manually notifies administrators to fix poisoning problems, then automation extent is reduced, but ease of repair is improved, but loss of time increases due to manual detection and response

Engineering Contradiction:
Improvedetection and response timeVSAvoidautomatic problem detection
Core Design Contradiction:
Loss of timeVSExtent of automation

Solution Approach 1:

The cluster performs self-service by automatically monitoring its own health, detecting poisoning problems through event message analysis, calculating probability of breach values, and initiating reprovisioning requests without administrator intervention. The system serves itself by identifying and responding to its own problems, eliminating the need for manual detection and notification.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements automated feedback loops where event messages from the cluster are continuously analyzed, probability of breach values are calculated and monitored, and reprovisioning actions are automatically triggered when thresholds are exceeded. This automated feedback mechanism eliminates manual detection delays and enables rapid response to emerging problems.

Inventive Principle:
Principle #23Feedback

3Reliability

If the cluster adds more redundant resources to handle failures, then reliability is improved, but loss of substance increases due to unused standby resources, and device complexity increases

Engineering Contradiction:
Improvefault toleranceVSAvoidunused resource capacity
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The system transitions from static resource allocation to dynamic resource management. Instead of maintaining fixed numbers of standby resources, the cluster dynamically calculates probability of breach values based on actual event messages and current system state, and requests reprovisioning resources only when and where needed. This dynamic approach optimizes resource utilization while maintaining reliability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter for resource allocation from fixed redundancy counts to variable probability of breach values. The system calculates probabilistic metrics based on event messages and uses these parameters to dynamically determine reprovisioning needs, allowing resource allocation to adapt to changing system conditions and actual failure rates rather than relying on conservative static estimates.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7464302B2Method and apparatus for expressing high availability cluster demand based on probability of breach
Publication Date: 2008.12.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US7464302B2 patent drawing
  • US7464302B2 patent drawing
  • US7464302B2 patent drawing

AI summary

A method, apparatus, and computer instructions are provided for expressing high availability (H/A) cluster demand based on probability of breach. When a failover occurs in the H/A cluster, event messages are sent to a provisioning manager server. The mechanism of embodiments of the present invention filters the event messages and translates the events into probability of breach data. The mechanism then updates the data model of the provision manager server and makes a recommendation to the provisioning manager server as to whether reprovisioning of new node should be performed. The provisioning manager server makes the decision and either reprovisions new nodes to the H/A cluster or notifies the administrator of detected poisoning problem.