Cloud Resiliency Scoring for Hardware Failure Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In cloud-based service-oriented architectures, ensuring high availability and disaster recovery while maintaining resiliency in the face of hardware and software failures is challenging, particularly due to the aging of physical drivers and components, which can lead to increased failure probabilities and disruptions.

Innovation Solution

A method that involves receiving time-date data sets and machine logic-based rules to determine a resiliency value for server computers, which is then used to recommend changes to hardware, firmware, high availability policies, and disaster recovery plans, thereby enhancing the system's ability to handle failures and maintain service integrity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If high availability policies and disaster recovery plans are implemented, then service availability is improved, but system complexity increases

Engineering Contradiction:
Improveservice availabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system dynamically adjusts resiliency parameters (such as monitoring intensity, backup frequency, and failover thresholds) based on the calculated resiliency score. When hardware components age or performance degrades, the system modifies operational parameters to maintain service availability without requiring complete system redesign or permanent complex configurations.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If proactive hardware replacement and policy adjustments are made, then resiliency is improved, but resource consumption increases

Engineering Contradiction:
ImproveresiliencyVSAvoidresource consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system performs preliminary assessments by calculating resiliency scores based on hardware age, performance metrics, and operational data. By identifying components that are likely to fail soon, the system enables proactive replacement or adjustment before actual failures occur, preventing service disruptions and reducing the need for emergency resource allocation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors hardware performance and updates the resiliency score based on feedback from operational data. This closed-loop approach allows the system to adjust resource allocation dynamically - increasing resources when resiliency is low and reducing resources when resiliency is adequate - thereby optimizing the balance between reliability and resource consumption.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If monitoring and analysis of hardware age and performance are intensified, then failure detection is improved, but operational overhead increases

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidoperational overhead
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The resiliency score calculation system serves multiple functions simultaneously: it assesses hardware health, predicts failures, guides replacement decisions, and informs policy adjustments. By consolidating these functions into a single unified metric, the system reduces operational overhead compared to maintaining separate monitoring systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11310276B2Adjusting resiliency policies for cloud services based on a resiliency score
Publication Date: 2022.04.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11310276B2 patent drawing
  • US11310276B2 patent drawing
  • US11310276B2 patent drawing

AI summary

Evaluating a plurality of computers hosting a cloud platform for effectiveness at operating through operational failures with minimal or no degradation to operations by identifying vulnerabilities in hardware, firmware, software and operational policy/plan aspects of the plurality of computers and managing the identified vulnerabilities by modifying hardware, firmware, software, and operational policy/plan aspects of the plurality of computers and the hosted cloud platform to improve effectiveness at operating through operational failures with minimal or no degradation to operations.