Machine Learning Module for Predictive Failure Risk Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computational systems face challenges in proactively identifying and mitigating potential failures, as existing methods lack effective predictive capabilities to anticipate and prevent malfunctions in storage systems before they occur.

Innovation Solution

A machine learning module, specifically a neural network, is employed to analyze attributes of a computing environment, generating a risk score that indicates the likelihood of potential malfunctions. Based on this score, proactive measures are taken to prevent failures, such as updating firmware or replacing faulty drives, thereby reducing the risk of failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional monitoring methods are used to detect failures, then the system can identify problems after they occur, but the system cannot proactively prevent failures before they happen

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddowntime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The machine learning module performs preliminary analysis of system attributes to predict potential failures before they occur. By evaluating multiple attributes simultaneously and generating risk scores in advance, the system can take preventive actions before actual failures happen, thereby improving reliability and reducing downtime.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors system attributes and uses the machine learning module to generate feedback in the form of risk scores. This feedback mechanism allows the system to adjust operations proactively based on predicted failure risks, creating a closed-loop control system that prevents failures rather than merely responding to them.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If comprehensive monitoring of all system attributes is implemented to improve failure prediction accuracy, then the precision of failure identification increases, but the complexity of the system increases

Engineering Contradiction:
Improvefailure prediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The machine learning module serves as a universal processing unit that handles multiple different attributes (firmware levels, drive temperatures, workload patterns, etc.) simultaneously. By using a single multi-functional module rather than separate monitoring systems for each attribute, the patent achieves comprehensive monitoring without proportionally increasing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transforms multiple diverse attributes into a unified risk score parameter. By changing the representation of various system attributes into a single comprehensive risk metric, the system maintains high measurement precision while simplifying the overall system architecture and making the output more actionable.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220075676A1Using a machine learning module to perform preemptive identification and reduction of risk of failure in computational systems
Publication Date: 2022.03.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20220075676A1 patent drawing
  • US20220075676A1 patent drawing
  • US20220075676A1 patent drawing

AI summary

Input on a plurality of attributes of a computing environment is provided to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within the computing environment. A determination is made as to whether the risk score exceeds a predetermined threshold. In response to determining that the risk score exceeds a predetermined threshold, an indication is transmitted to indicate that potential malfunctioning is likely to occur within the computing environment. A modification is made to the computing environment to prevent the potential malfunctioning from occurring.