Unified Process-IT Topology for Unseen Failure Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Enterprises face challenges in detecting and responding to previously unencountered IT failures due to separate silos between process owners and IT operations teams, leading to difficulties in creating unified views of IT systems and manually linking IT processes, which can result in undetected or ignored errors with significant revenue and reputation costs.

Innovation Solution

A method using computer hardware to detect IT failures by mapping unseen failures to seen ones based on similarity scores, generating IT failure impact predictions and recommendations through a unified process-IT topology created by aligning API calls with process steps, and employing machine learning models to forecast and mitigate the impact of IT failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual linkages are created among IT system facets, then understanding IT failures improves, but complexity and time consumption increase

Engineering Contradiction:
Improvefailure impact detection accuracyVSAvoidsystem linkage complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical linkage creation with automated machine learning algorithms. The system automatically learns and maps relationships between IT components and business processes through training on historical failure data, eliminating the need for manual documentation and linkage creation while achieving accurate failure impact detection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-learning and self-mapping of IT-component-to-process relationships through automated training on operational data. The machine learning models continuously improve their understanding of system topology and failure patterns without requiring manual intervention, enabling the system to serve itself in mapping and analysis tasks.

Inventive Principle:
Principle #25Self-service

2Reliability

If comprehensive monitoring of all IT components is implemented, then failure detection capability improves, but system complexity and resource consumption increase

Engineering Contradiction:
Improvefailure detection capabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and focuses monitoring on critical failure indicators rather than comprehensively monitoring all IT components. The machine learning models identify and prioritize the most relevant signals and components based on historical failure patterns, allowing the system to maintain high reliability by focusing resources on what matters most rather than uniformly monitoring everything.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system dynamically adjusts monitoring parameters and thresholds based on learned failure patterns and system state. The machine learning models adapt their detection criteria over time, changing parameters such as alert thresholds and monitoring priorities to balance comprehensive coverage with resource efficiency, improving reliability without linearly increasing complexity.

Inventive Principle:
Principle #35Parameter changes

3Speed

If automated failure mapping using machine learning is implemented, then response time to unseen failures improves, but computational resource consumption increases

Engineering Contradiction:
Improvefailure response speedVSAvoidcomputational resource consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary training and modeling of failure patterns in advance, building comprehensive knowledge bases before actual failures occur. By pre-training machine learning models on historical data and establishing baseline behavior patterns, the system can rapidly respond to new failures without consuming excessive computational resources during the critical response moment, as the heavy lifting of pattern recognition has already been performed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by using machine learning models that process only the most relevant failure indicators and components rather than analyzing every possible system parameter. The models are designed to focus computational effort on the most predictive features, achieving fast response times for unseen failures while minimizing overall computational resource consumption by avoiding exhaustive analysis of all system data.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240193023A1Predicting the impact of previously unseen computer system failures on the system using a unified topology
Publication Date: 2024.06.13 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20240193023A1 patent drawing
  • US20240193023A1 patent drawing
  • US20240193023A1 patent drawing

AI summary

Predicting the impact of an information technology (IT) failure includes detecting a computer-generated indication of the failure. Responsive to determining that the IT failure is a previously unseen IT failure, operations of a computer-implemented unseen event handler can be invoked. The unseen event handler can map the previously unseen IT failure to a previously seen IT failure based on a similarity score generated by a computer-implemented similarity scorer, wherein the similarity score is based on a unified process-IT topology. A machine learning model can generate an IT failure impact prediction and recommendation based on the mapping, wherein the machine learning model also is based on the unified process-IT topology. An output of the IT failure prediction and recommendation can be generated.