Failure Prediction Explainability for Data Processing Infrastructure

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems face challenges in managing failures due to the lack of transparency and trustworthiness in inference models, which are crucial for predicting component failures, leading to inefficiencies and potential system downtime.

Innovation Solution

Implementing a system that uses a large language model as a controller to interpret user inputs, retrieve structured knowledge attributes from a repository, and generate failure prediction responses, enhancing the interpretability and trustworthiness of failure predictions through explainable AI.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If inference models are used to predict failures in data processing systems, then prediction capability is improved, but transparency and trustworthiness deteriorate due to lack of interpretability

Engineering Contradiction:
Improvefailure prediction capabilityVSAvoidtransparency of prediction reasoning
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent introduces an explainable AI layer as an intermediary between the inference model and end users. This layer extracts hidden knowledge from the inference model's internal representations and transforms it into structured, interpretable explanations that preserve the original prediction capability while adding transparency about the reasoning process.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts hidden knowledge from the internal workings of the inference model through explainable AI techniques. By taking out and analyzing the model's internal representations, the system generates structured explanations that reveal the reasoning behind predictions without altering the model's core prediction function.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If manual validation of inference model predictions is performed, then trustworthiness is improved, but time consumption and operational efficiency deteriorate

Engineering Contradiction:
Improvetrustworthiness of predictionsVSAvoidvalidation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent enables the inference model to self-explain its predictions through integrated explainable AI capabilities. The model automatically generates structured explanations of its reasoning process, eliminating the need for external manual validation while maintaining trustworthiness through transparent, interpretable output.

Inventive Principle:
Principle #25Self-service

3Loss of information

If explainable AI is implemented to extract hidden knowledge, then transparency is improved, but system complexity increases

Engineering Contradiction:
Improvetransparency of model reasoningVSAvoidsystem architecture complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent implements a universal explainable AI layer that can extract and explain hidden knowledge from multiple different inference models through a single standardized interface. This multi-functional approach handles various model types and prediction scenarios without requiring separate explanation mechanisms for each, thereby managing complexity while maintaining transparency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12613764B2Managing data processing system failures using a predictive model as a controller and hidden knowledge from predictive models
Publication Date: 2026.04.28 DELL PROD LP
  • US12613764B2 patent drawing
  • US12613764B2 patent drawing
  • US12613764B2 patent drawing

AI summary

Methods and systems for managing data processing systems are disclosed. A data processing system may include and depend on the operation of hardware and/or software components. Inference models may be implemented to predict future system infrastructure outcomes (e.g., component failures) using information recorded in logs that reflect the operation of the components. However, the models may be complex “black boxes” and may generate critical outcome predictions for downstream consumers without explanations of how the predictions are determined, resulting in downstream consumers having low confidence in the predictions. Therefore, hidden knowledge (e.g., structured knowledge attributes) of the models may be extracted and/or used to understand the underlying processes that the models use to predict the system infrastructure outcomes. The hidden knowledge may be utilized interactively (e.g., using AI chatbots) to provide users with failure prediction responses that allow users to better remediate failures of the data processing systems.