Distributed System Support Manager for Common Component Failure Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed systems face challenges in efficiently identifying and correcting common component failures across multiple deployments, leading to reduced uptime and performance due to traditional approaches that often replace unnecessary components.

Innovation Solution

A system with a support manager that performs deployment-level monitoring, identifies common component failures, and initiates iterative outcome-driven corrective actions using a heuristically derived knowledge base to remediate issues dynamically, reducing the number of corrective actions needed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional component replacement approaches are used in distributed systems, then component failures can be addressed, but unnecessary components are replaced and system downtime increases

Engineering Contradiction:
Improvesystem uptimeVSAvoidsystem downtime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary monitoring and analysis of component failures across multiple deployments before initiating corrective actions. By identifying common component failures through deployment-level monitoring and analyzing failure patterns in advance, the system prepares remediation strategies proactively, reducing the need for unnecessary component replacements and minimizing system downtime when failures occur.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where component failure data from multiple deployments is collected, analyzed, and used to improve future failure responses. The monitoring system continuously gathers failure information, identifies common patterns, and uses this feedback to refine corrective actions, thereby reducing unnecessary replacements and improving system uptime over time through learned optimizations.

Inventive Principle:
Principle #23Feedback

2Reliability

If traditional component replacement approaches are used, then component failures can be addressed, but the number of corrective actions increases and cognitive load on users increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoidcomplexity of corrective actions
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges monitoring and analysis functions across multiple deployments into a unified system that identifies common component failures. By combining failure data from various deployments and applying centralized analysis, the system reduces the number of corrective actions needed and simplifies the overall process, decreasing cognitive load on users while maintaining system reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary layer between component failures and corrective actions. This intermediary monitoring and analysis system processes failure information, identifies common patterns, and determines optimal corrective actions, thereby reducing the complexity of direct failure-response relationships and minimizing the number of actions users must take while preserving system reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10795756B2System and method to predictively service and support the solution
Publication Date: 2020.10.06 EMC IP HLDG CO LLC
  • US10795756B2 patent drawing
  • US10795756B2 patent drawing
  • US10795756B2 patent drawing

AI summary

A system state monitor for managing a distributed system includes a persistent storage and a processor. The persistent storage includes a heuristically derived knowledge base. The processor performs deployment-level monitoring of deployments of the distributed system and identifies a common component failure of components of the deployments based on the deployment-level monitoring. In response to identifying the common component failure, the processor identifies impacted computing devices each hosting a respective component of the components; obtains deployment level state information from each of the impacted computing devices; identifies an iterative set of outcome driven corrective actions based on the obtained deployment level state information and the heuristically derived knowledge base; and initiates a computing device correction on an impacted computing device of the impacted computing devices using the iterative set of outcome driven corrective actions to obtain a corrected computing device.