Event Clustering System for Infrastructure Failure Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for managing and organizing large volumes of electronic messages and events from infrastructure are inefficient, lacking automated techniques for effective indexing and retrieval, leading to difficulties in finding relevant information due to the vast amount of data and the complexity of web-based communication systems.

Innovation Solution

An event clustering system with a collaborative interface that utilizes machine learning to group events from managed infrastructures, employing engines like NMF, k-means clustering, and topology proximity to identify common characteristics and create clusters related to failures or errors, thereby facilitating the organization and retrieval of relevant information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated techniques are implemented for indexing and retrieval of electronic messages and events, then information retrieval efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveinformation retrieval efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system employs machine learning models that automatically learn and adapt to the specific characteristics of the infrastructure events and messages. The models self-train on historical data, automatically improving their indexing and retrieval capabilities without requiring manual configuration or intervention, thus maintaining productivity gains while limiting complexity growth through autonomous operation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent utilizes neural network models with adjustable parameters that can be fine-tuned based on the specific infrastructure being monitored. The system dynamically adjusts model parameters such as embedding dimensions, learning rates, and threshold values to optimize retrieval efficiency for different infrastructure types, achieving high productivity across diverse scenarios without requiring completely different systems for each case

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If machine learning models are used to cluster events and identify causal relationships, then accuracy in identifying actionable problems is improved, but computational resources required increase

Engineering Contradiction:
Improveaccuracy in identifying actionable problemsVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system implements a two-stage processing approach where events are first filtered through lightweight preprocessing that identifies obvious patterns and discards clearly irrelevant events. Only events that pass this initial filter are subjected to the computationally intensive neural network analysis, reducing overall computational resource consumption while maintaining high accuracy for the events that require detailed analysis

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent divides the complex event analysis task into multiple specialized neural network components: embedding models that convert events to vectors, clustering models that group similar events, and causal inference models that identify relationships. This segmentation allows each component to be optimized independently and processed in parallel, improving accuracy through specialized processing while managing computational resources more efficiently than a monolithic approach

Inventive Principle:
Principle #1Segmentation

3Loss of information

If all events from infrastructure are stored and analyzed, then completeness of information is improved, but time required for analysis and retrieval increases

Engineering Contradiction:
Improvecompleteness of informationVSAvoidtime for analysis and retrieval
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary embedding of events into vector representations as they are ingested, and pre-clusters them using approximate nearest neighbor algorithms. This preliminary processing creates an organized structure that enables fast retrieval later, ensuring that when analysis is needed, the system can quickly access relevant events without having to scan all stored events, thus maintaining information completeness while reducing analysis time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential features and characteristics of events needed for causal analysis and clustering, storing these extracted features in optimized data structures rather than storing complete event data. This extraction approach maintains the completeness of information needed for accurate analysis while significantly reducing the time required to retrieve and process event data

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10050910B2Application of neural nets to determine the probability of an event being causal
Publication Date: 2018.08.14 DELL PROD LP
  • US10050910B2 patent drawing
  • US10050910B2 patent drawing
  • US10050910B2 patent drawing

AI summary

An event clustering system has an extraction engine in communication with a managed infrastructure. A signalizer engine includes one or more of an NMF engine, a k-means clustering engine and a topology proximity engine. The signalizer engine determines one or more common characteristics or features from events, the signalizer engine using the common features of events to produce clusters of events relating to the failure or errors in the managed infrastructure. Membership in a cluster indicates a common factor of the events that is a failure or an actionable problem in the physical hardware managed infrastructure directed to supporting the flow and processing of information. The system is configured to group two or more situations, where a situation is a collection of one or more events or alerts representative of a problem in the managed infrastructure.