Causality Mapping Model for Distributed System Fault Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As computer networks and distributed systems become increasingly complex, accurately modeling causality relationships between events and their sources becomes difficult due to the super-linear increase in network complexity and fault propagation, making it hard to manage problems effectively.

Innovation Solution

A method and apparatus for generating a causality mapping model that automatically maps dependencies between causing events and detectable events in a distributed system, using techniques like causality matrices and probabilistic models to represent system operations and predict the impact of node failures on application connections.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated event collection and reporting systems are implemented, then the load on human operators is reduced, but the complexity of managing and analyzing the increased volume of events increases

Engineering Contradiction:
Improveautomation of event collection and reportingVSAvoidcomplexity of event management system
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent introduces an event correlation engine as an intermediary component that automatically processes, filters, and prioritizes events before presenting them to operators. This mediator layer reduces the burden on human operators while managing the complexity of event analysis through automated correlation algorithms and predefined event patterns.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates simplified copies or representations of complex event data through standardized event formats and correlation models. By copying event information into a structured framework with defined relationships, the system makes large volumes of events manageable without requiring operators to directly handle the full complexity of raw event data.

Inventive Principle:
Principle #26Copying

2Productivity

If event correlation techniques are used to group distinct events, then the event stream is compressed into a more manageable form, but the accuracy of identifying underlying causes may be reduced

Engineering Contradiction:
Improveefficiency of problem managementVSAvoidaccuracy of cause identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts correlation parameters and thresholds based on the specific event types and system conditions. By changing parameters such as correlation time windows, event weightings, and priority levels, the system optimizes the balance between compressing event streams for efficiency and maintaining sufficient precision for accurate cause identification in different operational contexts.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The event correlation engine employs dynamic correlation rules that adapt to changing system states and event patterns. Rather than using static grouping criteria, the system dynamically adjusts correlation strength and scope based on real-time analysis of event relationships, allowing it to maintain accuracy while achieving compression of the event stream.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If the number of computer nodes in a network increases, then the system capacity and functionality are improved, but the network complexity and fault rate increase super-linearly

Engineering Contradiction:
Improvesystem capacityVSAvoidnetwork complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the distributed system into manageable domains or zones, each with its own event correlation rules and management policies. By dividing the large-scale network into smaller logical units, the system can handle increased node counts without experiencing super-linear complexity growth, as each segment can be independently analyzed and managed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces additional organizational dimensions beyond simple node count, such as hierarchical levels, functional domains, and geographic regions. By organizing nodes along multiple dimensions rather than treating them as a flat collection, the system manages complexity through structured categorization that scales more gracefully with system growth.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Adaptability or versatility

If fault propagation between machines and protocol layers is allowed, then system functionality and interoperability are improved, but the rate of generated events and difficulty of problem management increase

Engineering Contradiction:
Improvesystem interoperabilityVSAvoidcomplexity of problem management
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The event correlation engine acts as an intermediary that intercepts and analyzes fault propagation events as they travel between machines and protocol layers. By introducing this mediation layer, the system maintains full interoperability and fault propagation capabilities while the correlator automatically tracks and manages the resulting event chains, reducing the complexity burden on operators.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7949739B2Method and apparatus for determining causality mapping of distributed systems
Publication Date: 2011.05.24 VMWARE INC
  • US7949739B2 patent drawing
  • US7949739B2 patent drawing
  • US7949739B2 patent drawing

AI summary

A method and apparatus for determining causality mapping between causing events and detectable events among a plurality of nodes in a distributed system is disclosed. The method comprises the steps of automatically generating a causality mapping model of the dependences between causing events at the nodes of the distributed system and the detectable events in a subset of the nodes, the model suitable for representing the execution of at least one system operation. In one aspect the generation is perform by selecting nodes associated with each of the detectable events from the subset of the nodes and indicating the dependency between a causing event and at least one detectable event for each causing event at a node when the causing event node is a known distance from at least one node selected from the selected nodes. In still another aspect, the processing described herein is in the form of a computer-readable medium suitable for providing instruction to a computer or processing system for executing the processing claimed.