Causality Mapping for Distributed System Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional distributed system monitoring methods require a large number of monitoring entities, leading to high computation overhead and scalability issues, especially in networks with limited bandwidth, as they monitor all significant components, often detecting unnecessary events to identify root causes.

Innovation Solution

A method and system for determining the optimal number and location of monitoring entities by generating a causality mapping model that reduces the number of detectable events, allowing for efficient monitoring of only the necessary events to represent system operations, with monitoring entities placed at selected nodes associated with these events.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If monitors are placed at every significant component in the network to monitor all events, then system monitoring coverage is improved, but the number of monitors and computation overhead increases significantly

Engineering Contradiction:
Improvesystem monitoring coverageVSAvoidnumber of monitors
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes unnecessary monitoring events from the system by identifying and eliminating redundant events that do not contribute to root cause analysis. This reduces the number of monitors needed while maintaining effective system monitoring coverage through selective event monitoring based on causality relationships.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the monitoring system into critical and non-critical events by establishing causality relationships between events. This segmentation allows the system to focus monitoring resources on essential events that actually contribute to system operation understanding, reducing the total number of monitors required.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If all events are monitored to ensure complete system visibility, then measurement precision is improved, but bandwidth consumption increases

Engineering Contradiction:
Improvesystem visibilityVSAvoidbandwidth consumption
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent extracts only the essential events needed for system visibility by establishing causality relationships and identifying which events are necessary for understanding system operation. This extraction process eliminates redundant event data transmission, reducing bandwidth consumption while maintaining adequate system visibility for monitoring purposes.

Inventive Principle:
Principle #2Taking out (Extraction)

3Difficulty of detecting and measuring

If monitors are placed throughout the network to detect all events, then detection capability is improved, but scalability deteriorates

Engineering Contradiction:
Improveevent detection capabilityVSAvoidscalability
Core Design Contradiction:
Difficulty of detecting and measuringVSAdaptability or versatility

Solution Approach 1:

The patent extracts and eliminates unnecessary events from the monitoring scope by analyzing causality relationships. This reduction in monitored events improves scalability by decreasing the computational burden and resource requirements, allowing the monitoring system to scale more effectively as the network grows while maintaining adequate detection capability for essential events.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS7546609B2Method and apparatus for determining monitoring locations in distributed systems
Publication Date: 2009.06.09 VMWARE INC
  • US7546609B2 patent drawing
  • US7546609B2 patent drawing
  • US7546609B2 patent drawing

AI summary

A method and apparatus for determining the number and location of monitoring entities in a distributed system is disclosed. The method comprising the steps of automatically generating a causality mapping model of the dependences between causing events at the nodes of the distributed system and the detectable events associated with a subset of the nodes, the model suitable for representing the execution of at least one system operation, reducing the number of detectable events in the model, wherein the reduced number of detectable events is suitable for substantially representing the execution of the at least one system operation; and placing at least one of the at least one monitoring entities at selected ones of the nodes associated with the detectable events in the reduced model. In another aspect, the processing described herein is in the form of a computer-readable medium suitable for providing instruction to a computer or processing system for executing the processing claimed.