Causal Model for Distributed System Performance Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods fail to provide an automated and comprehensive approach to detect latent performance degradations and build dependency models for distributed systems, especially in complex network environments, as they are either computationally heavy, costly, or limited to detecting hard faults rather than performance issues.

Innovation Solution

A method and system that uses a controlled cloud environment with a sandbox to perturb resources, measure responses, identify correlations, and build causal models to detect and model performance degradations and failures in a distributed application, enabling automated testing and root cause analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated test environments and tools are used to stimulate distributed applications, then system dependency discovery and performance degradation detection improve, but computational complexity and implementation cost increase

Engineering Contradiction:
Improveautomated system dependency discoveryVSAvoidcomputational complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system segments the distributed application into individual nodes and their dependencies, analyzing each node separately through controlled stimuli application. This segmentation allows automated dependency discovery without requiring complex global analysis of the entire system at once, thereby reducing computational complexity while maintaining automation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by applying controlled stimuli to nodes before actual performance degradation occurs. This proactive approach allows the system to map dependencies and identify potential failure points in advance, enabling automated detection without requiring complex real-time analysis during actual failures.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If comprehensive behavioral modeling is performed to detect latent performance degradations, then detection precision improves, but measurement and analysis difficulty increase

Engineering Contradiction:
Improveperformance degradation detection precisionVSAvoidbehavioral analysis difficulty
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The system introduces an intermediary layer that observes and measures node behaviors indirectly through controlled stimuli responses. This intermediary approach allows precise detection of performance degradations by measuring how nodes respond to standardized inputs, avoiding the need for complex direct analysis of internal system states while maintaining high detection precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes parameters by applying controlled stimuli that modify node operating conditions temporarily. By observing how nodes respond to these parameter changes, the system can precisely detect performance degradations and map dependencies without requiring complex continuous monitoring of all system parameters under normal operation.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If controlled environment with stimuli application is used, then causal relationship identification improves, but system invasiveness increases

Engineering Contradiction:
Improvecausal relationship identification accuracyVSAvoidsystem invasiveness
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system applies partial actions by using controlled stimuli that affect only specific nodes or resources temporarily, rather than impacting the entire system continuously. This approach enables accurate causal relationship identification through targeted measurements while minimizing overall system invasiveness, as the stimuli are applied selectively and transiently rather than globally and permanently.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11461685B2Determining performance in a distributed application or system
Publication Date: 2022.10.04 ALCATEL LUCENT SA
  • US11461685B2 patent drawing
  • US11461685B2 patent drawing
  • US11461685B2 patent drawing

AI summary

In one embodiment, the method includes determining one or more nodes associated with a treatment of a query; generating one or more stimuli associated with the treatment of the query wherein the or each stimulus are likely to perturb one or more resources within a system; measuring data at the or each node relating to the resources to determine the effect of the or each stimuli at the or each node; identifying one or more pairs of nodes which have a correlation in the measured data; transforming the correlation into a causal relationships where the cause is a measuring device measuring the response and the consequences are the other correlated measuring devices; generating a list of causal relationships; and combining different causal relationships into a causal model so that a chain of causal propagations can be built.