Hybrid Observability for Distributed Microservice Root Cause Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems struggle to effectively monitor and manage highly distributed microservice architectures, leading to complex debugging and downtime issues due to interdependencies, data inconsistencies, and the inability to detect anomalies before they cause adverse events.

Innovation Solution

A hybrid observability system using application and network monitoring software, combined with machine learning and graph neural networks, to analyze and classify components, detect anomalies, and provide root cause analysis, while suggesting proactive deployment alternatives.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If microservice-based architecture is used to achieve scalability and flexibility, then system adaptability and productivity are improved, but device complexity and difficulty of detecting and measuring system state increase

Engineering Contradiction:
Improvesystem flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary system comprising a graph builder, anomaly detector, and root cause analyzer that mediates between the complex microservice architecture and the monitoring/analysis processes. This intermediary layer manages the complexity by automatically constructing service dependency graphs, detecting anomalies, and performing root cause analysis, thereby maintaining system flexibility while reducing the operational burden of complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The monitoring system performs multiple functions through a unified approach: it builds service dependency graphs, collects metrics from multiple sources, detects anomalies, and performs root cause analysis. This multi-functional system addresses various aspects of microservice monitoring within a single framework, reducing the need for separate specialized tools and simplifying the overall monitoring architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If manual monitoring processes are used for small applications, then measurement precision is maintained, but loss of time and productivity decrease when scaling to large deployments

Engineering Contradiction:
Improvemonitoring accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system implements self-service capabilities through automated anomaly detection and root cause analysis. The anomaly detector automatically identifies unusual patterns in service metrics, and the root cause analyzer autonomously determines the source of issues by traversing the service dependency graph. This automation eliminates the need for manual monitoring while maintaining high measurement precision and significantly reducing response time to incidents.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by pre-building service dependency graphs and establishing baseline metrics before anomalies occur. The graph builder proactively constructs the service topology map, and the system continuously collects and analyzes metrics to establish normal behavior patterns. This preliminary preparation enables rapid anomaly detection and root cause analysis when issues arise, reducing response time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If traditional monitoring systems are used without considering interdependencies, then ease of operation is maintained, but reliability and measurement precision deteriorate due to undetected cascading failures

Engineering Contradiction:
Improvemonitoring simplicityVSAvoidsystem reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the monitoring approach by first building a service dependency graph that divides the system into discrete service components and their relationships. This segmentation allows the system to analyze failures in the context of service dependencies rather than as isolated events. The graph structure enables the anomaly detector to trace issues through specific service paths, improving reliability while maintaining operational simplicity through structured analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The service dependency graph acts as an intermediary structure that connects monitoring operations with reliability analysis. The graph builder creates a visual and analytical representation of service relationships, which the root cause analyzer then uses to trace failure paths. This intermediary graph enables the system to maintain ease of operation through visualizable service maps while simultaneously improving reliability by considering interdependencies in failure analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250225056A1System and method for hybrid observability of highly distributed systems
Publication Date: 2025.07.10 DRAJVNETS LTD
  • US20250225056A1 patent drawing
  • US20250225056A1 patent drawing
  • US20250225056A1 patent drawing

AI summary

The invention uses application monitoring software running on individual machines, as well as network monitoring software to observe the operations of multiple machines and the networks connecting them, in order to achieve efficient understanding of highly distributed systems. Applications are monitored at multiple levels including cloud, edge and all levels in between, in order to detect the existence of and classify the type of network and application and infrastructure components, provide an efficient way to monitor these components in real time, provide an efficient way to find correlations between discrete parts (e.g. servers, containers, APIs) of the system, understand different deployment alternatives, and provide a means to provide Root Cause Analysis when a fault or degradation is detected in the service or application performance.