Multi-Cloud Path Fault Detection via Telemetry Endpoints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current solutions fail to accurately identify and isolate faults causing increased latency and poor performance in multi-cloud and hybrid-cloud environments, due to complex heterogeneous networks and proprietary instrumentation across different cloud providers.
Innovation Solution
The system activates cloud triage endpoints and agents to collect telemetry data through OAM packets, which include details such as cloud region, availability zones, path types, latency, and jitter, enabling the determination of problematic paths and potential rerouting to optimize performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If proprietary instrumentation from different cloud providers is used, then each cloud provider's specific performance metrics can be measured, but cloud-agnostic visibility and fault detection across multi-cloud environments cannot be achieved
Solution Approach 1:
The patent implements a universal telemetry collection framework that can gather performance metrics from multiple cloud providers (AWS, Azure, GCP) using a single standardized interface. The system defines common telemetry data structures and collection mechanisms that work across different cloud environments, enabling cloud-agnostic monitoring while maintaining the ability to measure provider-specific metrics through adapter patterns.
Solution Approach 2:
The patent introduces an intermediary telemetry collection layer between the diverse cloud provider instruments and the monitoring system. This intermediary standardizes the interface for collecting latency, throughput, and error metrics from various cloud services, translating proprietary instrumentation outputs into a unified format that enables cross-cloud comparison and fault detection.
2Adaptability or versatility
If complex heterogeneous multi-cloud networks are used, then cloud providers can offer diverse services and flexibility, but accurate identification and isolation of faults causing increased latency becomes difficult
Solution Approach 1:
The patent segments the complex multi-cloud network into discrete telemetry collection points and measurable path segments between interest points. By breaking down the end-to-end path into individual segments (e.g., cloud edge to cloud region, availability zone to workload), the system can isolate faults to specific segments and identify which cloud provider or network segment is causing latency issues.
Solution Approach 2:
The patent implements continuous feedback loops where telemetry data from multiple cloud services is collected, analyzed, and used to dynamically adjust monitoring and troubleshooting actions. The system uses feedback from latency measurements and error rates to automatically identify problematic paths and trigger fault isolation procedures, making the complex monitoring process manageable and adaptive.
3Ease of manufacture
If on-premise infrastructure is decommissioned in favor of public clouds, then maintenance costs and IT staff requirements are reduced, but visibility and control over network paths become limited
Solution Approach 1:
The patent implements self-service telemetry collection capabilities within the public cloud environment, where cloud agents and endpoints automatically gather and report performance metrics without requiring on-premise infrastructure or dedicated IT staff for manual monitoring. The system autonomously collects latency, throughput, and error data from cloud workloads and transmits it to the analysis platform.
Solution Approach 2:
The patent replaces mechanical on-premise monitoring infrastructure with software-based telemetry collection mechanisms deployed within the public cloud environment. Instead of physical monitoring equipment and on-site IT staff, the system uses virtualized cloud agents, automated data collection scripts, and cloud-native monitoring services to maintain visibility over network paths.
Data Source
AI summary
In one embodiment, a method includes identifying a problematic event between a first interest point and a second interest point of a network and activating, in response to identifying the problematic event between the first interest point and the second interest point, a first endpoint associated with the first interest point and a second endpoint associated with the second interest point. The method also includes receiving, from the first endpoint and the second endpoint, telemetry data associated with a problematic path between the first interest point and the second interest point. The method further includes determining the problematic path between the first interest point and the second interest point using the telemetry data received from the first endpoint and the second endpoint.


