Anomaly Detection Using Client Telemetry Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Service providers face challenges in identifying and addressing issues outside their domain, such as equipment failures at third-party ISPs, which can cause connection problems for online services, leading to increased recovery times and reduced reliability.

Innovation Solution

The use of telemetry data from client devices to automatically detect anomalies, allowing service providers to infer and remediate issues within the online service, including those occurring outside their domain, without the need for local monitoring at every network point.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If service providers implement local monitoring systems at every discrete point along the network topology, then measurement precision and anomaly detection capability improve, but device complexity and implementation cost increase significantly

Engineering Contradiction:
Improveanomaly detection capabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by using a single centralized monitoring system that performs multiple functions: collecting telemetry data from diverse sources (clients, servers, network devices), processing various types of data (performance metrics, error logs, connectivity status), and detecting anomalies across the entire service ecosystem. This eliminates the need for separate monitoring systems at each discrete point while maintaining comprehensive monitoring capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces telemetry data as an intermediary that carries information about system state and anomalies from distributed sources to the centralized monitoring system. This intermediary mechanism enables the centralized system to indirectly observe and detect issues throughout the network without requiring direct monitoring infrastructure at every location, thus reducing complexity while preserving detection precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If service providers actively monitor third-party ISP equipment and client-side configurations, then reliability of online service improves, but loss of time for identification and recovery increases due to lack of direct access

Engineering Contradiction:
Improveonline service reliabilityVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by continuously collecting and analyzing telemetry data before failures manifest as service outages. The system proactively identifies degradation patterns, predicts potential failures, and alerts operators in advance, enabling preventive maintenance and faster response when issues occur. This reduces both the frequency and duration of service disruptions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes a feedback loop where telemetry data from clients and network devices continuously informs the monitoring system about service health. When anomalies are detected, the system generates alerts and tracks resolution status, creating a closed-loop feedback mechanism that enables rapid identification and correction of issues, thereby improving reliability and reducing recovery time.

Inventive Principle:
Principle #23Feedback

3Loss of energy

If telemetry data is collected and analyzed centrally, then network bandwidth usage is reduced compared to distributed monitoring, but measurement precision may be compromised

Engineering Contradiction:
Improvenetwork bandwidth usageVSAvoidanomaly detection accuracy
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The patent extracts only the essential and relevant telemetry data from distributed sources for centralized analysis. Instead of transmitting all raw data from every device, the system selectively collects key performance indicators, error metrics, and anomaly signals that are most useful for detecting service-wide patterns. This extraction approach minimizes bandwidth consumption while preserving the precision needed for effective anomaly detection.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10511545B2Anomaly detection and classification using telemetry data
Publication Date: 2019.12.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10511545B2 patent drawing
  • US10511545B2 patent drawing
  • US10511545B2 patent drawing

AI summary

Historical telemetry data can be used to generate predictions for various classes of data at various aggregates of a system that implements an online service. An anomaly detection process can then be utilized to detect anomalies for a class of data at a selected aggregate. An example anomaly detection process includes receiving telemetry data originating from a plurality of client devices, selecting a class of data from the telemetry data, converting the class of data to a set of metrics, aggregating the set of metrics according to a component of interest to obtain values of aggregated metrics over time for the component of interest, determining a prediction error by comparing the values of the aggregated metrics to a prediction, detecting an anomaly based at least in part on the prediction error, and transmitting an alert message of the anomaly to a receiving entity.