Automated Causality Detection for Cloud Upgrade Regressions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud-based platforms face challenges in identifying whether user experience issues are caused by upgrades or localized environmental factors, as manual monitoring of telemetry data is time-consuming and ineffective, especially with rapidly growing server farms.

Innovation Solution

The system automates causality detection by computing upgrade-to-upgrade and upgrade unit-to-unit scores using telemetry data to determine if issues are related to recent upgrades, allowing for sequential deployment and early mitigation of problems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual monitoring of telemetry data is used to identify problems caused by upgrades, then analysts can detect anomalies and determine root causes, but the process becomes time-consuming and ineffective as server farms grow rapidly

Engineering Contradiction:
Improveproblem detection accuracyVSAvoidmonitoring time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of analysts reviewing telemetry data with an automated system that uses machine learning models to compute causality scores. The system automatically compares telemetry data before and after upgrades, calculates upgrade-to-upgrade and upgrade unit-to-unit scores, and determines root causes without human intervention, thereby eliminating time loss while maintaining detection precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an automated causality determination system as an intermediary between telemetry data collection and problem identification. This intermediary system processes large volumes of telemetry data, computes causality scores, and generates problem identifications, serving as a bridge that handles the complexity and scale of monitoring growing server farms without requiring proportional increases in manual analyst capacity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If upgrades are deployed frequently to improve user experience, then service functionality is enhanced, but the likelihood of introducing regressions increases

Engineering Contradiction:
Improveservice improvement rateVSAvoidregression risk
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the automated causality determination system continuously monitors telemetry data after upgrades are deployed. By computing causality scores and identifying regressions automatically, the system provides rapid feedback on whether upgrades are causing problems, enabling quick mitigation actions while maintaining frequent upgrade deployment to improve user experience.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent enables preliminary identification of regressions by automatically analyzing telemetry data and computing causality scores shortly after upgrade deployment. This preliminary action allows the system to detect potential regressions early in the deployment cycle, enabling proactive mitigation before regressions affect a large number of users, thus supporting frequent upgrades with reduced risk.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If analysts manually review endless logs of telemetry data to diagnose root causes, then problem identification may be achieved, but the process is largely ineffective particularly because cloud-based platforms are growing rapidly

Engineering Contradiction:
Improveroot cause identification accuracyVSAvoiddiagnosis efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the manual mechanical process of analysts reviewing endless telemetry logs with an automated system that efficiently processes and analyzes telemetry data using computer-executable instructions. The system automatically computes causality scores, compares telemetry patterns, and identifies root causes, thereby maintaining diagnostic accuracy while dramatically improving diagnosis efficiency to keep pace with rapid platform growth.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the approach to telemetry data analysis by changing parameters from manual review of raw logs to automated computation of causality scores. The system transforms telemetry data into meaningful metrics through automated score computation, enabling efficient root cause identification that scales with platform growth without requiring proportional increases in analyst productivity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4122163B1Causality determination of upgrade regressions via comparisons of telemetry data
Publication Date: 2024.12.25 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4122163B1 patent drawingFigure 1
  • EP4122163B1 patent drawingFigure 2A
  • EP4122163B1 patent drawingFigure 2B

AI summary

Disclosed herein is a system for automating the causality detection process when upgrades are deployed to different resources that provide a service. The resources can include physical and/or virtual resources (e.g., processing, storage, and/or networking resources) that are divided into different, geographically dispersed, resource units. To determine whether a root cause of a problem is associated with an upgrade event that has recently been deployed, a system is configured to use telemetry data to compute an upgrade-to-upgrade score that represents differences between two different upgrade events that are deployed to the same resource unit. The system is further configured to use telemetry data to compute an upgrade unit-to-unit score that represents differences between the same upgrade event being deployed to two different resource units. The scores can be used to output an alert, for an analyst, that signals whether a recently deployed upgrade event is the cause of a problem.