Automated Causality Detection for Cloud Upgrade Regressions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud-based platforms face challenges in identifying whether user experience issues are caused by upgrades or localized environmental factors, as manual monitoring of telemetry data is time-consuming and ineffective, especially with rapidly growing server farms.
Innovation Solution
The system automates causality detection by computing upgrade-to-upgrade and upgrade unit-to-unit scores using telemetry data to determine if issues are related to recent upgrades, allowing for sequential deployment and early mitigation of problems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual monitoring of telemetry data is used to identify problems caused by upgrades, then analysts can detect anomalies and determine root causes, but the process becomes time-consuming and ineffective as server farms grow rapidly
Solution Approach 1:
The patent replaces the manual mechanical process of analysts reviewing telemetry data with an automated system that uses machine learning models to compute causality scores. The system automatically compares telemetry data before and after upgrades, calculates upgrade-to-upgrade and upgrade unit-to-unit scores, and determines root causes without human intervention, thereby eliminating time loss while maintaining detection precision.
Solution Approach 2:
The patent introduces an automated causality determination system as an intermediary between telemetry data collection and problem identification. This intermediary system processes large volumes of telemetry data, computes causality scores, and generates problem identifications, serving as a bridge that handles the complexity and scale of monitoring growing server farms without requiring proportional increases in manual analyst capacity.
2Productivity
If upgrades are deployed frequently to improve user experience, then service functionality is enhanced, but the likelihood of introducing regressions increases
Solution Approach 1:
The patent implements a feedback mechanism where the automated causality determination system continuously monitors telemetry data after upgrades are deployed. By computing causality scores and identifying regressions automatically, the system provides rapid feedback on whether upgrades are causing problems, enabling quick mitigation actions while maintaining frequent upgrade deployment to improve user experience.
Solution Approach 2:
The patent enables preliminary identification of regressions by automatically analyzing telemetry data and computing causality scores shortly after upgrade deployment. This preliminary action allows the system to detect potential regressions early in the deployment cycle, enabling proactive mitigation before regressions affect a large number of users, thus supporting frequent upgrades with reduced risk.
3Measurement precision
If analysts manually review endless logs of telemetry data to diagnose root causes, then problem identification may be achieved, but the process is largely ineffective particularly because cloud-based platforms are growing rapidly
Solution Approach 1:
The patent replaces the manual mechanical process of analysts reviewing endless telemetry logs with an automated system that efficiently processes and analyzes telemetry data using computer-executable instructions. The system automatically computes causality scores, compares telemetry patterns, and identifies root causes, thereby maintaining diagnostic accuracy while dramatically improving diagnosis efficiency to keep pace with rapid platform growth.
Solution Approach 2:
The patent transforms the approach to telemetry data analysis by changing parameters from manual review of raw logs to automated computation of causality scores. The system transforms telemetry data into meaningful metrics through automated score computation, enabling efficient root cause identification that scales with platform growth without requiring proportional increases in analyst productivity.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Disclosed herein is a system for automating the causality detection process when upgrades are deployed to different resources that provide a service. The resources can include physical and/or virtual resources (e.g., processing, storage, and/or networking resources) that are divided into different, geographically dispersed, resource units. To determine whether a root cause of a problem is associated with an upgrade event that has recently been deployed, a system is configured to use telemetry data to compute an upgrade-to-upgrade score that represents differences between two different upgrade events that are deployed to the same resource unit. The system is further configured to use telemetry data to compute an upgrade unit-to-unit score that represents differences between the same upgrade event being deployed to two different resource units. The scores can be used to output an alert, for an analyst, that signals whether a recently deployed upgrade event is the cause of a problem.