Random Forest Root Cause Analysis for Cloud Data Shifts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud data warehouses, changes in data metrics can be difficult to detect and analyze due to the large volume of data and complex storage structures, making it challenging to identify the root causes of these changes.
Innovation Solution
The use of a random forest machine learning model to monitor data metrics, detect and rank data segments contributing to or counteracting shifts in data, and identify the root causes of these shifts, even in datasets with hundreds of dimensions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional data analysis methods are used in cloud data warehouses, then data storage capacity is maintained, but the ability to detect and analyze changes in data metrics deteriorates due to large volume and complex storage structures
Solution Approach 1:
The patent segments the large volume of data into manageable dimensions and hierarchies, allowing the system to analyze changes at different levels of granularity. By breaking down the data structure into segments that can be independently analyzed, the system can detect metric changes without being overwhelmed by the total data volume.
Solution Approach 2:
The patent introduces an intermediary layer of analysis that sits between the raw data storage and the detection process. This intermediary layer processes and structures the data before analysis, making it easier to detect changes even in large volumes of data.
2Loss of information
If comprehensive data storage is maintained to preserve all information, then data completeness is improved, but the complexity of identifying root causes of changes worsens
Solution Approach 1:
The patent segments the comprehensive data into organized dimensions and hierarchies, maintaining completeness while reducing the complexity of analysis. By structuring data into segments, the system can identify root causes by examining specific segments rather than the entire data set.
Solution Approach 2:
The patent introduces dimensional organization to the data storage structure, allowing analysis to proceed along multiple dimensions simultaneously. This dimensional approach transforms the complexity of root cause identification by providing structured pathways through the data.
3Device complexity
If traditional monitoring approaches are used, then system simplicity is maintained, but the precision of detecting data shifts and their causes deteriorates in high-dimensional datasets
Solution Approach 1:
The patent implements dynamic analysis capabilities that adapt to the specific characteristics of the data being analyzed. The system can dynamically adjust its monitoring approach based on the dimensionality and structure of the data, maintaining simplicity where possible while achieving high precision when needed.
Solution Approach 2:
The patent changes the parameters of analysis by introducing dimensional hierarchies and multiple levels of granularity. This allows the system to achieve high measurement precision by selecting appropriate parameter levels for different analysis scenarios.
Data Source
AI summary
Techniques described herein can monitor various data metrics. The techniques can select a subset of dimensions from a plurality of dimensions related to a data shift. The techniques including generating a plurality of decision tree graphs to classify a plurality of segments, each segment representing a combination of two or more dimensions of the subset of dimensions, and each decision tree graph including a different root node representing a respective dimension of the subset of dimensions.


