Auto-Insight Data Change Detection in Cloud Warehouses
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud data warehouses, changes in data metrics can be difficult to detect due to the large volume of data, and identifying the root causes of these changes is challenging, especially when they are hidden beneath other changes.
Innovation Solution
The implementation of auto-insight techniques that monitor data metrics, detect shifts, rank contributing data segments, and identify non-relevant factors, using pruning and ranking methods to correct quantification and uncover hidden shifts in composite business metrics such as ratios or probabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in large volumes in cloud data warehouses, then data storage capacity and analytical capability are improved, but detection difficulty and root cause identification become more challenging
Solution Approach 1:
The patent segments data into hierarchical levels (e.g., total revenue, revenue by region, revenue by product) and analyzes changes at different segmentation levels. This allows the system to drill down from aggregate data to detailed data segments, making change detection more manageable and root cause identification more precise despite large overall data volumes.
Solution Approach 2:
The patent extracts and isolates specific data segments that contribute to metric changes from the larger data set. By identifying and separating relevant data segments (e.g., extracting regional data when revenue change is detected), the system reduces the complexity of analyzing large data volumes while maintaining focus on changes that matter.
2Reliability
If data is organized in complex structures to maintain relationships, then data integrity and analytical depth are improved, but root cause discovery becomes more difficult
Solution Approach 1:
The patent divides complex data structures into manageable segments while preserving relationships through hierarchical organization. This segmentation maintains data integrity by preserving contextual relationships while simplifying the analysis process, allowing users to navigate from high-level summaries to detailed segments without losing data context.
Solution Approach 2:
The patent introduces intermediate summary layers between raw data and analysis queries. These intermediate layers act as mediators that pre-process and organize complex data structures, making root cause discovery easier while maintaining the integrity of the underlying data relationships.
3Measurement precision
If comprehensive data analysis is performed to identify all factors, then analysis completeness is improved, but processing time and computational resources increase
Solution Approach 1:
The patent applies partial analysis by focusing computational resources on data segments that are most likely to contain root causes of metric changes. Rather than analyzing all data comprehensively, the system identifies and analyzes only the relevant segments (e.g., only regional data when revenue changes are detected), achieving sufficient analysis completeness while reducing processing time.
Solution Approach 2:
The patent performs preliminary analysis by pre-calculating and storing summary statistics and segmentations before actual analysis queries are executed. This preliminary preparation allows the system to quickly identify relevant data segments during analysis without performing comprehensive calculations in real-time, thereby reducing processing time while maintaining analysis completeness.
Data Source
AI summary
Techniques described herein can monitor various data metrics. The auto-insight techniques can further detect and rank data segments that contributed to, or counteracted, shifts in data and detect when such shifts occurred. Thus, the techniques described herein can detect and identify root causes in shifts in different metrics. The techniques include pruning and ranking causes to identify the root causes and identify non-relevant factors, as well.


