Data Quality Anomaly Root Cause Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data quality check mechanisms fail to identify the root cause of data quality anomalies in data warehouses, limiting organizations' ability to understand and address the underlying issues effectively.

Innovation Solution

A method that monitors data parameters, detects disparities, identifies and scores events associated with the disparities, and takes preventive actions by forming a dependency graph and using machine learning models to determine the root cause, adjusting weights and recalculating scores iteratively to pinpoint the cause of data disparities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional data quality check mechanisms are used, then data quality monitoring is simple and fast, but the root cause of data quality anomalies cannot be identified

Engineering Contradiction:
Improveroot cause identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the data quality anomaly analysis into multiple independent components: event collection module, event scoring module, and root cause identification module. Each module handles a specific aspect of the analysis, allowing the system to identify root causes accurately without requiring a monolithic complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces event scores as an intermediary mechanism that bridges the gap between raw event data and root cause identification. By calculating and comparing event scores, the system can objectively determine the root cause without requiring direct complex analysis of all possible factors.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If detailed event analysis is performed to identify root causes, then understanding of underlying issues improves, but analysis time and computational resources increase

Engineering Contradiction:
Improveinformation completenessVSAvoidanalysis time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent applies partial action by focusing analysis only on relevant events associated with the data quality anomaly rather than analyzing all system events. The event scoring mechanism prioritizes likely causes, allowing the system to achieve sufficient understanding without exhaustive analysis of every possible factor.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameter of event evaluation from qualitative assessment to quantitative scoring. By assigning numerical scores to events based on predefined criteria, the system can rapidly compare and rank potential root causes, significantly reducing analysis time while maintaining information completeness.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If event scoring and weight adjustment is implemented, then root cause identification accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvecause identification accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent performs preliminary action by pre-defining event weights and scoring criteria before actual anomaly analysis occurs. This preparation work is done once and stored, so during runtime the system only needs to retrieve and apply these predefined values, reducing computational burden while maintaining high identification accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The event scoring system is self-service in that it automatically calculates scores and identifies root causes without requiring manual intervention or complex external computations. The system uses its own predefined weights and criteria to evaluate events, making the process computationally efficient while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12189608B2System and method of identifying event as root cause of data quality anomaly
Publication Date: 2025.01.07 VISA INTERNATIONAL SERVICE ASSOCIATION
  • US12189608B2 patent drawing
  • US12189608B2 patent drawing
  • US12189608B2 patent drawing

AI summary

Embodiments detect and predict data disparity issues in data warehouses. Embodiments derive meaningful insights about the events occurred prior to the data disparity and correlate the events to understand the root cause of the data disparity (or the root cause of an alert generated as a result of detecting the data disparity). Embodiments either take or recommend actionable measures to prevent further occurrences of the event identified as the root cause. According to various embodiments, when the monitored data is transaction data (e.g. transaction volume, transaction amount, transaction processing speed, etc.), internal events (e.g. data job failures, job delays, job server maintenances) or external events (e.g. seasonal holiday events, natural calamities) may cause a dip or spike in the transaction data resulting in a data quality anomaly (i.e. a data disparity).