Data Quality Lineage Assessment for Cross-Dataset Exception Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data quality assessment systems fail to identify and track data quality problems upstream and downstream among datasets, leading to inefficient and costly shutdowns of downstream processing tasks, and require duplicate rules for each dataset, increasing implementation effort.

Innovation Solution

A method and system for assessing data quality that identifies exceptional values in a first dataset, stores entries in an exception table, and links these to related datasets, using assessment rules to track and communicate impacts across datasets, allowing for targeted adjustments without halting entire processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data quality rules are applied to every dataset individually, then data quality issues can be identified in each dataset, but implementation effort and time consumption increase significantly

Engineering Contradiction:
Improvedata quality assessment completenessVSAvoidimplementation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates a universal data quality rule engine that can assess multiple datasets using a single set of rules. The system stores datasets and their relationships in a graph structure, allowing one rule set to evaluate data quality across the entire data ecosystem rather than requiring separate rule sets for each dataset. This multi-functional approach maintains comprehensive data quality assessment while significantly reducing implementation time.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the assessment of multiple datasets into a unified process by representing them as a graph of related datasets. Instead of processing datasets independently, the system combines them into a relational model where data quality issues in one dataset can be propagated to related datasets automatically, eliminating the need for duplicate rule execution across multiple datasets.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If data quality issues are identified in upstream datasets, then downstream datasets can be affected, but traditional systems cannot identify the precise impact on downstream datasets

Engineering Contradiction:
Improvedata quality issue tracking accuracyVSAvoiddownstream impact information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent implements feedback mechanisms that track data quality issues from upstream datasets through the graph structure to downstream datasets. When a data quality issue is detected in an upstream dataset, the system automatically propagates this information through the graph to identify which downstream datasets are affected and by what degree, providing continuous feedback about the impact of data quality issues across the entire data ecosystem.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent adds a relational dimension to data quality assessment by representing datasets and their relationships as a graph structure. This dimensional change from isolated dataset assessment to relational graph assessment enables the system to track and communicate the impact of data quality issues across multiple datasets, providing a holistic view of data quality propagation that traditional systems cannot achieve.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If downstream processing is halted to avoid data quality problems, then data quality issues can be addressed, but process delays and costs increase significantly

Engineering Contradiction:
Improvedata quality controlVSAvoiddownstream processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements dynamic data quality control that allows downstream processing to continue at varying levels of strictness based on the severity and propagation of data quality issues. Rather than a binary halt-or-continue approach, the system dynamically adjusts processing parameters and data quality thresholds based on real-time assessment of issue severity, enabling flexible control that maintains productivity while ensuring data quality.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality control by addressing data quality issues at specific locations in the data ecosystem rather than halting all downstream processing. The graph structure enables targeted intervention at specific datasets or relationships where issues exist, allowing other parts of the system to continue operating normally. This localized approach maintains overall productivity while addressing specific data quality problems.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12505088B2System and method for data quality assessment
Publication Date: 2025.12.23 TORANA INC
  • US12505088B2 patent drawing
  • US12505088B2 patent drawing
  • US12505088B2 patent drawing

AI summary

Methods and systems for providing data assessment across related datasets, including identifying exceptional values in datasets and assessing upstream and or downstream datasets that utilize the exceptional values. Data assessment rules use exceptional values found in a dataset and data lineage information to identify impacted data in upstream or downstream datasets.