Data Quality Management System for Multi-Source Inconsistency Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Inconsistent and incomplete data associated with objects across different data sources can lead to quality issues, causing issues in data consumption and usage.

Innovation Solution

A data quality management system that identifies and resolves data quality issues by comparing attribute values across multiple data sources, providing graphical user interfaces for users to address missing, inconsistent, and untranslated values, and implementing actions such as copying or replacing values, and translating attribute values as needed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is collected from multiple different data sources, then the quantity and variety of information associated with an object increases, but data quality issues such as inconsistency and incompleteness arise

Engineering Contradiction:
Improvequantity of informationVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system implements automated feedback loops where data quality issues are continuously detected through comparison algorithms, and resolution actions are automatically executed or prompted. The system monitors data consistency across sources, identifies discrepancies, and feeds this information back to users or automated processes for correction, creating a continuous improvement cycle for data quality.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system introduces an intermediary data quality management layer between multiple data sources and end users. This intermediary component compares data from different sources, identifies inconsistencies and completions, and presents unified, quality-assured information to users, mediating the conflict between data quantity and quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If data from multiple data sources is integrated, then completeness of object information improves, but inconsistency and data quality issues worsen

Engineering Contradiction:
Improvecompleteness of informationVSAvoiddata consistency
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system continuously compares data attributes across multiple sources and provides feedback on consistency and completeness. When inconsistencies are detected, the system flags them for review and can automatically resolve them based on predefined rules or user input, creating a feedback mechanism that maintains data quality while integrating multiple sources.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system applies different quality assurance strategies to different data attributes based on their specific characteristics and requirements. Critical attributes undergo stricter validation and comparison, while less critical attributes receive lighter processing, allowing the system to maintain high completeness without uniformly sacrificing consistency.

Inventive Principle:
Principle #3Local quality

3Reliability

If automated data comparison and resolution systems are implemented, then data quality improves, but system complexity increases

Engineering Contradiction:
Improvedata qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements self-service capabilities where data quality issues are automatically detected, analyzed, and resolved without requiring constant human intervention. Automated algorithms compare data sources, identify inconsistencies, and execute resolution actions based on predefined rules, allowing the system to maintain high data quality while minimizing operational complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary data validation and comparison actions before data is fully integrated or used. By pre-identifying and resolving potential quality issues before they propagate through the system, the need for complex post-processing and manual intervention is reduced, simplifying the overall system architecture.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If manual review and resolution of data quality issues is performed, then data accuracy improves, but time consumption and productivity decrease

Engineering Contradiction:
Improvedata accuracyVSAvoiddata processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system implements automated feedback mechanisms that identify and resolve the majority of data quality issues without manual intervention. Only complex or ambiguous cases are flagged for manual review, creating a tiered processing system that maintains high accuracy while maximizing automated efficiency and minimizing time-consuming manual operations.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service data quality management by automatically comparing data sources, identifying inconsistencies, and executing resolution actions based on predefined rules. This automation handles routine data quality tasks efficiently, reserving manual review only for exceptional cases, thereby maintaining high accuracy without sacrificing productivity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10268711B1Identifying and resolving data quality issues amongst information stored across multiple data sources
Publication Date: 2019.04.23 AMAZON TECH INC
  • US10268711B1 patent drawing
  • US10268711B1 patent drawing
  • US10268711B1 patent drawing

AI summary

The techniques described herein are directed to identifying data quality issues within information stored across multiple different data sources. For instance, the data quality issues can comprise missing values, inconsistent values, and un-translated values. Once identified, the techniques implement actions to resolve the data quality issues so that consumption or use of the information stored is improved. In at least one example, the identification and resolution of a data quality issue can be implemented in response to receiving a query that identifies an object. Based on the query, the system can collect values, from the multiple different sources, for attributes that have been defined for an item. The system can use algorithms (e.g., a comparison algorithm) to identify a data quality issue and can output a graphical user interface that visually distinguishes between attributes with a data quality issue and attributes without a data quality issue.