Knowledge-Based Data Quality Solution for Domain-Specific Cleansing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data quality solutions rely on generic approaches that fail to effectively address errors in data, leading to low-quality data that can compromise decision-making and trust, as they do not account for the specific characteristics and contexts of the data being processed.
Innovation Solution
A knowledge-based data quality solution that separates knowledge acquisition from data processing, utilizing a movable and extensible knowledge container to cleanse, profile, and perform semantic de-duplication, leveraging internal and external information sources for improved data quality management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If generic algorithms are applied to cleanse data, then data processing can be performed, but the data quality remains low because the specific problems are not addressed
Solution Approach 1:
The patent segments data quality issues into distinct domains (e.g., contact information, addresses, names) and applies specialized algorithms tailored to each domain's specific error patterns, rather than using a single generic algorithm for all data types
Solution Approach 2:
The patent implements domain-specific quality rules and validation criteria that are customized to the characteristics of each data type, ensuring that the cleansing process addresses the particular problems associated with each domain while maintaining overall processing efficiency
2Ease of operation
If a one-size-fits-all approach is used for data cleansing, then the system is simple to operate, but it cannot solve specific data quality problems
Solution Approach 1:
The patent creates a universal data quality system that automatically detects data types and applies appropriate domain-specific algorithms through a unified interface, maintaining ease of operation while delivering specialized quality improvement for each data domain
3Reliability
If domain-specific knowledge is integrated into data processing, then data quality improves, but the system complexity increases
Solution Approach 1:
The patent pre-configures domain-specific knowledge bases, validation rules, and cleansing algorithms for various data types before processing begins, allowing the system to automatically apply appropriate expertise without requiring complex real-time decision-making or user configuration
Data Source
AI summary
The subject disclosure relates to a knowledge-driven data quality solution that is based on a rich knowledge base. The data quality solution can provide continuous improvement and can be based on continuous (or on-going) knowledge acquisition. The data quality solution can be built once and can be reused for multiple data quality improvements, which can be for the same data or for similar data. The disclosed aspects are easy to use and focus on productivity and user experience. Further, the disclosed aspects are open and extendible and can be applied to cloud-based reference data (e.g., a third party data source) and/or user generated knowledge. According to some aspects, the disclosed aspects can be integrated with data integration services.


