Data Integration System for Automatic Inconsistency Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data integration systems face challenges in automatically detecting and adjusting data inconsistencies across different data types and qualities, leading to increased engineering man-hours and adverse effects on business applications as data volume and diversity grow.
Innovation Solution
A data integration method and system that automatically detects inconsistencies by calculating feature values, detecting inconsistencies, and adjusting data types and qualities based on user-defined specifications through a data integration system comprising schema mapping detection, inconsistency detection, map addition, query interpretation, and data adjustment units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data preparation processing is constructed assuming constant data type or quality, then business application can operate normally, but data inconsistency detection and correction require increased engineering man-hours
Solution Approach 1:
The system performs self-service by automatically detecting data inconsistencies and executing correction processing without requiring engineer intervention. The inconsistency detection unit automatically identifies data type or quality inconsistencies, and the correction processing unit automatically adjusts the data based on stored correction rules, enabling the system to serve itself rather than requiring external manual correction.
Solution Approach 2:
The system implements feedback by continuously monitoring data for inconsistencies and automatically triggering correction processes. The inconsistency detection unit provides feedback about detected inconsistencies to the correction processing unit, which then applies appropriate corrections and stores updated correction rules for future use, creating a closed-loop system that improves over time.
2Manufacturing precision
If data is adjusted according to type or quality assumed for each application, then data consistency is maintained, but engineering man-hours increase as data volume and diversity grow
Solution Approach 1:
The system automatically maintains data consistency through self-service mechanisms. The inconsistency detection unit continuously monitors data against assumed types and qualities, and the correction processing unit automatically adjusts data to match requirements, eliminating the need for manual data adjustment even as data volume and diversity increase.
Solution Approach 2:
The system performs preliminary action by pre-storing correction rules for various data inconsistency scenarios. When inconsistencies are detected, the system can immediately apply pre-prepared correction rules without requiring real-time engineering decisions, enabling rapid automatic correction of data consistency issues.
3Difficulty of detecting and measuring
If automatic inconsistency detection is implemented, then data inconsistencies can be promptly detected, but automatic adjustment based on user-defined specifications requires complex system architecture
Solution Approach 1:
The system applies segmentation by dividing the data integration process into distinct functional units: a feature value calculation unit that extracts data characteristics, an inconsistency detection unit that identifies problems, and a correction processing unit that applies fixes. This modular segmentation makes the complex detection and adjustment process manageable and maintainable.
Solution Approach 2:
The system uses an intermediary approach by introducing a correction rule storage unit that acts as a mediator between inconsistency detection and correction execution. Rather than directly implementing complex adjustment logic, the system stores predefined correction rules that bridge the gap between detecting inconsistencies and applying appropriate corrections, simplifying the overall system architecture.
Data Source
AI summary
In the present invention, there is provided a data integration method executed by a data integration system that corrects inconsistency in type or quality between first data and second data stored in a data lake. The method includes a feature value calculation step of calculating feature values of the first data and the second data, an inconsistency detection step of detecting the inconsistency between the first data and the second data based on the feature values, and a data adjustment step of correcting the inconsistency by using an integrated view according to a request to acquire the first data and the second data from a user.


