Data Observation Metrics for Multi-Platform Database Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data sets extracted from multiple data platforms contain errors such as incorrect values, missing data, and duplicates, leading to reduced accuracy and complexity in integrated databases due to fluctuating data amounts and types, making it difficult to maintain data accuracy and availability.
Innovation Solution
An information processing apparatus and method that extracts data sets from multiple platforms, analyzes quality using dynamic metrics, validates data items, and constructs metrics based on processing history to improve data accuracy and availability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If uniform and unchanged error evaluation criteria are used to filter errors in data sets, then system complexity is reduced and ease of operation is improved, but data accuracy deteriorates because the criteria cannot follow high-frequency changes in data sources
Solution Approach 1:
The patent applies dynamics by making the error evaluation criteria changeable and adaptable. The system allows error criteria to be dynamically adjusted based on data source characteristics and changes, transforming static uniform criteria into dynamic adaptive criteria that can follow high-frequency changes in data sources while maintaining ease of operation through automated adjustment mechanisms
Solution Approach 2:
The patent changes the parameters of error evaluation criteria to match different data sources and their characteristics. By allowing criteria parameters to be modified based on data source types, update frequencies, and error patterns, the system maintains data accuracy without requiring complex manual configuration for each scenario
2Manufacturing precision
If error evaluation criteria are dynamically adjusted to reflect changes in multiple data platforms, then data accuracy is improved, but system complexity increases and maintenance difficulty arises
Solution Approach 1:
The patent applies self-service by enabling the error evaluation system to automatically adapt to data source changes without requiring manual intervention. The system self-configures criteria based on observed data patterns, automatically updates evaluation parameters, and maintains accuracy through autonomous adjustment, reducing the complexity burden on operators
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously monitors data source changes and error patterns, then uses this feedback to automatically adjust error evaluation criteria. This closed-loop approach allows the system to maintain data accuracy by responding to changes while keeping complexity managed through automated feedback-driven adjustment
3Manufacturing precision
If comprehensive error filtering is applied to all data sets from multiple platforms, then data accuracy is improved, but processing time increases due to the need to monitor and maintain criteria for the entire integrated database
Solution Approach 1:
The patent applies segmentation by dividing the error filtering process into targeted segments based on data source characteristics, error types, and risk levels. Instead of uniformly applying comprehensive filtering to all data, the system segments filtering intensity and criteria application to specific data sources or error categories, maintaining data accuracy while reducing overall processing time through selective focused filtering
Data Source
AI summary
There is provided an information processing apparatus capable of making the accuracy and behavior of data observable when constructing and updating a database by aggregating data from multiple data sources so as to improve data accuracy and availability. The apparatus includes a data set extraction unit configured to extract a data set from a plurality of databases belonging to a plurality of platforms, respectively, a quality analysis unit configured to analyze quality of the data set per data set by applying a first metric to the data set; a data validation unit configured to validate a data value per data item in the data set by applying a second metric to the data set; and a metric construction unit configured to dynamically construct at least a part of the first and second metrics to be applied to the data set based on a processing history of the data set.


