Data Veracity Correction Using NLP and ML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current enterprise data management systems are unable to enhance data veracity scores, leading to wastage of storage and computational resources due to inadequate means to improve the quality of stored data for insights generation.
Innovation Solution
A system comprising a retriever, profiler, veracity generator, corrector, and recommender that analyzes data, identifies anomalies, and applies optimal correction techniques using machine learning models and natural language processing to enhance data veracity scores, thereby improving the usefulness of the data for insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is stored in the repository for future insights generation, then data availability is improved, but data veracity deteriorates over time leading to resource wastage
Solution Approach 1:
The system performs preliminary actions by continuously monitoring data veracity metrics and applying corrections before the data is used for insights generation. The automated correction system proactively identifies and rectifies veracity issues, preventing the data from becoming completely unusable and avoiding wastage of computational resources on low-quality data.
Solution Approach 2:
The system implements feedback mechanisms by continuously evaluating data veracity scores and using this information to trigger appropriate correction actions. The system monitors veracity metrics, compares them against thresholds, and automatically initiates correction processes when degradation is detected, creating a closed-loop system that maintains data quality over time.
2Measurement precision
If data veracity scoring is implemented to ensure data quality, then insights accuracy is improved, but system complexity increases due to lack of enhancement capabilities
Solution Approach 1:
The system segments the data quality management function into distinct modular components: veracity scoring modules that assess different aspects of data quality, correction modules that apply specific remediation techniques, and orchestration layers that coordinate their interaction. This modular architecture reduces overall system complexity by allowing each component to be independently developed, maintained, and optimized.
Solution Approach 2:
The system introduces intermediary components that act as mediators between raw data and the insights generation process. These intermediaries include veracity scorers that evaluate data quality and correction mechanisms that remediate issues, serving as a buffer layer that protects the core insights engine from dealing with raw, potentially problematic data directly.
3Quantity of substance
If data is stored without verification for resource efficiency, then storage costs are reduced, but data usability deteriorates leading to futile insights gathering
Solution Approach 1:
The system applies partial verification by selectively assessing data veracity based on predefined criteria and thresholds rather than performing exhaustive verification on all data. This approach balances storage efficiency with usability by verifying only the necessary aspects of data quality that impact insights generation, avoiding excessive verification overhead while preventing futile processing of severely degraded data.
Data Source
AI summary
Examples for enhancing veracity of data are described herein. Data from a repository may be received based on a data receiving rule. From the received data, a first dataset may be generated using statistical modeling. Also, a first data veracity score for the first dataset is generated which is indicative of a degree of usability of the dataset. Another aspect relates to identifying an anomaly in the first dataset, the corrector, for each anomaly, to identify a correction technique from amongst a plurality of correction techniques. Further, a second dataset is generated using the identified correction technique having second data veracity score higher than the first data veracity score.


