Neural Network Data Quality Prediction via Metadata Profiles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Poor data quality, often resulting from unstructured or inconsistent data, hinders the benefits of big data analytics and management, requiring efficient techniques to improve data quality without consuming excessive time and resources.
Innovation Solution
The use of artificial intelligence, specifically a trained neural network, to process incoming data, extract features, and generate metadata profiles, which are then used to predict data categories, thereby improving data organization and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional techniques are used to improve data quality, then data quality can be improved, but a large amount of time and resources are consumed
Solution Approach 1:
The patent replaces manual data quality assessment and categorization processes with an automated machine learning system. The system uses trained models to automatically evaluate data quality metrics and categorize data, substituting human mechanical processes with computational algorithms that operate faster and more consistently.
Solution Approach 2:
The system enables data to be automatically assessed and categorized without requiring manual intervention. The machine learning models self-evaluate data quality by analyzing metadata profiles and automatically assign categories, allowing the system to serve itself rather than relying on external human resources.
2Stability of the object's composition
If manual data categorization is performed, then data organization can be improved, but specialized human resources are required
Solution Approach 1:
The patent replaces manual data categorization performed by specialized human resources with automated machine learning algorithms. The system extracts features from data, generates metadata profiles, and uses trained models to automatically categorize data, eliminating the need for human experts while maintaining or improving categorization consistency.
Solution Approach 2:
The system creates metadata profiles that are simplified representations or copies of the actual data characteristics. These metadata profiles capture essential features and patterns, allowing the system to categorize data based on these copies rather than analyzing the full complexity of the original data, thereby reducing resource requirements.
Data Source
AI summary
Embodiments improve data quality using artificial intelligence. Incoming data that includes a plurality of rows of data and a trained neural network that is configured to predict a data category for the incoming data can be received, where the neural network has been trained with training data including training features, and the training data includes labeled data categories. The incoming data can be processed, where the processing extracts features about the plurality of rows of data to generate metadata profiles that represent the incoming data. Using the trained neural network, a data category for the incoming data can be predicted, where the prediction is based on the generated metadata profiles.


