Data Mapping via Machine Learning Classifier and User Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in analyzing data stored in legacy databases that are not compliant with common data models, leading to high costs and resource wastage due to the inability to infer the meaning of data effectively.
Innovation Solution
A machine-intelligent solution using machine learning and metadata analysis to infer the meaning of data, with user input for confirmation or modification when confidence is low, enabling data mapping into a compatible model for analytics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is stored in a legacy database with non-compliant schema, then data storage capacity is maintained, but data analysis capability is lost
Solution Approach 1:
The patent introduces an intermediary system comprising a data classifier and mapping specification generator that bridges legacy databases and modern data analytics platforms. The data classifier analyzes legacy schema structures and generates mapping specifications that translate non-compliant data into compliant formats, enabling both data retention and analytical capability without direct schema modification.
Solution Approach 2:
The system transforms the schema compatibility parameter by dynamically generating mapping specifications that convert legacy schema parameters into modern data model parameters. This parameter transformation enables data from diverse legacy systems with varying schema structures to be universally analyzed without losing either storage capacity or analytical capability.
2Speed
If automatic data classification is performed without user input, then processing speed is improved, but classification accuracy deteriorates
Solution Approach 1:
The system applies partial automation by performing automatic data classification for high-confidence cases while selectively invoking user input requests for low-confidence classifications. This partial action approach maintains high processing speed for straightforward cases while ensuring accuracy for ambiguous cases, avoiding the extremes of complete automation or manual classification.
Solution Approach 2:
The system incorporates feedback mechanisms where user inputs on classified data are captured and used to refine the data classifier's future classifications. This feedback loop continuously improves classification accuracy while maintaining processing speed, as the system learns from user corrections and reduces the frequency of manual interventions over time.
Data Source
AI summary
Methods, apparatuses, and systems for improving data mapping are provided. An example method may include retrieving a first plurality of data objects associated with a first database schema from a database, determining a first data classifier corresponding to the first database schema, generating a mapping specification based at least in part on the first data classifier and the first plurality of data objects, and generating a second plurality of data objects based at least in part on the first plurality of data objects and the mapping specification.


