Training Data Harmonization for Accurate Predictive Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive models face challenges in training due to inconsistencies and incoherencies in supervised and unsupervised training data, which affect the quality and accuracy of predictions.
Innovation Solution
A method that harmonizes a wide range of real-world training data into a single, error-free, uniformly formatted record file by cleaning and transforming data using algorithms to correct values, discern context, and remove inconsistencies, followed by building smart-agent predictive models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If real-world training data is used for predictive model training, then the quantity of training data increases, but inconsistencies and incoherencies in the data reduce prediction quality and accuracy
Solution Approach 1:
The patent applies preliminary data cleaning and harmonization actions before training the predictive model. The system identifies and corrects inconsistencies, incoherencies, and errors in the training data through preprocessing steps including data validation, normalization, and quality assessment, ensuring the data is ready for effective model training.
Solution Approach 2:
The patent introduces an intermediary data cleaning and harmonization layer between the raw real-world data and the predictive model training process. This intermediary component processes the data to remove inconsistencies and ensure quality, acting as a mediator that allows both large data quantity and high prediction quality to coexist.
2Reliability
If data cleaning and harmonization processes are applied to training data, then data quality and coherence improve, but the complexity of the data processing increases
Solution Approach 1:
The patent segments the data cleaning and harmonization process into distinct modular components including data validation, normalization, quality assessment, and inconsistency resolution. Each module handles a specific aspect of data processing, making the overall complex process more manageable and maintainable while improving data coherence.
Data Source
AI summary
A method that improves the training of predictive models. Better trained predictive models make better predictions, and can classify transactions with reduced levels of false positives and false negative. Included is an apparatus for executing a data clean-up algorithm that harmonizes a wide range of real world supervised and unsupervised training data into a single, error-free, uniformly formatted record file that has every field coherent and well populated with information.


