Training Data Harmonization for More Reliable Predictive Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive model training methods face challenges in harmonizing diverse and inconsistent supervised and unsupervised training data into a coherent, error-free format, which affects the quality and reliability of predictive models.
Innovation Solution
A data clean-up method harmonizes a wide range of real-world training data into a single, uniformly formatted record file by correcting data values and discerning context, using algorithms to ensure every field is coherent and well-populated, and employs smart-agents to build predictive models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If diverse supervised and unsupervised training data is used to improve model accuracy, then predictive model performance is enhanced, but data inconsistency and errors increase
Solution Approach 1:
The patent segments the data processing workflow into distinct phases: data collection from multiple sources, data cleaning and harmonization, and model training. By separating supervised and unsupervised data processing into distinct segments with specialized cleaning algorithms for each type, the system maintains high reliability while managing complexity of diverse data sources
Solution Approach 2:
The patent introduces an intermediary data harmonization layer between raw diverse data and the predictive model. This intermediary process includes algorithms that detect and correct inconsistencies, standardize formats, and validate data quality across different sources before feeding into the model, thus resolving the contradiction between data diversity and consistency
2Manufacturing precision
If extensive data cleaning and harmonization processes are applied to improve data quality, then data integrity is enhanced, but processing time and complexity increase
Solution Approach 1:
The patent applies preliminary data cleaning and validation actions during the data collection phase itself. By performing initial harmonization and error detection before full processing, the system reduces the burden of subsequent cleaning operations, thereby improving data integrity without proportionally increasing total processing time
Solution Approach 2:
The patent implements self-service mechanisms where the data cleaning algorithms automatically detect and correct their own errors without manual intervention. The system includes built-in validation rules and correction algorithms that operate autonomously, reducing both processing time and operational complexity while maintaining high data integrity standards
Data Source
AI summary
A method that improves the training of predictive models. Better trained predictive models make better predictions, and can classify transactions with reduced levels of false positives and false negative. Included is an apparatus for executing a data clean-up algorithm that harmonizes a wide range of real world supervised and unsupervised training data into a single, error-free, uniformly formatted record file that has every field coherent and well populated with information.


