Training Data Harmonization for Accurate Predictive Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive models face challenges in training due to inconsistencies and incoherencies in supervised and unsupervised training data, which affect the quality and accuracy of predictions.

Innovation Solution

A method that harmonizes a wide range of real-world training data into a single, error-free, uniformly formatted record file by cleaning and transforming data using algorithms to correct values, discern context, and remove inconsistencies, followed by building smart-agent predictive models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If real-world training data is used for predictive model training, then the quantity of training data increases, but inconsistencies and incoherencies in the data reduce prediction quality and accuracy

Engineering Contradiction:
Improvequantity of training dataVSAvoidprediction quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies preliminary data cleaning and harmonization actions before training the predictive model. The system identifies and corrects inconsistencies, incoherencies, and errors in the training data through preprocessing steps including data validation, normalization, and quality assessment, ensuring the data is ready for effective model training.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary data cleaning and harmonization layer between the raw real-world data and the predictive model training process. This intermediary component processes the data to remove inconsistencies and ensure quality, acting as a mediator that allows both large data quantity and high prediction quality to coexist.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If data cleaning and harmonization processes are applied to training data, then data quality and coherence improve, but the complexity of the data processing increases

Engineering Contradiction:
Improvedata coherenceVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the data cleaning and harmonization process into distinct modular components including data validation, normalization, quality assessment, and inconsistency resolution. Each module handles a specific aspect of data processing, making the overall complex process more manageable and maintainable while improving data coherence.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12412127B2Data clean-up method for improving predictive model training
Publication Date: 2025.09.09 BRIGHTERION INC
  • US12412127B2 patent drawing
  • US12412127B2 patent drawing
  • US12412127B2 patent drawing

AI summary

A method that improves the training of predictive models. Better trained predictive models make better predictions, and can classify transactions with reduced levels of false positives and false negative. Included is an apparatus for executing a data clean-up algorithm that harmonizes a wide range of real world supervised and unsupervised training data into a single, error-free, uniformly formatted record file that has every field coherent and well populated with information.