Automated Data Pre-Processing Templates for Adaptive ML Inputs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning-based prediction systems require manual intervention for data pre-processing, which is inefficient and cannot adapt to changing user inputs, leading to suboptimal performance.

Innovation Solution

An automated system that generates pre-processing templates based on historical data sets, calculates matching scores for input data, and selects the appropriate template for processing, ensuring consistent and adaptive data formatting for machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual data pre-processing is used, then data can be processed with human judgment and flexibility, but the process is inefficient and cannot adapt quickly to changing user inputs

Engineering Contradiction:
Improveadaptability to changing user inputsVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system dynamically adapts to changing user inputs by learning from historical pre-processing operations. The machine learning model continuously updates its understanding of pre-processing requirements based on new data patterns, enabling the system to adjust its behavior without manual reconfiguration while maintaining high processing efficiency through automated operations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs self-service by automatically learning pre-processing patterns from historical data and applying them to new inputs without requiring continuous manual intervention. The machine learning model autonomously improves its pre-processing capabilities by analyzing past operations and adapting to new data formats, reducing dependency on manual expert intervention while maintaining high adaptability.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If manual data pre-processing is used, then complex data transformations can be performed with expert judgment, but the process is time-consuming and labor-intensive

Engineering Contradiction:
Improvedata pre-processing qualityVSAvoidpre-processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by learning pre-processing patterns from historical data in advance. The machine learning model is trained on past pre-processing operations, capturing expert knowledge and transformation patterns beforehand. This allows the system to quickly apply learned patterns to new data without requiring real-time manual expert intervention, maintaining high pre-processing quality while significantly reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces the mechanical system of manual expert intervention with an automated machine learning-based approach. The ML model captures and replicates expert pre-processing knowledge, substituting human manual operations with automated algorithms that maintain high data transformation quality while eliminating the time-consuming nature of manual processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated pre-processing is implemented, then processing efficiency is improved, but the system may lack the flexibility to handle novel data formats without retraining

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidhandling of novel data formats
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system implements feedback mechanisms where the machine learning model continuously learns from new data patterns and pre-processing outcomes. By incorporating feedback from historical operations and performance metrics, the model adapts to novel data formats over time, maintaining both high processing efficiency and flexibility. The feedback loop enables the system to refine its understanding of new data types while preserving the benefits of automated processing.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12423937B2Automated data pre-processing for machine learning
Publication Date: 2025.09.23 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12423937B2 patent drawing
  • US12423937B2 patent drawing
  • US12423937B2 patent drawing

AI summary

Computer-implemented methods for performing automated pre-processing of data for a machine-learning based prediction system are provided. Aspects include receiving a plurality of raw data sets, receiving a plurality of processed data sets, wherein each of the plurality of processed data sets corresponds to one of the plurality of raw data sets, and generating a plurality of pre-processing templates based on the plurality of raw data sets and the processed data set. Aspects also include receiving an input data set, generating, for each of the plurality of pre-processing templates, a matching score for the input data set, and selecting one of the plurality of pre-processing templates based on the matching score. Aspects further include pre-processing the input data set using the selected pre-processing template and providing the pre-processed input data set to the machine learning based prediction system.