Automated Machine Learning Data Preparation and Model Tuning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in efficiently utilizing vast amounts of data for organizations to improve business practices through machine learning models, which often requires time-consuming and resource-intensive processes prone to manual errors and subjective distortions.
Innovation Solution
The system automates data preparation, training, and tuning of machine learning models by generating code blocks, labels, and feature records based on domain data and schema information, using ontologies and generative AI to reduce bias and variability, and segmenting data for training, evaluation, and testing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual operations are used to prepare and train machine learning models, then flexibility and adaptability are maintained, but time consumption and resource intensity increase significantly
Solution Approach 1:
The system enables self-service automation where the machine learning platform automatically performs data preparation, feature engineering, model training, and evaluation without requiring manual intervention. The automated machine learning system serves itself by autonomously selecting algorithms, tuning hyperparameters, and optimizing models based on predefined business objectives and data characteristics.
Solution Approach 2:
The system changes key parameters of the machine learning process from manual control to automated control. It transforms discrete manual operations into continuous automated processes, changing the state of model development from ad-hoc to systematic, thereby increasing productivity while managing complexity through standardized parameters.
2Reliability
If manual operations are used for model evaluation, then subjective judgment can be applied, but errors and distortions increase due to human bias
Solution Approach 1:
The system implements feedback mechanisms where model evaluation results automatically feed back into the model training process. Performance metrics are continuously monitored and used to adjust model parameters, select algorithms, and optimize features, creating a closed-loop system that improves reliability through iterative automated feedback rather than static manual evaluation.
Solution Approach 2:
The system replaces manual mechanical evaluation processes with automated computational evaluation mechanisms. Instead of human analysts manually assessing model performance, the system uses automated algorithms to evaluate models against predefined criteria, eliminating human bias while maintaining evaluation accuracy through consistent, reproducible computational assessments.
3Quantity of substance
If vast amounts of data are collected, then more comprehensive insights can be gained, but the difficulty of effective utilization increases
Solution Approach 1:
The system segments vast amounts of data into manageable, meaningful portions through automated feature engineering and data categorization. It divides raw data into relevant features, groups related information, and organizes data streams into structured formats that automated algorithms can process efficiently, reducing the difficulty of detecting and measuring patterns in large datasets.
Solution Approach 2:
The system introduces automated intermediaries between raw data and model training. These intermediaries include automated data cleaning modules, feature extraction algorithms, and data transformation layers that process and prepare vast amounts of data, making the data more accessible and easier to utilize effectively without requiring manual intervention.
Data Source
AI summary
Embodiments are directed to managing machine learning models. Domain items may be determined based on domain data and schema information. Labels that correspond to a predicted outcome may be generated based on the domain data. A model may be trained based on a portion of a plurality of feature records and the labels such that each feature record may be associated with an observance of a domain item. The trained model may be disqualified based on evaluation metrics that may be below a threshold value causing further actions, including: submitting other portions of the feature records to the disqualified model; determining erroneous feature fields in the feature records based on metrics associated with the submission of the other portions of feature records; updating the feature records to exclude the erroneous feature fields; retraining the disqualified model based on the updated feature records.


