ML Model Selection via Data Redundancy Elimination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning (ML) models are often rigidly tied to specific cloud platforms, making data integration and predictive analysis across hybrid or multi-cloud deployments challenging, and require expertise to choose the right ML algorithm, which can lead to sub-optimal choices due to complexity and data variability issues.
Innovation Solution
A system comprising a processor, data lake, data analyzer, and model selector and evaluator that ingests data from a data lake, processes it to eliminate redundant attributes, and selects an optimal ML model for predictive analysis, allowing for continuous training and validation based on performance metrics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple ML models are evaluated and selected based on performance metrics, then prediction accuracy is improved, but system complexity and computation time increase
Solution Approach 1:
The system performs preliminary evaluation of multiple ML models using performance metrics before deployment. The model evaluator assesses candidate models in advance, selecting the optimal model based on predetermined criteria, which resolves the complexity issue by pre-determining the best model rather than requiring complex real-time selection
Solution Approach 2:
The system automatically evaluates and selects ML models based on performance metrics without requiring manual intervention. The model evaluator self-manages the selection process by comparing models against performance criteria and automatically identifying the optimal model, reducing operational complexity
2Measurement precision
If data is transformed and redundant attributes are eliminated, then ML model performance is improved, but data processing time increases
Solution Approach 1:
The data transformer extracts and removes redundant attributes from the input data, keeping only the essential features needed for ML model training. This selective extraction improves model performance by eliminating noise while the automated nature of the process manages the time trade-off efficiently
Solution Approach 2:
The transformation process applies different processing operations to different attributes based on their specific characteristics. Rather than uniformly processing all data, the system identifies and applies appropriate transformations only where needed, optimizing the balance between performance improvement and processing time
3Measurement precision
If ML models are continuously trained and validated, then model accuracy is improved, but computational resources and time are consumed
Solution Approach 1:
The system implements periodic training and validation of ML models rather than continuous operation. Models are retrained at scheduled intervals or when triggered by specific conditions, allowing computational resources to be used efficiently while maintaining model accuracy through regular updates
Solution Approach 2:
The model evaluator provides feedback on model performance, which triggers retraining only when performance degradation is detected or improvement is potential. This feedback-driven approach optimizes computational resource usage by avoiding unnecessary training cycles while maintaining model accuracy
4Quantity of substance
If data from multiple cloud platforms is integrated, then data comprehensiveness is improved, but integration complexity increases
Solution Approach 1:
The data lake implements a universal data storage and management interface that can accommodate data from multiple cloud platforms through standardized protocols. This universal interface simplifies integration by providing a common access point regardless of the source platform, reducing integration complexity while maintaining data comprehensiveness
Data Source
AI summary
The present disclosure relates to systems and methods for carrying out predictive analysis where a plurality of data sets may be ingested from a data lake. A data analyzer may tag the ingested data sets, detect redundant occurrence of multiple attributes such as, a row, a column, and a list in the tagged data set. The data analyzer may eliminate the detected redundant multiple attributes. Further, a model selector and evaluator may execute a machine learning (ML) model to conduct predictive analysis on the data set. The execution may be done based on a predefined set of instructions stored in a database. The executed ML model may be validated upon determining that the predictive analysis yields a positive response for the transformed data set.


