Relevant Variable Selection Using Genetic Subset Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data analysis tools are time-consuming, non-automated, and lack a unified framework for preprocessing and statistical analysis, making it difficult to identify relevant variables in heterogeneous data sets, which can lead to inconsistent results and reduced responsiveness in industrial processes.
Innovation Solution
A method and system using a variable selection module that applies genetic algorithms and selection techniques to generate and evaluate subsets of variables, ensuring relevance and reliability, and dynamically adapts to changes in data sets, suitable for large volumes of heterogeneous data in industrial processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data preprocessing and analysis tools are used, then data cleaning and exploration can be performed, but the process becomes extremely time-consuming and cumbersome
Solution Approach 1:
The system enables automated self-service through genetic algorithms that automatically perform variable selection and subset generation without requiring manual Data Scientist intervention for each preprocessing step, thereby reducing time consumption while maintaining analysis reliability
Solution Approach 2:
The invention changes the parameter of automation level from manual to automated by implementing genetic algorithms and iterative optimization processes that systematically explore variable subsets, transforming the preprocessing workflow from artisanal to systematic and time-efficient
2Adaptability or versatility
If multiple analysis techniques are applied to heterogeneous data sets, then comprehensive analysis results can be obtained, but the results become inconsistent and difficult to compare
Solution Approach 1:
The system provides a universal framework that can handle multiple analysis techniques and heterogeneous data types through a common genetic algorithm interface, enabling consistent comparison across different techniques while maintaining versatility in the types of data and methods that can be analyzed
Solution Approach 2:
The invention segments the analysis process into standardized components (variable selection, subset generation, performance evaluation) that can be consistently applied across different data sets and techniques, improving result comparability while maintaining analytical versatility
3Measurement precision
If manual variable selection and preprocessing are performed, then relevant variables can be identified, but the process lacks automation and systematic comparison
Solution Approach 1:
The system implements feedback mechanisms through iterative genetic algorithms that evaluate variable subsets based on performance metrics, automatically refining variable selections through systematic comparison and feedback loops, thereby achieving both high automation and precise variable identification
Solution Approach 2:
The invention performs preliminary automated actions by systematically generating and evaluating variable subsets before final analysis, using genetic algorithms to pre-process and rank variables, which reduces manual intervention while maintaining identification precision
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method for identifying relevant variables for a dataset, said variables being derived from a plurality of variables involved during the processing of the dataset, said method comprising: - a step of generating a subset of variables from the plurality of variables, - a step of assigning a quantification value to each variable in the generated subset of variables, - a step of selecting a relevant variable, - a further step of generating a new subset of variables when the quantitative value of the selected variable is less than a predetermined threshold value.