Relevant Variable Selection Using Genetic Subset Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data analysis tools are time-consuming, non-automated, and lack a unified framework for preprocessing and statistical analysis, making it difficult to identify relevant variables in heterogeneous data sets, which can lead to inconsistent results and reduced responsiveness in industrial processes.

Innovation Solution

A method and system using a variable selection module that applies genetic algorithms and selection techniques to generate and evaluate subsets of variables, ensuring relevance and reliability, and dynamically adapts to changes in data sets, suitable for large volumes of heterogeneous data in industrial processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data preprocessing and analysis tools are used, then data cleaning and exploration can be performed, but the process becomes extremely time-consuming and cumbersome

Engineering Contradiction:
Improvedata analysis reliabilityVSAvoidpreprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables automated self-service through genetic algorithms that automatically perform variable selection and subset generation without requiring manual Data Scientist intervention for each preprocessing step, thereby reducing time consumption while maintaining analysis reliability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention changes the parameter of automation level from manual to automated by implementing genetic algorithms and iterative optimization processes that systematically explore variable subsets, transforming the preprocessing workflow from artisanal to systematic and time-efficient

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If multiple analysis techniques are applied to heterogeneous data sets, then comprehensive analysis results can be obtained, but the results become inconsistent and difficult to compare

Engineering Contradiction:
Improveanalysis technique versatilityVSAvoidresult consistency
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system provides a universal framework that can handle multiple analysis techniques and heterogeneous data types through a common genetic algorithm interface, enabling consistent comparison across different techniques while maintaining versatility in the types of data and methods that can be analyzed

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The invention segments the analysis process into standardized components (variable selection, subset generation, performance evaluation) that can be consistently applied across different data sets and techniques, improving result comparability while maintaining analytical versatility

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If manual variable selection and preprocessing are performed, then relevant variables can be identified, but the process lacks automation and systematic comparison

Engineering Contradiction:
Improvevariable relevance identificationVSAvoidpreprocessing automation
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system implements feedback mechanisms through iterative genetic algorithms that evaluate variable subsets based on performance metrics, automatically refining variable selections through systematic comparison and feedback loops, thereby achieving both high automation and precise variable identification

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The invention performs preliminary automated actions by systematically generating and evaluating variable subsets before final analysis, using genetic algorithms to pre-process and rank variables, which reduces manual intervention while maintaining identification precision

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3846047A1Method and system for identifying relevant variables
Publication Date: 2021.07.07 BULL SA
  • EP3846047A1 patent drawingFigure 1
  • EP3846047A1 patent drawingFigure 2
  • EP3846047A1 patent drawingFigure 3

AI summary

The invention relates to a method for identifying relevant variables for a dataset, said variables being derived from a plurality of variables involved during the processing of the dataset, said method comprising: - a step of generating a subset of variables from the plurality of variables, - a step of assigning a quantification value to each variable in the generated subset of variables, - a step of selecting a relevant variable, - a further step of generating a new subset of variables when the quantitative value of the selected variable is less than a predetermined threshold value.