GLL-PC Causal Discovery for Feature Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for local causal learning and Markov blanket discovery in large datasets, particularly in biomedical fields, face challenges in scalability and accuracy, making it difficult to identify direct causes and effects, and transforming datasets into minimal forms for optimal prediction.

Innovation Solution

The development of a generative method, GLL-PC, which includes an inclusion heuristic, elimination strategy, and symmetry check to identify the set of direct causes and effects, along with its variants GLL-PC-nonsym and GLL-MB, to systematically determine the Markov boundary and direct causal relationships.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If prior art methods are used for local causal learning and Markov blanket discovery, then variable/feature selection can be performed, but scalability to large datasets and accuracy in identifying direct causes and effects deteriorate

Engineering Contradiction:
Improveaccuracy in identifying direct causes and effectsVSAvoidscalability to large datasets
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the large dataset processing into local causal learning tasks focused on identifying the Markov blanket of a target variable. Instead of analyzing all variable relationships globally, the method segments the problem into finding direct causes and effects locally, which improves both accuracy in identifying causal relationships and scalability to large datasets by reducing computational complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by first identifying the Markov blanket (set of direct causes and effects) before performing classification or further analysis. This preliminary causal discovery step filters the large dataset to only relevant variables, enabling subsequent tasks to scale efficiently while maintaining high accuracy in identifying true causal relationships

Inventive Principle:
Principle #10Preliminary action

2Productivity

If datasets with many variables are transformed into minimal reduced datasets, then feature selection efficiency improves, but loss of information about potential causal relationships increases

Engineering Contradiction:
Improvefeature selection efficiencyVSAvoidinformation about potential causal relationships
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent replaces mechanical feature selection methods (which randomly or heuristically reduce features) with a causal inference-based system that uses statistical tests and causal diagrams to identify the Markov blanket. This substitution ensures that feature selection efficiency is improved while minimizing information loss about causal relationships, because the reduction is guided by causal theory rather than arbitrary criteria

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent incorporates feedback mechanisms where the identified Markov blanket is validated and refined through iterative causal inference processes. The method uses statistical feedback from the data to confirm or reject potential causal relationships, ensuring that information about true causal structures is preserved while achieving efficient feature reduction

Inventive Principle:
Principle #23Feedback

3Loss of information

If comprehensive causal relationships are discovered in large datasets, then understanding of system causation improves, but computational complexity and time requirements increase

Engineering Contradiction:
Improveunderstanding of system causationVSAvoidcomputational time for causal discovery
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent applies local quality by focusing computational resources on discovering causal relationships locally around the target variable rather than attempting to map all causal relationships in the dataset. By concentrating on the Markov blanket (direct causes and effects), the method achieves deep understanding of local causation with significantly reduced computational time compared to global causal discovery approaches

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8655821B2Local causal and Markov blanket induction method for causal discovery and feature selection from data
Publication Date: 2014.02.18 ALIFERIS KONSTANTINOS CONSTANTIN F
  • US8655821B2 patent drawing
  • US8655821B2 patent drawing
  • US8655821B2 patent drawing

AI summary

Methods for discovery of local causes/effects and of Markov blankets enable discovery of causal relationships from large data sets and provide principled solutions to the variable/feature selection problem, an integral part of predictive modeling. The present invention provides a generative method for learning local causal structure around target variables of interest in the form of direct causes/effects and Markov blankets applicable to very large real world datasets even with small samples. The selected feature sets can be used for causal discovery, classification, and regression. The generative method GLL can be instantiated in many ways giving rise to novel method variants. The method transforms a dataset with many variables into either a minimal reduced dataset where all variables are needed for optimal prediction of the response variable, or a dataset where all variables are direct causes and direct effects or the Markov blanket of the response variable.