Causal Discovery in Mixed Datasets via Hybrid Discretization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing algorithms struggle to determine causal relationships in mixed datasets containing both continuous and discrete variables, as they either require homogenous data or are inefficient in processing non-homogeneous data, leading to loss of information and errors.

Innovation Solution

A multi-phase hybrid approach that uses constraint-based algorithms to establish dependency, followed by data-driven discretization of continuous variables, and then employs a scoring function to identify a directed graph that preserves causal relationships, utilizing algorithms like Fast Conditional Independence Test and Fast Greedy Equivalence Search.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If statistical approaches are used to determine causality in mixed datasets, then causality can be determined while controlling for confounding influences, but the approaches struggle to properly determine causality amongst variables when the underlying dataset contains data related to continuous variables and discrete variables

Engineering Contradiction:
Improvecausality determination accuracyVSAvoidhandling mixed data types
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the causal discovery process into distinct phases: constraint-based phase for structure learning and score-based phase for parameter estimation. This segmentation allows each phase to be optimized for its specific task, with the constraint-based phase handling structural relationships and the score-based phase refining causal directions, thereby improving overall reliability for mixed datasets

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation by introducing a hybrid scoring function that combines constraints and scores, and by using different probability distributions for continuous and discrete variables. This parameter adaptation enables the system to handle mixed data types effectively while maintaining causality determination accuracy

Inventive Principle:
Principle #35Parameter changes

2Productivity

If existing algorithms are applied to mixed datasets, then processing can be performed, but information loss and errors occur due to the algorithms being designed for homogenous data

Engineering Contradiction:
Improveprocessing capabilityVSAvoidcausal relationship accuracy
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent creates a universal causal discovery framework that can handle both homogeneous and heterogeneous datasets. The hybrid approach unifies constraint-based and score-based methods, and the discretization module adapts to work with both continuous and discrete variables, making the system multi-functional without losing causal relationship accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary discretization step that transforms continuous variables into discrete representations. This intermediary process enables existing discrete-variable algorithms to process mixed datasets effectively while preserving causal relationships, thereby maintaining both productivity and information integrity

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If discretization is applied to continuous variables, then the data can be processed by algorithms designed for discrete variables, but information loss occurs during the discretization process

Engineering Contradiction:
Improvealgorithm compatibilityVSAvoidcontinuous variable detail
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent applies preliminary discretization to continuous variables before the main causal discovery process. This preliminary action transforms continuous data into discrete categories that are compatible with constraint-based algorithms, enabling algorithm compatibility while minimizing information loss through careful discretization strategy selection

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12008589B2Discovering causal relationships in mixed datasets
Publication Date: 2024.06.11 ADOBE INC
  • US12008589B2 patent drawing
  • US12008589B2 patent drawing
  • US12008589B2 patent drawing

AI summary

Introduced here are approaches to determining causal relationships in mixed datasets containing data related to continuous variables and discrete variables. To accomplish this, a marketing insight and intelligence platform may employ a multi-phase approach in which dependency is established before the data related to continuous variables is discretized. Such an approach ensures that information regarding dependence is not lost through discretization.