Phylogenetic Generative Models for Biological Association Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Biological data with phylogenetic structure poses challenges for statistical analysis, as existing methods assume independence and identically distributed data, leading to confounding effects and inaccurate identification of associations.

Innovation Solution

The use of generative models that account for phylogenetic structure, such as conditional and directed joint models, to determine associations between variables, which include learning the phylogenetic tree simultaneously with statistical models, and employing frequentist, Bayesian, and cross-validation techniques to assess the strength of associations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If standard statistical methods assuming independence are used, then the analysis is simple, but the identification of associations is inaccurate due to confounding by phylogeny

Engineering Contradiction:
Improveaccuracy of association identificationVSAvoidcomplexity of statistical model
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a phylogenetic tree as an intermediary structure that mediates between the observed data and the independence assumption. By modeling the data generation process through the phylogenetic tree, the method accounts for evolutionary relationships while maintaining a tractable statistical framework. This resolves the contradiction by providing accurate association identification without requiring overly complex models.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts and removes the phylogenetic confounding effect from the statistical analysis by explicitly modeling it separately. Through techniques such as phylogenetic correction and conditional independence testing, the method isolates the phylogenetic signal from the association signal, allowing accurate identification of true associations while simplifying the overall analytical approach.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If phylogenetic structure is accounted for using generative models, then false discovery rates are well-calibrated and discriminatory power increases, but the computational complexity increases

Engineering Contradiction:
Improvecalibration of false discovery ratesVSAvoidcomplexity of generative models
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex generative modeling process into distinct, manageable components: (1) phylogenetic tree construction, (2) parameter estimation for each variable, (3) association testing through conditional models. This segmentation allows the method to achieve well-calibrated false discovery rates through systematic modeling while keeping computational complexity tractable through modular implementation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs dynamic modeling approaches where the phylogenetic tree and model parameters are estimated simultaneously with the association structure. This dynamic, iterative process allows the model to adapt to the data while maintaining computational efficiency through algorithms such as expectation-maximization and cross-validation, resolving the contradiction between reliability and complexity.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If multiple predictor variables are analyzed with directional constraints, then the network of dependencies can be represented as a directed acyclic graph, but the model fitting becomes more complex

Engineering Contradiction:
Improveability to represent directional relationshipsVSAvoidcomplexity of directed acyclic graph modeling
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary constraints to define the directionality of relationships before full model fitting. By specifying which variables are predictors and which are targets a priori, and by constraining the model to produce a directed acyclic graph structure, the method enables versatile representation of directional relationships while simplifying the model fitting process through reduced search space and pre-defined structural constraints.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8050870B2Identifying associations using graphical models
Publication Date: 2011.11.01 ZHIGU HLDG
  • US8050870B2 patent drawing
  • US8050870B2 patent drawing
  • US8050870B2 patent drawing

AI summary

Statistical models for identifying associations are described herein. By way of example, a system for identifying associations between variables can include a model builder and an association identifier. The model builder can receive observations about the variables and generate a null model and a non-null model. The association identifier can assess the strength of the association between the variables by determining how much the non-null model better explains the observed data than the null model. Additionally or alternatively, the structure of the observed data can be inferred simultaneously with the statistical model.