Probabilistic Predictor for Phenotype Association Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Early genome-wide association studies (GWAS) often miss associations where multiple single-nucleotide polymorphisms (SNPs) have a mild influence on phenotypes, as finding a robust aggregation function to quantify the relationship between sets of SNPs and phenotypes has been elusive.

Innovation Solution

A probabilistic predictor is used to summarize the relationship between sets of biological predictors and phenotypes, employing functions like L1-regularized logistic regression, L1-regularized softmax, or L1-regularized linear regression based on the phenotype type, and trained on a portion of the data for application to another portion, facilitating genome-wide association analysis and gene-set enrichment analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional GWAS methods are used to focus on one or a small number of SNPs, then the analysis is simple and computationally efficient, but associations where multiple SNPs have mild influence are missed

Engineering Contradiction:
Improvedetection capabilityVSAvoidaggregation function complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent changes the mathematical parameters by using L1-regularized regression to create a sparse aggregation function that selects important SNPs while handling multiple predictors. This allows the system to detect associations where multiple SNPs have mild influence without overwhelming computational complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary aggregation function that summarizes the relationship between sets of SNPs and phenotypes. This mediator layer transforms complex multi-SNP relationships into interpretable probability distributions, enabling detection of subtle associations while maintaining computational efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a robust aggregation function is developed to quantify relationships between sets of SNPs and phenotypes, then detection precision improves, but the complexity of the system increases

Engineering Contradiction:
Improveassociation detection precisionVSAvoidpredictor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies L1-regularization to change the optimization parameters, creating a sparse solution that identifies important SNPs while keeping the model interpretable. This parameter change enables high measurement precision without proportional increases in system complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the analysis by using gene sets as predefined groups of SNPs, allowing the aggregation function to operate on biologically meaningful units rather than individual SNPs. This segmentation improves detection precision while managing complexity through biological prior knowledge

Inventive Principle:
Principle #1Segmentation

3Reliability

If the probabilistic predictor is trained on a portion of the data and applied to another portion, then the reliability of phenotype prediction improves, but the time required for analysis increases

Engineering Contradiction:
Improveprediction reliabilityVSAvoidtraining and analysis time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by training the probabilistic predictor on a portion of the data before applying it to the target population. This preliminary training establishes reliable prediction models that can be efficiently applied to new data, improving reliability while minimizing repeated computation time

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8315957B2Predicting phenotypes using a probabilistic predictor
Publication Date: 2012.11.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8315957B2 patent drawing
  • US8315957B2 patent drawing
  • US8315957B2 patent drawing

AI summary

Aspects of the subject matter described herein relate to predicting phenotypes. In aspects, a probabilistic predictor is used to summarize a relationship between a set of biological predictors and a phenotype. The probabilistic predictor may use a function that is selected based on the type of the phenotype (e.g., binary, multi-state, or continuous). The probabilistic predictor may use genetic and/or epigenetic information. The probabilistic predictor may be trained on a portion of the data in conjunction with predicting phenotypes in another portion of the data. The probabilistic predictor may be used for various analyses including genome-wide association analysis and gene-set enrichment analysis.