Biomolecule Sequence Coevolution Analysis for Variant Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting the phenotypic effects of genetic variants in biomolecules, such as proteins, are limited by their inability to distinguish between evolutionary relationships due to functional adaptation versus non-functional genetic drift, making it difficult to accurately predict the functional relationships between residues and the phenotypic consequences of variants.

Innovation Solution

A computer-implemented method that identifies functionally-related residues in biomolecules by using multiple sequence alignment data, computing covariation values, defining clades based on phylogenetic cutoffs, building covariation matrices, and applying dimensionality reduction techniques to predict the phenotypic consequences of genetic variants through machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional coevolution methods are used to measure relationships between equivalent positions in biomolecules, then evolutionary relationships can be detected, but the methods cannot distinguish between functional adaptation and non-functional genetic drift, resulting in inaccurate prediction of phenotypic effects

Engineering Contradiction:
Improveaccuracy of phenotypic effect predictionVSAvoidfunctional relationship information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The method segments the evolutionary signal by dividing it into two distinct components: intra-clade covariation (functional adaptation) and inter-clade covariation (genetic drift). By computing these separately using phylogenetic tree structure, the method isolates the functional signal from noise, enabling accurate prediction of phenotypic effects while preserving functional relationship information.

Inventive Principle:
Principle #1Segmentation

2Reliability

If multiple sequence alignment data is analyzed to predict phenotypic effects, then evolutionary relationships can be identified, but the complexity of distinguishing functional vs. non-functional relationships increases computational difficulty

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational method complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The phylogenetic tree serves as an intermediary structure that mediates the analysis of multiple sequence alignment data. By using the tree topology to guide the computation of intra-clade and inter-clade covariation, the method simplifies the complex task of distinguishing functional from non-functional relationships into a structured computational process based on phylogenetic distances.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If coevolution analysis is performed without considering phylogenetic structure, then computational simplicity is maintained, but the ability to distinguish functional adaptation from genetic drift is lost

Engineering Contradiction:
Improvecomputational simplicityVSAvoidfunctional relationship detection accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The method introduces dynamic consideration of phylogenetic structure by adapting the analysis to the specific topology and branch lengths of the phylogenetic tree. Rather than using a static, one-size-fits-all approach, the computational method dynamically adjusts to the evolutionary relationships in the data, maintaining simplicity while improving functional relationship detection accuracy.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10886007B2Methods and systems for identification of biomolecule sequence coevolution and applications thereof
Publication Date: 2021.01.05 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • US10886007B2 patent drawing
  • US10886007B2 patent drawing
  • US10886007B2 patent drawing

AI summary

Generation of biomolecule sequence coevolution data structures, matrices, scores, and sectors are described. Generally, the generated coevolution data removes covariant noise due to phylogenetic drift and can reveal coevolution of residue positions in multiple phylogenetic distances. Scores can be built upon the data structures and matrices to reveal sectors of residue positions that function and evolve together. Furthermore, the coevolution data structures, matrices, scores, and sectors can be used to predict structure or function of residue variants.