Phylogenetic Tree Vector Representations for Protein Sequence Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current bioinformatics techniques face challenges in accurately predicting nucleic acid or protein sequences and inferring evolutionary relationships using multi-sequence alignments and phylogenetic trees, as they lack effective methods for generating predictive models from phylogenetic data.

Innovation Solution

A computer-implemented method that creates a multi-sequence alignment and generates a phylogenetic tree, which is then used to train machine learning models to produce vector representations of nucleic acid or protein sequences, enabling the prediction of evolution, regression, and sibling sequences based on the phylogenetic tree.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional bioinformatics techniques are used to predict nucleic acid or protein sequences from multi-sequence alignments and phylogenetic trees, then the methods are computationally feasible, but the prediction accuracy and ability to infer evolutionary relationships are insufficient

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces vector representations as an intermediary between the phylogenetic tree data and the sequence prediction task. The machine learning models first convert phylogenetic tree data into vector representations, which then serve as input for predicting nucleic acid or protein sequences. This intermediary representation enables more accurate capture of evolutionary relationships while maintaining computational feasibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the structural phylogenetic tree data into continuous vector space representations, changing the parameter space from discrete tree structures to continuous vectors. This parameter transformation allows the application of sophisticated machine learning models that can capture nuanced evolutionary relationships, thereby improving prediction accuracy without prohibitive computational cost.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If sophisticated machine learning models are employed to generate vector representations from phylogenetic trees, then the predictive capability for evolutionary relationships improves, but the computational complexity and resource requirements increase

Engineering Contradiction:
Improveevolutionary relationship inferenceVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent segments the complex task of sequence prediction into distinct components: (1) generating vector representations from phylogenetic trees, (2) using these vectors to predict sequences, and (3) evaluating evolutionary relationships. This segmentation allows each component to be optimized independently, improving reliability while managing computational resources more efficiently.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If deep generative modeling is used to learn sequences from phylogenetic data, then the understanding of structural properties and evolutionary relationships is enhanced, but the complexity of the analysis pipeline increases

Engineering Contradiction:
Improveevolutionary information retentionVSAvoidanalysis pipeline complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical bioinformatics analysis pipelines with data-driven machine learning models. Instead of using conventional sequence alignment and phylogenetic analysis methods, the system uses neural network-based generative models that automatically learn evolutionary patterns from data, thereby reducing information loss while managing pipeline complexity through automated feature learning.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250006306A1Generative modeling and representational learning from multi-sequence alignment and phylogenetic tree data
Publication Date: 2025.01.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250006306A1 patent drawing
  • US20250006306A1 patent drawing
  • US20250006306A1 patent drawing

AI summary

Generative modeling from phylogenetic data is provided. The method comprises creating a multi-sequence alignment (MSA) based on a nucleic acid or protein sequence and generating a phylogenetic tree based on the MSA. The phylogenetic tree is fed into a number of machine learning models, which generate vector representations of the nucleic acid or protein sequences based on the phylogenetic tree. The machine learning models generate from the vector representation predicted nucleic acid or protein sequences for at least one of an evolution sequence, regression sequence, or sibling sequences of nucleic acids or proteins according to the phylogenetic tree.