Taxonomy-Independent Phylogenetic Placement via De Novo Tree Binning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for harmonizing microbiological data across independent studies lack robust, generalizable, and biologically meaningful features, particularly in phylogenetic placement, which is not fine-grained enough for predictive clinical models or post hoc harmonization.

Innovation Solution

A computer-implemented method and system for generating taxonomy-independent, generalizable features by receiving amplicon sequence variants, generating a de novo phylogenetic tree, and using a divide-and-conquer strategy to bin amplicon sequence variants into phylogenetically-binned groups (phylotypes) based on phylogenetic distances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional phylogenetic placement methods are used, then microbiological data can be harmonized across studies, but the precision and granularity are insufficient for predictive clinical models

Engineering Contradiction:
Improvephylogenetic placement precisionVSAvoidfeature generation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the phylogenetic tree into hierarchical levels (e.g., phylum, class, order, family, genus, species) and assigns features at each level independently. This segmentation allows the system to capture fine-grained phylogenetic relationships without overwhelming complexity, as each hierarchical level can be processed and validated separately, improving precision while managing computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimensional approach by creating a multi-level hierarchical feature structure that extends traditional flat phylogenetic classification. By adding hierarchical depth and creating nested feature levels, the system achieves finer granularity in phylogenetic placement without linearly increasing overall system complexity, as the hierarchical structure organizes complexity in a manageable way.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If reference-set-dependent methods like cOTUs and taxonomy are used, then data harmonization is simplified, but the features lack biological meaningfulness and generalizability

Engineering Contradiction:
Improvefeature generalizabilityVSAvoidde novo phylogenetic tree construction
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary construction of a de novo phylogenetic tree from raw sequence data before any classification or analysis. This preliminary action creates a reference-free phylogenetic framework that is not biased by existing reference sets, enabling better biological meaningfulness and generalizability. The tree construction is done once and can be reused across multiple studies, managing the initial complexity investment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system generates its own phylogenetic features independently without relying on external reference sets. The de novo phylogenetic tree and associated features are self-generated from the input sequences themselves, making the method adaptable to any microbial community without being constrained by the scope or quality of reference databases. This self-service approach enhances generalizability across different studies and organisms.

Inventive Principle:
Principle #25Self-service

3Reliability

If fine-grained phylogenetic features are generated, then predictive clinical models can be developed, but computational processing complexity increases

Engineering Contradiction:
Improvepredictive model reliabilityVSAvoidcomputational processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the feature generation process into discrete hierarchical levels, where features are extracted and validated at each level (phylum, class, order, etc.). This segmentation allows computational processing to be distributed across multiple manageable stages rather than handling all features simultaneously, reducing the burden on computational resources while maintaining fine-grained precision for reliable predictive modeling.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of phylogenetic resolution by allowing users to adjust the granularity level of feature extraction based on computational resources available. The system can operate at different hierarchical depths (coarser or finer granularity), enabling flexible adaptation between model reliability and computational complexity requirements through parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250037793A1Phylogenetic placement using taxonomy-independent feature generation
Publication Date: 2025.01.30 THE RGT UNIV OF MICHIGAN
  • US20250037793A1 patent drawing
  • US20250037793A1 patent drawing
  • US20250037793A1 patent drawing

AI summary

Methods and systems for generating improved taxonomy-independent, generalizable features of alleles. The systems and methods may include (1) receiving a plurality of amplicon sequence variants corresponding to one or more microorganism communities; (2) generating a de novo phylogenetic tree representing a plurality of full-length and non-clustered alleles and the plurality of amplicon sequence variants; (3) generating a set of one or more phylogenetically-binned amplicon sequence variants (phylotypes) by a divide-and-conquer strategy; and/or (4) storing the set of phylotypes in one or more computer memories.