Taxonomy-Independent Phylogenetic Placement via De Novo Tree Binning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for harmonizing microbiological data across independent studies lack robust, generalizable, and biologically meaningful features, particularly in phylogenetic placement, which is not fine-grained enough for predictive clinical models or post hoc harmonization.
Innovation Solution
A computer-implemented method and system for generating taxonomy-independent, generalizable features by receiving amplicon sequence variants, generating a de novo phylogenetic tree, and using a divide-and-conquer strategy to bin amplicon sequence variants into phylogenetically-binned groups (phylotypes) based on phylogenetic distances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional phylogenetic placement methods are used, then microbiological data can be harmonized across studies, but the precision and granularity are insufficient for predictive clinical models
Solution Approach 1:
The patent segments the phylogenetic tree into hierarchical levels (e.g., phylum, class, order, family, genus, species) and assigns features at each level independently. This segmentation allows the system to capture fine-grained phylogenetic relationships without overwhelming complexity, as each hierarchical level can be processed and validated separately, improving precision while managing computational complexity.
Solution Approach 2:
The patent introduces a new dimensional approach by creating a multi-level hierarchical feature structure that extends traditional flat phylogenetic classification. By adding hierarchical depth and creating nested feature levels, the system achieves finer granularity in phylogenetic placement without linearly increasing overall system complexity, as the hierarchical structure organizes complexity in a manageable way.
2Adaptability or versatility
If reference-set-dependent methods like cOTUs and taxonomy are used, then data harmonization is simplified, but the features lack biological meaningfulness and generalizability
Solution Approach 1:
The patent performs preliminary construction of a de novo phylogenetic tree from raw sequence data before any classification or analysis. This preliminary action creates a reference-free phylogenetic framework that is not biased by existing reference sets, enabling better biological meaningfulness and generalizability. The tree construction is done once and can be reused across multiple studies, managing the initial complexity investment.
Solution Approach 2:
The system generates its own phylogenetic features independently without relying on external reference sets. The de novo phylogenetic tree and associated features are self-generated from the input sequences themselves, making the method adaptable to any microbial community without being constrained by the scope or quality of reference databases. This self-service approach enhances generalizability across different studies and organisms.
3Reliability
If fine-grained phylogenetic features are generated, then predictive clinical models can be developed, but computational processing complexity increases
Solution Approach 1:
The patent segments the feature generation process into discrete hierarchical levels, where features are extracted and validated at each level (phylum, class, order, etc.). This segmentation allows computational processing to be distributed across multiple manageable stages rather than handling all features simultaneously, reducing the burden on computational resources while maintaining fine-grained precision for reliable predictive modeling.
Solution Approach 2:
The patent changes the parameter of phylogenetic resolution by allowing users to adjust the granularity level of feature extraction based on computational resources available. The system can operate at different hierarchical depths (coarser or finer granularity), enabling flexible adaptation between model reliability and computational complexity requirements through parameter adjustment.
Data Source
AI summary
Methods and systems for generating improved taxonomy-independent, generalizable features of alleles. The systems and methods may include (1) receiving a plurality of amplicon sequence variants corresponding to one or more microorganism communities; (2) generating a de novo phylogenetic tree representing a plurality of full-length and non-clustered alleles and the plurality of amplicon sequence variants; (3) generating a set of one or more phylogenetically-binned amplicon sequence variants (phylotypes) by a divide-and-conquer strategy; and/or (4) storing the set of phylotypes in one or more computer memories.


