Identifying microbiome constituents
A computer-implemented method for sequencing and analyzing vaginal samples addresses the limitations of current diagnostics by providing a comprehensive, data-driven approach for accurate identification and classification of vaginal infections through taxonomic and metabolomic analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-04-02
AI Technical Summary
Current diagnostic methods for vaginal infections, such as bacterial vaginosis, are subjective, technician-dependent, and lack comprehensive analysis of microbial ecology and metabolomic context, leading to inconsistent results and underutilization of network and metabolomic features.
A computer-implemented method for sequencing biological samples, including 16S rRNA sequencing or shotgun metagenomic sequencing, to identify microorganisms and perform taxonomic clustering, cross-correlation analysis, and metabolomics to generate a report for diagnosing vaginal infections, incorporating consensus-based methods and bias correction for improved accuracy.
Provides a data-driven, clinically compatible diagnostic workflow that flexibly accommodates different sequencing strategies, addressing compositional biases and enhancing species-level resolution, thereby improving the accuracy and consistency of vaginal infection diagnoses.
Smart Images

Figure US2025048559_02042026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTIDENTIFYING MICROBIOME CONSTITUENTSFIELD
[0001] The present disclosure is directed generally to microbiome-based diagnostics, and more particularly, to laboratory and diagnostic techniques for profiling the composition, metabolome, and microbiome network of a microbiome such as vaginal microbiome to identify biomarkers associated with a particular disease state such as vaginal infections.BACKGROUND
[0002] The human microbiome is a collection of microorganisms (composed of bacteria, bacteriophages, fungi, protozoa, archaea, and viruses) that live inside and on the human body. These microorganisms reside in various parts of the body, including the gastrointestinal tract, skin, mammary glands, seminal fluid, uterus, ovarian follicles, lung, saliva, oral mucosa, conjunctiva, and the biliary tract. A microbiome environment of particular concern is the vaginal microbiome as it plays a crucial role in maintaining vaginal health, protecting against infections, and influencing overall reproductive health. The composition of the vaginal microbiome can vary significantly among individuals and is influenced by factors such as age, hormonal changes, sexual activity, and hygiene practices. In healthy women, the vaginal microbiome is typically dominated by Lactobacillus species, which produce lactic acid that helps to maintain an acidic pH, creating an environment that inhibits the growth of pathogenic microorganisms. Moreover, a balanced vaginal microbiome provides a protective barrier against pathogens, including bacteria, viruses, and fungi, thereby reducing the risk of infections such as bacterial vaginosis, yeast infections, and sexually transmitted infections (STIs). With respect to reproductive health, the vaginal microbiome has functional roles in fertility, pregnancy outcomes, and the risk of preterm birth. Dysbiosis, or an imbalance in the vaginal microbiome, is linked to adverse reproductive health outcomes.
[0003] One example of dysbiosis is bacterial vaginosis (BV), which is a type of vaginal inflammation caused by bacterial overgrowth, upsetting the healthy microbiome of the vagina and is the most common vaginal infection worldwide. BV and related vaginal dysbioses are common, frequently intermittent, and in many women asymptomatic, yet they have been associated with adverse reproductive and infectious disease outcomes, underscoring a longstanding need for more reliable, objective, and comprehensive diagnostic approaches. Traditional clinic-based methods such as Amsel’s criteria and Nugent scoring are widely usedAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT but can be technician-dependent and subjective, leading to variability across settings and inconsistent results that complicate clinical decision-making. Molecular PCR panels (e.g., assays targeting Atopobium vaginae, BVAB-2, and Megasphaera-1) offer greater analytical sensitivity and specificity, but many symptomatic patients still receive negative or indeterminate results, and these targeted constructs may not capture the broader microbial ecology or alternate etiologies that present similarly.
[0004] Over the past decade, sequencing-based characterizations of the vaginal microbiome have illuminated distinct community state types (CSTs) and diversity patterns, including Lactobacillus-deficient states that correlate with BV positivity; however, translating these insights into routine diagnostics remains challenging. While short-amplicon 16S rRNA gene sequencing (e.g., V3-V4) can reproduce certain PCR-based findings and stratify CSTs, it may not consistently resolve taxa at the species level across all genera, potentially limiting clinical interpretability in borderline or mixed presentations. In turn, full-length 16S rRNA gene sequencing has been shown to enhance species-level resolution and reduce misclassification, but it introduces new practical considerations around platform choice, workflow integration, and data processing.
[0005] Analytical biases further complicate microbial abundance comparisons: compositional data, sampling fraction effects, and library size differences can yield false negatives or false positives unless appropriately normalized or corrected, which is not uniformly addressed across commonly used pipelines. Moreover, conventional assays rarely assess higher-order features (such as cross-taxonomy co-occurrence networks or modular correlation structures) that may provide clinically meaningful context when single-marker signals are weak or discordant. Parallel work suggests that BV-positive states are associated with conserved shifts in predicted metabolic pathways and short-chain fatty acid (SCFA) signatures across CSTs, yet most frontline diagnostics do not incorporate metabolomic inference or aggregated pathway-level signals.
[0006] Collectively, these limitations (subjectivity in legacy microscopy -based criteria, targeted PCR panels that may not capture expanded biomarker universes, amplicon resolution constraints, compositional biases in abundance analyses, and underutilization of network and metabolomic context) reflect broader gaps in existing solutions. In view of these gaps, there is an ongoing need for clinically compatible, data-driven workflows that can flexibly accommodate different sequencing strategies, address compositional and correlationAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT structures, and synthesize taxonomic, network, and functional signals into decision-support outputs, while maintaining practical throughput and robustness for real-world use.BRIEF SUMMARY
[0007] In various embodiments, a computer-implemented method comprises: performing a sequencing method on a biological sample collected from a subject, wherein the sequencing method generates read data for microorganisms within the biological sample; identifying, using a clinical assay, sequence variants from the read data, wherein the sequence variants distinguish between different species or strains of microorganisms within the biological sample, and wherein the clinical assay generates a relative abundance value for each identified sequence variant; performing, using the clinical assay, taxonomic clustering of the microorganisms based on their sequencing similarities to identify populations of species; clustering, using a consensus-based method, the species within the populations of species identified in the biological sample based on their co-occurrence patterns to generate clusters; performing a cross-correlation analysis by comparing the clusters identified in the biological sample to clusters belonging to a healthy biological sample to identify significant relationships between the microorganisms; and generating a report that comprises (i) a list of the microorganisms identified in the biological sample and (ii) the cross-correlation between the microorganisms in the biological sample compared to the healthy biological sample, wherein (i) and (ii) are used to determine if the biological sample is positive or negative for a disease.
[0008] In some embodiments, the sequencing method is 16S rRNA sequencing, shotgun metagenomic sequencing, or both.
[0009] In some embodiments, Sanger sequencing, next-generation sequencing (NGS), Illumina MiSeq, Illumina NextSeq, Illumina NovaSeq, 454 pyrosequencing, SMRT sequencing, or PacBio sequencing are used to perform 16S rRNA sequencing, shotgun metagenomic sequencing, or both.
[0010] In some embodiments, the sequencing method is 16S rRNA sequencing.
[0011] In some embodiments, the read data are amplicon sequencing read data for the 16S rRNA gene.
[0012] In some embodiments, the 16S rRNA gene is sequenced entirely, and the read data comprises reads aligning to the entire 16S rRNA gene.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0013] In some embodiments, a portion of the 16S rRNA gene is sequenced and the read data comprises reads aligning to the portion of the 16S rRNA gene.
[0014] In some embodiments, the portion of the 16S rRNA gene sequenced comprises one or more variable regions within the 16S rRNA gene and the read data comprises reads aligning to the one or more variable regions.
[0015] In some embodiments, the portion of the 16S rRNA gene sequenced comprises variable regions 3 and 4 of the 16S rRNA gene and the read data comprises reads aligning to variable regions 3 and 4.
[0016] In some embodiments, the biological sample is a vaginal swab collected from the subject.
[0017] In some embodiments, the vaginal swab comprises the microorganisms comprising the population of species in the vaginal swab.
[0018] In some embodiments, the microorganisms comprise bacteria, bacteriophages, fungi, protozoa, archaea, viruses, or any combination thereof.
[0019] In some embodiments, the subject is suspected of having, symptomatic, asymptomatic, diagnosed with, or receiving treatment for one or more vaginal infection.
[0020] In some embodiments, the subject is symptomatic for one or more vaginal infections.
[0021] In some embodiments, the one or more vaginal infections comprises: bacterial vaginosis, a yeast infection, one or more sexually transmitted infections (STIs), urinary tract infection (UTI), mycoplasma infection, or any combination thereof.
[0022] In some embodiments, at least one of the one or more vaginal infections is bacterial vaginosis.
[0023] In some embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Campylobacteria, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes, Saccharimonadia, or any combination thereof.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0024] In some embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes or any combination thereof.
[0025] In some embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic genera [Eubacterium] brachy group, [Ruminococcus] gnavus group, Acidaminococcus, Acinetobacter, Actinomyces, Aerococcus, Agathobacter, Alloscardovia, Anaerococcus, Anaeroglobus, Aquabacterium, Arcanobacterium, Atopobium, Bacteroides, Bifidobacterium, Brevibacterium, Bulleidia, Butyri cicoccus, Campylobacter, Clostridium sensu stricto 1, Corynebacterium, Criibacterium, Cryptobacterium, Dermabacter, Dialister, DNF00809, Enterococcus, Eremococcus, Escherichia-Shigella, Ezakiella, Facklamia, Fasti diosipila, Fenollaria, Finegoldia, Fusobacterium, Gallicola, Gardnerella, Gemella, Granulicatella, Haemophilus, Helcococcus, Howardella, HT002, Klebsiella, Lachnospiraceae FE2018 group, Lacticaseibacillus, Lactobacillus, Lawsonella, Limosilactobacillus, Listeria, Mageeibacillus, Megasphaera, Mobiluncus, Mogibacterium, Moryella, Murdochiella, Mycoplasma, Olsenella, Parvimonas, Pelomonas, Peptococcus, Peptoniphilus, Peptostreptococcus, Porphyromonas, Prevotella, Prevotella_7, Pseudomonas, Pseudoramibacter, Ralstonia, Rikenellaceae RC9 gut group, Roseburia, S5-A14a, Shuttleworthia, Sneathia, Solobacterium, Sphingobium, Sphingomonas, Staphylococcus, Streptococcus, Sutterella, Ureaplasma, Varibaculum, Veillonella, or any combination thereof.
[0026] In some embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic genera Aerococcus, Anaerococcus, Atopobium, Bacteroides, Bulleidia, Corynebacterium, Criibacterium, Dialister, DNF00809, Escherichia-Shigella, Ezakiella, Fastidiosipila, Finegoldia, Gardnerella, Gemella, Howardella, HT002, Lactobacillus, Limosilactobacillus, Mageeibacillus, Megasphaera, Parvimonas, Prevotella, Pseudomonas, Shuttleworthia, Sneathia, Staphylococcus, Streptococcus, Ureaplasma or any combination thereof.
[0027] In some embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic species Acidaminococcus intestine, Acinetobacter johnsonii, Actinomyces ihuae, Actinomyces turicensis, Aerococcus christensenii, Alloscardovia omnicolens, Anaerococcus hydrogenalis,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTAnaerococcus lactolyticus, Anaerococcus obesiensis, Anaerococcus provencensis, Anaerococcus senegalensis, Anaeroglobus geminatus, Aquabacterium parvum, Arcanobacterium urinimassiliense, Atopobium deltae, Atopobium parvulum, Atopobium vaginae, Bacteroides fragilis, Bacteroides vulgatus, Bifidobacterium bifidum, Bifidobacterium breve, Brevibacterium ravenspurgense, Butyricicoccus faecihominis, BVAB-1, BVAB-2, BVAB-3, Campylobacter faecalis, Campylobacter ureolyticus, Corynebacterium aurimucosum, Corynebacterium coyleae, Corynebacterium mycetoides, Corynebacterium pyruviciproducens, Corynebacterium sundsvallense, Criibacterium bergeronii, Cryptobacterium curtum, Drmabacter jinjuensis, Dialister propionicifaciens, Eremococcus coleocola, Facklamia hominis, Facklamia ignava, Fasti diosipila sanguinis, Finegoldia magna, Fusobacterium nucleatum, Gardnerella vaginalis, Gemella asaccharolytica, Granulicatella elegans, Helcococcus sueciensis, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus iners, Lactobacillus jensenii, Lawsonella clevelandensis, Megasphaera-1, Megasphaera-2, Mobiluncus mulieris, Mogibacterium timidum, Moryella indoligenes, Murdochiella asaccharolytica, Mycoplasma girerdii, Mycoplasma hominis, Olsenella phocaeensis, Parvimonas micra, Pelomonas aquatica, Peptococcus niger, Peptoniphilus coxii, Peptoniphilus duerdenii, Peptoniphilus koenoeneniae, Peptoniphilus lacrimalis, Peptoniphilus obesi, Peptostreptococcus anaerobius, Peptostreptococcus stomatis, Porphyromonas asaccharolytica, Porphyromonas bennonis, Porphyromonas somerae, Porphyromonas uenonis, Prevotella amnii, Prevotella bivia, Prevotella buccalis, Prevotella corporis, Prevotella disiens, Prevotella timonensis, Prevotella ? denticola, Prevotella ? melaninogenica, Pseudoramibacter alactolyticus, Roseburia intestinalis, Sneathia amnii, Sneathia sanguinegens, Solobacterium moorei, Staphylococcus lugdunensis, Ureaplasma urealyticum, Varibaculum cambriense, Veillonella atypica, Veillonella dispar, and Veillonella montpellierensis, or any combination thereof.
[0028] In some embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic species Aerococcus christensenii, Anaerococcus obesiensis, Atopobium vaginae, BVAB-1, BVAB-2, BVAB-3, Dialister propionicifaciens, Finegoldia magna, Gardnerella vaginalis, Gemella asaccharolytica, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus jensenii, Megasphaera-1, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens, or any combination thereof.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0029] In some embodiments, the computer-implemented method further comprises performing differential abundance analysis on the sequence variants to identify microorganisms that enriched or depleted in the biological sample.
[0030] In some embodiments, the microorganisms belonging to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes or any combination thereof are differentially abundant.
[0031] In some embodiments, the microorganisms belonging to taxonomic genera Aerococcus, Anaerococcus, Atopobium, Bacteroides, Bulleidia, Corynebacterium, Criibacterium, Dialister, DNF00809, Escherichia-Shigella, Ezakiella, Fastidiosipila, Finegoldia, Gardnerella, Gemella, Howardella, HT002, Lactobacillus, Limosilactobacillus, Mageeibacillus, Megasphaera, Parvimonas, Prevotella, Pseudomonas, Shuttleworthia, Sneathia, Staphylococcus, Streptococcus, Ureaplasma or any combination thereof are differentially abundant.
[0032] In some embodiments, the microorganisms belonging to taxonomic species Aerococcus christensenii, Anaerococcus obesiensis, Atopobium vaginae, BVAB-1, BVAB-2, BVAB-3, Dialister propionicifaciens, Finegoldia magna, Gardnerella vaginalis, Gemella asaccharolytica, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus jensenii, Megasphaera- 1, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens, or any combination thereof are differentially abundant.
[0033] In some embodiments, the computer-implemented method further comprises performing a metabolomics analysis on the sequence variants to identify metabolomic pathways, metabolites, or both that are dysregulated in the biological sample as compared to the healthy biological sample.
[0034] In various embodiments, a computer-implemented method comprises: extracting microbial DNA from the biological sample; performing full-length 16S rRNA gene sequencing spanning V1-V9 on the extracted microbial DNA to generate long-read sequencing reads; basecalling and demultiplexing the long-read sequencing reads to produce per-sample read sets; performing error correction and consensus generation on the per-sample read sets to obtain full-length 16S amplicon sequences; detecting and removing PCR chimeras and dereplicating the full-length 16S amplicon sequences to construct a featureAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT table; assigning species-level taxonomy to features in the feature table using a curated full- length 16S reference database refined by alignment to a custom panel of bacterial vaginosis- associated phylotypes; aggregating feature counts of the features to a lowest confident taxonomic rank to form a taxonomically aggregated matrix; determining a Community State Type based on a most abundant Lactobacillus species from the taxonomically aggregated matrix; and generating a structured report comprising a list of microorganisms with their relative abundances and the Community State Type.
[0035] In some embodiments, the computer-implemented method further comprises performing a bias-corrective differential abundance analysis on the taxonomically aggregated matrix to estimate per-taxon log2 fold-change between bacterial vaginosis-positive and bacterial vaginosis-negative cohorts under false discovery rate control, and persisting per- taxon statistics for visualization in the structured report.
[0036] In some embodiments, the computer-implemented method further comprises transforming abundances in the taxonomically aggregated matrix by a log-ratio method, computing pairwise correlations using Sparse Correlations for Compositional data separately within the bacterial vaginosis-positive and bacterial vaginosis-negative cohorts, and aggregating multiple base clusterings into consensus clusters, with node annotations referencing the persisted per-taxon statistics from the bias-corrective differential abundance analysis.
[0037] In some embodiments, the computer-implemented method further comprises crosscorrelating the consensus clusters derived from the biological sample to consensus clusters belonging to a healthy reference cohort to identify preserved or rewired modules, and retaining correlation edges above a prespecified magnitude threshold for inclusion in the structured report.
[0038] In some embodiments, the computer-implemented method further comprises predicting pathway-level functional profiles by placing features on a phylogenetic reference, inferring gene family content by hidden-state prediction, assembling gene families into MetaCyc pathways, and performing compositional differential pathway analysis across the bacterial vaginosis-positive and bacterial vaginosis-negative cohorts, wherein pathway interpretations leverage the consensus clusters and the persisted per-taxon statistics.
[0039] In some embodiments, the computer-implemented method further comprises computing cross-method concordance by mapping and reconciling taxon identifiers andAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT names across the species-level taxonomy assignments and the Community State Type, and at least one additional modality selected from short-amplicon 16S, shotgun metagenomics, or panel PCR outputs, and producing concordance statistics for taxon, Community State Type, consensus cluster, and pathway levels.
[0040] In some embodiments, the computer-implemented method further comprises applying quality control and acceptance criteria to the long-read sequencing reads and the full-length 16S amplicon sequences, including per-sample read depth, read-length distribution centered on full-length 16S, platform-specific quality metrics, negative-control background limits, mock-community recovery metrics, and chimera rates, and gating the downstream bias-corrective differential abundance analysis, consensus clustering, crosscorrelation, pathway prediction, and cross-method concordance steps based on the quality control and acceptance criteria. In some embodiments, the computer-implemented method is executed on a throughput-oriented compute architecture comprising streaming or chunked ingestion of long-read files, hardware-accelerated basecalling, multithreaded consensus generation and chimera detection, caching and indexing of full-length 16S reference databases, containerized analytics modules, asynchronous job orchestration, and persistent audit logs that record parameters, software versions, reference hashes, and provenance for the basecalling, consensus generation, differential abundance analysis, consensus clustering, cross-correlation, pathway prediction, and cross-method concordance.
[0041] In some embodiments, the computer-implemented method further comprises reflexively invoking a full-length 16S workflow when the cross-method concordance step identifies insufficient species-level resolution, discordant results across modalities, or an indeterminate bacterial vaginosis classification, and reusing the taxonomically aggregated matrix, persisted per-taxon statistics, consensus clusters, cross-correlation outputs, and pathway analysis results to update the structured report.
[0042] In some embodiments, the computer-implemented method further comprises computing, for inclusion in the structured report, a bacterial vaginosis classification according to predefined or composite rules that consider: presence or absence and relative abundances from the species-level taxonomy assignments, the Community State Type, the persisted per-taxon statistics from the bias-corrective differential abundance analysis, membership of bacterial vaginosis-associated taxa in preserved or rewired consensus clustersAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT identified by the cross-correlation analysis, and enrichment of bacterial vaginosis-associated pathway signatures identified by the pathway-level functional profiles.
[0043] Also described herein is a computer-implemented method of diagnosing a subject with bacterial vaginosis comprising: (i) performing any one of the methods described herein to identify BV-associated biomarkers present in the biological sample from the subject, wherein the BV-associated biomarkers comprise Aerococcus christensenii, Atopobium vaginae, BVAB-1, BVAB-2, BVAB-3, Gardnerella vaginalis, Gemella asaccharolytica, Megasphaera-1, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens; (ii) diagnosing the subject as: BV negative when Atopobium vaginae, BVAB-2, and Megasphaera-1 are not detected, or when either Atopobium vaginae, BVAB-2, or Megasphaera-1 is detected and additional BV biomarkers are not detected, BV positive when at least two of Atopobium vaginae, BVAB-2, or Megasphaera-1 are detected, BV indeterminate when either Atopobium vaginae, BVAB-2, or Megasphaera-1 is detected and at least one additional BV biomarker is detected; and (iii) treating the bacterial vaginosis by administering a therapeutic agent to the subject.
[0044] In some embodiments, the additional BV biomarkers comprise bacterial species among Atopobium vaginae, BVAB-2, and Megasphaera-1.
[0045] In some embodiments, the therapeutic agent is clindamycin oral suppositories or metronidazole vaginal gel.
[0046] In some embodiments, a system is provided that includes one or more processors and a non-transitory computer readable medium containing instructions which, when executed on the one or more processors, cause the one or more processors to perform part or all actions or operations in one or more methods or processes disclosed herein.
[0047] In some embodiments, a non-transitory computer readable storage medium is provided comprising computer program instructions that, when executed by one or more processors, cause the one or more processors to perform part or all actions or operations in one or more methods or processes disclosed herein.
[0048] In some embodiments, a computer-program product is provided that is tangibly embodied in a non-transitory machine-readable medium and that includes instructions configured to cause one or more data processors to perform part or all actions or operations in one or more methods or processes disclosed herein.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0049] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The present application includes the following figures. The figures are intended to illustrate certain embodiments and / or features of the compositions and methods, and to supplement any description(s) of the compositions and methods. The figures do not limit the scope of the compositions and methods, unless the written description expressly indicates that such is the case.
[0051] FIG. 1 illustrates an exemplary clinical assay used to characterize the composition of microbiomes, in accordance with various embodiments.
[0052] FIG. 2 shows a flowchart for detecting biomarkers in accordance with various embodiments.
[0053] FIG. 3 shows a flowchart for characterizing a microbiome for clinical decision support in accordance with various embodiments.
[0054] FIG. 4 illustrates an exemplary clinical assay used to characterize the composition of microbiomes, in accordance with various embodiments.
[0055] FIG. 5 shows a flowchart for characterizing a microbiome for clinical decision support in accordance with various embodiments
[0056] FIGS. 6A and 6B show 16S V3-V4 rRNA sequencing of remnant clinician- collected vaginal swabs previously analyzed via Labcorp NuSwab® BV PCR test. FIG. 6A illustrates an example process of determining Bacterial Vaginosis (BV) status of a vaginal swab using the NuSwab® three microbe panel composed of Atopobium vaginae, B VAB-2, and Megasphaera-1. Each microbe is quantified and given a score of 0 (low), 1 (moderate), or 2 (high). The total score is then interpreted as BV-POS (3-6), BV-IND (2), or BV-NEG (0-1).Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTFIG. 6B shows the 16S relative abundances of NuSwab® panel microbes stratified by their scores. Statistical comparisons of 16S relative abundances were made between samples with NuSwab® scores of 2 and those with scores of 0 or 1 using Wilcoxon rank-sum tests with FDR control.
[0057] FIGS. 7A-7C illustrates the taxonomic and phylogenetic analysis of 16S V3-V4 Amplicon Sequence Variants (ASVs). FIG. 7A are stacked bar plots of class relative abundances colored consistently with FIG. 7B and with samples stratified by BV status. FIG. 4B shows maximum likelihood (ML) phylogeny of all ASVs, with branches colored by class. The Bacilli class was highlighted and constructed a phylogeny. FIG. 4C shows ML phylogeny of ASVs within the Bacilli class shown in FIG. 7B with branches colored by genus and tips labeled with the lowest taxonomic classifications. Neighboring tips with the same label were aggregated into a single label.
[0058] FIGS. 8A-8C Illustrate the community state type (CST) analysis of vaginal microbiomes. FIG. 8 A are stacked bar plots of key genera and species relative abundances across CSTs I, II, III, IV, and V. Samples are stratified by CST (solid vertical lines), then by BV status (dashed vertical lines). FIG. 8B shows multi-dimensional scaling scatter plot of vaginal microbiome Bray-Curtis Dissimilarities with data points colored by CST and shaped by BV status. FIG. 8C shows boxplots of vaginal microbiome diversities as measured by Shannon index, stratified first by CST, then by BV status with data points shaped and colored consistently with those in FIG. 8B.
[0059] FIGS. 9A-9C shows differential abundance (DA) analysis and modularized cooccurrence network analysis of BV-POS and BV-NEG samples. FIG. 9A is a volcano plot of the differential abundance results, depicting the log2 fold-change (L2FC) (x-axis, BV-POS vs BV-NEG), and the -loglO(q-value) (y-axis) for each taxa assessed. FDR control for multiple testing was used to calculate q-values. Data point colors represent statistical significance of taxa, with red representing enriched taxa (q-value < 0.05, L2FC >= 1), blue representing depleted taxa (q-value < 0.05, L2FC <= -1), and yellow representing all other taxa without significant changes. FIG. 9B shows boxplots comparing the relative abundances of select taxa between BV-POS and BV-NEG samples with statistical significance indicated to the right (*** q<0.001). Unique identification of taxa is represented by only a single boxplot shown (e.g. Parvimonas exclusively detected in BV-NEG samples). FIG. 9C shows the clustering results from C3NA with the first column representing clustering among the BV-NEG samplesAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT(n=25 clusters) followed by the modular preservation ribbon linked to the colored BV-POS clusters (n=19). The arc on the right represents the spearman correlations above 0.2 among each BV-POS cluster, with each node colored by the DA result (enriched, neutral, depleted) and the edge colored by the correlation magnitude. All testing was performed on the species level with any non-speciated taxa labeled at the highest resolved taxonomic levels, (i.e., "g_DNF000809" represents the ASVs that resolved to the DNF000809 genus without species assignment).
[0060] FIGS. 10A-10B show a heatmap of differentially abundant (DA) predicted metabolic pathways between BV-POS and BV-NEG samples found in CST-I, CST-III, and CST-IV. Rows represent MetaCyc metabolic pathways and are ordered via hierarchical clustering, while columns (samples) are stratified first by CST (solid vertical lines) then by BV status (dashed vertical lines). The heatmap combines the results of four pairwise comparisons, including: 1) CST-I BV-NEG vs CST-III BV-NEG, 2) CST-I BV-NEG vs CST-IV BV-POS, 3) CST-III BV-NEG vs CST-IV BV-POS, and 4) CST-IV BV-POS vs CST-IV BV-NEG. The top 20 differentially abundant pathways of each comparison with the greatest variance were selected and merged into the list of significant pathways across all pairs of sample groups (combined set of 45 pathways shown).
[0061] FIG. 11 indicates the 16S rRNA gene hypervariable region configurations for which data has been analyzed and provided in this document, where the inside black box indicates the V3-V4 regions and the outside black box indicates the V1-V9 regions (i.e., full-length 16S region). Each individual hypervariable region is indicated by a patterned box (Nine total: VI, V2, V3, etc.), and non-exhaustive list of available PCR primers for amplification of hypervariable region(s) is shown below as indicated by black arrows.
[0062] FIG. 12 shows a table summarizing the mapping statistics for five samples that were sequencing via both 16S V1-V9 and V3-V4 rRNA gene sequencing with marginal totals. For each sequencing sample, the number of genus-mapped reads is shown along with the number and percentage of ASV reads that were shared between the two methods. Briefly, V3-V4 ASVs were aligned to the V1-V9 ASVs and were considered “shared” if the V3-V4 ASV aligned with 100% coverage and identity to one or more V1-V9 ASVs. Note that this is a one-to-many relationship, as VI -V9 spans additional hypervariable regions. This notion of “shared ASVs” indicates the exact sequencing concordance of the two methods, where theAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT table shows on average >98% reads derived from the shared AS Vs in both sequencing methods.
[0063] FIG. 13 shows a table summarizing the detection of the three NuSwab microbes (Atopobium vaginae, BVAB-2, and Megasphaera-1) in each sample from FIG. 12. In addition to 16S V3-V4 and V1-V9 rRNA gene sequencing, these samples were also previously evaluated by shotgun metagenomics (MGx) and NuSwab. Rows highlighted with a black box indicate the expected positives with NuSwab score of 1 or 2, where all three sequencing methods (V1-V9, V3-V4, MGx) have non-zero relative abundance, i.e. detection. Rows in white without black box are expected negatives with NuSwab score of 0, where all three methods correctly show a relative abundance result of 0%.
[0064] FIG. 14 shows a stacked bar plot comparing species / genus relative abundances between 16S V3-V4 and V1-V9 rRNA gene sequenced samples first described in FIG. 12. Relative abundances are only shown for the “shared ASVs”, which constitute >98% of the original relative abundance on average (see FIG. 12). This enabled relative abundances to be compared thrice: 1) V3-V4 ASV-derived abundances, 2) V1-V9 ASV-derived abundances using the shared V3-V4 ASV assigned taxonomy (V1V9_V3V4), and 3) V1-V9 ASV- derived abundances using the full V1-V9 ASV assigned taxonomy (V1V9_V1V9). Comparison of (1) with (2) highlights the concordance of the methods, focusing on abundances while ignoring improved taxonomic resolution afforded by the additional hypervariable regions covered by VI -V9. Comparison of (2) with (3) highlights the improvements in taxonomic resolution when considering the additional hypervariable regions in V1-V9. For example, Sample 4 had similar ASV abundances between V3-V4 and V1-V9, while the VI -V9 method improved the resolution of a Lactobacillus ASV (no species resolution) to Lactobacillus jensenii-resolved ASV.
[0065] FIG. 15 shows a concordance table summary of the five samples originally described in FIG. 12. Sample relative abundances were compared between 16S V3-V4 and VI -V9 rRNA gene sequencing methods at both genus and species resolution. This was evaluated in a similar fashion to FIG.14, where V1-V9 was compared with V3-V4 using either the shared V3-V4 taxonomy (VI V9_V3V4) and using the VI V9 taxonomy (V1V9 V1V9). Two metrics were used: 1) Bray-Curtis Dissimilarity (BCD; ranging from 0- 1), which measures the similarity of composition with a value of 0 indicating completely identical composition and a value of 1 indicating completely dissimilar compositions (i.e., noAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT overlap) and 2) Spearman correlation (ranging from -1 to 1), which measures that rank correlation with a value of -1 representing a perfect negative monotonic relationship and a value of 1 representing a perfect positive monotonic relationship. A low BCD (close to 0) and high Spearman correlation (close to 1) are features of strong concordance. V3-V4 was strongly concordant with both V1-V9 taxonomy comparisons (VI V9 V3 V4 and V1V9_V3V4) for both genus and species resolutions. As expected V1V9_V1V9 had lower species concordances with V3-V4, since the improved resolution afforded by the additional hypervariable regions in V1-V9 altered the species taxonomy, similar to as shown in FIG. 14.
[0066] FIG. 16 shows a table summarizing the increases in bacterial detections by VI -V9 hypervariable regions compared to V3-V4 across the five samples first described in FIG. 12. The black boxes show the percentage increases in alpha diversity (measured by Shannon index) and number of genera detected. Number of AS Vs from selected key vaginal florae (Gardnerella, Lactobacillus, Prevotella) are also shown. The CST for each sample is also shown, showing the increased diversity and detections span multiple CSTs.
[0067] FIG. 17 shows rarefaction curves for both 16S V3-V4 (black) and V1-V9 (gray) rRNA gene sequencing across the five samples first described in FIG. 12. These curves are generated by downsampling the read counts for each sample to a specified number of reads (e.g., 5,000 reads) and enumerating the ASVs detected. These curves indicate that despite sample-to-sample variation in the number of ASVs detected (expected, as vaginal microbiomes vary in diversity and these samples span multiple CSTs — see FIG. 16), targeting the V1-V9 hypervariable regions consistently increases the number of ASVs with the gap present even at low numbers of sequenced reads. The gap increases up until approximately 10,000 reads and then persists at higher sequencing depths.
[0068] FIG. 18 illustrates an exemplary clinical assay workflow used to characterize the composition of microbiomes, in accordance with various embodimentsTERMS
[0069] As used herein, the articles ‘a’ and ‘an’ are used herein to refer to one or to more than one (i.e. at least one) of the grammatical object of the article. By way of example, an element means at least one element and can include more than one element.
[0070] As used herein, the terms “about,” “similarly,” “substantially,” and “approximately” are defined as being largely but not necessarily wholly what is specified (and include wholly what is specified) as understood by one of ordinary skill in the art. In any disclosedAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT embodiment, the term “about,” “similarly,” “substantially,” or “approximately” may be substituted with “within [a percentage] of’ what is specified, where the percentage includes 0.1 percent, 1 percent, 5 percent, and 10 percent, etc. Moreover, the term terms “about,” “similarly,” “substantially,” and “approximately” are used to provide flexibility to a numerical range endpoint by providing that a given value may be slightly above or slightly below the endpoint without affecting the desired result.
[0071] As used herein, the term “absolute abundance,” when used in the context of describing the presence of a particular bacterial species, refers to the amount of DNA derived from the bacterial species out of the amount of all DNA in a sample. For instance, the absolute abundance of one bacterium can be determined by comparing the quantity of DNA specific for this bacterial species (e.g., determined by quantitative PCR) in one given sample with the quantity of all DNA in the same sample.
[0072] As used herein, when an action is “based on” something, this means the action can be based at least in part on at least a part of the something.
[0073] The use herein of the terms including, comprising, or having, and variations thereof, is meant to encompass the elements listed thereafter and equivalents thereof as well as additional elements. Embodiments recited as including, comprising, or having certain elements are also contemplated as consisting essentially of and consisting of those certain elements. As used herein, and / or, refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations were interpreted in the alternative (or).
[0074] As used herein, the term “percentage relative abundance,” when used in the context of describing the presence of a particular bacterial species in relation to all bacterial species present in the same environment, refers to the relative amount of the bacterial species out of the amount of all bacterial species as expressed in a percentage form. For instance, the percentage relative abundance of one particular bacterial species can be determined by comparing the quantity of DNA specific for this species (e.g., determined by quantitative polymerase chain reaction) in one given sample with the quantity of all bacterial DNA (e.g., determined by quantitative polymerase chain reaction (PCR) and sequencing based on the 16s rRNA sequence) in the same sample.
[0075] Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unlessAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. For example, if a range is stated as 1 to 50, it is intended that values such as 2 to 40, 10 to 30, or 1 to 3, etc., are expressly enumerated in this specification. These are only examples of what is specifically intended, and all possible combinations of numerical values (e.g., integer, whole number, decimal, fraction, and the like) between and including the lowest value and the highest value enumerated are to be considered to be expressly stated in this disclosure.DETAILED DESCRIPTION
[0076] The ensuing description provides preferred exemplary embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the ensuing description of the preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It is understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.
[0077] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0078] Also, it is noted that individual embodiments may be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart or diagram may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTI. INTRODUCTIONTechnical Problem
[0079] Diagnostics for vaginal infections, including bacterial vaginosis (BV), remain fragmented across subjective microscopy -based criteria, targeted PCR panels, and researchgrade sequencing, resulting in variability of clinical calls, indeterminate outcomes, and limited visibility into the broader microbial ecology underlying symptomatic presentations. Legacy approaches such as Amsel’s criteria and the Nugent score are technician-dependent and can yield inconsistent results across settings, constraining reliability and downstream care decisions. PCR-based assays that quantify a narrow group of BV-associated microbes (e.g., Atopobium vaginae, BVAB-2, Megasphaera-1) improve analytical sensitivity and specificity, yet a substantial fraction of symptomatic patients test negative or indeterminate, indicating that targeted constructs may under-sample relevant taxa or miss network-level signatures that present similarly.
[0080] Sequencing-based analyses have clarified community state types (CSTs) and diversity shifts associated with BV (e.g., Lactobacillus-deficient CST-IV), but clinical translation is often limited by amplicon-region choices (such as V3-V4) that may insufficiently resolve species across certain genera and by workflows that do not surface decision-grade reports aligned to clinical practice. While short-amplicon 16S rRNA gene sequencing can reproduce certain PCR-based findings and stratify CSTs, reduced specieslevel resolution in some genera (e.g., Gardnerella) can limit interpretability in borderline or mixed presentations, and inconsistency across primer sets and pipelines reduces comparability across studies and cohorts. Full-length 16S rRNA gene sequencing can increase species-level resolution and reduce misclassification relative to short amplicons, but many pipelines lack configurable ways to incorporate full-length reads alongside established workflows, constraining adoption and interpretability.
[0081] Independent of sequencing method, microbial abundance comparisons are vulnerable to compositional and sampling-fraction biases (e.g., variable sampling fractions, library sizes), which can yield false negatives or false positives unless addressed by biascorrective methods suitable for microbiome compositions. Conventional pipelines also underutilize higher-order analytic constructs — such as consensus clustering and crosstaxonomy co-occurrence networks — that capture coordinated microbial signatures and may be more robust than single-marker strategies in borderline or mixed presentations. Further,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT most frontline diagnostics do not incorporate pathway-level inferences (e.g., SCFA-linked fermentative signatures) or fused functional signals that can augment specificity. From a systems perspective, clinical laboratories face operational burdens processing high- throughput sequencing files (e.g., paired-end FASTQ), training error models, performing denoising, taxonomic assignment, multi-branch analytics (metabolomics prediction, differential abundance testing, network modeling), and summarizing results into interpretable reports, all of which can be resource-intensive, latency-prone, and difficult to audit or reproduce at scale when implemented as ad hoc toolchains. Heterogeneous hardware, variable sample quality, disparate reference databases, and limited parallelization or caching further complicate standardized processing, increase turnaround time (TAT), and inflate cost structures, creating practical barriers to routine adoption.Technical Solution
[0082] In order to address these challenges and others, the disclosed computer- implemented clinical assay provides an integrated, end-to-end workflow that ingests clinician-collected biological samples and supports configurable sequencing modalities including short-amplicon 16S rRNA gene profiling (e.g., V3-V4), optional shotgun metagenomics, and an alternative full-length 16S rRNA gene workflow, which are harmonized into a unified analytics suite comprising high-resolution ASV calling, taxonomic assignment, consensus clustering, cross-correlation to healthy-reference clusters, differential abundance testing, and metabolomics pathway prediction to produce decision-grade reports suitable for clinical use. In some embodiments, the short-amplicon configuration is processed via a DAD A2 -based pipeline that performs read filtering, trimming, denoising, merging, chimera removal, ASV table construction, and taxonomic assignment using curated reference sets (e.g., SILVA, Greengenes, RDP, and custom BV sequences), thereby yielding high- resolution microbial profiles aligned to clinical samples with consistent reporting across modalities. In complementary embodiments, the assay integrates a full-length 16S rRNA gene sequencing workflow (e.g., PacBio SMRT or Oxford Nanopore), including library preparation spanning V1-V9, platform-appropriate error correction / consensus calling, species-level classification (e.g., Emu for Nanopore), and harmonized downstream processing and reporting to increase species-level resolution and reduce misclassification in genera where short amplicons are limiting.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0083] To mitigate compositional and sampling-fraction biases, the assay applies biascorrective normalization and statistically principled differential abundance methods tailored to microbiome data (e.g., ANCOM-BC), thereby reducing false inferences and enhancing robustness across cohorts. The assay further implements correlation-driven, consensus-based cross-taxonomy network analysis (e.g., C3NA with log-ratio transformations and SparCC correlations) that identifies coordinated microbial modules preserved across BV-positive and BV-negative groups, supporting resilient, multi-target diagnostic strategies when singlemarker signals are discordant or borderline. In addition, the assay predicts MetaCyc pathways from ASV profiles (e.g., PICRUSt2) and applies compositional differential pathway analysis (e.g., ALDEx2) to surface conserved BV-associated metabolic signatures (including SCFA- linked fermentative pathways), which can be incorporated into composite scoring frameworks or used to adjudicate indeterminate cases.
[0084] From an implementation standpoint, the system is architected to improve computational efficiency, reduce latency, and enhance reproducibility and auditability: streaming and chunked ingestion of large FASTQ files with multithreaded preprocessing (e.g., quality filtering and denoising) to reduce wall-clock time; caching and indexing of reference databases to minimize repeated I / O and accelerate taxonomic assignment; and asynchronous job orchestration to isolate long-running steps (e.g., error-rate learning, consensus clustering, network construction), thereby smoothing burst loads and improving throughput on shared infrastructure. Memory usage is managed via sparse matrix representations for ASV count tables and correlation graphs, while numerical stability is enhanced by consistent log-ratio transformations, with containerized analytics modules and persistent audit logs capturing parameters, versions, and provenance to support regulated laboratory environments and reproducible runs. The reporting services serialize taxonomic, network, and pathway outputs into structured summaries — including cross-correlation to healthy-reference clusters, normal-range contextualization, and optional CST stratification — to enable faster interpretation and actionable decision support in routine practice.
[0085] The alternative full-length 16S configuration can be invoked when species-level resolution is clinically material, with library preparation designed to span V1-V9, platform options including PacBio SMRT (for high-accuracy circular consensus reads) and Oxford Nanopore (for real-time long reads), and data processing that includes platform-specific error correction, consensus generation, and species-level taxonomic assignment. These full-length reads are integrated into the same analytics and reporting framework as short-amplicon data,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT enabling side-by-side interpretability, cross-method concordance checks, and configuration- driven selection of sequencing modality based on clinical context and laboratory constraints. Taken together, the configurable assay fuses taxonomic, network, and functional signals, improves analytical rigor, and addresses practical system-level constraints — computational efficiency, latency, auditability, and cross-method harmonization — associated with deploying sequencing-based diagnostics for BV and related vaginal infections.II. CLINICAL ASSAY CONFIGURED TO CHARACTERIZE VAGINAL MICROBIOMES
[0086] The ensuing description details methods directed towards diagnosing and treating subjects with BV; however, this is not intended to limit the scope of this disclosure and the methods may also be reasonably applied to other vaginal and bacterial infections such as aerobic vaginosis, yeast infections (Candidiasis), sexually transmitted infections (STIs), urinary tract infections (UTIs), genital Mycoplasma infections, Haemophilus influenzae, and the like. These methods are suitable for use in both symptomatic and asymptomatic presentations and can serve as reflex testing when panel PCR results are positive or negative but clinical suspicion remains high, thereby aligning sequencing-based outputs with familiar diagnostic contexts. In addition, Community State Type (CST) stratification, based on the most abundant Lactobacillus species detected per sample, provides clinically useful context for differentiating BV from other dysbiotic or infectious states by distinguishing Lactobacillus-dominant communities (e.g., CST-I, CST-II, CST-V) from diverse anaeroberich communities (e.g., CST-IV).
[0087] Current methods of diagnosing BV (e.g., Nugent score, Amsel’s criteria, PCR-based methods) fail to identify BV symptomatic women (roughly 50% in some instances), which underscores the clinical need for improved tools. Accordingly, a clinical assay that combines experimental data collection with bioinformatic analysis to identify BV biomarkers from the vaginal microbiome of subjects who are symptomatic or asymptomatic for BV is described herein, with assay configurations focused on short-amplicon 16S rRNA gene sequencing (e.g., V3-V4) and, in some cases, optional shotgun metagenomics to broaden coverage beyond bacterial marker genes. In validation studies, 16S V3-V4 sequencing reproduced NuSwab BV target quantification and shotgun metagenomics-based CST assignments with high concordance, providing a robust basis for expanded biomarker detection beyond the conventional three-microbe panel.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0088] As described herein, “bacterial vaginosis,” abbreviated as “BV,” is to be understood in the broad sense as the alterations of vaginal microbiome composition in a woman, as compared to a baseline, reference or “normal” vaginal microbiome composition. In other words, the term “bacterial vaginosis” is not limited to a specific vaginal microbiome composition, or any particular symptoms observed in a particular woman or population of women, and may encompass Lactobacillus-deficient CST-IV states with increased taxonomic diversity and BV-associated taxa enrichment. While BV can manifest through a variety of symptoms (e.g., changes in discharge color / odor / amount, vaginal itching or irritation, pain during intercourse, painful urination, light bleeding), BV can also be asymptomatic; the assay methods described herein are generally probative for vaginal microbiome alterations and can be used to determine such alterations in both symptomatic and asymptomatic females. Moreover, “vaginal infections” are also to be understood broadly as alterations of vaginal microbiome composition relative to baseline / reference compositions.
[0089] Alterations or changes to the vaginal microbiome composition are due to shifts in the microorganism populations that comprise the vaginal microbiome. This collection of microorganisms can include bacteria, bacteriophages, fungi, protozoa, and viruses. In various embodiments, bacterial shifts are measured by the clinical assay 100, which, through the microbiome pipeline 110 and downstream analytics, supports detection bacteria at various taxonomic ranks (for example, classes, genera, and species) associated with BV and other infections (non-limiting enumerations given below) and contextualization via CST stratification and differential / compositional analyses. Non-limiting examples of bacteria (e.g., BV biomarkers) that make up the vaginal microbiome composition of healthy and BV infected vaginas include bacteria belonging to the taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Campylobacteria, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes, Saccharimonadia, or any combination thereof. More specifically, the bacteria (e.g., BV biomarkers) can belong to the taxonomic genus [Eubacterium] brachy group, [Ruminococcus] gnavus group, Acidaminococcus, Acinetobacter, Actinomyces, Aerococcus, Agathobacter, Alloscardovia, Anaerococcus, Anaeroglobus, Aquabacterium, Arcanobacterium, Atopobium, Bacteroides, Bifidobacterium, Brevibacterium, Bulleidia, Butyri cicoccus, Campylobacter, Clostridium sensu stricto 1, Corynebacterium, Criibacterium, Cryptobacterium, Dermabacter, Dialister, DNF00809, Enterococcus, Eremococcus, Escherichia-Shigella, Ezakiella, Facklamia, Fasti diosipila, Fenollaria, Finegoldia, Fusobacterium, Gallicola, Gardnerella, Gemella, Granulicatella,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTHaemophilus, Helcococcus, Howardella, HT002, Klebsiella, Lachnospiraceae FE2018 group, Lacticaseibacillus, Lactobacillus, Lawsonella, Limosilactobacillus, Listeria, Mageeibacillus, Megasphaera, Mobiluncus, Mogibacterium, Moryella, Murdochiella, Mycoplasma, Olsenella, Parvimonas, Pelomonas, Peptococcus, Peptoniphilus, Peptostreptococcus, Porphyromonas, Prevotella, Prevotella ?, Pseudomonas, Pseudoramibacter, Ralstonia, Rikenellaceae RC9 gut group, Roseburia, S5-A14a, Shuttleworthia, Sneathia, Solobacterium, Sphingobium, Sphingomonas, Staphylococcus, Streptococcus, Sutterella, Ureaplasma, Varibaculum, Veillonella, or any combination thereof. Examples of bacterial species (e.g., BV biomarkers) that can make up the vaginal microbiome composition include: Acidaminococcus intestine, Acinetobacter johnsonii, Actinomyces ihuae, Actinomyces turicensis, Aerococcus christensenii, Alloscardovia omnicolens, Anaerococcus hydrogenalis, Anaerococcus lactolyticus, Anaerococcus obesiensis, Anaerococcus provencensis, Anaerococcus senegalensis, Anaeroglobus geminatus, Aquabacterium parvum, Arcanobacterium urinimassiliense, Atopobium deltae, Atopobium parvulum, Atopobium vaginae (also known as Fannyhessea vaginae), Bacteroides fragilis, Bacteroides vulgatus, Bifidobacterium bifidum, Bifidobacterium breve, Brevibacterium ravenspurgense, Butyricicoccus faecihominis, BVAB-1 (also known as Candidates Lachnocurva vaginae), BVAB-2 (also known as Amygdalobacter indicium), BVAB-3 (also known as Mageeibacillus indolicum), Campylobacter faecalis, Campylobacter ureolyticus, Corynebacterium aurimucosum, Corynebacterium coyleae, Corynebacterium mycetoides, Corynebacterium pyruviciproducens, Corynebacterium sundsvallense, Criibacterium bergeronii, Cryptobacterium curtum, Drmabacter jinjuensis, Dialister propionicifaciens, Eremococcus coleocola, Facklamia hominis, Facklamia ignava, Fasti diosipila sanguinis, Finegoldia magna, Fusobacterium nucleatum, Gardnerella vaginalis, Gemella asaccharolytica, Granulicatella elegans, Helcococcus sueciensis, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus iners, Lactobacillus jensenii, Lawsonella clevelandensis, Megasphaera- 1, Megasphaera-2, Mobiluncus mulieris, Mogibacterium timidum, Moryella indoligenes, Murdochiella asaccharolytica, Mycoplasma girerdii, Mycoplasma hominis, Olsenella phocaeensis, Parvimonas micra, Pelomonas aquatica, Peptococcus niger, Peptoniphilus coxii, Peptoniphilus duerdenii, Peptoniphilus koenoeneniae, Peptoniphilus lacrimalis, Peptoniphilus obesi, Peptostreptococcus anaerobius, Peptostreptococcus stomatis, Porphyromonas asaccharolytica, Porphyromonas bennonis, Porphyromonas somerae, Porphyromonas uenonis, Prevotella amnii, Prevotella bivia, Prevotella buccalis, Prevotella corporis, Prevotella disiens, Prevotella timonensis, Prevotella ? denticola, Prevotella ?Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT melaninogenica, Pseudoramibacter alactolyticus, Roseburia intestinalis, Sneathia amnii, Sneathia sanguinegens, Solobacterium moorei, Staphylococcus lugdunensis, Ureaplasma urealyticum, Varibaculum cambriense, Veillonella atypica, Veillonella dispar, Veillonella montpellierensis or any combination thereof.Overview
[0090] FIG. 1 illustrates an exemplary computing environment for performing clinical assay 100 configured to characterize vaginal microbiomes from bacterial vaginosis (BV) diagnosed samples, and more generally, vaginal infection diagnosed samples, by integrating laboratory processing with bioinformatic analytics to derive decision-support biomarkers and contextual microbial signals, using short-amplicon 16S rRNA gene profiling (e.g., V3-V4) as a primary modality and, in some embodiments, optional shotgun metagenomics for broader microbial profiling when clinically useful. The exemplary environment includes a database management platform 101, a network 102, an assay platform 103, and end devices 104. The database management platform 101 and the assay platform 103 can be implemented using software only (e.g., each module of the platform is a digital entity implemented using programs, code, or instructions executable by one or more processors), using hardware (e.g., a medical tool to perform testing, a sequencer, a GPU, a CPU, or the like), or using a combination of hardware and software. Although FIG. 1 illustrates a particular set and arrangement of the components, it should be understood that any suitable number or configuration of components may be included in the environment. Additional components such as various sequencing systems, cloud-based data repositories, or parallel computing resources may also be integrated as appropriate for specific implementations. Security measures, such as encrypted data transmission and user authentication, can be implemented to protect sensitive clinical and genomic information during processing and data sharing.
[0091] Clinical assay 100 begins with clinician-collected biological samples (e.g., vaginal swabs) that undergo DNA extraction and amplicon-based 16S rRNA gene sequencing to generate sequencing reads 105, which are ingested for downstream analysis. The database management platform 101 is a core component of the environment that provides the infrastructure and computational tools necessary for organizing, maintaining, and providing data such as the generated sequencing read data 105 for downstream analysis. The database management platform 101 is configured for storing sequencing, microbiome, and subject information, e.g., collections of sequencing read data 105 for various subjects, biomarkers asAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT reference databases, reference cohorts, reference genomes, and the like. The data stores or devices with the database management platform 101 can be deployed using local servers, network-attached storage devices, or cloud-based data warehousing solutions. Various database management systems may be used, including relational databases such as PostgreSQL and MySQL, or scalable NoSQL architectures, selected based on requirements for performance, scalability, and data accessibility. The database management platform 101 is routinely updated to incorporate newly validated references and improvements from ongoing curation, thereby enabling the system to adapt to new developments and expansions in publicly available sequence repositories.
[0092] The database management platform 101 is configured to interact with other components of environment, including the network 102, assay platform 103, and end devices 104. Through these interactions, the database management platform 101 supplies reference datasets and annotation metadata for the microbiome analysis, receives updates and new sequence data, and supports distributed, cloud-based, or hybrid deployments. Data exchange between modules can occur over standard network protocols, high-speed data buses, or cloud application programming interfaces, depending on the system architecture. The database management platform 101 may be deployed as a combination of software applications, such as Python scripts, Docker or Conda environments, or as integrated systems that combine software with dedicated hardware resources including CPUs, GPUs, or high-performance storage appliances. This modular and scalable architecture enables the database management platform 101 to support analysis for a range of microbiomes, accommodate changes in reference data as new bacteria, viruses, etc. and subtypes with various microbiomes are discovered, and integrate with high-throughput sequencing instruments or automated update mechanisms.
[0093] Network 102 is configured to provide robust, high-speed, and secure data communications among the components of environment, including the database management platform 101, the assay platform 103, and end devices 104. Network 102 supports a range of modern networking protocols and architectures, enabling reliable connectivity and efficient data transfer to facilitate the workflows illustrated in FIGS. 1-5. Network 102 may be implemented as a local area network (LAN), a wide-area network (WAN), or a combination of public and private networks, including the Internet, virtual private networks (VPNs), or dedicated research networks. Contemporary network protocols such as TCP / IP, Ethernet, and advanced wireless standards (for example, Wi-Fi 6, Wi-Fi 7, or 5G cellular networks) areAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT supported to provide high bandwidth, low latency, and secure transmission of large volumes of genomic data and analytical results. The network 102 can also integrate optical fiber links and high-throughput backbone connections where ultra-fast data movement is required, such as for distributed storage clusters or remote laboratory facilities.
[0094] To connect the database management platform 101, the assay platform 103, and the end devices 104 to network 102, various types of physical and wireless links may be employed. Wireline connections such as Ethernet, fiber optic cables, or DOCSIS cable modems can deliver reliable and scalable connectivity. Wireless solutions including Wi-Fi, 5G, and Bluetooth can enable mobility, remote access, and ease of installation. For geographically distributed deployments, network 102 may further integrate cloud-based networking services, software-defined networking (SDN), and edge computing nodes to optimize data routing and processing efficiency. The integration of these diverse connection methods ensures a resilient and high-performance data communication framework. Network 102 enables seamless data exchange and real-time interaction among all components of environment, supporting complex workflows and large-scale analysis tasks required for microbiome diagnostics. Security features such as encrypted data transmission, multi-factor authentication, and firewall protections may be incorporated to safeguard sensitive patient and genomic data during transfer and remote access.
[0095] The assay platform 103 is configured to process samples and profile the composition, metabolome, and microbiome network of a microbiome such as vaginal microbiome to identify biomarkers associated with a particular disease state such as vaginal infections. As depicted in FIG. 1, the assay platform 103 comprises several specialized pipelines, tools, and modules. These pipelines, tools, and modules may be realized in software, hardware, or a hybrid configuration, and are designed to support high-throughput, accurate, and automated workflows. More specifically, the sequencing reads 105 are processed by a microbiome pipeline 110 (e.g., a DAD A2 -based workflow) that performs quality filtering, denoising, read merging, chimera removal, and ASV table construction, followed by taxonomic assignment against curated references (e.g., SILVA, RDP, Greengenes, and custom BV phylotype sequences) to produce high-resolution amplicon sequencing variants (AS Vs) for each sample. The microbiome pipeline 110 further supports CST stratification by assessing the most abundant Lactobacillus species per sample (CST-I: L. crispatus; CST-II: L. gasseri; CST-III: L. iners; CST-IV: diverse; CST-V: L. jensenii), following Ravel et al., and is designed to benchmark sequencing-derived quantificationsAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT against PCR-based BV status calls (e.g., NuSwab BV panel targets Atopobium vaginae, BVAB-2, and Megasphaera-1) to facilitate cross-method concordance. In validation cohorts with paired 16S V3-V4 and metagenomics (MGx), CST calls were concordant in the substantial majority of samples, supporting the reliability of 16S-derived CST assignments used in the assay.
[0096] Downstream of the microbiome pipeline 110, the AS Vs are independently processed by a metabolomic prediction pipeline 115 to infer pathway-level functional profiles and to identify BV-associated metabolomic pathways 125 via differential analysis, thereby capturing conserved signatures (e.g., short-chain fatty acid-linked fermentative pathways) that can complement taxonomic markers. In parallel, a microbial network pipeline 120 computes compositional correlations and consensus clustering to derive BV-associated networks 130, revealing modular co-occurrence patterns that expand the BV biomarker universe beyond single targets and improve interpretability in indeterminate or mixed presentations. Outputs from the microbiome pipeline 110, metabolomic prediction pipeline 115, and microbial network pipeline 120 are synthesized into consolidated summaries that align CST stratification and cross-method concordance with expanded biomarker sets and pathway-level signals, enabling clinical reporting consistent with BV diagnostic contexts and supporting more general vaginal infection assessments.
[0097] To enable seamless integration with laboratory operations, the assay platform 103 is designed to interact directly with a variety of sequencing modules or sequencers. This interaction may be realized through data transfer protocols and integration with the output systems of next-generation sequencing (NGS) instruments, Sanger sequencers, or other automated sequencing platforms. The assay platform 103 can automatically retrieve sequencing files and associated metadata using direct USB, Ethernet, Wi-Fi, or through connections established with laboratory information management systems (LIMS) via the network 102 that aggregate data across multiple instruments and data stores or device (e.g., those data stores and device that are part of the database management platform 101). In some embodiments, sequencing platforms may be physically linked to dedicated processing servers or high-performance workstations where the assay platform 103 is installed, while in other scenarios, data may be routed through secure cloud-based storage or network-attached servers for centralized access and processing.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0098] Each of the one or more end devices 104 is an electronic device comprising hardware, software, embedded logic components, or a combination of these elements, and is configured to interact with the database management platform 101, network 102, and the assay platform 103. The end device 104 may include a range of contemporary computing systems such as desktop computers, laptops, workstation computers, tablets, smartphones, portable handheld devices, wearable computing devices, thin clients, or other specialized laboratory or clinical terminals. These computing devices can run various operating systems and application environments, including Windows, macOS, Linux distributions, Android, iOS, or other modern or embedded operating systems.
[0099] The end device 104 may be designed to execute a variety of client-side or webbased applications that support user interaction with the virus subtyping pipeline. For example, the end device 104 may run specialized software for submitting sequence data, reviewing analysis reports, managing database updates, or accessing curated reference datasets. The end device 104 may also support secure user authentication, audit trails, and role-based access control, ensuring that only authorized users can access sensitive genomic and diagnostic information.
[0100] The end device 104 comprises an interface, such as a graphical user interface (GUI), which enables users to interact intuitively with the environment. Through the interface, users can upload sequence data, initiate new microbiome analyses, visualize results, monitor workflow status, or configure system parameters. The interface may support advanced visualization tools for exploring profiles, confidence scores, or analysis, and may also integrate with LIMS or electronic medical records for seamless data exchange.
[0101] The end device 104 is capable of both inputting and receiving data over the network 102. For example, a laboratory technician, clinician, or researcher may use the end device 104 to submit nucleotide sequence data or analysis requests to the assay platform 103. The end device 104 may also be used to retrieve results, download analytical reports, or access the latest updates to reference databases maintained by the database management platform 101. Data transmission between the end device 104 and other components may occur via wired connections (such as Ethernet or USB) or via wireless protocols (such as Wi-Fi, Bluetooth, or cellular networks), depending on the deployment scenario.
[0102] In some embodiments, the end device 104 may also support integration with cloudbased services or distributed computing environments. This enables remote access to theAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT virus subtyping system, supports telemedicine or distributed research collaborations, and allows authorized users to interact with the environment from virtually any location. Security features, including encrypted data transfer, multi-factor authentication, and digital certificates, may be implemented to protect sensitive patient, sample, and genomic data during transmission and remote access.
[0103] The end device 104 may further incorporate notification systems, audit logs, and automated reporting tools, enabling users to receive alerts about workflow completion, system updates, or quality control events. Advanced deployments may allow the end device 104 to interface with laboratory automation platforms, robotic sample handlers, or sequencing instruments for fully automated, end-to-end workflows.
[0104] In some embodiments, the environment may be further augmented with additional components designed to enhance performance, scalability, and adaptability for advanced virus subtyping workflows. For example, high-throughput sequencing systems can be incorporated to generate large volumes of raw sequence data from clinical, environmental, or research samples. These sequencing systems may be directly connected to the assay platform 104, enabling automated transfer of sequencing outputs into the assay platform 104 for immediate downstream processing. Sequencing instruments may be physically located in laboratory environments and interfaced with the network 102 for seamless data integration and real-time analysis.
[0105] Parallel computation resources, such as GPU clusters, multicore CPU servers, or cloud-based high-performance computing environments, may also be deployed within environment to accelerate computationally intensive steps. These resources can be allocated to the database management platform 101 and the assay platform 103, enabling rapid analysis even when processing large sample batches or extensive reference databases. Parallel computing capabilities may also support reference set generation and data management within the database management platform 101, expediting the construction and updating of references.
[0106] The modular architecture of the environment enables flexible scaling and adaptation to a variety of laboratory, clinical, or research settings. Components such as the database management platform 101, network 102, and the assay platform 103 are designed to support distributed processing, remote access, and collaborative workflows, allowing laboratories to accommodate increasing data volumes, support geographically dispersed teams, or integrateAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT with external diagnostic networks. End devices 04, equipped with user interfaces, provide access points for laboratory personnel, clinicians, or researchers to monitor system operations, initiate analyses, review results, and manage database updates from local or remote locations.Sample Acquisition
[0107] Sequencing read data 105 are generated from clinician-collected biological samples, preferably vaginal swabs obtained under standard-of-care procedures, followed by DNA extraction and amplicon-based 16S rRNA gene library preparation to capture the vaginal microbiome for downstream analysis within clinical assay 100. The workflow is designed for BV and more general vaginal infection contexts, and may incorporate contemporaneous PCR-based BV status labels (e.g., NuSwab BV panel targets) for cross-method concordance during analytics, while the sequencing read data 105 output serves as the primary source of ASV-resolved taxonomic, CST, network, and metabolomic inference.
[0108] In various embodiments, biomarkers correspond to one or more vaginal or bacterial infections, including without limitation BV, yeast infections (Candidiasis), sexually transmitted infections (STIs), urinary tract infections (UTIs), genital mycoplasma infections, or any combination thereof. Biological samples may be collected from subjects who are suspected of having, symptomatic for, asymptomatic for, diagnosed with, or receiving treatment for BV or other vaginal infections. For BV-specific testing, commonly used clinical assays quantify Atopobium vaginae (also known as Fannyhessea vaginae), BVAB-2 (also known as Amygdalobacter indicium), and Megasphaera-1, which can be leveraged as contextual labels in the integrated workflow while sequencing read data 105 are used to profile a broader microbial community. By way of example, without limitation, a female subject may be tested for BV using a BV test (e.g., NuSwab® VG) to test for the presence of predictive marker organisms (e.g., Atopobium vaginae (also known as Fannyhessea vaginae), BVAB-2 (also known as Amygdalobacter indicium), and Megasphaera-1).
[0109] The sample can be a cell-containing liquid or tissue, and may include, without limitation, cells from a vaginal swab, amniotic fluid, biopsies, blood or blood cells, peritoneal fluid, plasma, pleural fluid, saliva, semen, serum, tissue sections, or tissue homogenates. Samples are obtained by standard procedures, including swabs (e.g., vaginal swabs), biofilm sampling, aspirations, tissue sections, venipuncture, and surgical or needle biopsies, and may be used immediately or stored under appropriate conditions for later processing. In preferredAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT embodiments, a clinician collects a vaginal swab (e.g., from the lower third of the vaginal wall) to capture vaginal fluid and the associated microbiome for sequencing read generation 105. Different clinically acceptable sampling strategies and anatomical sites may yield comparable microbiota profiles for downstream 16S-based analyses when other technical variables are controlled.
[0110] BV-associated biomarkers and other vaginal infection biomarkers are exemplified by bacterial nucleic acids (DNA and / or RNA), including sequences specific to a particular taxonomic group, such as class, genus, and / or species. Sequencing-derived ASVs and taxonomic assignments can identify targeted BV panel organisms and, where present, an expanded set of BV-associated taxa (e.g., Gardnerella, Prevotella, Sneathia, Dialister, BVAB- 1, BVAB-3, Megasphaera-2), as well as Lactobacillus species relevant to CST stratification. The biomarkers are not limited to currently enumerated taxa; newly recognized or yet- uncultivated organisms may be detected and incorporated as biomarkers as reference databases and empirical evidence evolve. The methods may be adapted to detect a variety of BV (and / or vaginal infection) associated microorganisms, with the biomarker composition and reporting configurable per clinical use case.DNA Isolation[OHl] Nucleic acids (e.g., DNA and / or RNA) associated with bacteria, bacteriophages, fungi, protozoa, archaea, viruses, or combinations thereof may be isolated from biological samples collected as described herein to enable downstream characterization of the vaginal microbiome within clinical assay 100. In some embodiments, microbial DNA is preferentially isolated for amplicon-based 16S rRNA gene profiling, while in other embodiments total DNA is extracted to preserve broad community representation across microorganisms present in the sample. As used herein, “nucleic acid” encompasses deoxyribonucleic acids (DNA) or ribonucleic acids (RNA), in single- or double-stranded forms, including naturally occurring molecules, synthetic analogues, and structural variants such as hairpins or stem-loop configurations, without limitation.
[0112] Nucleic acid isolation may be performed using any suitable technique, including solid-phase or magnetic bead workflows, with process steps that can include mechanical and / or chemical lysis of cells, removal of proteins and inhibitors, and recovery of purified nucleic acids compatible with downstream amplification and sequencing. In various embodiments, DNA extraction from clinician-collected vaginal swabs is performed using aAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT validated kit (e.g., ZymoBIOMICS Magbead DNA isolation kit) under conditions that balance robust lysis of diverse taxa (including Gram-positive organisms) with minimization of extraction bias, followed by elution suitable for library preparation. Optional process controls may include negative extraction blanks and mock-community references to monitor contamination and assess extraction consistency; quality control checkpoints can include spectrophotometric or fluorometric quantification, fragment integrity assessment, and evaluation of inhibitor carryover to support reproducible amplification of marker genes (e.g., the 16S rRNA gene). Where appropriate, host-DNA depletion and / or additional inhibitor removal steps may be employed to increase the proportion of microbial reads and improve downstream taxonomic resolution; purified nucleic acids may be used immediately for library preparation or stored at appropriate temperatures (e.g., -20°C to -80°C for DNA) with chain- of-custody tracking and sample barcoding to maintain traceability through the clinical assay 100 workflow.Methods of Detecting Nucleic Acid Molecules
[0113] In various embodiments, the presence or relative abundance of vaginal infection biomarkers (e.g., BV-associated taxa) is detected by amplifying extracted nucleic acids from a biological sample using polymerase chain reaction (PCR) techniques, including endpoint PCR, quantitative PCR (qPCR), reverse-transcriptase PCR (RT-PCR), multiplex PCR, nested PCR, and high-fidelity PCR, with primers designed to target genes of interest relevant to bacterial identification and quantification. In some contexts, targeted PCR assays for BV- associated organisms (e.g., Atopobium vaginae, BVAB-2, Megasphaera-1) may be used as clinical comparators or labels alongside sequencing-based analytics in clinical assay 100.
[0114] Primers may target conserved loci that enable broad bacterial detection with species-level discrimination via flanking variable regions. In certain embodiments, primers target the 16S ribosomal RNA (rRNA) gene, with forward and reverse oligonucleotides hybridizing to conserved regions that flank one or more hypervariable regions (V1-V9), thereby producing amplicons suitable for downstream sequencing and taxonomic classification in microbiome pipeline 110. Primer length and composition are selected to achieve robust amplification under conditions suitable for polymerase activity, with optimization based on sample quality, target region (e.g., V3-V4), and platform requirements.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0115] In other embodiments, biomarker detection is performed using sequencing methods that generate sequencing read data 105 from amplified regions or from total extracted nucleic acids. Suitable sequencing approaches include chain-termination (Sanger), sequencing by synthesis, sequencing by ligation, mass-spectrometry-based sequencing, and next-generation sequencing (NGS) platforms (e.g., Illumina MiSeq, NextSeq, NovaSeq), among others, which may be used alone or in combination to produce read data for taxonomic profiling via microbiome pipeline 110.
[0116] In certain embodiments for clinical assay 100, short-amplicon 16S rRNA gene sequencing (e.g., V3-V4) is used to detect and classify bacteria based on conserved and variable regions within the 16S gene. The nine hypervariable regions (V1-V9) provide species-informative sequence diversity interspersed with conserved anchors that facilitate primer binding and amplification Sequencing of 16S rRNA gene amplicons enables comparison against curated databases (e.g., SILVA, Greengenes, RDP) to assign taxonomy and quantify relative abundances, and has been demonstrated to reproduce PCR-based BV target quantifications and stratify vaginal microbiomes into Community State Types (CSTs).
[0117] To generate 16S-targeted sequencing read data 105, extracted DNA from a biological sample (e.g., clinician-collected vaginal swab) is PCR-amplified using primers targeting the 16S rRNA gene. In some embodiments, primers amplify the V3-V4 hypervariable regions with priming sequences located in adjacent conserved regions, yielding amplicons that are purified and sequenced on NGS platforms (e.g., Illumina MiSeq) to produce paired-end reads suitable for downstream ASV inference and taxonomic assignment in microbiome pipeline 110. Alternative primer sets may target different hypervariable regions or combinations thereof, with selection informed by performance characteristics for taxa of interest, while maintaining compatibility with downstream modules 132, 134, and 136.
[0118] In certain embodiments, shotgun metagenomic sequencing is used to generate sequencing read data 105 that encompass bacterial 16S rRNA genes and additional microbial genomes and markers (e.g., bacteriophages, fungi, protozoa, archaea, viruses). Taxonomic classification of shotgun reads may leverage sequence similarity against reference databases (e.g., NCBI, SILVA, Greengenes) using algorithms such as BLAST, Kraken, and Centrifuge, optionally refined with tools like Bracken, and may incorporate binning strategies based on sequence composition (e.g., GC content, k-mer profiles) and coverage patterns, therebyAttomey Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT enabling broader microbiome profiling beyond marker-gene amplicons when clinically warranted.
[0119] Sequencing workflows (e.g., 16S amplicon and shotgun metagenomics) generate large numbers of single-end or paired-end reads stored with quality scores in FASTQ (or FASTA) files. For amplicon-based analyses within clinical assay 100, reads are processed via microbiome pipeline 110 to infer high-resolution AS Vs and assign taxonomy, while shotgun data are aligned or classified to references to determine taxonomic composition and, in some cases, functional potential. These processed outputs form the basis for downstream CST stratification, differential abundance testing, network analyses, and metabolomic pathway prediction via metabolomic prediction pipeline 115 and microbial network pipeline 120.Processing of 16S rRNA Gene Sequencing Using a Microbiome Pipeline
[0120] Following sequencing, sequencing read data 105 are analyzed using the computational pipelines integrated into clinical assay 100 to produce high-resolution, clinically interpretable profiles of the vaginal microbiome. In embodiments configured for 16S rRNA gene sequencing targeting V3-V4, the read data 105 primarily map to the 16S rRNA gene and enable amplicon sequence variant (ASV) inference and taxonomic assignment suitable for downstream stratification and reporting, including CST calls, differential abundance (DA) testing, compositional correlation analyses, and pathway-level functional inference.
[0121] The sequencing read data 105 from BV-negative and BV-positive cohorts are input into the microbiome pipeline 110 designed to process high-throughput microbial community data. In various embodiments, the microbiome pipeline 110 infers AS Vs from amplicon sequence reads by generating a run-specific error model and distinguishing true biological sequences from artifacts introduced during amplification and sequencing. Exemplary ASV inference methods include DADA2, Deblur, MED, and UNOISE; in preferred configurations, the pipeline employs a DAD A2 -based workflow to produce high-resolution ASV profiles for each sample, with learned per-cycle / per-base error rates and consensus merging supporting accuracy and reproducibility.
[0122] To achieve robust ASV detection, the microbiome pipeline 110 employs built-in tools 132, 134, and 136 that collectively perform quality assurance, statistical error modeling, artifact removal, and quantification, and may be implemented as software modules (e.g., R-Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT or Python-based microservices, containerized workflows) executing on general-purpose compute (on-premises servers or cloud infrastructure) with persistent storage, audit logging, and optional hardware acceleration. These tools can operate sequentially or in parallel and are orchestrated to produce reproducible, high-quality amplicon sequence outputs for downstream analysis.
[0123] Tool 132 (read ingestion and quality-control module) is configured to import singleend or paired-end FASTQ files, normalize read orientation, and perform quality filtering and end trimming (e.g., removal of adapters / primers and low-quality tails) using platform- appropriate parameters (such as DADA2 filterAndTrim with optimally chosen trimRight values), while generating QC metrics (read counts, Phred score distributions, base composition) to gate downstream steps. Tool 132 may be a software service that can be multithreaded for performance and integrated with laboratory information management systems (LIMS) for sample traceability, supporting batch and streaming modes to reduce latency in high-throughput contexts.
[0124] Tool 134 (error-rate learning, denoising, and read-merging module) learns a runspecific nucleotide error model from the observed quality scores (e.g., DADA2’s per- cycle / per-base error estimation), applies statistical denoising to correct base-calling errors, and merges overlapping paired-end reads into consensus sequences, persisting learned parameters and processing metadata for auditability. Tool 134 may be implemented in software (e.g., DADA2 vl.22.0 in R), supports parallel execution to reduce latency, and caches intermediate artifacts to minimize repeated I / O; merging parameters are selected to balance overlap length and mismatch tolerance consistent with V3-V4 amplicon sizes.
[0125] Tool 136 (chimera removal, ASV table construction, and taxonomic assignment module) detects and removes PCR chimeras (e.g., removeBimeraDenovo), constructs an ASV count matrix across all samples, and assigns taxonomy using curated reference databases (e.g., SILVA vl38 via the RDP classifier) with confidence thresholds suitable for clinical interpretability. Tool 136 may further perform species-refined speciation by direct alignment (e.g., VSEARCH) to custom reference sets for BV-associated phylotypes (BVAB1 / 2 / 3, Megasphaeral / 2) and cluster-based refinement (e.g., CD-HIT for Lactobacillus speciation), producing a taxonomically aggregated matrix for downstream CST calling, differential abundance testing, network analysis, and metabolomic prediction across pipelinesAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT115 and 120. Tool 136 may be software with optional hardware-assisted indexing / caching of large reference databases to improve throughput and reproducibility across runs.
[0126] After the ASV table and taxonomic assignments are generated, the AS Vs are independently processed by different arms of clinical assay 100 for additional data interpretation. One arm is a metabolomic prediction pipeline 115 that identifies metabolomic pathways potentially dysregulated in association with BV. The metabolomic prediction pipeline 115 ingests the ASV sequences and their relative abundances output by the microbiome pipeline 110 and computes predicted gene family and pathway profiles per sample, followed by statistical differential analysis to identify metabolic pathways potentially dysregulated in association with BV. The pipeline is implemented as a modular software workflow (e.g., containerized R / Python microservices) executing on general -purpose compute (on-premises or cloud), with audit logging of parameters, versions, and provenance to support reproducibility. Inputs typically include: (i) an ASV fasta / abundance table, (ii) taxonomic assignments, and (iii) sample-level metadata (e.g., BV-positive / negative labels derived from NuSwab BV panel results and / or sequencing-based biomarker quantification, and optional CST stratum), enabling group-wise comparisons aligned to clinical contexts.
[0127] Metabolic pathway tool 138 (e.g., PICRUSt2) first performs phylogenetic placement of the observed ASVs onto a reference tree (e.g., via SEPP or internally supported placement methods), then runs hidden-state prediction (HSP) to infer per-ASV gene family abundances (e.g., KEGG orthologs or equivalent gene family units) based on the phylogenetic relatedness of ASVs to sequenced reference genomes. From predicted gene family profiles, tool 138 infers MetaCyc pathway abundances using established pathway inference steps (e.g., MinPath-like logic to assemble genes into pathways), yielding persample predictions of pathway-level functional capacity, with optional exclusion of superpathways / engineered pathways from visualization layers to maintain biological interpretability.
[0128] Metabolic pathway differential analysis tool 140 (e.g., ALDEx2) then evaluates group-wise differences in predicted pathway relative abundances using compositional data methods. ALDEx2 models pathway counts via a Dirichlet-multinomial framework to generate multiple Monte Carlo instances per feature, applies centered log-ratio (clr) transformation to address compositional constraints, and computes effect sizes alongside nonparametric and / or parametric tests (e.g., Wilcoxon rank-sum and Welch’s t-tests), withAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT multiplicity control (e.g., Benjamini-Hochberg FDR) to produce q-values for significance assessment. In preferred configurations, tool 140 performs pairwise comparisons aligned to the clinical assay study design (e.g, CST-I BV-NEG vs CST-IV BV-POS; CST-III BV-NEG vs CST-IV BV-POS; CST-I BV-NEG vs CST-III BV-NEG; and CST-IV BV-POS vs CST- IV BV-NEG), thereby identifying conserved BV-associated pathway perturbations across CSTs and within CST-IV specifically..
[0129] The metabolomic prediction pipeline 115 is configured to surface functional signatures that complement taxonomic biomarkers and network modules, including shortchain fatty acid (SCFA)-linked fermentative pathways consistently enriched in BV-positive samples (e.g, acetyl-CoA fermentation to butanoate, succinate fermentation to butanoate, pyruvate fermentation to propanoate, and L-glutamate / L-glutamine biosynthesis), which imply elevated acetate, succinate, butanoate, and propanoate concentrations in BV states. These conserved enrichments are observed across CSTs and, in some cases, in CST-IV BV- negative samples exhibiting BV-like pathway profiles, supporting utility for early detection or adjudication of indeterminate findings; integrated heatmap visualizations ordered via hierarchical clustering across CST / BV strata are consistent with FIGS. 10A-10B.
[0130] The microbial network pipeline 120 is configured as a modular analytics suite that ingests the ASV table and associated taxonomic assignments output by the microbiome pipeline 110, and then processes these data along two complementary arms: (i) a microbial taxa differential analysis tool 145 and (ii) a microbial taxa correlation analysis tool 150, to jointly quantify condition-associated abundance shifts and coordinated co-occurrence structures across taxa. Inputs include an ASV count matrix; a corresponding taxonomy map aggregated to a lowest common rank (e.g, species); sample-level metadata (e.g, BV- positive / negative labels derived from NuSwab BV panel results and / or sequencing-derived biomarker quantification); and optional stratification variables such as CST.
[0131] Microbial taxa differential analysis tool 145 (e.g, an ANCOM-BC-based module) operates on ASV counts aggregated to the species rank and applies bias-corrective normalization designed for compositional microbiome data, explicitly mitigating distortions arising from variable sampling fractions, library sizes, and zero inflation that can otherwise yield false negative and false positive conclusions. The tool 145 estimates per-taxon log2 fold-change (L2FC) between groups (e.g, BV-POS vs. BV-NEG) using a log-linear modeling framework with sample-specific offset terms and bias correction, and computesAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT multiplicity-adjusted q-values via false discovery rate (FDR) control; enriched taxa may be defined by q-value < 0.05 and L2FC > 1, depleted taxa by q-value < 0.05 and L2FC < -1, and neutral taxa otherwise, consistent with reporting thresholds.
[0132] The microbial taxa correlation analysis tool 150 is designed to infer co-occurrence and coordination among taxa in the presence of compositional constraints by using correlation estimators tailored to relative abundance data. In some embodiments, tool 150 first transforms the taxon abundance vectors with log-ratio methods (e.g., centered log-ratio, additive log-ratio, or isometric log-ratio) to improve numerical stability and interpretability, and then applies Sparse Correlations for Compositional data (SparCC) to estimate pairwise correlations separately within BV-positive and BV-negative groups; bootstrapping or permutation schemes may assess edge robustness, and optional regularization (e.g., LASSO or Elastic Net) may down-weight spurious or weak edges, retaining only correlations above a prespecified magnitude (e.g., |r| > 0.2) for network construction.
[0133] To derive robust modules of co-occurring taxa, pipeline 120 implements a consensus-based clustering procedure (e.g., C3NA) that aggregates multiple base clustering results into a co-association matrix, where each entry (i, j) encodes the frequency with which taxa i and j co-cluster across runs / parameterizations. A final clustering (e.g., hierarchical or spectral) applied to this co-association matrix yields consensus clusters that are less sensitive to any single algorithm or parameter choice; optimal cluster numbers may be determined separately for BV-POS and BV-NEG cohorts to facilitate modular preservation analysis and cross-condition comparisons, with ribbons / arcs linking preserved modules and nodes colored by DA status (enriched / neutral / depleted) and edges colored by correlation magnitude, consistent with FIGS. 9A-9C visualizations.
[0134] Operationally, the microbial network pipeline 120 includes preprocessing steps (e.g., filtering low-prevalence taxa; handling zeros with appropriate pseudo-counts; optional CST stratification), computes summary statistics (cluster membership, within-cluster density, inter-cluster correlation, edge robustness scores), and stores all intermediate artifacts for auditability. Outputs, including differential abundance tables from tool 145, correlation edge lists and consensus clusters from tool 150, and preservation summaries across BV groups, are integrated into reporting layers that accompany taxonomic and CST-contextual results from the microbiome pipeline 110 and functional pathway signals from the metabolomics prediction pipeline 115, enabling clinically interpretable decision support.Attorney Docket No.: 057618-1521428 Client Reference No.: LC 2024-14-WO-PCT
[0135] By way of example, Table 1 provides an example for a false negative conclusion due to failure to account for sample fraction bias. As illustrated, the ecosystem composition of microbiome A includes 4 blue microbes and 4 red microbes (total microbial load of 8), while microbiome B includes 6 blue and 6 red microbes (load of 12). When sampling from microbiome A and B, both samples may include 2 microbes from the blue and red taxa, indicating that the observed abundance for both microbiome A and B are the same and relative abundances are equal, incorrectly suggesting no difference in taxon abundance when sampling fractions differ (A: 1 / 2 vs B: 1 / 3).Table 1: False negative due to failure to account for sample fraction bias
[0136] As illustrated in Table 1, equal observed counts and equal relative abundances can mask true differences in underlying ecosystems when sampling fractions differ, underscoring the importance of bias-corrective normalization methods (e.g., ANCOM-BC) that estimate and adjust for sampling fraction and library size effects in between-group comparisons.
[0137] By way of example, Table 2 provides an example for a false positive scenario due to failure to account for sample fraction bias. The ecosystem composition of microbiome A includes 4 blue and 4 red microbes (load 8), whereas microbiome B includes 12 blue and 4 red microbes (load 16). Sampling from microbiome A may include 2 blue and 2 red microbes (relative abundances of 1 / 2 each), while sampling from microbiome B may include 3 blue and 1 red microbe (relative abundances of 3 / 4 and 1 / 4), incorrectly suggesting both taxa differ when only the blue taxon is different in the underlying ecosystems.Table 2: False positive due to failure to account for sample fraction biasAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0138] As illustrated in Table 2, uncorrected sampling biases may yield spurious differential abundance calls, and bias-corrective normalization methods suitable for compositional microbiome data mitigate this artifact by modeling sampling fractions and library size differences explicitly.
[0139] To avoid the false negative and false positive scenarios described above, data normalization techniques may be used to account for differences in sampling fractions and library sizes. Examples of data normalization methods can include Analysis of Compositions of Microbiomes with Bias Correction (ANCOM-BC), cumulative sum scaling (CSS), Median (MED), Trimmed Mean of M-values (TMM), total-sum scaling (TSS), Upper Quartile (UQ), effective library size corrections (ELib-TMM; ELib-UQ), Wrench, or other appropriate methods known in the art. In various embodiments, the microbial taxa differential analysis tool 145 uses ANCOM-BC normalization methods on the ASV counts to determine if there are differential abundances in taxa between samples.
[0140] In addition to, or as an alternative of using the microbial taxa differential analysis tool 145, microbial network pipeline 120 is used to analyze the relative abundances of ASVs aggregated at different taxonomic ranks to determine relationships across different microbial taxa of a single sample and / or between different microbial taxa across samples. In various embodiments, the microbial network pipeline 120 uses computational tools (e.g., Correlation and Consensus-based Cross-taxonomy Network Analysis (C3NA)) to investigate microbial sequencing data (e.g., 16S rRNA sequencing read data and / or ASVs) to identify cooccurrence patterns across different taxonomic levels within a biological sample (e.g., a BV negative sample and / or a BV positive sample). As used herein, co-occurrence refers to microbes / bacteria occurring in the same sample, and optionally to identify relationships between the microbes / bacteria occurring in the same sample. In various embodiments, the taxa (e.g., class, genus, species) identified within a single sample may be clustered using a consensus-based approach. A consensus-based approach for clustering involves combining multiple clustering results to produce a single, consolidated clustering solution that is more robust and reliable than any individual clustering result. Multiple clusters may be generatedAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT by applying various clustering algorithms (e.g., K-means, hierarchical clustering, DBSCAN) to the AS Vs. Alternatively, the same clustering algorithm may be iteratively used with different parameters or on a different subset of the ASVs to generate multiple clusters. Regardless of how the multiple clusters are generated, each cluster may be converted into an association matrix to capture the co-occurrence of data points in the same cluster across different clustering results. To capture the consensus or agreement between different clustering results, the association matrixes are combined into a co-association matrix, where the entry at position (i, j) indicates how often data points i and j are clustered together across all the generated clusters. A final set of consensus clusters are generated by applying a clustering algorithm (e.g., hierarchical clustering, spectral clustering, or any suitable algorithm) to the co-association matrix. In various embodiments, the clustering method described herein is independently applied to a BV-positive vaginal sample and to a BV- negative vaginal sample.
[0141] To determine relationships between different microbial taxa across samples (e.g., BV-positive versus BV-negative samples as determined by the NuSwab BV panel and / or the 16S rRNA BV biomarker quantification), the microbial network pipeline 120 may implement the microbial taxa correlation analysis tool 150 (e.g., correlations for compositional data, SparCC) for cross-correlation analysis within independent groups of samples (e.g., BV- positive and BV-negative). The goal of cross-correlation analysis is to measure how well the relative abundances of ASVs across samples resemble each other under compositional constraints; log-ratio transformations (e.g., centered / additive / isometric) enhance numerical stability, and regularization (e.g., LASSOZElastic Net) may be applied to retain only significant microbe-microbe correlations. Correlation networks can be visualized where nodes represent components (e.g., taxa / clusters) and edges represent significant correlations (e.g., |r| > 0.2), enabling observation of microbial interaction roles in health and disease and how the microbial community changes in response to environmental changes or B V status, consistent with the arc / ribbon visualizations in FIGS. 9A-9C.III. IDENTIFYING BACTERIAL VAGINOSIS CONSTITUENTS
[0142] FIG. 2 is a flowchart illustrating process 200 for identifying vaginal infection- associated biomarkers (e.g., BV-biomarkers) to determine a diagnosis and / or a treatment plan for a subject. The processing depicted in FIG. 2 may be implemented in software (e.g., code, instructions, program) executed by one or more processing units (e.g., processors, cores) ofAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT the respective systems, hardware, or combinations thereof (e.g., the intelligent selection machine). The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 2 and described below is intended to be illustrative and non-limiting. Although FIG. 2 depicts the various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In certain alternative embodiments, the steps may be performed in some different orders, or some steps may also be performed in parallel. In some embodiments, such as the embodiments depicted in FIG. 1, the processing depicted in FIG. 2 may be performed by the components of the computing environment for clinical assay 100 described with respect to FIG. 1.
[0143] At step 205, a sequencing method is performed on a biological sample collected from a subject. The sequencing method generates read data for the microorganisms within the biological sample. For example, the sequencing method may be 16S rRNA sequencing where the sequencing read data comprises reads from bacterial microorganisms within the biological sample. In other instances, the sequencing method may be shotgun metagenomic sequencing where the sequencing read data comprises reads from bacteria, bacteriophage, fungi, protozoa, archaea, viral, and other microorganisms that may be present in the biological sample. Moreover, any sequencing method known in the art, for example Sanger sequencing, next-generation sequencing (NGS), Illumina sequencing, 454 pyrosequencing, SMRT sequencing, PacBio sequencing, and the like can be used to perform 16S rRNA sequencing, shotgun metagenomic sequencing, or both. To identify BV-associated biomarkers, 16S rRNA sequencing is the sequencing method used. Accordingly, the read data generated from the 16S rRNA sequencing method is amplicon sequencing read data, as the sequencing method includes PCR amplification of the 16S rRNA gene.
[0144] The biological sample collected from the subject can comprise cells, tissues or fluid obtained from a subject suspected of having, is symptomatic, is asymptomatic, diagnosed with, or receiving treatment for one or more vaginal and / or bacterial infections. The one or more vaginal and / or bacterial infections can comprise bacterial vaginosis (BV), a yeast infection, one or more sexually transmitted infections (STIs), a urinary tract infection (UTI), a mycoplasma infection, or any combination thereof. In various embodiments, at least one of the one or more vaginal and / or bacterial infections is bacterial vaginosis (BV). A subject suspected of having BV has not been formally diagnosed, thus the biological sample collected from such a subject is unknown to be BV negative or BV positive. Accordingly, there are some instances in which the biological sample is negative for BV. Additionally,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT there are some instances in which the biological sample is positive for BV and the subject is asymptomatic (or symptomatic) for BV. The biological sample may be cells, tissues, or fluids that comprise nucleic acid molecules (e.g., DNA and / or RNA). In various embodiments the biological sample is a vaginal swab that comprises a collection / population of microorganisms (e.g., bacteria, bacteriophages, fungi, protozoa, archaea, viruses, or any combination thereof) that make up the composition of subject’s vaginal microbiome.
[0145] As described herein, a population of microorganisms that make up the vaginal microbiome can average between 107to 109microorganisms per gram of vaginal fluid. The total number of microorganisms in the vaginal microbiome can vary significantly between individuals and over time due to factors such as hormonal changes, menstrual cycles, sexual activity, and overall health. Accordingly, one of skill in the art can appreciate that (i) the number of microorganisms that are detected using the methods described herein are beyond the processing capabilities of the human mind, (ii) it is unreasonable to interpret a collection or population in this scenario by its simplest definition of two microorganisms as the average vaginal microbiome can comprise between 107to 109different microorganisms, and (iii) the composition of microorganisms from one subject to another are not identical or necessarily consistent, thus any determined proportions, percentages, relative abundance values, etc. of the microorganisms is subject dependent.
[0146] In various embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Campylobacteria, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes, Saccharimonadia, or any combination thereof. More specifically, the microorganisms belong to the taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes, or any combination thereof. More specifically, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes or any combination thereof.
[0147] In various embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic genera [Eubacterium] brachy group, [Ruminococcus] gnavus group, Acidaminococcus,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTAcinetobacter, Actinomyces, Aerococcus, Agathobacter, Alloscardovia, Anaerococcus, Anaeroglobus, Aquabacterium, Arcanobacterium, Atopobium, Bacteroides, Bifidobacterium, Brevibacterium, Bulleidia, Butyri cicoccus, Campylobacter, Clostridium sensu stricto 1, Corynebacterium, Criibacterium, Cryptobacterium, Dermabacter, Dialister, DNF00809, Enterococcus, Eremococcus, Escherichia-Shigella, Ezakiella, Facklamia, Fasti diosipila, Fenollaria, Finegoldia, Fusobacterium, Gallicola, Gardnerella, Gemella, Granulicatella, Haemophilus, Helcococcus, Howardella, HT002, Klebsiella, Lachnospiraceae FE2018 group, Lacticaseibacillus, Lactobacillus, Lawsonella, Limosilactobacillus, Listeria, Mageeibacillus, Megasphaera, Mobiluncus, Mogibacterium, Moryella, Murdochiella, Mycoplasma, Olsenella, Parvimonas, Pelomonas, Peptococcus, Peptoniphilus, Peptostreptococcus, Porphyromonas, Prevotella, Prevotella ?, Pseudomonas, Pseudoramibacter, Ralstonia, Rikenellaceae RC9 gut group, Roseburia, S5-A14a, Shuttleworthia, Sneathia, Solobacterium, Sphingobium, Sphingomonas, Staphylococcus, Streptococcus, Sutterella, Ureaplasma, Varibaculum, Veillonella, or any combination thereof. More specifically, the microorganisms belong to the taxonomic genera Aerococcus, Anaerococcus, Atopobium, Bacteroides, Bulleidia, Corynebacterium, Criibacterium, Dialister, DNF00809, Escherichia-Shigella, Ezakiella, Fasti diosipila, Finegoldia, Gardnerella, Gemella, Howardella, HT002, Lactobacillus, Limosilactobacillus, Mageeibacillus, Megasphaera, Parvimonas, Prevotella, Pseudomonas, Shuttleworthia, Sneathia, Staphylococcus, Streptococcus, Ureaplasma, or any combination thereof. More specifically, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic genera Aerococcus, Anaerococcus, Atopobium, Bacteroides, Bulleidia, Corynebacterium, Criibacterium, Dialister, DNF00809, Escherichia-Shigella, Ezakiella, Fasti diosipila, Finegoldia, Gardnerella, Gemella, Howardella, HT002, Lactobacillus, Limosilactobacillus, Mageeibacillus, Megasphaera, Parvimonas, Prevotella, Pseudomonas, Shuttleworthia, Sneathia, Staphylococcus, Streptococcus, Ureaplasma or any combination thereof.
[0148] In various embodiments, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to the taxonomic species Acidaminococcus intestine, Acinetobacter johnsonii, Actinomyces ihuae, Actinomyces turicensis, Aerococcus christensenii, Alloscardovia omnicolens, Anaerococcus hydrogenalis, Anaerococcus lactolyticus, Anaerococcus obesiensis, Anaerococcus provencensis, Anaerococcus senegalensis, Anaeroglobus geminatus, Aquabacterium parvum, Arcanobacterium urinimassiliense, Atopobium deltae, Atopobium parvulum, AtopobiumAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT vaginae (also known as Fannyhessea vaginae), Bacteroides fragilis, Bacteroides vulgatus, Bifidobacterium bifidum, Bifidobacterium breve, Brevibacterium ravenspurgense, Butyricicoccus faecihominis, BVAB-1 (also known as Candidatus Lachnocurva vaginae), BVAB-2 (also known as Amygdalobacter indicium), BVAB-3 (also known as Mageeibacillus indolicus), Campylobacter faecalis, Campylobacter ureolyticus, Corynebacterium aurimucosum, Corynebacterium coyleae, Corynebacterium mycetoides, Corynebacterium pyruviciproducens, Corynebacterium sundsvallense, Criibacterium bergeronii, Cryptobacterium curtum, Drmabacter jinjuensis, Dialister propionicifaciens, Eremococcus coleocola, Facklamia hominis, Facklamia ignava, Fasti diosipila sanguinis, Finegoldia magna, Fusobacterium nucleatum, Gardnerella vaginalis, Gemella asaccharolytica, Granulicatella elegans, Helcococcus sueciensis, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus iners, Lactobacillus jensenii, Lawsonella clevelandensis, Megasphaera-1, Megasphaera-2, Mobiluncus mulieris, Mogibacterium timidum, Moryella indoligenes, Murdochiella asaccharolytica, Mycoplasma girerdii, Mycoplasma hominis, Olsenella phocaeensis, Parvimonas micra, Pelomonas aquatica, Peptococcus niger, Peptoniphilus coxii, Peptoniphilus duerdenii, Peptoniphilus koenoeneniae, Peptoniphilus lacrimalis, Peptoniphilus obesi, Peptostreptococcus anaerobius, Peptostreptococcus stomatis, Porphyromonas asaccharolytica, Porphyromonas bennonis, Porphyromonas somerae, Porphyromonas uenonis, Prevotella amnii, Prevotella bivia, Prevotella buccalis, Prevotella corporis, Prevotella disiens, Prevotella timonensis, Prevotella ? denticola, Prevotella ? melaninogenica, Pseudoramibacter alactolyticus, Roseburia intestinalis, Sneathia amnii, Sneathia sanguinegens, Solobacterium moorei, Staphylococcus lugdunensis, Ureaplasma urealyticum, Varibaculum cambriense, Veillonella atypica, Veillonella dispar, and Veillonella montpellierensis, or any combination thereof. More specifically, in response to the at least one of the one or more vaginal infections being bacterial vaginosis, the microorganisms belong to taxonomic species Aerococcus christensenii, Anaerococcus obesiensis, Atopobium vaginae (also known as Fannyhessea vaginae), BVAB-1 (also known as Candidatus Lachnocurva vaginae), BVAB-2 (also known as Amygdalobacter indicium), BVAB-3 (also known as Mageeibacillus indolicus), Dialister propionicifaciens, Finegoldia magna, Gardnerella vaginalis, Gemella asaccharolytica, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus jensenii, Megasphaera-1, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens, or any combination thereof.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0149] To determine the bacterial microorganism composition of the biological sample, 16S rRNA sequencing is used to generate bacterial read data (e.g., amplicon sequencing read data) for the 16S ribosomal RNA (16S rRNA) gene. The 16S rRNA gene comprises conserved and variable regions, where the variable regions are used to distinguish between different bacterial species and / or strains. In other words, the variable regions comprise the sequence variants provide species-specific signatures. The 16S rRNA gene has nine hypervariable regions (VI -V9) that exhibit sequence diversity among different bacterial species. Accordingly, the targeted sequencing method can generate read data that align to the entirety of the 16S rRNA gene or a portion of the 16S rRNA gene.
[0150] To generate read data that targets the 16S rRNA gene, PCR primers targeting the 16S rRNA gene are used to amplify the sequence from isolated DNA obtained from the biological samples using any of the DNA isolation methods described herein. In some embodiments, the primers used amplify the entirety of the 16S rRNA gene, thus the read data comprises reads aligning to the entire 16S rRNA gene. In some embodiments, the primers used amplify a portion or a target region of the 16S rRNA gene. As such, the read data comprises reads aligning to the portion or targeted region of the 16S rRNA gene. In other instances, the portion of the 16S rRNA comprises one or more variable regions within the 16S rRNA gene and the read data comprises reads aligning to the one or more variable regions. In some embodiments, the primers used amplify variable regions 3 and 4 of the 16S rRNA gene, where the priming sequences (e.g., sequence where primers can bind to the DNA) are located in the adjacent conserved region surrounding the targeted variable regions. In such an instance, the read data comprises reads aligning to variable regions 3 and 4. The PCR product is purified using known methods in the art and sequenced using DNA sequencing methods described herein to generate the read data.
[0151] At step 210, a clinical assay is used to process the read data (e.g., amplicon sequencing read data) corresponding to the 16S rRNA gene. The clinical assay is responsible for identifying sequence variants (e.g., amplicon sequencing variants (AS Vs)) and obtaining a count and / or relative abundance value (e.g., the relative amount of a bacterial species out of the amount of all bacterial species in a sample) for each identified sequence variant. As described herein, sequence variants, or AS Vs, refer to non-conserved DNA sequences that distinguish species / strains from one another despite being from the same taxonomic genus and class. The non-conserved DNA sequences may affect a single nucleotide (e.g., SNP), 2-5 nucleotides in a row (small insertions / deletions, nucleotide sequence alterations, etc.), and / orAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT larger regions of nucleotides generating variable regions that give a species-specific DNA signature among the conserved and highly conserved regions of the DNA sequence. Once identified, the microbiome pipeline constructs an ASV table that contains counts of each unique ASV across all samples.
[0152] At step 215, in addition to identifying ASVs, the clinical assay also performs clustering and taxonomic assignment of the microorganism based on their read data and identified sequence variants. Taxonomic clustering may be performed using a variety of methods, such as operational taxonomic units (OTUs) or by the ASVs. The OTU clustering strategy “genericizes” the sequences of the microorganisms and groups them based on a sequence similarity threshold such as a sequence similarity of 95%, 96%, 97%, 98%, 99%, or 100%. On the other hand, clustering by ASVs is based purely on the exact sequence similarity between microorganisms, providing a higher resolution and more precise identification of species and strains. In various embodiments, taxonomic clustering uses sequence similarities (e.g., sequence variants and read data) to cluster the same species of microorganisms within the biological sample and then give each cluster a taxonomic assignment. The taxonomic assignment may be at the class, genus, species, or any combination thereof. In various embodiments, taxonomic assignments are given at the species level.
[0153] At step 220, the different species of microorganisms identified in the biological sample are clustered based on their co-occurrence patterns. Co-occurrence refers to microbes / bacteria occurring in the same sample or having a connection with one another. To determine how often different species of microorganisms co-occur together, a consensusbased approach is used to cluster the microorganisms. Consensus clustering involves combining multiple clustering results to produce a single, consolidated clustering solution that is more robust and reliable than any individual clustering result. The consolidated clusters are converted into an association matrix to capture the co-occurrence of data points in the same cluster across different clustering results. For example, the numeric value in the cell at the intersection of microorganism 1 and microorganism 2 indicates how often the two microorganisms co-occur.
[0154] At step 225, a cross-correlation analysis is performed by comparing the clusters identified in the biological sample to clusters belonging to a healthy biological sample to identify significant relationships between the microorganisms. The goal of cross-correlationAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT is to measure how microorganisms in any cluster of the biological sample changes or interacts relative to the microorganisms in any cluster of the healthy biological sample. For example, in a healthy vaginal microbiome, the relative abundance of Lactobacillus species is very high compared to other non-dominant bacterial species. However, the vaginal microbiome of a subject with BV, will have a significant decrease in Lactobacillus species and an increase in Atopobium vaginae (also known as Fannyhessea vaginae), Bacterial Vaginosis Associated Bacterium (BVAB)-2, and Megasphaera-1. Accordingly, the crosscorrelation relationship between Lactobacillus species and any one of Atopobium vaginae (also known as Fannyhessea vaginae), BVAB-2 (also known as Amygdalobacter indicium), or Megasphaera-1 is a negative relationship. Various types of cross-correlation relationships can exist between two biological samples, such as strong / weak positive / negative crosscorrelations, no (zero) correlation, lagging correlation (i.e., the relationship between two variables is not simultaneous but occurs with a delay), periodic or cyclic correlations (abundance increases and decreases in a regular pattern over time), and nonlinear relationships (e.g., quadratic, cubic, exponential, logarithmic, cosine), etc.
[0155] As described herein, a healthy biological sample is to be understood in the broad sense where the vaginal microbiome composition is considered at baseline or in reference to a “normal” vaginal microbiome composition as defined by clinical guidelines. A biological sample may be considered healthy when collected from a woman with no self-reported symptoms, no clinical symptoms, and received negative outcomes for Amsel’s criteria (including Nugent score) testing.
[0156] As described herein, a disease / infection-negative (e.g., a BV-negative) biological sample refers to a sample where key biomarkers are absent and / or do not fall significantly outside of their normal range. For example, a BV-negative sample refers to a sample wherein Atopobium vaginae, BVAB-2, and Megasphaera-1 are not detected, or when either Atopobium vaginae, BVAB-2, or Megasphaera-1 is detected and additional BV biomarkers (e.g., bacterial species among Atopobium vaginae, BVAB-2, and Megasphaera-1) are not detected. In some instances, a BV-negative biological sample may be positive for biomarkers associated with other, non-BV-associated, vaginal infections.
[0157] At step 230, process 200 for identifying BV-associated biomarkers can optionally include a step to perform differential abundance analysis on the sequence variants to identify microorganisms that are enriched or depleted in the biological sample as compared to aAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT healthy biological sample. In various instances, box 230 may be optionally performed or performed simultaneously with the above-described methods in boxes 205-225 to determine whether differentially abundant microbes (e.g., microorganisms) coordinate with each other based on BV status. A taxa (e.g., class, genus, or species) of microorganisms are said to be significantly enriched if the difference in abundance between the biological sample as compared to a healthy biological sample has an FDR < 0.05 and Log2 fold change > 1. Similarly, a taxa (e.g., class, genus, or species) of microorganisms are said to be significantly depleted if the difference in abundance between the biological sample as compared to a healthy biological sample has an FDR < 0.05 and Log2 fold change < -1. Taxa that do not fall within these criteria are identified as neutral (FDR >0.05 or -1 < Log2 fold change < 1).
[0158] In various embodiments, in response to the at least one of the one or more vaginal infections being bacterial vaginosis, the microorganisms belonging to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes or any combination thereof are differentially abundant.
[0159] In various embodiments, in response to the at least one of the one or more vaginal infections being bacterial vaginosis, the microorganisms belonging to taxonomic genera Aerococcus, Anaerococcus, Atopobium, Bacteroides, Bulleidia, Corynebacterium, Criibacterium, Dialister, DNF00809, Escherichia-Shigella, Ezakiella, Fastidiosipila, Finegoldia, Gardnerella, Gemella, Howardella, HT002, Lactobacillus, Limosilactobacillus, Mageeibacillus, Megasphaera, Parvimonas, Prevotella, Pseudomonas, Shuttleworthia, Sneathia, Staphylococcus, Streptococcus, Ureaplasma or any combination thereof are differentially abundant.
[0160] In various embodiments, in response to the at least one of the one or more vaginal infections being bacterial vaginosis, the microorganisms belonging to taxonomic species Aerococcus christensenii, Anaerococcus obesiensis, Atopobium vaginae (also known as Fannyhessea vaginae), BVAB-1 (also known as Candidatus Lachnocurva vaginae), BVAB-2 (also known as Amygdalobacter indicium), BVAB-3 (also known as Mageeibacillus indolicus), Dialister propionicifaciens, Finegoldia magna, Gardnerella vaginalis, Gemella asaccharolytica, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus jensenii, Megasphaera- 1, Megasphaera-2, Prevotella amnii, PrevotellaAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT timonensis, Sneathia amnii, Sneathia sanguinegens, or any combination thereof are differentially abundant.
[0161] At step 235, process 200 can further include the optional step of performing metabolomic analysis on the sequence variants. The goal of this is to identify metabolic pathways, metabolites, or both that are dysregulated in the biological sample as compared to a healthy biological sample. Metabolomic analysis involves identifying metabolic pathways based on marker gene sequencing profiles (e.g., sequence variants / ASVs). To achieve this, ASV sequences and their relative abundances are input into a metabolomics database (MetaCyc, NIH’s metabolomics workbench, etc.) that comprise robust and experimentally validated metabolic pathways. Based on the ASV sequences and their relative abundances, metabolic pathways and metabolites predicted to be reflective of the biological sample’s sequencing profile are output. Examples of metabolic pathways can include, without limitation, acetyl-CoA fermentation to butanoate, L-glutamate and L-glutamine biosynthesis, succinate fermentation to butanoate, and / or pyruvate fermentation to propanoate. In addition to metabolic pathway identification, metabolomic analysis may optionally include performing differential analysis on the pathways / metabolites to determine which of the identified pathways are enriched and associated with BV.
[0162] At step 240, a report is generated comprising (i) a list of the microorganisms identified in the biological sample with their relative abundance and (ii) the cross-correlation between the microorganisms in the biological sample compared to the healthy biological sample, wherein (i) and (ii) are used to determine if the biological sample is positive or negative for a disease (e.g., bacterial vaginosis). The list of microorganisms and their relative abundance may be illustrated on a reference range chart that indicates values of low, normal, and high ranges that are specific to the microorganism. For example, a normal range for a microorganism refers to the expected range observed from a population of healthy biological samples. Low and high ranges indicate values that are significantly lower than or higher than the normal range, where the term significantly refers to a p-value of less than or equal to 0.05. With respect to the term “normal range,” in this context, it is not to be limited to a generalized population of women. For example, vaginal microbiomes are known to have different microbial compositions based on geographical location (e.g., United States versus Africa, versus Europe, etc.). As such, the “normal range” for a specific microorganism for a woman who lives is Africa may be defined by comparing to a population of healthy biological samples collected from women who also live in Africa. Furthermore, “normalAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT range” may also be influenced by the age of the women, for example a 30-year-old female as compared to a 70-year-old female. As such, normal ranges may be defined based on age groups, for example and without limitation, a population of healthy biological samples from women 25-35 years of age, 36-45 years of age, 46-55 years of age, 56-65 years of age, 66-75 years of age, etc. Additional considerations that may be taken into account when defining “normal range” can also include pre / postmenarchal (related to age), pre / postmenopausal (related to age), pregnancy status, oral contraceptive use, and the like. In addition to considering the relative abundance patterns of the microorganisms, their relationships with other microorganisms identified in the biological sample also considered. For example, if the relative abundance of Lactobacillus species is found in the low range and the relative abundances of any one of Atopobium vaginae (also known as Fannyhessea vaginae), BVAB- 2 (also known as Amygdalobacter indicium), or Megasphaera-1 are found in the high range, knowing that the relationship between these bacterial species is negative, this information is considered when making a diagnosis.
[0163] For a microorganism to be considered as a BV-associated biomarker, the presence, relative abundance, and its relationship with other microbes in the biological sample are considered. For many of the identified microorganisms, they naturally are present in a healthy vaginal microbiome, thus a simple binary “yes” or “no” for presence is not sufficient alone to make a diagnosis. Similarly, considering the relative abundance alone of any microorganisms is also insufficient to make a diagnosis as a healthy vaginal microbiome is naturally highly enriched and depleted for specific bacterial species. As such, a combination of at least 2 of the criteria (e.g., presence, abundance, and / or relationship) is considered when making a diagnosis. More preferably, all 3 criteria are used to make a diagnosis. In so doing, the detection rate and treatment of BV in women significantly improves compared to only relying on one of the criteria for a diagnosis.
[0164] Additionally, the report may also include information related to the differential abundance of microorganisms, indicating which are significantly enriched or depleted and which are neutral. Also include may be information related to metabolic pathways and metabolites. In the event these optional tests are ran, the results may also be included in the report.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTIV. CHARACTERIZING A MICROBIOME FOR CLINICAL DECISION SUPPORT
[0165] FIG. 3 is a flowchart illustrating process 300 for characterizing a microbiome for clinical decision support. The processing depicted in FIG. 3 may be implemented in software (e.g., code, instructions, program) executed by one or more processing units (e.g., processors, cores) of the respective systems, hardware, or combinations thereof (e.g., the intelligent selection machine). The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 3 and described below is intended to be illustrative and non-limiting. Although FIG. 3 depicts the various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In certain alternative embodiments, the steps may be performed in some different orders, or some steps may also be performed in parallel. In some embodiments, such as the embodiments depicted in FIG. 1, the processing depicted in FIG. 3 may be performed by the components of the computing environment for clinical assay 100 described with respect to FIG. 1.
[0166] At step 305, a sequencing method is performed on a clinician-collected biological sample from a subject, preferably a vaginal swab obtained under standard-of-care procedures and chain-of-custody, to generate read data for microorganisms present in the sample. In some embodiments, the primary sequencing modality is short-amplicon 16S rRNA gene sequencing targeting V3-V4 with primers annealing in conserved regions flanking the hypervariable regions, producing paired-end reads compatible with downstream ASV inference and taxonomic assignment. In certain embodiments, shotgun metagenomic sequencing is optionally performed to capture additional microbial genomes and markers (e.g., bacteriophages, fungi, protozoa, archaea, viruses) for broader profiling and concordance checks with the amplicon modality. Suitable sequencing platforms include, without limitation, Illumina MiSeq for amplicons and Illumina NextSeq / NovaSeq for shotgun, with read data stored in FASTQ (or FASTA) files along with quality scores. Sequencing read data generated in step 305 constitute “sequencing reads” 105 that are ingested by the clinical assay pipeline depicted in FIG. 1.
[0167] At step 310, the sequencing reads 105 are processed by a microbiome pipeline 110 to identify sequence variants and to construct an ASV table with counts and relative abundances per sample, using modules that include quality filtering, trimming, statistical denoising (e.g., a learned per-cycle / per-base error model), paired-end read merging, andAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT chimera removal. The microbiome pipeline 110 then assigns taxonomy to AS Vs against curated references (e.g., SILVA, RDP, Greengenes, and custom BV phylotype sequences), and, where resolvable, refines species-level assignments with direct alignment to custom panels for BV-associated organisms (e.g., BVAB-1 / 2 / 3; Megasphaera-1 / 2) and clusteringbased speciation for Lactobacillus, yielding a taxonomically aggregated count matrix suitable for downstream analytics. The pipeline 110 further determines the Community State Type (CST) by the most abundant Lactobacillus species detected in each sample (e.g., CST-I: L. crispatus; CST-IL L. gasseri; CST-IIL L. iners; CST-IV: diverse anaerobe-rich; CST-V: L. jensenii), which provides clinically useful context for distinguishing Lactobacillus-dominant communities from diverse anaerobic communities typical of BV-associated dysbiosis. In some embodiments, the pipeline associates the sample with clinical labels (e.g., BV-positive, BV-negative, BV-indeterminate) derived from a panel PCR assay (e.g., NuSwab) for crossmethod concordance and for use as comparators in later steps. Implementation details may include streaming ingestion, multithreaded preprocessing, and caching / indexing of reference databases to reduce latency and improve throughput, consistent with clinical assay 100 in FIG. 1.
[0168] At step 315, the microorganisms identified at step 310 are grouped into taxonomic clusters (e.g., species-level aggregation) to generate a taxonomically aggregated matrix, which serves as a standardized input for differential abundance testing, correlation estimation, consensus clustering, pathway prediction, and reporting. In various embodiments, microbial features (classes, genera, species) relevant to BV — including Lactobacillus, Gardnerella, Prevotella, Sneathia, Dialister, Veillonellaceae members (e.g., Megasphaera), and BVAB phylotypes — are represented to enable robust multi-marker analytics beyond single-target detection.
[0169] At step 320, the species-level matrix is analyzed to derive co-occurrence patterns using a consensus-based clustering approach that aggregates multiple base clusterings into a co-association matrix, with clustering executed separately within BV-positive and BV- negative cohorts to identify within-group modules of co-occurring taxa. In certain embodiments, correlation estimation tailored to compositional data (e.g., SparCC) is used to inform the clustering inputs, following appropriate log-ratio transformations to improve numerical stability and interpretability. The resulting consensus clusters reveal modular signatures, such as distinct clusters that contain canonical BV targets (e.g., Atopobium vaginae, BVAB-2, Megasphaera- 1) and expanded BV-associated taxa (e.g., BVAB-1 / 3,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTMegasphaera-2, Sneathia spp., Prevotella timonensis), supporting multi-target diagnostics and increasing robustness when single markers are borderline.
[0170] At step 325, a cross-correlation analysis compares the consensus clusters derived from the sample (or cohort) to clusters belonging to a healthy reference sample (or cohort) to identify significant positive or negative relationships, modular preservation, and reorganization of microbial communities associated with disease states. A healthy biological sample is understood as a sample whose vaginal microbiome composition is within a baseline or “normal” range and, for example, has no self-reported symptoms, no clinical symptoms, and negative outcomes under Amsel / Nugent testing, providing a benchmark for crosscondition comparison. In some embodiments, correlations above a prespecified magnitude (e.g., |r| > 0.2) are retained for visualization and interpretation, with ribbons / arcs linking preserved modules and node / edge annotations reflecting differential abundance status and correlation strength.
[0171] At step 330, a bias-corrective differential abundance analysis is performed on the species-level matrix to identify microorganisms enriched or depleted in the subject’s cohort relative to healthy or comparator cohorts, using methods designed for compositional microbiome data (e.g., ANCOM-BC) to mitigate distortions from variable sampling fractions, library sizes, and zero inflation. In some embodiments, microorganisms are classified as significantly enriched (e.g., false discovery rate (FDR) < 0.05 and log2 fold-change > 1), significantly depleted (e.g., FDR < 0.05 and log2 fold-change < -1), or neutral (otherwise), with per-taxon statistics (e.g., effect size, confidence intervals, q-values) persisted for subsequent visualization (e.g., volcano plots; boxplots of select taxa). This step may be executed in parallel with step 320 to quantify abundance shifts and to color code nodes within the network modules, enabling integrated interpretation of differential abundance findings and co-occurrence structures.
[0172] At step 335, a metabolomic prediction pipeline 115 infers pathway-level functional profiles from ASV sequences and abundances, typically by phylogenetic placement and hidden-state prediction of gene families followed by assembly into MetaCyc pathways, yielding per-sample predictions of functional capacity. A compositional differential pathway analysis (e.g., ALDEx2) then evaluates pathway-level differences across CST / BV strata and between BV-positive and BV-negative cohorts, applying multiplicity control (e.g., Benjamini-Hochberg) to identify significantly enriched or depleted pathways (e.g., FDR <Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT0.05). In some embodiments, BV-positive samples exhibit conserved enrichment of shortchain fatty acid (SCFA)-linked fermentative pathways, including acetyl-CoA fermentation to butanoate, succinate fermentation to butanoate, pyruvate fermentation to propanoate, and L- glutamate / L-glutamine biosynthesis, implying elevated acetate, succinate, butanoate, and propanoate in BV states and providing complementary signals for adjudicating indeterminate cases.
[0173] At step 340, the system generates a structured report containing: (i) a list of microorganisms identified in the biological sample with taxonomy and relative abundance, (ii) CST stratification derived from the most abundant Lactobacillus species detected, (iii) cross-correlation and consensus-module context indicating preserved / disrupted modules relative to healthy reference clusters, (iv) differential abundance tables with statistical thresholds and vol cano / b oxplot-ready outputs, and (v) predicted pathway-level signatures with statistical measures and optional heatmap-ready dataln some embodiments, relative abundances are contextualized against “normal ranges” derived from healthy population cohorts that may be stratified by geography, age, menopausal status, pregnancy status, and other clinically relevant covariates, with low / normal / high ranges mapped to statistical criteria (e.g., p < 0.05) for each microorganism. Where a PCR panel result exists (e.g., NuSwab targets Atopobium vaginae, BVAB-2, Megasphaera-1), the report optionally provides crossmethod concordance by juxtaposing panel scores with sequencing-derived quantifications for those targets and with expanded biomarkers and pathway signatures
[0174] At step 345, and optionally as part of the report, a BV classification may be determined using predefined rules or a composite decision framework, for example: BV- negative when Atopobium vaginae, BVAB-2, and Megasphaera-1 are not detected, or when one of these is detected without additional BV biomarkers; BV-positive when at least two of Atopobium vaginae, BVAB-2, or Megasphaera-1 are detected; and BV-indeterminate when one of Atopobium vaginae, BVAB-2, or Megasphaera-1 is detected and at least one additional BV biomarker is detectedln certain embodiments, composite criteria may also incorporate consensus-module membership of BV-associated taxa and enrichment of BV- associated pathway signatures to improve specificity and to adjudicate borderline or mixed presentations, consistent with the modular and pathway-level outputs produced in steps 320, 325, 330, and 335.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0175] Cross-correlation networks may display strong / weak positive / negative relationships, no correlation, or patterns suggesting lag or periodicity where sufficient longitudinal data exist; in typical cross-sectional clinical contexts, stable positive / negative associations and modular preservation / rewiring are emphasized for interpretability, with node coloring to indicate DA class and edge coloring to indicate correlation magnitude, facilitating clinical review.
[0176] In various embodiments, computational implementation of process 300 includes streaming / chunked ingestion of FASTQ files; multithreaded quality filtering and denoising; caching / indexing of large reference databases to reduce repeated I / O; sparse matrix representations for ASV count tables and correlation graphs; asynchronous job orchestration for long-running steps (e.g., error-rate learning, consensus clustering, network construction); and containerized modules with persistent audit logs for reproducibility and compliance.
[0177] The structured report produced in step 340 may serialize taxonomic, CST, DA, network, and pathway outputs with versioned parameters and provenance; results can be integrated with laboratory information management systems (LIMS) and surfaced alongside panel PCR outputs for cross-method concordance and clinician-facing decision support.
[0178] In some embodiments, steps 320, 325, 330, and 335 can be executed in parallel or in a pipelined fashion to minimize total turnaround time, and the normal-range contextualization can be cohort- and geography-aware to improve specificity across diverse patient populations, with age and pregnancy status optionally influencing normal-range boundaries.
[0179] As with process 200, the objective of process 300 is to combine at least two independent lines of evidence (presence / absence and relative abundance patterns of multiple BV-associated microorganisms, their co-occurrence relationships and modular context, and pathway-level functional signatures) when making or supporting a BV-related call, with the entire output organized to facilitate clinical interpretation and downstream care decisions.
[0180] Although process 300 is illustrated in a particular sequence, one of ordinary skill will recognize that certain steps can be rearranged, executed concurrently, repeated, or omitted depending on the clinical use case, data quality, or laboratory constraints without departing from the scope of the described methods. Termination of the process occurs upon report generation and, where applicable, computation of a BV classification, though subsequent post-therapeutic monitoring can re-invoke the same pipeline to assess response.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0181] The foregoing description of process 300 is intended to be illustrative and nonlimiting, and can be implemented in software code stored on a non-transitory machine- readable medium and executed by one or more processors; module boundaries in FIG. 1 (e.g., microbiome pipeline 110, metabolomic prediction pipeline 115, microbial network pipeline 120, and tools 132 / 134 / 136 / 138 / 140 / 145 / 150) are exemplary and may be fused or subdivided while preserving the functional characteristics and decision-grade reporting described herein.V. CLINICAL ASSAY CONFIGURED TO CHARACTERIZE VAGINAL MICROBIOMES VIA FULL-LENGTH 16S RRNA GENE SEQUENCING
[0182] As shown in FIG. 4, clinical assay 400 provides an alternative, end-to-end workflow that ingests clinician-collected biological samples (e.g., vaginal swabs) and performs full- length 16S rRNA gene sequencing spanning V1-V9 to increase species-level resolution and reduce misclassification for taxa where short-amplicon configurations (e.g., V3-V4) can be limiting, while harmonizing outputs with the short-amplicon assay for side-by-side interpretability and decision support consistent with the architecture of FIG. 1. In preferred configurations, clinical assay 400 can plug into the same downstream analytics and reporting stack used by clinical assay 100 (namely microbiome pipeline 410 (for CST stratification and taxonomic aggregation), metabolomic prediction pipeline 415 (for pathway inference and compositional differential analysis), and microbial network pipeline 420 (for bias-corrective differential abundance and consensus-based co-occurrence networks)) so laboratories may select sequencing modality based on clinical context, resolution needs, and operational constraints without disrupting validation artifacts or reporting formats.
[0183] FIG.4 illustrates an exemplary computing environment for performing clinical assay 400 (can be same as computing environment used for clinical assay 100 with additional or alternative components, subsystem, modules, etc. or can be a stand alone computing environment from that used to implement clinical assay 100) configured to characterize vaginal microbiomes from bacterial vaginosis (BV) diagnosed samples, and more generally, vaginal infection diagnosed samples, by integrating laboratory processing with bioinformatic analytics to derive decision-support biomarkers and contextual microbial signals, using full- length 16S rRNA gene sequencing spanning V1-V9 as a primary modality. The exemplary environment includes a database management platform 401, a network 402, an assay platform 403, and end devices 404. The database management platform 401 and the assay platform 403 can be implemented using software only (e.g., each module of the platform is a digitalAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT entity implemented using programs, code, or instructions executable by one or more processors), using hardware (e.g., a medical tool to perform testing, a sequencer, a GPU, a CPU, or the like), or using a combination of hardware and software. Although FIG. 4 illustrates a particular set and arrangement of the components, it should be understood that any suitable number or configuration of components may be included in the environment. Additional components such as various sequencing systems, cloud-based data repositories, or parallel computing resources may also be integrated as appropriate for specific implementations. Security measures, such as encrypted data transmission and user authentication, can be implemented to protect sensitive clinical and genomic information during processing and data sharing.
[0184] The database management platform 401 is configured for storing sequencing, microbiome, and subject information, e.g., collections of sequencing read data 405 for various subjects, biomarkers as reference databases, reference cohorts, reference genomes, and the like. The data stores or devices with the database management platform 401 can be deployed using local servers, network-attached storage devices, or cloud-based data warehousing solutions. Various database management systems may be used, including relational databases such as PostgreSQL and MySQL, or scalable NoSQL architectures, selected based on requirements for performance, scalability, and data accessibility. The database management platform 401 is routinely updated to incorporate newly validated references and improvements from ongoing curation, thereby enabling the system to adapt to new developments and expansions in publicly available sequence repositories.
[0185] The database management platform 401 is configured to interact with other components of environment, including the network 102, assay platform 403, and end devices 404. Through these interactions, the database management platform 401 supplies reference datasets and annotation metadata for the microbiome analysis, receives updates and new sequence data, and supports distributed, cloud-based, or hybrid deployments. Data exchange between modules can occur over standard network protocols, high-speed data buses, or cloud application programming interfaces, depending on the system architecture. The database management platform 401 may be deployed as a combination of software applications, such as Python scripts, Docker or Conda environments, or as integrated systems that combine software with dedicated hardware resources including CPUs, GPUs, or high-performance storage appliances. This modular and scalable architecture enables the database management platform 101 to support analysis for a range of microbiomes, accommodate changes inAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT reference data as new bacteria, viruses, etc. and subtypes with various microbiomes are discovered, and integrate with high-throughput sequencing instruments or automated update mechanisms.
[0186] Network 402 is configured to provide robust, high-speed, and secure data communications among the components of environment, including the database management platform 401, the assay platform 403, and end devices 404. Network 402 supports a range of modern networking protocols and architectures, enabling reliable connectivity and efficient data transfer to facilitate the workflows illustrated in FIGS. 1-5. Network 402 may be implemented as a local area network (LAN), a wide-area network (WAN), or a combination of public and private networks, including the Internet, virtual private networks (VPNs), or dedicated research networks. Contemporary network protocols such as TCP / IP, Ethernet, and advanced wireless standards (for example, Wi-Fi 6, Wi-Fi 7, or 5G cellular networks) are supported to provide high bandwidth, low latency, and secure transmission of large volumes of genomic data and analytical results. The network 402 can also integrate optical fiber links and high-throughput backbone connections where ultra-fast data movement is required, such as for distributed storage clusters or remote laboratory facilities.
[0187] To connect the database management platform 401, the assay platform 403, and the end devices 404 to network 402, various types of physical and wireless links may be employed. Wireline connections such as Ethernet, fiber optic cables, or DOCSIS cable modems can deliver reliable and scalable connectivity. Wireless solutions including Wi-Fi, 5G, and Bluetooth can enable mobility, remote access, and ease of installation. For geographically distributed deployments, network 402 may further integrate cloud-based networking services, software-defined networking (SDN), and edge computing nodes to optimize data routing and processing efficiency. The integration of these diverse connection methods ensures a resilient and high-performance data communication framework. Network 402 enables seamless data exchange and real-time interaction among all components of environment, supporting complex workflows and large-scale analysis tasks required for microbiome diagnostics. Security features such as encrypted data transmission, multi-factor authentication, and firewall protections may be incorporated to safeguard sensitive patient and genomic data during transfer and remote access.
[0188] The assay platform 403 is configured to process samples and profile the composition, metabolome, and microbiome network of a microbiome such as vaginalAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT microbiome to identify biomarkers associated with a particular disease state such as vaginal infections. As depicted in FIG. 4, the assay platform 403 comprises several specialized pipelines, tools, and modules described in detail below. These pipelines, tools, and modules may be realized in software, hardware, or a hybrid configuration, and are designed to support high-throughput, accurate, and automated workflows.
[0189] To enable seamless integration with laboratory operations, the assay platform 403 is designed to interact directly with a variety of sequencing modules or sequencers. This interaction may be realized through data transfer protocols and integration with the output systems of next-generation sequencing (NGS) instruments, Sanger sequencers, or other automated sequencing platforms. The assay platform 403 can automatically retrieve sequencing files and associated metadata using direct USB, Ethernet, Wi-Fi, or through connections established with laboratory information management systems (LIMS) via the network 402 that aggregate data across multiple instruments and data stores or device (e.g., those data stores and device that are part of the database management platform 401). In some embodiments, sequencing platforms may be physically linked to dedicated processing servers or high-performance workstations where the assay platform 403 is installed, while in other scenarios, data may be routed through secure cloud-based storage or network-attached servers for centralized access and processing.
[0190] Each of the one or more end devices 404 is an electronic device comprising hardware, software, embedded logic components, or a combination of these elements, and is configured to interact with the database management platform 401, network 402, and the assay platform 403. The end device 404 may include a range of contemporary computing systems such as desktop computers, laptops, workstation computers, tablets, smartphones, portable handheld devices, wearable computing devices, thin clients, or other specialized laboratory or clinical terminals. These computing devices can run various operating systems and application environments, including Windows, macOS, Linux distributions, Android, iOS, or other modern or embedded operating systems.
[0191] The end device 404 may be designed to execute a variety of client-side or webbased applications that support user interaction with the virus subtyping pipeline. For example, the end device 404 may run specialized software for submitting sequence data, reviewing analysis reports, managing database updates, or accessing curated reference datasets. The end device 404 may also support secure user authentication, audit trails, andAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT role-based access control, ensuring that only authorized users can access sensitive genomic and diagnostic information.
[0192] The end device 404 comprises an interface, such as a graphical user interface (GUI), which enables users to interact intuitively with the environment. Through the interface, users can upload sequence data, initiate new microbiome analyses, visualize results, monitor workflow status, or configure system parameters. The interface may support advanced visualization tools for exploring profiles, confidence scores, or analysis, and may also integrate with LIMS or electronic medical records for seamless data exchange.
[0193] The end device 404 is capable of both inputting and receiving data over the network 402. For example, a laboratory technician, clinician, or researcher may use the end device 404 to submit nucleotide sequence data or analysis requests to the assay platform 403. The end device 404 may also be used to retrieve results, download analytical reports, or access the latest updates to reference databases maintained by the database management platform 401. Data transmission between the end device 404 and other components may occur via wired connections (such as Ethernet or USB) or via wireless protocols (such as Wi-Fi, Bluetooth, or cellular networks), depending on the deployment scenario.
[0194] In some embodiments, the end device 404 may also support integration with cloudbased services or distributed computing environments. This enables remote access to the virus subtyping system, supports telemedicine or distributed research collaborations, and allows authorized users to interact with the environment from virtually any location. Security features, including encrypted data transfer, multi-factor authentication, and digital certificates, may be implemented to protect sensitive patient, sample, and genomic data during transmission and remote access.
[0195] The end device 450 may further incorporate notification systems, audit logs, and automated reporting tools, enabling users to receive alerts about workflow completion, system updates, or quality control events. Advanced deployments may allow the end device 404 to interface with laboratory automation platforms, robotic sample handlers, or sequencing instruments for fully automated, end-to-end workflows.
[0196] In some embodiments, the environment may be further augmented with additional components designed to enhance performance, scalability, and adaptability for advanced virus subtyping workflows. For example, high-throughput sequencing systems can be incorporated to generate large volumes of raw sequence data from clinical, environmental, orAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT research samples. These sequencing systems may be directly connected to the assay platform 404, enabling automated transfer of sequencing outputs into the assay platform 404 for immediate downstream processing. Sequencing instruments may be physically located in laboratory environments and interfaced with the network 402 for seamless data integration and real-time analysis.
[0197] Parallel computation resources, such as GPU clusters, multicore CPU servers, or cloud-based high-performance computing environments, may also be deployed within environment to accelerate computationally intensive steps. These resources can be allocated to the database management platform 401 and the assay platform 403, enabling rapid analysis even when processing large sample batches or extensive reference databases. Parallel computing capabilities may also support reference set generation and data management within the database management platform 401, expediting the construction and updating of references.
[0198] The modular architecture of the environment enables flexible scaling and adaptation to a variety of laboratory, clinical, or research settings. Components such as the database management platform 401, network 402, and the assay platform 403 are designed to support distributed processing, remote access, and collaborative workflows, allowing laboratories to accommodate increasing data volumes, support geographically dispersed teams, or integrate with external diagnostic networks. End devices 404, equipped with user interfaces, provide access points for laboratory personnel, clinicians, or researchers to monitor system operations, initiate analyses, review results, and manage database updates from local or remote locations.
[0199] As with clinical assay 100, clinical assay 400 begins with clinician-collected biological samples, preferably vaginal swabs obtained under standard-of-care procedures and chain-of-custody, followed by DNA extraction under conditions that balance robust lysis of diverse taxa (including Gram-positive organisms) with minimization of extraction bias; optional negative extraction blanks and mock-community references may be included for contamination monitoring and process verification. In some embodiments, host-DNA depletion and inhibitor-removal steps are employed to improve microbial signal prior to library preparation for full-length amplification, and the same LIMS integration, sample barcoding, and traceability practices described for clinical assay 100 apply to clinical assay 400 to enable unified tracking and audit across modalities.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0200] Library preparation for full-length 16S rRNA gene sequencing is designed to amplify V1-V9 from extracted microbial DNA using primers annealing to conserved flanking regions and producing ~1.5 kb amplicons suitable for long-read platforms, with platform-compatible barcoding and adapters applied per vendor guidance. In some embodiments, PacBio SMRT sequencing generates high-accuracy circular consensus (HiFi) reads with per-molecule error correction and minimum pass / quality thresholds enforced to ensure species-resolved accuracy; in other embodiments, Oxford Nanopore sequencing produces real-time long reads with platform-appropriate basecalling and polishing (e.g., sup basecalling and Medaka / Racon consensus) to achieve target accuracy for species-level assignmentsSize selection and amplicon QC (e.g., Bioanalyzer / Tapestation traces and qPCR quantification) may be performed prior to sequencing to confirm fragment integrity and library concentrations across multiplexed samples.
[0201] Sequencing generates long-read sequencing reads 405 that serve as inputs to microbiome pipeline 410 comprising platform-aware processing tools as follows: Tool 437 (long-read basecalling, demultiplexing, and quality-control module) imports raw instrument outputs, normalizes read orientation, performs demultiplexing and basecalling with platform- appropriate parameters (e.g., PacBio CCS filters or Nanopore high-accuracy models), and computes QC metrics (read counts, read-length distributions centered on full-length 16S, accuracy / quality estimates) to gate downstream processing. Tool 439 (platform-specific consensus generation and error-correction module) applies per-platform consensus calling and polishing (e.g., CCS for PacBio; Medaka / Racon for Nanopore) to produce high-fidelity amplicon sequences, persisting learned parameters and processing metadata for auditability, with filters selected to balance consensus depth and accuracy consistent with full-length amplicon sizes. Tool 441 (long-read chimera detection and dereplication module) detects and removes PCR chimeras using long-read-appropriate heuristics and constructs a dereplicated feature table of full-length amplicon sequences across all samples, retaining length- and quality-filtered features for downstream taxonomic classification and aggregation. Tool 443 (species-level taxonomic assignment module for full-length 16S) assigns taxonomy using curated full-length reference databases (e.g., SILVA / Greengenes2 / RDP / GTDB) with confidence thresholds suitable for clinical interpretability and refines species-level assignments by direct alignment (e.g., BLASTn / VSEARCH) to a custom BV phylotype panel (e.g., BVAB-1 / 2 / 3; Megasphaera phylotypes), producing a taxonomically aggregated count matrix keyed at the lowest confident rank (preferably species).Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0202] To support cross-method harmonization, clinical assay 400 optionally invokes Tool 455 (cross-method harmonization and concordance module) to map and reconcile taxon identifiers and naming conventions across full-length 16S outputs, short-amplicon outputs, shotgun metagenomics classifications, and panel PCR targets (e.g., NuSwab BV), and to compute concordance summaries at taxon, CST, network, and pathway levels for adjudication in borderline or mixed presentations. In preferred configurations, Tool 455 also serializes harmonization parameters and reference hashes for auditability and methodcomparison studies.
[0203] Quality control and acceptance criteria for clinical assay 400 include per-sample read depth; read-length distributions centered on full-length 16S; platform-specific per-read and per-sample quality thresholds (e.g., QV metrics for PacBio HiFi, minimum estimated accuracies for polished Nanopore consensus); negative-control background levels; mockcommunity recovery metrics (composition and abundance concordance); and chimera rates, with failure modes triggering reprocessing or resequencing per laboratory SOPs. Tool 460 (long-read QC acceptance aggregator) collects run-specific statistics (e.g., SMRT pass counts, barcode balance, on-target percentages, read N50, taxonomic assignment confidence distributions) and persists them for auditable review and longitudinal performance monitoring.
[0204] The compute architecture for clinical assay 400 is designed to maximize throughput and reproducibility while minimizing latency and cost by employing streaming / chunked ingestion of long-read files, GPU- or ASIC-accelerated basecalling where available, multithreaded consensus generation and chimera detection, and caching / indexing of large full-length reference databases to reduce repeated I / O and accelerate taxonomic assignments. As with clinical assay 100, containerized analytics modules, asynchronous job orchestration, sparse matrix representations for count tables and correlation graphs, and persistent audit logs (parameters, software versions, reference hashes, and provenance) support regulated laboratory environments and reproducible runs at scale. In preferred configurations, clinical assay 100 and clinical assay 400 share common microservices for LIMS integration, authentication / authorization, configuration management, and reporting serialization, enabling laboratories to switch modalities without duplicating validation and IT integration efforts.
[0205] Downstream analytics for clinical assay 400 conform to the multi -branch approach used in assay 100. Microbiome pipeline 410 performs CST stratification based on the mostAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT abundant Lactobacillus species detected per sample (e.g., CST-I: L. crispatus; CST-II: L. gasseri; CST-III: L. iners; CST-IV: diverse anaerobe-rich; CST-V: L. jensenii), providing clinically useful context for distinguishing Lactobacillus-dominant communities from BV- associated dysbioses. Downstream of the microbiome pipeline 410, the ASVs are independently processed by a metabolomic prediction pipeline 415 to infer pathway -level functional profiles and to identify BV-associated metabolomic pathways 425 via differential analysis, thereby capturing conserved signatures (e.g., short-chain fatty acid-linked fermentative pathways) that can complement taxonomic markers. In parallel, a microbial network pipeline 420 computes compositional correlations and consensus clustering to derive BV-associated networks 430, revealing modular co-occurrence patterns that expand the BV biomarker universe beyond single targets and improve interpretability in indeterminate or mixed presentations. Outputs from the microbiome pipeline 410, metabolomic prediction pipeline 415, and microbial network pipeline 420 are synthesized into consolidated summaries that align CST stratification and cross-method concordance with expanded biomarker sets and pathway-level signals, enabling clinical reporting consistent with BV diagnostic contexts and supporting more general vaginal infection assessments.
[0206] Metabolomic prediction pipeline 415 can be invoked to infer MetaCyc pathwaylevel functional profiles from full-length 16S taxonomic outputs via phylogenetic placement and hidden-state prediction (metabolic pathway tool 438, e.g., PICRUSt2), followed by compositional differential pathway analysis (metabolic pathway differential analysis tool 440, e.g., ALDEx2) across CST / BV strata to identify significantly enriched or depleted pathways (e.g., FDR < 0.05), with conserved BV-associated signatures including short-chain fatty acid- linked fermentative pathways (e.g., acetyl-CoA fermentation to butanoate, succinate fermentation to butanoate, pyruvate fermentation to propanoate, and L-glutamate / L- glutamine biosynthesis). In some contexts, full-length reads may improve phylogenetic placement certainty and pathway inference stability for closely related taxa, subject to validation.
[0207] Microbial network pipeline 420 can be invoked and includes microbial taxa differential analysis tool 445 (e.g., ANCOM-BC) operating on species-level matrices with bias-corrective normalization to mitigate distortions from variable sampling fractions, library sizes, and zero inflation, classifying taxa as significantly enriched, significantly depleted, or neutral under multiplicity control consistent with reporting thresholds. Microbial taxa correlation analysis tool 450 estimates co-occurrence structures under compositionalAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT constraints using log-ratio transformations and SparCC-based correlation within BV-positive and BV-negative groups, followed by consensus-based clustering (e.g., C3NA) and modular preservation analysis, with node coloring to indicate differential abundance status and edge coloring to indicate correlation magnitude, thereby surfacing modules containing canonical and expanded BV-associated taxa.
[0208] The reporting layer for clinical assay 400 serializes taxonomic, CST, differential abundance, network, and pathway outputs into structured summaries with versioned parameters and provenance, providing: (i) a list of microorganisms identified with specieslevel assignments and relative abundances; (ii) normal-range contextualization by geography, age, and other clinical covariates; (iii) cross-correlation to healthy-reference clusters and visualization of preserved / rewired modules; and (iv) optional cross-method concordance against PCR-panel outputs, all aligned to clinician-facing formats used for assay 100 to facilitate routine interpretation. BV classification may be determined using predefined rules or composite criteria incorporating presence / absence and relative abundance of canonical and expanded BV-associated taxa, CST state, consensus-module context, and pathway-level signatures to adjudicate borderline or mixed presentations, with transparency of decision logic in the signed-out report.
[0209] Operationally, laboratories may select between clinical assay 100 and clinical assay 400 (or run them side-by-side) based on clinical indications (e.g., suspected mixed infections, indeterminate PCR or short-amplicon calls, or cases where species-level resolution impacts care), specimen quality / quantity, turnaround time requirements, and platform availability, with both modalities feeding into the common analytics and reporting backbone to standardize results interpretation, audit, and regulatory documentation. Tool 455 may further support method-comparison studies to quantify agreement across modalities and define reflex criteria and performance specifications for clinical deployment and quality management.
[0210] From a systems perspective, clinical assay 400 extends the configurable, data-driven diagnostic framework by enabling a full-length 16S option that leverages platform-specific strengths (e.g., PacBio HiFi accuracy; Nanopore real-time sequencing) while preserving compositional -bias correction, consensus network modeling, healthy-reference crosscorrelation, and composite BV decision constructs that underpin assay 100, thereby improving analytical rigor and clinical utility across diverse patient populations and laboratory environments.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTVI. CHARACTERIZING A MICROBIOME FOR CLINICAL DECISION SUPPORT
[0211] FIG. 5 is a flowchart illustrating process 500 for identifying characterizing a microbiome for clinical decision support. The processing depicted in FIG. 5 may be implemented in software (e.g., code, instructions, program) executed by one or more processing units (e.g., processors, cores) of the respective systems, hardware, or combinations thereof (e.g., the intelligent selection machine). The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 5 and described below is intended to be illustrative and non-limiting. Although FIG. 5 depicts the various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In certain alternative embodiments, the steps may be performed in some different orders, or some steps may also be performed in parallel. In some embodiments, such as the embodiments depicted in FIG. 4, the processing depicted in FIG. 5 may be performed by the components of the computing environment for clinical assay 400 described with respect to FIG. 4.
[0212] At step 505, a sequencing method is performed on a clinician-collected biological sample from a subject, preferably a vaginal swab obtained under standard-of-care procedures and chain-of-custody, to generate long-read sequencing reads for microorganisms present in the sample. In embodiments directed to clinical assay 400, the primary sequencing modality is full-length 16S rRNA gene sequencing spanning V1-V9, with primers annealing in conserved regions flanking the hypervariable regions to produce ~1.5 kb amplicons compatible with long-read platforms (e.g., PacBio SMRT for high-accuracy circular consensus reads or Oxford Nanopore for real-time long reads), thereby enabling improved species-level resolution relative to short amplicons and facilitating cross-method harmonization with assay 100 outputs.
[0213] At step 510, the long-read sequencing reads 405 are processed to identify full-length amplicon features and construct a species-level matrix suitable for downstream analytics. Platform-aware modules may be invoked in sequence: Tool 437 (long-read basecalling, demultiplexing, and quality-control) to import raw instrument outputs, normalize orientation, demultiplex barcodes, perform basecalling with platform-appropriate models, and compute QC metrics; Tool 439 (platform-specific consensus generation and error correction) to generate high-fidelity amplicon sequences (e.g., CCS for PacBio; Medaka / Racon polishingAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT for Nanopore), persisting parameters and metadata for auditability; Tool 441 (long-read chimera detection and dereplication) to remove PCR chimeras and construct a dereplicated feature table across samples; and Tool 443 (species-level taxonomic assignment for full- length 16S) to assign taxonomy against curated full-length references (e.g., SILVA / Greengenes2 / RDP / GTDB) with refinement via direct alignment to a custom BV phylotype panel (e.g., BVAB- 1 / 2 / 3; Megasphaera phylotypes), yielding a taxonomically aggregated count matrix keyed at the lowest confident rank, preferably species.
[0214] At step 515, the microorganisms identified at step 510 are grouped into taxonomic clusters (e.g., species-level aggregation) to generate a standardized input for downstream CST stratification, differential abundance, correlation estimation, consensus clustering, pathway prediction, and reporting. In some embodiments, CST determinations, bias thresholds, and taxon naming conventions are harmonized with those used by clinical assay 100 to support side-by-side interpretability and method-comparison analyses in regulated laboratory environments.
[0215] At step 520, co-occurrence patterns are derived using a consensus-based clustering approach tailored to compositional data, executed separately within bacterial vaginosis (BV)- positive and BV-negative cohorts to identify within-group modules of coordinated taxa. Microbial taxa correlation analysis tool 450 transforms abundances (e.g., centered / additive / isometric log-ratio) and estimates pairwise correlations using Sparse Correlations for Compositional data (SparCC), with bootstrapping or permutation to assess edge robustness, optional regularization to down-weight spurious edges, and retention of correlations above a prespecified magnitude (e.g., |r| > 0.2) for network construction. The resulting co-association matrix is clustered via a consensus procedure (e.g., C3NA), revealing modular signatures that may contain canonical BV targets (e.g., Atopobium vaginae, BVAB- 2, Megasphaera- 1) and expanded BV-associated taxa (e.g., BVAB-1 / 3, Megasphaera-2, Sneathia, Prevotella timonensis), with node coloring to indicate differential abundance status and edge coloring to indicate correlation magnitude.
[0216] At step 525, a cross-correlation analysis compares the consensus clusters derived from the subject’s sample (or cohort) to clusters belonging to a healthy reference cohort to identify significant positive or negative relationships, modular preservation, and rewiring of microbial communities associated with disease states. Healthy biological samples are understood as baseline or “normal” vaginal microbiomes (e.g., no symptoms, negativeAttomey Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTAmsel / Nugent), providing a benchmark for cross-condition comparison; ribbons / arcs may link preserved modules with visualization retaining only edges above a magnitude threshold (e.g., |r| > 0.2), while node / edge annotations reflect differential abundance status and correlation strength for clinical review.
[0217] At step 530, a bias-corrective differential abundance analysis is performed on the species-level matrix to identify microorganisms enriched or depleted in the subject’s cohort relative to healthy or comparator cohorts, using methods designed for compositional microbiome data (e.g., microbial taxa differential analysis tool 445, ANCOM-BC). This step mitigates distortions arising from variable sampling fractions, library sizes, and zero inflation and classifies taxa as significantly enriched (e.g., FDR < 0.05 and log2 fold-change > 1), significantly depleted (e.g., FDR < 0.05 and log2 fold-change < -1), or neutral (otherwise), with per-taxon statistics persisted for subsequent visualization (e.g., volcano plots and boxplots of select taxa), consistent with assay 100 reporting conventions and thresholds.
[0218] At step 535, the metabolomic prediction pipeline 415 infers pathway-level functional profiles from full-length 16S features and abundances, optionally improving phylogenetic placement certainty for closely related taxa. Metabolic pathway tool 438 (e.g., PICRUSt2) performs phylogenetic placement and hidden-state prediction of gene families, assembles predicted families into MetaCyc pathways, and produces per-sample pathway profiles. Metabolic pathway differential analysis tool 440 (e.g., ALDEx2) then evaluates pathway-level differences across CST / BV strata and between BV-positive and BV-negative cohorts under multiplicity control (e.g., FDR < 0.05), surfacing conserved BV-associated metabolic signatures (e.g., short-chain fatty acid-linked fermentations such as acetyl-CoA fermentation to butanoate, succinate fermentation to butanoate, pyruvate fermentation to propanoate, and L-glutamate / L-glutamine biosynthesis).
[0219] At step 540, the system generates a structured report containing: (i) a list of microorganisms identified in the biological sample at species rank (where resolvable) with relative abundances, (ii) CST stratification derived from the most abundant Lactobacillus species detected (e.g., CST-I: L. crispatus; CST-II: L. gasseri; CST-III: L. iners; CST-IV: diverse anaerobe-rich; CST-V: L. jensenii), (iii) cross-correlation and consensus-module context indicating preserved / disrupted modules relative to healthy reference clusters, (iv) differential abundance tables with statistical thresholds and volcano / boxplot-ready outputs, and (v) predicted pathway-level signatures with statistical measures and optional heatmap-Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT ready data. Relative abundances are contextualized against “normal ranges” derived from healthy population cohorts and stratified by geography, age, menopausal / pregnancy status, or other covariates, with low / normal / high ranges mapped to statistical criteria (e.g., p < 0.05) for each microorganism; where panel PCR results exist (e.g., NuSwab BV), the report optionally provides cross-method concordance by juxtaposing panel scores with full-length 16S quantifications and expanded biomarkers and pathway signatures to support clinician-facing decision-making.
[0220] At step 545, a BV classification may be determined using predefined rules or a composite decision framework, for example: BV-negative when Atopobium vaginae, BVAB- 2, and Megasphaera-1 are not detected, or when one of these is detected without additional BV biomarkers; BV-positive when at least two of Atopobium vaginae, BVAB-2, or Megasphaera-1 are detected; and BV-indeterminate when one of Atopobium vaginae, BVAB-2, or Megasphaera-1 is detected and at least one additional BV biomarker is detected. In certain embodiments, composite criteria incorporate consensus-module membership of BV-associated taxa and enrichment of BV-associated pathway signatures to improve specificity and adjudicate borderline or mixed presentations, consistent with outputs produced in steps 520, 525, 530, and 535.
[0221] Computational implementation includes streaming / chunked ingestion of long-read files; hardware-accelerated basecalling where available (e.g., GPU / ASIC for platform basecalling); multithreaded consensus generation, chimera detection, and dereplication; caching and indexing of large full-length reference databases to reduce repeated I / O; sparse matrix representations for feature tables and correlation graphs; asynchronous job orchestration for long-running steps (e.g., consensus generation, consensus clustering, network construction); and containerized modules with persistent audit logs capturing parameters, software versions, reference hashes, and provenance to support reproducibility and compliance in regulated laboratory environments. Clinical assay 100 and clinical assay 400 share common microservices for LIMS integration, configuration management, authentication / authorization, and reporting serialization to streamline IT integration and validation.
[0222] Quality control and acceptance criteria include per-sample read depth, read-length distributions centered on full-length 16S (~1.5 kb), platform-specific per-read and per-sample quality metrics (e.g., PacBio HiFi QV thresholds; minimum estimated accuracies for polishedAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTNanopore consensus), negative-control background limits, mock-community recovery metrics (composition and abundance concordance), and chimera rates. Tool 460 (long-read QC acceptance aggregator) collects run-specific statistics (e.g., number of passes per molecule for SMRT, barcode balance, on-target percentages, read N50, and taxonomic assignment confidence distributions) and persists them for auditable review and longitudinal performance monitoring; failure modes automatically trigger reprocessing or resequencing according to laboratory SOPs.
[0223] Cross-method harmonization and concordance analysis is optionally performed via Tool 455 (cross-method harmonization and concordance module) to map and reconcile taxon identifiers and naming conventions across full-length 16S outputs, short-amplicon 16S outputs, shotgun metagenomics classifications, and panel PCR targets (e.g., NuSwab BV), producing concordance statistics at one or more of taxon, CST, network module, and pathway levels to inform reflex criteria, adjudicate discordant findings, and support methodcomparison studies.
[0224] Steps 520, 525, 530, and 535 can be executed in parallel or in a pipelined fashion to minimize total turnaround time, and normal-range contextualization can be cohort- and geography-aware to improve specificity across diverse patient populations, with age and pregnancy status optionally influencing normal-range boundaries. The system architecture permits selective invocation of full-length 16S (clinical assay 400) as a reflex to shortamplicon indeterminate calls or to address clinically material species-level distinctions (e.g., Gardnerella spp. or Lactobacillus species), while maintaining unified reporting outputs.
[0225] The structured report produced in step 540 serializes taxonomic, CST, differential abundance, network, and pathway outputs with versioned parameters and provenance; results are integrated with LIMS and surfaced alongside panel PCR outputs for cross-method concordance and clinician-facing decision support. Cross-correlation networks may display strong / weak positive / negative relationships, no correlation, or — where longitudinal data exist — patterns suggesting lag or periodicity; in typical cross-sectional clinical contexts, stable positive / negative associations and modular preservation / rewiring are emphasized to facilitate.
[0226] The objective of process 500 is to combine at least two independent lines of evidence (presence / absence and relative abundance patterns of multiple BV-associated microorganisms, their co-occurrence relationships and modular context, and pathway-levelAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT functional signatures) when making or supporting a BV-related call, with the entire output organized to facilitate clinical interpretation and downstream care decisions. Although process 500 is illustrated in a particular sequence, one of ordinary skill will recognize that certain steps can be rearranged, executed concurrently, repeated, or omitted depending on the clinical use case, data quality, or laboratory constraints without departing from the scope of the described methods. Termination occurs upon report generation and, where applicable, computation of a BV classification, though subsequent post-therapeutic monitoring can reinvoke the same pipeline to assess response.VII. METHODS OF DIAGNOSING AND TREATING
[0227] Any of the methods of detecting BV-associated bacteria (e.g., BV biomarkers) described herein can be used to diagnose a subject with bacterial vaginosis. In some cases, the female subject is suspected of having, symptomatic, asymptomatic, diagnosed with, or being treated for bacterial vaginosis. As used herein the terms diagnose, diagnosis or diagnosing, refer to distinguishing or identifying a disease, syndrome or condition or distinguishing or identifying a person having a particular disease, syndrome or condition. In some embodiments, detection of one or more B V biomarkers in a biological sample collected from a subject is used to diagnose bacterial vaginosis in the subject. The method of diagnosing a subject with bacterial vaginosis may further include (i) the relative abundance of a plurality of BV biomarkers as compared to a normal or expected abundance and / or (ii) the relationships between detected B V biomarkers. In some embodiments, any combination of detecting (e.g., presence or absence), comparing relative abundances, and / or considering the relationships of microorganisms found in the biological sample are used to diagnose the subject with bacterial vaginosis. Any of the methods provided herein can include one or more controls, for example, a positive and / or a negative control. Any of the methods described herein can be used in combination with, for example, a subject’s medical history, a pelvic examination, detection of clue cells, wet mount, and / or a vaginal pH test to diagnose bacterial vaginosis.
[0228] In various embodiments, a BV diagnosis is based on the BV-biomarker composition of the vaginal microbiome of a subject. The BV-biomarker composition is determined by the NuSwab® BV panel, the 16S rRNA BV biomarker quantification via ASV analysis, or both.
[0229] In various embodiments, a BV positive diagnosis is made when the composition of the vaginal microbiome of a subject is: (i) deficient in Lactobacillus species relative to aAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT healthy vaginal microbiome, (ii) enriched with two or more BV-biomarkers, such as Gardnerella, Prevotella, Sneathia, Dialister, BVAB-1, BVAB-3, veillonellaceae, DNF00809, Aerococcus christensenii, Parvimonas, and Bulleidia, (iii) enriched with two or more BV- associated bacteria (e.g., from the NuSwab® BV panel), such as Atopobium vaginae (also known as Fannyhessea vaginae), BVAB-2 (also known as Amygdalobacter indicium), or Megasphaera-1, or (iv) any combination of (i), (ii), and / or (iii).
[0230] In various embodiments, a BV negative diagnosis is made when the composition of the vaginal microbiome of a subject: (i) is depleted for Atopobium vaginae, BVAB-2, and Megasphaera-1, (ii) is depleted for BV-biomarkers not included among Atopobium vaginae, BVAB-2, and Megasphaera-1, such as, without limitation, Aerococcus christensenii, BVAB- 1, BVAB-3, Gardnerella vaginalis, Gemella asaccharolytica, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens, or both (i) and (ii).
[0231] In various embodiments, a BV indeterminate diagnosis is made when the composition of the vaginal microbiome of a subject: (i) either Atopobium vaginae, BVAB-2, or Megasphaera-1 is detected, (ii) at least one additional BV biomarker is detected, or (iii) both (i) and (ii).
[0232] Also provided herein are methods of treating a subject diagnosed with BV. The treatment method includes performing the detection methods described herein to identify B V biomarkers and selecting and administering treatment based on the results of method. Such treatment can be provided in a symptomatic or asymptomatic subject. Optionally, the BV biomarker detection method is repeated after treatment to track progression or improvement based on therapeutic intervention. Treatment refers to ameliorating the bacterial vaginosis in the subject. The term ameliorating refers to any therapeutically beneficial result in the treatment of a bacterial vaginosis, lessening in the severity or progression, delaying or preventing recurrence of the disease, or curing thereof. Thus, treating or treatment includes ameliorating at least one physical parameter or symptom. Treating or treatment includes modulating the disease or disorder, either physically (e.g., stabilization of a discernible symptom) or physiologically (e.g., stabilization of a physical parameter) or both.
[0233] Treatment can include providing to the subject an effective amount of a therapeutic agent such as, antibiotic(s), anti-bacterial medications, and the like. B V may be treated by anti-bacterial medications, including the antibiotics metronidazole, clindamycin, and tinidazole. In various embodiments, the therapeutic agent may be clindamycin oralAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT suppositories or a metronidazole vaginal gel. Some forms of vaginitis are a result of irritation of vaginal lining by common consumer products, such as soaps or lubricants, or particular types of undergarments, and the treatment may be as simple as stopping the use of an irritating product.
[0234] The term effective amount, as used throughout, is defined as any amount necessary to produce a desired physiologic response, for example, reducing or delaying one or more effects or symptoms of a disease or disorder. Effective amounts and schedules for administering the therapeutic agent can be determined empirically, making such determinations within the skill of one in the art. The dosage ranges for administration are those large enough to produce the desired effect in which one or more symptoms of the disease or disorder are affected (e.g., reduced or delayed). The dosage should not be so large as to cause substantial adverse side effects, such as unwanted cross-reactions, unwanted cell death, and the like. Generally, the dosage will vary with the species, age, body weight, general health, sex and diet of the subject, the mode and time of administration and severity of the condition and can be determined by one of skill in the art. The dosage can be adjusted by the individual physician in the event of any contraindications. Dosages can vary and can be administered in one or more doses.
[0235] The therapeutic agents described herein are administered in a number of ways depending on whether local or systemic treatment is desired. The compositions are administered via any of several routes of administration, including intraparenchymal injection, intravenously, intrathecally, intramuscularly, intracisternally, transdermally, or a combination thereof. Effective doses for any of the administration methods described herein can be extrapolated from dose-response curves derived from in vitro or animal model test systems.VIII. EXAMPLES
[0236] The examples below are intended to further illustrate certain aspects of the methods and compositions described herein and are not intended to limit the scope of the claims.Example 1 :Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTImportance and Summary of Key Results
[0237] Bacterial vaginosis (BV) poses a significant health burden for women during reproductive years and onward. Current BV diagnostics rely on either physical and microscopic evaluations by technicians or panels of select microbes.
[0238] In this study, the vaginal microbiomes of 75 clinician-collected remnant NuSwab® vaginal swabs via 16S V3-V4 rRNA gene sequencing were characterized. The rich diversity and CSTs of these vaginal microbiomes with amplicon sequence variant (ASV) resolution was elucidated, showing that BV typically manifests in Lactobacillus-deficient states. Also shown was that amplicon-based sequencing accurately identifies the three microbes targeted by the NuSwab® BV PCR assay, as well as identifying additional BV-associated microbes. Using metabolic predictions from these ASVs, metabolic signatures significantly associated with BV positivity across multiple CSTs were identified. These findings show that amplicon sequencing can accurately reproduce PCR-based testing results and provide an expanded array of biomarkers that may enhance currently available diagnostic tests.
[0239] Remnant clinician-collected NuSwab® vaginal swabs underwent DNA extraction and 16S V3-V4 rRNA gene sequencing to profile microbes in addition to those included in the Labcorp NuSwab® test. Community State Types (CSTs) were determined using the most abundant taxon detected in each sample. PCR results for NuSwab® panel microbial targets were compared against the corresponding microbiome profiles. Metabolic pathway abundances were characterized via metagenomic prediction from amplicon sequence variants (ASVs).
[0240] Sequencing of 75 remnant vaginal swabs yielded 492 unique 16S V3-V4 ASVs, identifying 83 unique genera. NuSwab® assay microbe quantification was strongly concordant with quantification by sequencing (p « 0.01). Samples in CST-I (18 of 18, 100%), CST-II (3 of 3, 100%), CST-III (15 of 17, 88%), and CST-V (1 of 1, 100%) were largely categorized as BV-negative via the NuSwab® panel, while most CST-IV samples (28 of 36, 78%) were BV-positive or BV-indeterminate. BV-associated microbial and predicted metabolic signatures were shared across multiple CSTs.
[0241] These findings show that 16S V3-V4 rRNA gene sequencing robustly reproduces PCR-based BV diagnostic testing results, accurately discriminates vaginal microbiome CSTsAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT dominated by distinct Lactobacilli, and further elucidates BV-associated bacterial and metabolic signatures.Results1. 16S rRNA gene sequencing accurately reproduces PCR-based BV diagnostic testing results
[0242] Seventy-five (75) remnant clinician-collected vaginal swabs were previously analyzed via the Labcorp NuSwab® test, a PCR-based diagnostic method that detects bacterial vaginosis (BV) by scoring three key BV-associated microbes: Atopobium vaginae, BVAB-2, and Megasphaera-1 (FIG. 6A). PCR analysis of these vaginal swabs yielded BV-positive (BV-POS, 27 of 75, 36%), BV-indeterminate (BV-IND, 3 of 75, 4%), and BV-negative (BV- NEG, 45 of 75, 60%) diagnostic results. To further characterize these vaginal swabs, their microbiomes were profiled via 16S V3-V4 rRNA gene sequencing analysis. Sequencing yielded 492 unique 16S V3-V4 Amplicon Sequence Variants (ASVs) that mapped to 83 unique genera, resulting in a mean of 93.6K genus mapped reads per sample. All three NuSwab®-tested microbes were detected among these ASVs, with significant enrichment of 16S relative abundances (RAs) among samples with scores of 2 (FIG. 6B, all p-values « 0.001, Wilcoxon rank-sum tests). This strong corroboration of BV diagnostic results by 16S sequencing analysis highlights the robustness of the NuSwab® test.
[0243] Next, to identify additional characteristics of these vaginal swabs, their ASV- resolved microbiomes were investigated. Juxtaposition of samples by BV status using classlevel taxonomy revealed a strong propensity for Bacilli detection among BV-NEG samples (80.6% mean RA) compared to BV-IND (31%) and BV-POS (11.5%) samples (FIG. 7A). Of the other 10 bacterial classes identified, six had an average RA of 5% or greater in BV-POS and BV-IND samples, while only one of those classes (Actinobacteria) was detected at similar levels in BV-NEG samples (FIG. 7A). Phylogenetic analysis of ASVs confirmed the robustness of these taxonomic classifications, with ASVs closely clustering with those assigned to the same class (FIG. 7B). Bacilli ASVs, consistent with their strong prevalence, constituted the plurality of ASVs (97 of 492, -20%) in the phylogeny and formed one large cluster of those predominantly in the Lactobacillus genus, while two other smaller clusters formed mostly comprising Ureaplasma, Mycoplasma, and Bulleidia (FIGS. 7B-7C). A diverse population of 51 Lactobacillus ASVs was detected, most of which were speciated (FIG. 7C, 38 of 51, 75%). In total, 13 L. iners, 11 L. crispatus, 7 L. jensenii, 5 L. gasseri, andAttorney Docket No.: 057618-1521428 Client Reference No.: LC 2024-14-WO-PCT2 L. hominis ASVs were found, each forming close subgroups in the phylogeny (FIG. 7C). These ASV-resolved, deeply sequenced vaginal microbiomes enable broad assessment of healthy and BV microbiome characteristics, expanding on BV diagnostic insights.2. Characterization of community state types
[0244] Next, samples were further characterized using community state type (CST) analysis, which differentiates samples by their detection of key Lactobacillus species. Specifically, a CST analysis was performed, wherein five CSTs were considered based on the most abundant species detected in each sample (CST-I: L. crispatus, CST-II: L. gasseri, CST-III: L. iners, CST-IV: diverse communities, CST-V: L. jensenii). All five CSTs were observed among the 75 NuSwab® samples in this study with a strongly significant association between CST and BV status (Fisher’s exact test, p < 0.001) (Table 3).Table 3: Community State Type (CST) classifications of NuSwab® vaginal samples. Each table entry indicates the number of samples classified in each CST from each of the NuSwab® result categories (BV-NEG, BV-IND, BV-POS). The marginal totals are shown on the right and bottom of the table.
[0245] While most BV-NEG samples were in Lactobacillus dominated CSTs (37 of 45, 82%), BV-POS and BV-IND samples were largely classified in CST-IV (28 of 30, 93%) (Table 3, FIG. 8A). When clustering sample microbiome profiles using multi-dimensional scaling of their pairwise Bray-Curtis Dissimilarities, significant separation of samples by their CST classification was observed (FIG. 8B, PERMANOVA, p < 0.001).
[0246] It was also observed that BV-POS samples were notably more diverse than BV- NEG samples (FIG. 8C), consistent with their tendency to be within CST-IV (Table 3, FIG. 8A). Interestingly, within CST-IV a significantly higher microbiome diversity among BV- POS and BV-IND samples compared to BV-NEG samples was still observed (FIG. 8C, p < 0.001, Wilcoxon rank-sum test). This suggests that vaginal microbiome classification in CST- IV may represent other forms of dysbiosis not necessarily caused by BV. Together theseAttorney Docket No.: 057618-1521428 Client Reference No.: LC 2024-14-WO-PCT findings indicate that among these NuSwab® samples, BV is associated with diverse, heterogenous microbiomes that generally lack the essential vaginal Lactobacilli.
[0247] To confirm the accuracy of the 16S rRNA sequencing of the V3-V4 hypervariable region for key vaginal Lactobacillus species (e.g. L. iners), shotgun metagenomic sequencing (MGx) was performed on the CST classifications. MGx is used to investigate species-refined vaginal microbiome diversity at the population level. A separate group of 54 remnant NuSwab® vaginal samples were sequenced via both 16S V3-V4 rRNA gene sequencing and MGx, and CST results were generated using the same ruleset as above for both sequencing methods (Table 4).Table 4. Comparison of Community State Type (CST) classifications between 16S rRNA gene V3-V4 sequencing and shotgun metagenomic sequencing (MGx). Rows and columns indicate CST calls determined from MGx and 16S rRNA gene V3-V4 sequencing results, respectively.
[0248] Fifty-three (53) of 54 (98%) samples had concordant CST calls across sequencing methods spanning all five CSTs. Closer inspection of the discrepant CST classification revealed only a difference in L. crispatus (16S: 37.1%, MGX: 61.4%) and L. iners (16S: 52.8%, MGx: 38.2%) relative abundances. Since L. crispatus and L. iners are phylogenetically distinct (FIG. 7C), this relative abundance difference was likely due to biases from the sequencing methods rather than 16S misclassification. These results show that the 16S V3-V4 vaginal Lactobacillus speciation is robust and capable of yielding accurate CST classifications for vaginal swabs.1. BV-POS samples are enriched with distinct bacterial networks
[0249] Since vaginal microbiomes exhibited clear differences in CST classifications and microbial diversity based on BV status (FIGS. 8A-8C, Table 3), the microbial differential abundances between BV-POS and BV-NEG samples were systematically analyzed. It wasAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT found that all three key NuSwab® BV-associated microbes significantly enriched and multiple Lactobacilli species significantly depleted among BV-POS samples (FIG. 9A). Also detected were multiple other BV-associated microbes enriched in BV-POS samples, including BVAB-1, BVAB-3, Megasphaera-2, Sneathia, DNF00809, and Parvimonas (FIGS. 9A-9B).
[0250] The modularity of these microbes was evaluated using a correlation-based analysis that assessed microbe-microbe correlations separately for BV-POS and BV-NEG samples (Methods). This approach enabled simultaneous evaluation of correlation network conservation and whether differentially abundant microbes coordinate with each other based on BV status. Interestingly, each NuSwab® BV-associated microbe was found in a separate correlation-based cluster, suggesting that they broadly capture BV microbial signatures (FIG. 9C). Atopobium vaginae formed a small cluster with only Sneathia sanguinegens (Cluster 10), while BVAB-2 clustered with BVAB-3 (Cluster 13) and Megasphaera-1 formed a broad cluster (Cluster 01) with multiple microbes enriched in BV-POS samples (FIG. 9C). Also identified was a fourth cluster of BV-associated microbes (Cluster 18), including Anaerococcus, Megasphaera-2, Prevotella timonensis, and Sneathia amnii (FIG. 9C). These results indicate that multiple coordinated microbial networks may drive BV pathogenesis and further support the use of multiple diagnostic targets for BV detection.2. Prediction of a BV-associated metabolic signature
[0251] To better understand the functional capacity of BV microbial signatures, the metagenomic predictions of ASVs were determined and differential abundance analysis of the MetaCyc metabolic pathways detected using PICRUSt2 was performed. Since only few samples were classified as CST-II and CST-V (Table 3), the analysis was constrained to CST-I, CST-III, and CST-IV, and it was investigated whether metabolic pathways were significantly enriched or depleted (FDR < 0.05) based on BV positivity within and between CSTs. Since all but two BV-POS samples were in CST-IV, only CST-IV BV-POS samples were considered. Comparisons of CST-I and CST-III BV-NEG samples versus the BV-POS samples separately revealed 270 and 232 differentially abundant (DA) metabolic pathways, respectively (Table 5).Table 5. Differential abundance analysis of predicted metabolic pathways between BV- NEG and BV-POS samples. Each row shows the number of significantly differentially abundant (DA, FDR < 0.05) metabolic pathways between BV-NEG and BV-POS samplesAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT within and between CSTs. In all comparisons, the BV-POS samples were from CST-IV, hence only the BV-NEG CST is shown in the table. The total number of DA pathways is shown on the left followed by the number of pathways that were also DA between BV-NEG samples compared across CST-I and CST-III. On the right, the remaining pathways are shown with the total followed by the number shared with other comparisons and the number unique to the comparison.
[0252] When comparing the CST-IV BV-NEG samples, 36 DA metabolic pathways were detected, all of which were detected in the other two comparisons (Table 5). Furthermore, a separate comparison of BV-NEG samples between CST-I and CST-III revealed only 39 DA metabolic pathways, which only represented a small fraction of the pathways in the other comparisons (Table 5). This indicated that the BV-associated metabolic pathways were largely not due to a deficiency of Lactobacilli and were more likely driven by the microbes enriched in BV-POS samples (FIGS. 9A-9C).
[0253] To visually examine the conservation of the four sample comparisons made above (Table 5, including the BV-NEG CST-I versus CST-III comparison), the most variable DA pathways from each comparison were selected and merged into a single heatmap (FIGS. 10A-10B). Consistent with our DA results (Table 5), BV-POS samples showed highly similar patterns of enriched and depleted pathways regardless of CST (FIGS. 10A-10B).Interestingly, the three BV-NEG samples within CST-IV (the three left most samples within BV-NEG CST-IV in FIGS. 10A-10B) exhibit a similar pattern of pathways with BV-POS samples. This similarity suggests a potential transition from a healthy state to BV, offering a plausible explanation for some patients experiencing B V symptoms while testing negative. Among the pathways significantly enriched in BV-POS samples in all comparisons with BV- NEG samples, four noteworthy pathways were observed: acetyl-CoA fermentation to butanoate, L-glutamate and L-glutamine biosynthesis, succinate fermentation to butanoate, and pyruvate fermentation to propanoate. These enrichments were indicative of higherAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT concentrations of four known BV-related metabolites, namely acetate, succinate, butanoate, and propanoate. Overall, these findings indicate that BV induces large metabolic disruptions in the vaginal microenvironment.Discussion
[0254] In this study, the microbial and metabolic signatures of remnant clinician-collected, PCR tested healthy and BV sample microbiomes were comprehensively profiled. It was demonstrated that 16S V3-V4 rRNA gene sequencing can 1) reproduce Labcorp NuSwab® BV PCR testing results; 2) accurately classify vaginal microbiome CSTs via Lactobacillus speciation; and 3) identify additional microbial and metabolic correlates of BV positivity.
[0255] By comprehensively profiling the vaginal microbiome with ASV-level resolution, multiple features of BV positivity were elucidated. The study found that BV-POS samples tended to be deficient in Lactobacillus species relative to BV-NEG samples, and instead were significantly more diverse and enriched with multiple BV-associated bacteria, such as Gardnerella, Prevotella, Sneathia, Dialister, BVAB-1, BVAB-3, veillonellaceae, DNF00809, Aerococcus christensenii, Parvimonas, and Bulleidia. The three NuSwab® BV-associated microbes (Atopobium vaginae, BVAB-2, and Megasphaera-1) were also among the enriched bacteria in BV-POS samples. Interestingly, a weak enrichment of Gemella asaccharolytica in BV-POS samples was observed, whereas Gemella spp. has been previously found to be less frequently detected in BV-POS samples. The correlation networks of BV-associated bacteria using C3NA were uniquely evaluated, highlighting key modular differences between the BV- POS and BV-NEG samples in this study. More specifically, the unique evaluation has to do with (i) the integration of the computational tools implemented into the pipeline for identifying disease associated constitutes, (ii) a priori knowledge of the NuSwab BV panel microbes and where they cluster within the networks, and (iii) the use of BV-positive and - negative status per the NuSwab BV panel assay. Four clusters of BV-associated microbes were uncovered, and intriguingly the three NuSwab® BV-associated microbes were identified in three separate clusters, indicating that those microbes collectively capture a significant portion of BV microbial signatures. One of those clusters was comprised of Gardnerella vaginalis, Megasphaera-1, DNF00809, and several other taxa.
[0256] Additionally, the metabolic signatures that distinguish vaginal microbiomes based on their CST and / or BV diagnostic status were interrogated. Four noteworthy pathways were observed: acetyl-CoA fermentation to butanoate, L-glutamate and L-glutamine biosynthesis,Attomey Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT succinate fermentation to butanoate, pyruvate fermentation to propanoate. These enrichments indicated higher concentrations of four known BV-related metabolites, namely acetate, succinate, butanoate, and propanoate. Acetate and butanoate are known to play a role in immune responses. Also observed were three pathways significantly depleted within CST-I, namely D-galacturonate degradation, lactose and galactose degradation and 6- hydroxymethyl -dihydropterin diphosphate biosynthesis, and some pathways significantly depleted in CST-III, such as sucrose degradation and L-lysine biosynthesis. The observed depletion of L-lysine biosynthesis in CST-III corresponds with findings from prior studies, which have indicated that / .. iners, the dominant species in CST-III, has a considerably smaller core genome for the biosynthesis of lysine compared to L. crispalus. the dominant species in CST-I. These metabolomic insights represent a complementary suite of biomarkers indirectly detectable by 16S rRNA gene sequencing.
[0257] Multiple commercially available tests exist for diagnosing BV in symptomatic women with no recommendations for those who are asymptomatic. Amsel’s criteria and Nugent score are mainstays of BV diagnosis but may vary in clinical settings due to their subjective nature and technician dependency. Multiplex PCR tests offer stronger sensitivity and specificity at greater expense than Nugent score and Amsel’s criteria but differ in the microbes that they target. This variability in BV diagnostics poses a challenge for uncovering novel BV biomarkers, which can be overcome using sequencing-based technologies. In this study the robustness of 16S rRNA gene sequencing of BV tested remnant vaginal swabs (Labcorp NuSwab® test) was confirmed.
[0258] In this study the vaginal microbiomes of remnant vaginal swabs clinically tested using the Labcorp NuSwab® BV PCR test was profiled. The findings support the accuracy of NuSwab® in identifying BV and provide valuable insights for advancing the diagnostic and treatment options available to patients. These results highlight the potential for 16S rRNA gene profiling in the diagnosis of BV.Materials and Methods1. Sample collection, 16S V3-V4 rRNA gene sequencing, and shotgun metagenomic sequencing
[0259] Seventy-five (75) clinician-collected remnant NuSwab® samples underwent DNA extraction using the ZymoBIOMICS Magbead DNA isolation kit. The 16S rRNA gene V3-Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTV4 hypervariable region was amplified using the commercially available forward and reverse primers. Library preparation was performed using KAPA HyperPrep and KAPA Library amplification kits, and the resulting libraries were sequenced on an Illumina MiSeq® platform using 3OObp paired end reads. A separate set of 54 remnant vaginal swabs were processed through the aforementioned sequencing clinical assay and through a shotgun metagenomic assay comprised of total microbial DNA extraction via the ZymoBIOMICS Magbead DNA isolation kit, library preparation using the KAPA HyperPlus Library Preparation Kit, and sequencing on the Illumina NextSeq2000® platform using 150bp paired end reads.2. DADA2 processing and taxonomic analysis
[0260] 16S rRNA gene V3-V4 sequencing reads were processed using R v4.1.1 by first reorienting reads to the same strand and then trimming using the DADA2 vl.22.0 filterAndTrim function with default parameters and trimRight =c(68, 52) (optimally chosen using in-house R function). Quality filtered trimmed reads were then processed through the standard DADA2 process, denoising and merging reads to generate Amplicon Sequence Variants (ASVs). ASVs were then filtered by removing singletons and those with lengths shorter than 350bp. The resulting ASVs were taxonomically classified using the SILVA vl38 database and DADA2 implemented RDP classifier with a bootstrap confidence threshold of 80. ASVs without at least a Phylum rank classification were removed from analysis.
[0261] Additional ASV speciation was performed by aligning ASVs with a genus rank but no species rank classification directly to the SILVA vl38 database using VSEARCH with parameters “—id 0.99 —strand both —maxgaps 0 — minwordmatches 0 — maxaccepts 0”. Aligned ASVs were assigned to a species if all optimal alignments matched to a single species. Lactobacillus-iocusQ< speciation was then performed by 99% clustering of all ASVs classified in the Lactobacillus genus using CD-HIT. ASVs found in clusters with ASVs previously taxonomically assigned to only a single Lactobacillus species (i.e. via RDP classification or direct alignment), were assigned to the same Lactobacillus species. Since the NuSwab® BV assay targets microbes not in the SILVA nomenclature (B VAB-2) and with phylotype resolution (Megasphaera-1), ASVs with genus rank classification were further taxonomically analyzed by direct alignment using VSEARCH as described above to a custom database of 16S sequences obtained from the Bacterial and Viral Bioinformatics Resource Center (BV-BRC) and additional BVAB-2 and Megasphaera phylotype reference sequences.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTAS Vs with uniquely optimal alignments to BVAB-1, BVAB-2, BVAB-3, Megasphaera-1, Megasphaera-2 were assigned accordingly at the species rank. A final taxonomically aggregated count matrix was then generated by aggregating AS Vs at the lowest assigned taxonomic rank.3. Phylogenetic, diversity, and Community State Type analysis
[0262] Phylogenetic analysis was performed by generating a multiple sequence alignment of all ASVs using Clustal Omega vl .2.4 and processing it in R using the phangorn package. A neighboring joining tree was constructed using the dist.ml and NJ functions. Then, a Jukes- Cantor optimized maximum likelihood tree was generated using the pml and optim.pml functions with optNni = T. The ggtree and dendextend R packages were used to visualize the resulting phylogenies.
[0263] Sample alpha diversity and Community State Types (CST) were determined using the final taxonomically aggregated count matrix (see section ‘DADA2 processing and taxonomic analysis’). Alpha diversity was calculated using the diversity function in the vegan R package. CSTs were determined using the most abundant taxon detected, following the criteria used by Ravel et al (CST-I: L. crispatus, CST-II: L. gasseri, CST-III: L. iners, CST- IV: diverse communities, CST-V: L. jensenii).4. Statistical analysis
[0264] The relationship between sample BV status and CST classification was examined using Fisher’s Exact tests using the fisher.exact function in R59. Differences in microbiome alpha diversity based on BV status and CST were examined using Wilcoxon rank-sum tests using the wilcox.test function in R59.5. Confirmation of 16S rRNA gene-resolved CSTs via shotgun metagenomic sequencing
[0265] To validate the 16S-based CST characterization, an additional set of 54 remnant vaginal swabs were sequenced via 16S rRNA gene V3-V4 sequencing as described above and via shotgun metagenomics (MGx). The CST classification obtained from the 16S and MGx were compared to validate the reliability of CST types from the 16S pipeline. The MGx reads were processed using Illumina bcl2fastq v2.2 and then filtered using BBMAP v38.9873 to ensure high-quality reads with a minimal length of 100 bp. Taxonomic assignments wereAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT performed using Kraken v2.1.274, followed by Bracken v2.775 using the pre-built standard plus Refseq protozoa and Fungi (PlusPF) database. CSTs were determined as described above in section ‘Phylogenetic, diversity, and Community State Type analysis’ using the most abundant species detected from the Bracken read re-assignment output (MGx) and the most abundant taxon from the 16S data processing.6. Differential abundance analysis of microbes
[0266] The differential abundance analysis of microbes between BV-POS and BV-NEG samples was performed by ANCOM-BC R package. The three BV-indeterminant samples were excluded. Significantly enriched taxa were identified with FDR < 0.05 and L2FC >= 1. Significantly depleted taxa were identified with FDR < 0.05 and L2FC <= -1. The remaining taxa were identified as neutral (FDR >= 0.05 or -1 < L2FC < 1).7. Modularized co-occurrence network analysis
[0267] BV-POS and BV-NEG samples were processed using C3NA38 to perform modularized co-occurrence network analysis. To account for the cross-taxonomy design and the unique microbial compositional structure, all cross-taxonomy correlations were obtained using Sparse Correlation for Compositional data (SparCC). Then, the optimal number of taxa clusters was separately determined for BV-POS and BV-NEG samples using a consensus approach implemented by C3NA.8. Prediction and differential abundance analysis of metabolic pathways
[0268] Two CSTs without enough number of samples, CST-II with three samples and CST-IV with one sample, were excluded in the metabolic pathway prediction. One outlier BV-NEG sample ID sample 43 in CST-IV and three BV-indeterminant samples were excluded. The remaining 67 samples from three major CSTs, CST-I, CST-III, and CST-IV, were used to perform metabolic pathway prediction. The MetaCyc metabolic pathways of microbes in samples were predicted using their 16S rRNA gene sequences using PICRUSt2 run with default options. Then, ALDEx2 was run with default options to identify pathways at differential relative abundance between different groups of samples. The difference abundance analysis of metabolic pathways was performed on four pairs of groups of samples, including BV-NEG samples within CST-I and BV-POS samples within CST-IV, BV-NEG samples within CST-III vs BV-POS samples within CST-IV, BV-NEG samples within CST-IAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT vs BV-POS samples within CST-IV, BV-NEG samples within CST-I vs BV-NEG samples within CST-III, and BV-NEG samples within CST-IV vs BV-POS samples within CST-IV. Significantly differentially abundant features among different groups of samples were identified using FDR < 0.05. The heatmap depicts the differential abundances of pathways that varied among different groups was generated based on centered log-ratio transformed score and plotted by pheatmap R package. Any pathways described as “superpathways” or “engineered” were excluded from the heatmap.Example 2:
[0269] Full-length 16S rRNA gene workflow (e.g., see clinical assay 400 and detailed description of FIG. 4) and cross-method harmonization for vaginal microbiome diagnosticsImportance and Summary of Key Results
[0270] Full-length 16S rRNA gene sequencing spanning V1-V9 is designed to increase species-level resolution and reduce misclassification in taxa where short-amplicon configurations (e.g., V3-V4) can be limiting, thereby strengthening diagnostic calls for bacterial vaginosis (BV) and related vaginal infections while maintaining side-by-side interpretability with established workflows and reports (FIG. 11). The clinical assay 400 workflow integrates platform-aware long-read processing (e.g., PacBio SMRT HiFi consensus or Oxford Nanopore high-accuracy basecalling / polishing), species-level taxonomic assignment against curated full-length references supplemented by BV phylotypes, and harmonized downstream analytics (CST stratification, bias-corrective DA testing, consensus network modules, and MetaCyc pathway prediction), consistent with clinical assay 100’s analytics stack. Validation studies (see FIGS. 12-17) indicate that full-length 16S outputs can be reconciled against short-amplicon 16S, shotgun metagenomics, and panel PCR results, with high-level concordance at CST and taxon ranks and improved resolution within genera that are challenging for short amplicons (e.g., Gardnerella and certain Lactobacillus species), supporting reflex testing and adjudication of borderline or indeterminate cases.Overview and Objectives
[0271] The objective of Example 2 is to demonstrate a clinically compatible, end-to-end full-length 16S workflow (clinical assay 400) that (i) ingests clinician-collected vaginal swabs under chain-of-custody; (ii) performs long-read sequencing of V1-V9; (iii) executes platform-specific consensus / error correction; (iv) infers high-fidelity full-length ampliconAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT features and assigns species-level taxonomy; (v) synthesizes outputs via the same CST, DA, network, and pathway analytics used in clinical assay 100; and (vi) serializes decision-grade reports integrated with LIMS, with optional cross-method harmonization to short-amplicon and PCR outputs.Methods Summary
[0272] Remnant clinician-collected vaginal swabs were processed via full-length 16S VIVO amplification, size selection, and barcoded library preparation, followed by long-read sequencing on platform-specific instruments (e.g., SMRT for HiFi reads or Nanopore for real-time long reads), with run-time QC of read-length distributions and per-read accuracy estimates. Long-read basecalling / demultiplexing, consensus generation (e.g., CCS for SMRT; polishing for Nanopore), chimera detection, dereplication, and species-level classification against curated full-length references supplemented with BV phylotypes produced a taxonomically aggregated species-level matrix suitable for downstream analytics. As in clinical assay 100 and Example 1, CST stratification based on the most abundant Lactobacillus species, bias-corrective DA testing (e.g., ANCOM-BC), consensus cooccurrence network analysis (log-ratio transforms, SparCC correlations, C3NA clustering), and MetaCyc pathway prediction / differential analysis (e.g., PICRUSt2 and ALDEx2) were applied to the full-length outputs, enabling harmonized reporting and cross-condition comparisons. Cross-method harmonization mapped and reconciled taxon names / identifiers across full-length outputs, short-amplicon 16S, shotgun metagenomics classifications, and panel PCR targets, producing concordance summaries at taxon, CST, network, and pathway levels for clinician-facing adjudication.Results (see FIGS. 12-17)1. Full-length 16S improves species-level resolution and supports CST stratification
[0273] Species-level assignments derived from full-length reads demonstrated enhanced resolution within clinically material genera (e.g., Lactobacillus speciation across L. crispatus, L. gasseri, L. iners, and L. jensenii; Gardnerella species distinctions were supported by curated references), facilitating robust CST determinations under the same ruleset used in clinical assay 100.Attorney Docket No.: 057618-1521428 Client Reference No.: LC 2024-14-WO-PCT2. Bias-corrective differential abundance and network modules retain diagnostic signal
[0274] Using the taxonomically aggregated species-level matrix, DA testing identified enriched BV-associated taxa and depleted Lactobacilli with thresholds that mitigate compositional biases (sampling fractions, library sizes, zero inflation), consistent with ANCOM-BC reporting conventions described for clinical assay 100. Consensus cooccurrence networks revealed coordinated modules containing canonical panel targets (Atopobium vaginae, BVAB-2, Megasphaera-1) as well as expanded BV taxa (e.g., BVAB- 1 / 3, Megasphaera-2, Sneathia spp., Prevotella timonensis), with modular preservation / rewiring across BV-positive and BV-negative cohorts visualized via arcs / ribbons and node / edge annotations.3. Predicted pathway signatures align with BV-associated metabolic profiles
[0275] Pathway-level functional profiles inferred from full-length features exhibited conserved enrichment of short-chain fatty acid-linked fermentative pathways (e.g., acetyl- CoA fermentation to butanoate, succinate fermentation to butanoate, pyruvate fermentation to propanoate, and L-glutamate / L-glutamine biosynthesis) in BV-positive states across CST strata, consistent with patterns described for assay 100. In select cases, full-length phylogenetic placement appeared to stabilize pathway inference for closely related taxa, subject to validation and reference coverage, with representative heatmaps and pathway lists.4. Cross-method concordance and reflex utility
[0276] Harmonization output demonstrated side-by-side interpretability of full-length 16S, short-amplicon 16S (e.g., V3-V4), shotgun metagenomics, and panel PCR targets (NuSwab), with concordance at CST and taxon ranks and adjudication logic that incorporates networkmodule context and pathway signatures for borderline or indeterminate cases. Cross-method summaries, including reconciliation logic, identifier mapping, and concordance metrics, which support reflex criteria for invoking full-length 16S when short-amplicon resolution is insufficient, results are discordant, or BV calls are indeterminate.Compute architecture, QC, and auditabilityAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT
[0277] The long-read pipeline employs streaming / chunked ingestion, hardware-accelerated basecalling where available, multithreaded consensus generation and chimera detection, caching / indexing of large full-length references, containerized analytics, asynchronous orchestration, sparse data structures, and persistent audit logs for reproducible, regulated laboratory operation. QC acceptance criteria encompass per-sample read depth, read-length distributions centered on full-length 16S (~1.5 kb), platform-specific quality thresholds (e.g., SMRT HiFi QV metrics; Nanopore accuracy estimates), negative-control background limits, mock-community recovery metrics, and chimera rates, with failure modes triggering reprocessing or resequencing per SOPs.Discussion
[0278] These findings indicate that a configurable, auditable full-length 16S clinical assay can be deployed alongside short-amplicon and complementary modalities to increase specieslevel resolution in clinically material genera, preserve diagnostic signal across CST / DA / network / pathway layers, and improve adjudication in borderline or mixed presentations while maintaining common reporting and LIMS integration. Taken together, clinical assay 400 (Example 2) extends the configurable diagnostic framework described for clinical assay 100 (Example 1) and addresses operational needs (computational efficiency, latency, auditability, and cross-method harmonization) associated with sequencing-based diagnostics for BV and related vaginal infections.Materials and Methods (see FIG. 18)1. Sample collection and full-length 16S library preparation
[0279] Clinician-collected remnant vaginal swabs were processed under chain-of-custody, with DNA extraction tuned to robustly lyse diverse taxa and optional host-DNA depletion / inhibitor removal to improve microbial signal. Full-length V1-V9 amplification, size selection, and barcoded library preparation were performed per platform guidance, with pre-sequencing QC (Bioanalyzer / T apestation traces and qPCR quantification) confirming fragment integrity and library concentration.2. Long-read sequencing and primary processing
[0280] Sequencing generated long-read datasets processed by (i) basecalling / demultiplexing with platform-appropriate models; (ii) consensus / error correctionAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT(e.g., CCS for HiFi SMRT; Medaka / Racon polishing for Nanopore); (iii) chimera detection; (iv) dereplication of full-length amplicon features; and (v) species-level classification against curated full-length references supplemented with BV phylotypes, yielding a taxonomically aggregated species-level matrix.3. CST, differential abundance, network, and pathway analytics
[0281] CST stratification was performed using the most abundant Lactobacillus species ruleset consistent with clinical assay 100 (Example 1). Bias-corrective DA testing (e.g., ANCOM-BC), log-ratio transforms with SparCC correlations, consensus clustering via C3NA, and pathway prediction / differential analysis (e.g., PICRUSt2 and ALDEx2) were applied to the species-level matrix, following the same thresholds and visualization conventions (volcano plots, boxplots, network arcs / ribbons, and heatmaps) as clinical assay 100 (Example 1) for harmonized reporting.4. Cross-method harmonization and reporting
[0282] Cross-method harmonization reconciled taxa across full-length outputs and other modalities (short-amplicon 16S, shotgun metagenomics, panel PCR) and produced concordance summaries at multiple levels; reporting serialized taxonomy, CST, DA, network, and pathway outputs with versioned parameters and provenance and optionally juxtaposed PCR panel scores with full-length quantifications and expanded biomarkers / pathways to facilitate clinician decision-making.5. Statistical analysis, QC, and audit
[0283] Multiplicity control (e.g., FDR) and effect size estimates were applied for DA and pathway analyses; QC acceptance criteria gated downstream analytics; and persistent audit logs captured parameters, versions, references, and provenance across pipeline stages for reproducibility and compliance, consistent with clinical assay 100’s auditable compute architecture.IX. ADDITIONAL CONSIDERATIONS
[0284] Implementation of the techniques, blocks, steps and means described above can be done in various ways. For example, these techniques, blocks, steps and means can be implemented in hardware, software, or a combination thereof. For a hardware implementation, the processing units can be implemented within one or more applicationAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, other electronic units designed to perform the functions described above, and / or a combination thereof.
[0285] Also, it is noted that the embodiments can be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart can describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations can be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to a return of the function to the calling function or the main function.
[0286] Furthermore, embodiments can be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting language, and / or microcode, the program code or code segments to perform the necessary tasks can be stored in a machine-readable medium such as a storage medium. A code segment or machine-executable instruction can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a script, a class, or any combination of instructions, data structures, and / or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, ticket passing, network transmission, etc.
[0287] For a firmware and / or software implementation, the methodologies can be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. Any machine-readable medium tangibly embodying instructions can be used in implementing the methodologies described herein. For example, software codes can be stored in a memory. Memory can be implemented within the processor or external to the processor. As used herein the term “memory” refers to any type of long term, short term,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT volatile, nonvolatile, or other storage medium and is not to be limited to any particular type of memory or number of memories, or type of media upon which memory is stored.
[0288] Moreover, as disclosed herein, the term “storage medium,” “storage” or “memory” can represent one or more memories for storing data, including read only memory (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage mediums, optical storage mediums, flash memory devices and / or other machine-readable mediums for storing information. The term "machine-readable medium" includes but is not limited to portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage mediums capable of storing that contain or carry instruction(s) and / or data.
[0289] While the principles of the disclosure have been described above in connection with specific apparatuses and methods, it is to be clearly understood that this description is made only by way of example and not as limitation on the scope of the disclosure.
Claims
Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTWHAT IS CLAIMED IS:
1. A computer-implemented method comprising: performing a sequencing method on a biological sample collected from a subject, wherein the sequencing method generates read data for microorganisms within the biological sample; identifying, using a clinical assay, sequence variants from the read data, wherein the sequence variants distinguish between different species or strains of microorganisms within the biological sample, and wherein the clinical assay generates a relative abundance value for each identified sequence variant; performing, using the clinical assay, taxonomic clustering of the microorganisms based on their sequencing similarities to identify populations of species; clustering, using a consensus-based method, the species within the populations of species identified in the biological sample based on their co-occurrence patterns to generate clusters; performing a cross-correlation analysis by comparing the clusters identified in the biological sample to clusters belonging to a healthy biological sample to identify significant relationships between the microorganisms; and generating a report that comprises (i) a list of the microorganisms identified in the biological sample and (ii) the cross-correlation between the microorganisms in the biological sample compared to the healthy biological sample, wherein (i) and (ii) are used to determine if the biological sample is positive or negative for a disease.
2. The computer-implemented method of claim 1, wherein the sequencing method is 16S rRNA sequencing, shotgun metagenomic sequencing, or both.
3. The computer-implemented method of claim 2, wherein Sanger sequencing, next-generation sequencing (NGS), Illumina MiSeq, Illumina NextSeq, Illumina NovaSeq, 454 pyrosequencing, SMRT sequencing, or PacBio sequencing are used to perform 16S rRNA sequencing, shotgun metagenomic sequencing, or both.
4. The computer-implemented method of claim 2, wherein the sequencing method is 16S rRNA sequencing.
5. The computer-implemented method of claim 1, wherein the read data are amplicon sequencing read data for the 16S rRNA gene.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT6. The computer-implemented method of claim 5, wherein the 16S rRNA gene is sequenced entirely, and the read data comprises reads aligning to the entire 16S rRNA gene.
7. The computer-implemented method of claim 5, wherein a portion of the 16S rRNA gene is sequenced and the read data comprises reads aligning to the portion of the 16S rRNA gene.
8. The computer-implemented method of claim 7, wherein the portion of the 16S rRNA gene sequenced comprises one or more variable regions within the 16S rRNA gene and the read data comprises reads aligning to the one or more variable regions.
9. The computer-implemented method of claim 8, wherein the portion of the 16S rRNA gene sequenced comprises variable regions 3 and 4 of the 16S rRNA gene and the read data comprises reads aligning to variable regions 3 and 4.
10. The computer-implemented method of claim 1, wherein the biological sample is a vaginal swab collected from the subject.
11. The computer-implemented method of claim 10, wherein the vaginal swab comprises the microorganisms comprising the population of species in the vaginal swab.
12. The computer-implemented method of claim 11, wherein the microorganisms comprise bacteria, bacteriophages, fungi, protozoa, archaea, viruses, or any combination thereof.
13. The computer-implemented method of claim 1, wherein the subject is suspected of having, symptomatic, asymptomatic, diagnosed with, or receiving treatment for one or more vaginal infection.
14. The computer-implemented method of claim 13, wherein the subject is symptomatic for one or more vaginal infections.
15. The computer-implemented method of claim 13, wherein the one or more vaginal infections comprises: bacterial vaginosis, a yeast infection, one or more sexually transmitted infections (STIs), urinary tract infection (UTI), mycoplasma infection, or any combination thereof.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT16. The computer-implemented method of claim 15, wherein at least one of the one or more vaginal infections is bacterial vaginosis.
17. The computer-implemented method of claim 16, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Campylobacteria, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes, Saccharimonadia, or any combination thereof.
18. The computer-implemented method of claim 16, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes or any combination thereof.
19. The computer-implemented method of claim 16, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic genera [Eubacterium] brachy group, [Ruminococcus] gnavus group, Acidaminococcus, Acinetobacter, Actinomyces, Aerococcus, Agathobacter, Alloscardovia, Anaerococcus, Anaeroglobus, Aquabacterium, Arcanobacterium, Atopobium, Bacteroides, Bifidobacterium, Brevibacterium, Bulleidia, Butyricicoccus, Campylobacter, Clostridium sensu stricto 1, Corynebacterium, Criibacterium, Cryptobacterium, Dermabacter, Dialister, DNF00809, Enterococcus, Eremococcus, Escherichia-Shigella, Ezakiella, Facklamia,Fasti diosipila, Fenollaria, Finegoldia, Fusobacterium, Gallicola, Gardnerella, Gemella, Granulicatella, Haemophilus, Helcococcus, Howardella, HT002, Klebsiella, Lachnospiraceae FE2018 group, Lacticaseibacillus, Lactobacillus, Lawsonella, Limosilactobacillus, Listeria, Mageeibacillus, Megasphaera, Mobiluncus, Mogibacterium, Moryella, Murdochiella, Mycoplasma, Olsenella, Parvimonas, Pelomonas, Peptococcus, Peptoniphilus, Peptostreptococcus, Porphyromonas, Prevotella, Prevotella_7, Pseudomonas, Pseudoramibacter, Ralstonia, Rikenellaceae RC9 gut group, Roseburia, S5-A14a, Shuttleworthia, Sneathia, Solobacterium, Sphingobium, Sphingomonas, Staphylococcus, Streptococcus, Sutterella, Ureaplasma, Varibaculum, Veillonella, or any combination thereof.
20. The computer-implemented method of claim 16, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong toAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT taxonomic genera Aerococcus, Anaerococcus, Atopobium, Bacteroides, Bulleidia, Corynebacterium, Criibacterium, Dialister, DNF00809, Escherichia-Shigella, Ezakiella, Fasti diosipila, Finegoldia, Gardnerella, Gemella, Howardella, HT002, Lactobacillus, Limosilactobacillus, Mageeibacillus, Megasphaera, Parvimonas, Prevotella, Pseudomonas, Shuttleworthia, Sneathia, Staphylococcus, Streptococcus, Ureaplasma or any combination thereof.
21. The computer-implemented method of claim 16, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic species Acidaminococcus intestine, Acinetobacter johnsonii, Actinomyces ihuae, Actinomyces turicensis, Aerococcus christensenii, Alloscardovia omnicolens, Anaerococcus hydrogenalis, Anaerococcus lactolyticus, Anaerococcus obesiensis, Anaerococcus provencensis, Anaerococcus senegalensis, Anaeroglobus geminatus, Aquabacterium parvum, Arcanobacterium urinimassiliense, Atopobium deltae, Atopobium parvulum, Atopobium vaginae, Bacteroides fragilis, Bacteroides vulgatus, Bifidobacterium bifidum, Bifidobacterium breve, Brevibacterium ravenspurgense, Butyricicoccus faecihominis, BVAB-1, BVAB-2, BVAB-3, Campylobacter faecalis, Campylobacter ureolyticus, Corynebacterium aurimucosum, Corynebacterium coyleae, Corynebacterium mycetoides, Corynebacterium pyruviciproducens, Corynebacterium sundsvallense, Criibacterium bergeronii, Cryptobacterium curtum, Drmabacter jinjuensis, Dialister propionicifaciens, Eremococcus coleocola, Facklamia hominis, Facklamia ignava, Fasti diosipila sanguinis, Finegoldia magna, Fusobacterium nucleatum, Gardnerella vaginalis, Gemella asaccharolytica, Granulicatella elegans, Helcococcus sueciensis, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus iners, Lactobacillus jensenii, Lawsonella clevelandensis, Megasphaera- 1, Megasphaera-2, Mobiluncus mulieris, Mogibacterium timidum, Moryella indoligenes, Murdochiella asaccharolytica, Mycoplasma girerdii, Mycoplasma hominis, Olsenella phocaeensis, Parvimonas micra, Pelomonas aquatica, Peptococcus niger, Peptoniphilus coxii, Peptoniphilus duerdenii, Peptoniphilus koenoeneniae, Peptoniphilus lacrimalis, Peptoniphilus obesi, Peptostreptococcus anaerobius, Peptostreptococcus stomatis, Porphyromonas asaccharolytica, Porphyromonas bennonis, Porphyromonas somerae, Porphyromonas uenonis, Prevotella amnii, Prevotella bivia, Prevotella buccalis, Prevotella corporis, Prevotella disiens, Prevotella timonensis, Prevotella ? denticola, Prevotella ? melaninogenica, Pseudoramibacter alactolyticus, Roseburia intestinalis, Sneathia amnii, Sneathia sanguinegens, Solobacterium moorei,Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCTStaphylococcus lugdunensis, Ureaplasma urealyticum, Varibaculum cambriense, Veillonella atypica, Veillonella dispar, and Veillonella montpellierensis, or any combination thereof.
22. The computer-implemented method of claim 16, in response to the at least one of the one or more vaginal infection being bacterial vaginosis, the microorganisms belong to taxonomic species Aerococcus christensenii, Anaerococcus obesiensis, Atopobium vaginae, BVAB-1, BVAB-2, BVAB-3, Dialister propionicifaciens, Finegoldia magna, Gardnerella vaginalis, Gemella asaccharolytica, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus jensenii, Megasphaera-1, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens, or any combination thereof.
23. The computer-implemented method of claim 1, further comprises performing differential abundance analysis on the sequence variants to identify microorganisms that enriched or depleted in the biological sample.
24. The computer-implemented method of claim 23, wherein the microorganisms belonging to taxonomic classes Actinobacteria, Alphaproteobacteria, Bacilli, Bacteroidia, Clostridia, Coriobacteriia, Fusobacteriia, Gammaproteobacteria, Negativicutes or any combination thereof are differentially abundant.
25. The computer-implemented method of claim 23, wherein the microorganisms belonging to taxonomic genera Aerococcus, Anaerococcus, Atopobium, Bacteroides, Bulleidia, Corynebacterium, Criibacterium, Dialister, DNF00809, Escherichia-Shigella, Ezakiella, Fasti diosipila, Finegoldia, Gardnerella, Gemella, Howardella, HT002, Lactobacillus, Limosilactobacillus, Mageeibacillus, Megasphaera, Parvimonas, Prevotella, Pseudomonas, Shuttleworthia, Sneathia, Staphylococcus, Streptococcus, Ureaplasma or any combination thereof are differentially abundant.
26. The computer-implemented method of claim 23, wherein the microorganisms belonging to taxonomic species Aerococcus christensenii, Anaerococcus obesiensis, Atopobium vaginae, BVAB-1, BVAB-2, BVAB-3, Dialister propionicifaciens, Finegoldia magna, Gardnerella vaginalis, Gemella asaccharolytica, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus hominis, Lactobacillus jensenii, Megasphaera-1, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens, or any combination thereof are differentially abundant.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT27. The computer-implemented method of claim 1, further comprises performing a metabolomics analysis on the sequence variants to identify metabolomic pathways, metabolites, or both that are dysregulated in the biological sample as compared to the healthy biological sample.
28. A computer-implemented method of diagnosing a subject with bacterial vaginosis comprising:(i) performing the method of any one of claims 1-27 to identify BV-associated biomarkers present in the biological sample from the subject, wherein the BV-associated biomarkers comprise Aerococcus christensenii, Atopobium vaginae, BVAB-1, BVAB-2, BVAB-3, Gardnerella vaginalis, Gemella asaccharolytica, Megasphaera-1, Megasphaera-2, Prevotella amnii, Prevotella timonensis, Sneathia amnii, Sneathia sanguinegens;(ii) diagnosing the subject as:BV negative when Atopobium vaginae, BVAB-2, and Megasphaera-1 are not detected, or when either Atopobium vaginae, BVAB-2, or Megasphaera-1 is detected and additional BV biomarkers are not detected,BV positive when at least two of Atopobium vaginae, BVAB-2, or Megasphaera-1 are detected,BV indeterminate when either Atopobium vaginae, BVAB-2, or Megasphaera- 1 is detected and at least one additional BV biomarker is detected; and(iii) treating the bacterial vaginosis by administering a therapeutic agent to the subject.
29. The computer-implemented method of claim 28, wherein the additional BV biomarkers comprise bacterial species among Atopobium vaginae, BVAB-2, and Megasphaera-1.
30. The computer-implemented method of claim 28, wherein the therapeutic agent is clindamycin oral suppositories or metronidazole vaginal gel.
31. A computer-implemented method comprising: extracting microbial DNA from the biological sample;Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT performing full-length 16S rRNA gene sequencing spanning V1-V9 on the extracted microbial DNA to generate long-read sequencing reads; basecalling and demultiplexing the long-read sequencing reads to produce per-sample read sets; performing error correction and consensus generation on the per-sample read sets to obtain full-length 16S amplicon sequences; detecting and removing PCR chimeras and dereplicating the full-length 16S amplicon sequences to construct a feature table; assigning species-level taxonomy to features in the feature table using a curated full- length 16S reference database refined by alignment to a custom panel of bacterial vaginosis- associated phylotypes; aggregating feature counts of the features to a lowest confident taxonomic rank to form a taxonomically aggregated matrix; determining a Community State Type based on a most abundant Lactobacillus species from the taxonomically aggregated matrix; and generating a structured report comprising a list of microorganisms with their relative abundances and the Community State Type.
32. The computer-implemented method of claim 31, further comprising performing a bias-corrective differential abundance analysis on the taxonomically aggregated matrix to estimate per-taxon log2 fold-change between bacterial vaginosis-positive and bacterial vaginosis-negative cohorts under false discovery rate control, and persisting per-taxon statistics for visualization in the structured report.
33. The computer-implemented method of claim 32, further comprising transforming abundances in the taxonomically aggregated matrix by a log-ratio method, computing pairwise correlations using Sparse Correlations for Compositional data separately within the bacterial vaginosis-positive and bacterial vaginosis-negative cohorts, and aggregating multiple base clusterings into consensus clusters, with node annotations referencing the persisted per-taxon statistics from the bias-corrective differential abundance analysis.
34. The computer-implemented method of claim 33, further comprising crosscorrelating the consensus clusters derived from the biological sample to consensus clusters belonging to a healthy reference cohort to identify preserved or rewired modules, andAttorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT retaining correlation edges above a prespecified magnitude threshold for inclusion in the structured report.
35. The computer-implemented method of claim 34, further comprising predicting pathway-level functional profiles by placing features on a phylogenetic reference, inferring gene family content by hidden-state prediction, assembling gene families into MetaCyc pathways, and performing compositional differential pathway analysis across the bacterial vaginosis-positive and bacterial vaginosis-negative cohorts, wherein pathway interpretations leverage the consensus clusters and the persisted per-taxon statistics.
36. The computer-implemented method of claim 35, further comprising computing cross-method concordance by mapping and reconciling taxon identifiers and names across the species-level taxonomy assignments and the Community State Type, and at least one additional modality selected from short-amplicon 16S, shotgun metagenomics, or panel PCR outputs, and producing concordance statistics for taxon, Community State Type, consensus cluster, and pathway levels.
37. The computer-implemented method of claim 36, further comprising applying quality control and acceptance criteria to the long-read sequencing reads and the full-length 16S amplicon sequences, including per-sample read depth, read-length distribution centered on full-length 16S, platform-specific quality metrics, negative-control background limits, mock-community recovery metrics, and chimera rates, and gating the downstream biascorrective differential abundance analysis, consensus clustering, cross-correlation, pathway prediction, and cross-method concordance steps based on the quality control and acceptance criteria.
38. The computer-implemented method of claim 37, wherein the method is executed on a throughput-oriented compute architecture comprising streaming or chunked ingestion of long-read files, hardware-accelerated basecalling, multithreaded consensus generation and chimera detection, caching and indexing of full-length 16S reference databases, containerized analytics modules, asynchronous job orchestration, and persistent audit logs that record parameters, software versions, reference hashes, and provenance for the basecalling, consensus generation, differential abundance analysis, consensus clustering, crosscorrelation, pathway prediction, and cross-method concordance.Attorney Docket No.: 057618-1521428Client Reference No.: LC 2024-14-WO-PCT39. The computer-implemented method of claim 37, further comprising reflexively invoking a full-length 16S workflow when the cross-method concordance step identifies insufficient species-level resolution, discordant results across modalities, or an indeterminate bacterial vaginosis classification, and reusing the taxonomically aggregated matrix, persisted per-taxon statistics, consensus clusters, cross-correlation outputs, and pathway analysis results to update the structured report.
40. The computer-implemented method of claim 39, further comprising computing, for inclusion in the structured report, a bacterial vaginosis classification according to predefined or composite rules that consider: presence or absence and relative abundances from the species-level taxonomy assignments, the Community State Type, the persisted per- taxon statistics from the bias-corrective differential abundance analysis, membership of bacterial vaginosis-associated taxa in preserved or rewired consensus clusters identified by the cross-correlation analysis, and enrichment of bacterial vaginosis-associated pathway signatures identified by the pathway-level functional profiles.
41. A system comprising: one or more processors; and a non-transitory computer readable medium containing instructions which, when executed on the one or more processors, cause the one or more processors to perform the actions or operations of the method in any one of Claim 1-40.
42. A non-transitory computer readable storage medium comprising computer program instructions that, when executed by one or more processors, cause the one or more processors to perform the actions or operations of the method in any one of Claim 1-40.