Methods of pathogen biomarker detection

WO2026198349A1PCT designated stage Publication Date: 2026-09-24LIU CINDY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/019074
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-18
Filing Date
2026-03-13
Publication Date
2026-09-24

Smart Images

  • Figure US2026019074_24092026_PF_FP_ABST
    Figure US2026019074_24092026_PF_FP_ABST
Patent Text Reader

Abstract

A method of determining a genomic biomarker for detecting an organism of interest, such as bacteria, fungi, or a virus, is disclosed. The method involves obtaining comprehensive genomic strain data from the organism of interest; removing at least one subset of genomic strain data from the comprehensive genomic strain data; and filtering the data obtained by requiring the resultant genomic strain data to have at least one pre-selected genomic feature that is common to the identified genomic biomarker.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Ref.: 70851.3W001METHODS OF PATHOGEN BIOMARKER DETECTION CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 773,580 filed March 18, 2025, and entitled “METHODS OF PATHOGEN BIOMARKER DETECTION,” the disclosure of which is incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY-SPONSORED RESEARCH

[0002] This invention was made with U.S. government support (NIH: R01AH25562). The U.S. government has certain rights in the invention.FIELD

[0003] This Application relates to methods for detecting organisms, and particularly for detecting pathogen biomarkers associated with, for example, bacteria, fungi, or a virus.BACKGROUND

[0004] Organisms, such as bacteria, fungi or a vims, can be resident in a tissue, an organ, or a biological sample. Epidemiologically, some of these organisms can be associated with the healthy state of the microbiome in a particular tissue or organ. However, certain organisms, or certain organisms at certain growth levels, can be associated with pathological conditions or disease states. The development of rapid and cost-effective methods for identifying such organisms, such as pathogenic bacteria or fungi, have numerous benefits, including for case-appropriate diagnostics and treatment.SUMMARY

[0005] This Summary is provided to introduce devices and methods for detection using bacterial genome that are further described below in the Detailed Description. This Summary is not intended to identify key aspects or essential aspects of the claimed subject matter.

[0006] All features of exemplary embodiments which are described in this disclosure and are not mutually exclusive can be combined with one another. Elements of one embodiment can be utilized in the other embodiments without further mention. Other aspects and features of the present invention will become apparent to those ordinarily skilled in the art upon review of the following description of specific embodiments in conjunction with any accompanying Figures.Attorney Ref.: 70851.3W001

[0007] In an aspect, a method of determining a genomic biomarker for detecting an organism of interest is disclosed. The method involves obtaining comprehensive genomic strain data from the organism of interest; removing at least one subset of genomic strain data from the comprehensive genomic strain data; and filtering the obtained data by requiring the resultant genomic strain data to have at least one pre-selected genomic feature that is common to the identified genomic biomarker.

[0008] In embodiments, the at least one subset of genomic strain data comprises ribosomal genomic data or homolog genomic data. In embodiments, the method further involves removing at least two subsets of genomic strain data. In embodiments, the at least two subsets of genomic strain data comprise ribosomal genomic data and homolog genomic data.

[0009] In embodiments, the at least one pre-selected genomic feature comprises a requirement that the identified genomic biomarker: is present in all genomes of the organism of interest; has at least 70% sequence identity with taxa members associated with the organism of interest; or contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.

[0010] In embodiments, the method further involves at least two pre-selected genomic features. In embodiments, the at least two pre-selected genomic features are selected from a requirement that the identified genomic biomarker: is present in all genomes of the organism of interest; has at least 70% sequence identity with taxa members associated with the organism of interest; and contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.

[0011] In embodiments, the method further involves at least three pre-selected genomic features. In embodiments, the at least three pre-selected genomic features comprise a requirement that the identified genomic biomarker: is present in all genomes of the organism of interest; has at least 70% sequence identity with taxa members associated with the organism of interest; and contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.Attorney Ref.: 70851.3W001

[0012] In certain embodiments, the comprehensive genomic strain data is associated with a pathogenic organism. In certain embodiments, the comprehensive genomic strain data is associated with a non-pathogenic organism. In certain embodiments, the comprehensive genomic strain data is associated with a bacterium. In a non-limiting example, the bacterium comprises Dolosigranulum pigrum or variants thereof. In a non-limiting example, the bacterium comprises Prevotella bivia or variants thereof. In a non-limiting example, the bacterium comprises Peptostreptococcus anaerobius or variants thereof. In a non-limiting example, the bacterium comprises Dialister micraerophilus or variants thereof. In certain embodiments, the comprehensive genomic strain data is associated with a fungus. In certain embodiments, the comprehensive genomic strain data is associated with a vims.BRIEF DESCRIPTION OF THE FIGURES

[0013] In this Application:

[0014] FIG. 1 is a schematic of a core genome -based approach for an assay design in accordance with embodiments of the disclosure.

[0015] FIG.2A illustrates an aspect of D. pigrum mur] phylogeny and sequence alignment, as a Neighbor joining tree constructed using full length mur] gene sequences from 21 D. pigrum isolates using Jalview 2.11 and ordered by branch lengths, highlighting that rnur] is part of the conserved core genome but is also phylogenetically informative, in accordance with aspects of the disclosure.

[0016] FIG.2B illustrates an aspect of D. pigrum mur] phylogeny and sequence alignment, as multiple sequence alignment of mur] amplicon region, where the forward primer is located at 1234-1255 bp and the reverse primer is located at 1436-1457 bp, in accordance with aspects of the disclosure.

[0017] FIG. 3 illustrates an aspect of D. pigrum core genome-based phylogeny in accordance with aspects of the disclosure.

[0018] FIG. 4 illustrates an aspect of D. pigrum pan-genome, where Pangenome of 21 D. pigrum genomes displayed with the maximum likelihood tree on the left using PHANDANGO in accordance with aspects of the disclosure.

[0019] FIG. 5A illustrates an aspect of quantitative assay validation for P. bivia, P. anaerobius, and D. micraerophilus assays, in accordance with aspects of the disclosure.Attorney Ref.: 70851.3W001

[0020] FIG. 5B illustrates an aspect of D. micraerophilus quantitative validation.DETAILED DESCRIPTION

[0021] A detailed description of one or more embodiments of the invention is provided below along with accompanying figures that illustrate the principles of the invention. The invention is described in connection with such embodiments, but the invention is not limited to any embodiment. The scope of the invention is limited only by the claims. Numerous specific details are set forth in the following description in order to provide a thorough understanding of the invention. These details are provided for the purpose of non-limiting examples and the invention may be practiced according to the claims without some or all of these specific details. For the purpose of clarity, technical material that is known in the technical fields related to the invention has not been described in detail so that the invention is not unnecessarily obscured.

[0022] Overview of Disclosure

[0023] In an aspect, a method of determining a genomic biomarker for detecting an organism of interest is disclosed. The organism may be a bacterium, a fungus, or a virus. The method involves obtaining comprehensive genomic strain data from the organism of interest; removing at least one subset of genomic strain data from the comprehensive genomic strain data; and filtering the obtained data by requiring the resultant genomic strain data to have at least one pre-selected genomic feature that is common to the identified genomic biomarker.

[0024] Definitions and Interpretation

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention pertains. As used herein, and unless stated otherwise or required otherwise by context, each of the following terms shall have the definition set forth below.

[0026] Other examples of implementations will become apparent to the person skilled in the art in view of the teachings of the present description and as such, will not be further described here.

[0027] Note that titles or subtitles may be used throughout the present disclosure for the convenience of the reader, but in no way should these limit the scope of the invention. Moreover, certain theories may be proposed and disclosed herein; however, in no way should they, whether they are right or wrong, limit the scope of the invention so long as the inventionAttorney Ref.: 70851.3W001is practiced according to the present disclosure without regard for any particular theory or scheme of action.

[0028] Any and all references cited throughout the specification are hereby incorporated by reference in their entirety for all purposes.

[0029] It will be understood by those of skill in the art that throughout the present specification, the term “a” used before a term encompasses embodiments containing one or more of what the term refers to. It will also be understood by those of skill in the art that throughout the present specification, the term “comprising”, which is synonymous with “including,” “containing,” or “characterized by,” is inclusive or open-ended and does not exclude additional, un-recited elements or method steps.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. In the case of conflict, the present document, including definitions, will control.

[0031] As used in the present disclosure, the terms “around”, “about” or “approximately” shall generally mean within the error margin generally accepted in the art. Hence, numerical quantities given herein generally include such error margin such that the terms “around”, “about” or “approximately” can be inferred if not expressly stated.

[0032] As used herein, the term “Dolosigranulum pigrum”, which is also referred to herein as “D. pigrum”, refers to all forms of this bacterium, including all variants and mutants thereof.

[0033] As used herein, the term “Prevotella bivia”, which is also referred to herein as “P. bivia”, refers to all forms of this bacterium, including all variants and mutants thereof.

[0034] As used herein, the term “Peptostreptococcus anaerobius”, which is also referred to herein as “P. anaerobius”, refers to all forms of this bacterium, including all variants and mutants thereof.

[0035] As used herein, the term “Dialister micraerophilus” , which is also referred to herein as “D. micraerophilus” , refers to all forms of this bacterium, including all variants and mutants thereof.

[0036] Detailed Description of Aspects and Embodiments of the DisclosureAttorney Ref.: 70851.3W001

[0037] In an aspect of the present disclosure, a method of determining a genomic biomarker for detecting an organism of interest is disclosed. The method involves obtaining comprehensive genomic strain data from the organism of interest; removing at least one subset of genomic strain data from the comprehensive genomic strain data; and filtering the obtained data by requiring the resultant genomic strain data to have at least one pre-selected genomic feature that is common to the identified genomic biomarker.

[0038] In embodiments, the at least one subset of genomic strain data comprises ribosomal genomic data or homolog genomic data. In embodiments, the method further involves removing at least two subsets of genomic strain data. In embodiments, the at least two subsets of genomic strain data comprise ribosomal genomic data and homolog genomic data.

[0039] In embodiments, the at least one pre-selected genomic feature comprises a requirement that the identified genomic biomarker: is present in all genomes of the organism of interest; has at least 70%, or 71%, or 72%, or 73%, or 74%, or 75%, or 76%, or 77%, or 78%, or 79%, or 80%, or 81%, or 82%, or 83%, or 84%, or 85%, or 86%, or 87%, or 88%, or 89%, or 90%, or 91% , or 92%, or 93%, or 94%, or 95%, or greater than 95% sequence identity with taxa members associated with the organism of interest; or contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 70%, or 69%, or 68%, or 67%, or 66%, or 65%, or 64%, or 63%, or 62%, or 61%, or 60%, or 59%, or 58%, or 57%, or 56%, or 55%, or 54%, or 53%, or 52%, or 51%, or 50%, or 49%, or 48%, or 47%, or 46%, or 45%, or 44%, or 43%, or 42%, or 41%, or 40%, or 39%, or 38%, or 37%, or 36%, or 35%, or 34%, or 33%, or 32%, or 31%, or 30%, or below 30% sequence identity and cover against sequences from taxa members associated with the organism of interest.

[0040] In embodiments, the method further involves at least two pre-selected genomic features. In embodiments, the at least two pre-selected genomic features are selected from a requirement that the identified genomic biomarker: is present in all genomes of the organism of interest; has at least 70% sequence identity with taxa members associated with the organism of interest; and contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.

[0041] In embodiments, the method further involves at least three pre-selected genomic features. In embodiments, the at least three pre-selected genomic features comprise aAttorney Ref.: 70851.3W001requirement that the identified genomic biomarker: is present in all genomes of the organism of interest; has at least 70% sequence identity with taxa members associated with the organism of interest; and contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.

[0042] In certain embodiments, the comprehensive genomic strain data is associated with a pathogenic organism. In certain embodiments, the comprehensive genomic strain data is associated with a non-pathogenic organism. In certain embodiments, the comprehensive genomic strain data is associated with a bacterium. In a non-limiting example, the bacterium comprises Dolosigranulurn pigrum or variants thereof. In a non-limiting example, the bacterium comprises Prevotella bivia or variants thereof. In a non-limiting example, the bacterium comprises Peptostreptococcus anaerobius or variants thereof. In a non-limiting example, the bacterium comprises Dialister micraerophilus or variants thereof. In certain embodiments, the comprehensive genomic strain data is associated with a fungus. In certain embodiments, the comprehensive genomic strain data is associated with a virus.

[0043] Rapid and cost-effective methods for the identification of some of these organism or cell species types are needed to facilitate future clinical and in vitro and in vivo studies. Standard biochemical methods are expensive and time-consuming, as are sequencing-based methods.

[0044] For example, MALDI-TOF analysis is cost-effective, but cannot be used to detect certain organisms directly from clinical samples. Accordingly, as embodied and broadly described herein, an aspect of the present disclosure relates to using a core-genome based approach, we designed and validated a PCR-based assay that can be used to confirm isolates of certain organisms, such as I), pigrum. As embodied and broadly described herein, an aspect of the present disclosure relates to using a core-genome based approach, we designed and validated a PCR-based assay that can be used to confirm isolates of certain organisms, such as P. bi via. As embodied and broadly described herein, an aspect of the present disclosure relates to using a core-genome based approach, we designed and validated a PCR-based assay that can be used to confirm isolates of certain organisms, such as P. anaerobius. As embodied and broadly described herein, an aspect of the present disclosure relates to using a core-genome based approach, we designed and validated a PCR-based assay that can be used to confirm isolates of certain organisms, such as D. micraerophilus.Attorney Ref.: 70851.3W001

[0045] As described herein, in an aspect, a method of detecting the presence of D. pigruin in a biological sample is provided. The method involves extracting DNA from the biological sample; and performing a molecular hybridization reaction on the extracted DNA to detect a gene or gene fragment that is unique to D. pigrum. In embodiments, the gene or gene fragment comprises a core genome-based candidate gene for D. pigrum. In certain embodiments, the gene or gene fragment comprises a gene selected from Table 5. In embodiments, the gene or gene fragment is unique to D. pigrum comprises D. pigrum inuri. In embodiments, the molecular hybridization process comprises introducing at least one oligonucleotide primer optimized to detect D. pigrum murJ.

[0046] Without limiting any other disclosure herein, the methods disclosed herein identify novel, sensitive, and specific microbial targets for bacteria and fungi at the specified taxonomic levels (e.g., genus, species, sequence type). We can further filter the microbial targets using bioinformatics tools to identify microbial targets that meet specific criteria, such as transmembrane proteins, resistance-conferring genes, and secreted proteins. Accordingly, these methods can identify novel microbial targets that could be used for:

[0047] PCR-based diagnostic assays for bacterial and fungal identification'. In this case, we design primers and probe sequences that will bind to regions of the novel targets using our method, which can be further developed into rapid PCR / qPCR / RT-qPCR. After identifying novel markers using our methods, we have successfully designed and validated PCR (2 primers; presence / absence detection) and real-time PCR (2 primers and 1 probe; presence / absence detection and quantification).

[0048] Immunoassay-based diagnostic assays for bacterial and fungal identification'. After identifying novel markers using our methods, these markers or portions of the markers can be used as epitopes in the experimental development, selection, and production of custom immunoassay reagents, including polyclonal antibodies, monoclonal antibodies, recombinant antibodies, camelid nanobodies, and single chain variable fragments. The markers, or portions of the markers can also be used as protein / polypeptide targets for the in silico design of immunoassay reagents such as antibody fragments and single chain variable fragments, with high affinity against regions of the novel markers. The resultant immunoassay reagents by additional in silico modifications that inserts specific amino acid sequences to permit site-directed conjugation of detectors or other site-directed modifications that improve the sensitivity, specificity, and / or dynamics range of the immunoassay reagent.Attorney Ref.: 70851.3W001

[0049] Additionally, because of the nature of the microbial targets that we are able to identify using our methods, such methods can be employed to develop taxon-specific anti-infectives and can be used to develop a reference library for sequencing-based diagnostics.EXAMPLES

[0050] The following methods were employed in the Examples described herein.

[0051] Example 1: D. pigrum Experimental Protocols

[0052] D. pigrum core genome analysis

[0053] A local D. pigrum genome database was curated by downloading publicly available genomes from NCB1 RefSeq and adding in-house sequenced and assembled D. pigrum genomes (see, for e.g.: Table 1). DNA from the inhouse D. pigrum isolates was extracted using a DNeasy Blood and Tissue kit (Qiagen) or MagNA Pure LC DNA Isolation Kit (Roche) and libraries were generated with a Nextera XT DNA Library kit (Illumina) according to manufacturer’s instructions for paired-end sequencing on an Illumina NextSeq 500 (Illumina, Inc., San Diego, CA) with a read length of 150 bp. We assembled Illumina short read sequences from inhouse D. pigrum isolates into contigs using the SPADES assembler (v.3.5) (see, for e.g.: Nurk et al., (2013) J Comput Biol.). Quality of the assembly was assessed using metrics generated by QUAST (v.2.3) (27) and all genomes were annotated with Prokka (v. 1.13) (see, for e.g. Seemann (2014) Bioinformatics). To maximize assay sensitivity for I), pigrum detection we focused on the core genome. The GFF files from the Prokka annotation step were used as input for the pan-genome analysis with Roary (v.3.12.0) (see, for e.g.: Page el al. (2015) Bioinformalics) [blastp identity=90%, gene presence in isolates to be core=99%]. We generated a maximum likelihood tree from core genome SNPs to assess relatedness of the D. pigrum. isolates using previously described methods (see, for e.g.: Price et al. (2012) mBio.,' Reid et al. (2019) Microb. Genom.). Briefly, Illumina short reads from inhouse D. pigrum isolates were mapped to the chromosome of the published D. pigrum reference genome (strain 83VPs-KB5; GenBank accession no. CP041626.1) using the NASP pipeline that uses BWA-MEM (v.0.7.12) (see, for e.g. : Durbin et al. (2009) Bioinformatics) to align and GATK (v.3.5) (see, for e.g.: McKenna et al. (2010) Genome Res.) to call SNPs. Publicly available genomes downloaded from NCBI RefSeq were aligned to the reference using MUMMER and SNPs were identified. The resultant SNP matrix was processed with Gubbins (see, for e.g. : Croucher et al. (2015) Nucleic Acids Res.) to remove recombinant regions. A Phylogenetic tree was constructed from the core SNPs in PhyML with Smart Model selection (v.3.0) (see, for e.g.:Attorney Ref.: 70851.3W001Guindon et al., (2010). Syst. Biol. .. The maximum likelihood phylogeny was visualized alongside the pangenome using PHANDANGO (see, for e.g. Hadfield et al. (2018) Bioinformatics) (see, for e.g.: FIG.4). Uniprot IDs of the core genes wherever available, were extracted from the GFF files using an inhouse script and were used to retrieve Gene Ontology terms from UniProt database (see, for e.g. C. UniProt (2021) Nucleic Acids Res.), (see: Table 2). The GO terms were analyzed and summarized using GAOTools (2018) Sci. Rep.).Table 1. D. pigrum whole genome sequences analyzedStrain Source Collection Host Country Total Total Totalyear contigs bases genes NCBT Accession KPT 1914 Nasal swab 2010 Human USA 76 1726398 1692 GCA_003263915.2 KPL1931_CDC4294-98 Blood 1998 Human USA 82 2014679 1999 GCF_003264085 1 KPL1937_CDC4199-99 Blood 1999 Human USA 65 1976602 1884 GCF_003264005 1 KPL1933_CDC4545-98 Nasopharyngeal 1998 Human USA 19 1861299 1787 GCF 003264045.1 KPL1939_CDC4792-99 Nasopharyngeal 1999 Human USA 47 1893917 1822 GCF 003263965.1 KPL1934_CDC4709-98 Eye 1998 Human USA 82 1912682 1805 GCA_0032640152 KPL1922_CDC39-95 Sinus 1995 Human USA 75 1859258 1794 GCF 003264145 1 87UNt-Sm4 Nose 2013 Human Germany 65 1954981 1874 SRR19918654 63VAs-B3 Nose 2012 Human Germany 51 1964942 1862 SRR19918661 63VAs-Sml Nose 2012 Human Germany 50 1969732 1867 SRR19918660 9VPs-B5 Nose 2011 Human Germany 55 1903686 1819 SRR 19918653 83VAs-Sm8 Nose 2012 Human Germany 29 1918043 1893 SRR19918656 83VPs-KB5 Nose 2012 Human Germany 39 1917955 1891 CP041626.1 44MNt_B4 Nose 2012 Human Germany 48 1891984 1890 SRR19918652 88MNs-Sm2 Nose 2013 Human Germany 20 1910927 1933 SRR19918655 88VPs-Sni9 Nose 2013 Human Germany 22 1862240 1849 SRR19918663 90VAs-B6 Nose Human Germany 44 1858197 18132013 SRR19918662 90VAs-Sm9 Nose Human Germany 38 1863257 18312013 SRR19918651 68VAs-B3 Nose Human Germany 20 1954390 18862012 SRR19918659 68VPs-B6 Nose 2012 Human Germany 50 1958590 1896 SRR19918658 81UNt-Sm4 Nose 2012 Human Germany 1 1876539 1793 SRR19918657Table 2. Summary of Gene Ontology termsGO Category Go Terms GO ID CountAttorney Ref.: 70851.3W001BP Cellular process G0:0009987 375 BP Metabolic process G0:0008152 338 BP Localization G0:0051179 63 BP Biological regulation G0:0065007 49 BP Response to stimulus G0:0050896 40 BP Biological process involved in interspecies interaction between G0:0044419 5organismsBP Developmental process G0:0032502 4 BP Detoxification G0:0098754 3 BP Reproductive process G0:0022414 1 BP Biological adhesion G0:0022610 1 BP Signaling G0:0023052 1 BP Multi-organism process G0:0051704 1 BP Immune system process GG:0002376 1 CC Protein-containing complex G0:0032991 45 CC Cellular anatomical entity GO:0110165 27 MF Catalytic activity G0:0003824 533 MF Binding G0:0005488 99 MF Transporter activity G0:0005215 87 MF ATPase G0:0016887 34 MF Antioxidant activity G0:0016209 5 MF Transcription regulator activity G0:0140110 5 MF Translation regulator activity G0:0045182 3 MF Molecular c arrier activity G0:0140104 2 MF Structural molecule activity G0:0005198 1 MF Molecular Uansducer activity GG:0060089 1 MF Small molecule sensor activity G0:0140299 1 MF Enzyme activator activity GG:0008047 1 MF Kinase regulator activity G0:0019207 1 MF Molecular function regulator G0:0098772 1 MF Molecular adaptor activity GG:0060090 1

[0054] D. pigrum assay target identification

[0055] The core genome was filtered and only SCSG were retained. An in-silico search for homology against non- / ). pigrum species was performed using BLAST (see, for e.g. Camacho et al. (2009) BMC Bioinformatics') against a local copy of the NT database (updated:Attorney Ref.: 70851.3W0012019-03-31). Gene targets with 70% similarity to non-D. pigruin species were removed. A final set of homologous single -copy core genes was used as the candidate pool for targets to design D. pigrum specific assay.

[0056] D. pigrum assay design

[0057] Primer3 (see, for e.g.: Untergasser et al. (2012) Nucleic A ci ds Res.) was used with default settings to identify candidate forward and reverse primers which were first compared to the D. pigrum gene alignment file then checked for similarity against other nasal bacteria, including Staphylococcus aureus, Staphylococcus epidermidis, Corynebacterium spp., Cutibacterium spp., Moraxella spp., Escherichia coli, Klebsiella spp., Citrobacter spp., Proteus spp., and Alloiococcus spp. Primers were excluded if 5 or more matching bases were found at the 3 ’-end of the primer.

[0058] D. pigrum assay validation

[0059] To assess the sensitivity of the primers, we tested the mur] assay against 12 D. pigrum isolates. These isolates had been previously verified to be D. pigrum by MALDI-TOF and their genomes were sequenced using Illumina HiSeq system (Illumina, San Diego, CA). Furthermore, we screened mur] primers against 110 clinical samples characterized by 16S rRNA gene-based sequencing as described previously (see, for e.g. : Liu et al. (2013) mBio). A non-D. pigrum control collection that included Moraxella catarrhalis, Staphylococcus aureus, Staphylococcus epidermidis, Corynebacterium pseudodiphtheriticum, Corynebacterium propinquum, Corynebacterium Corynebacterium accolens species was used to evaluate specificity of our primers.

[0060] Human subject research

[0061] This study and the protocols used were approved by the George Washington University Institutional Review Board and The Office of Human Research (IRB#: NCR191444).

[0062] Human nasal swab collection

[0063] Healthy community-dwelling adults were enrolled in a nasal microbiome study in Washington, DC (IRB#: NCR 191444). At enrollment, nasal specimens were self-collected using Puritan HydraFlock swabs (Puritan Medical Products, Guilford, ME) with staff instructions. Samples were placed immediately into transport media and stored at 4°C until processing. Samples were processed within 4 hours then transferred in 100 pl aliquots into labeled 2 mL cryovials and stored at -80°C.Attorney Ref.: 70851.3W001

[0064] DNA Isolation and purification

[0065] DNA from human nasal swabs were extracted using MagMax DNA Ultra 2.0 Kit with enzyme and chemical lysis. DNA from bacterial isolates were extracted through heat soak (I). pigrum, S. aureus, and S. epidermidis) or using the DNeasy Blood & Tissue Kit (Qiagen, Valencia, CA) (C. propinquum and C. pseudodiphtheriticum) according to manufacturer instructions.

[0066] mur] PCR amplificationEach mur] PCR was performed in a 20 pl reaction volume containing 1 pl of template DNA added to 19 pl of PCR reaction mix containing 0.4 pM of forward (5’-CAACAGCGTCCAGCAATCTA-‘3) (SEQ ID NO: 1) and reverse (5’-ATCGCTGTAATCCCGATGAG-‘3) (SEQ ID NO: 2) primers, IX Phusion High-Fidelity PCR Master Mix (ThermoFisher), and molecular-grade water. Amplification was performed on a Cl 000 Touch Thermocycler (Bio-Rad, Hercules, CA) using the following conditions: 98°C for 30s for denaturing, 54°C for 30s for annealing, and 72°C for 1 min for extension x 35 cycles. Amplified DNA was run on a 2% agarose E-gel (ThermoFisher) to assess amplification of D. pigrum DNA. Gels were imaged using a ChemiDoc-It2 (Analytik Jena US, Upland, CA). Presence of a visible band near the 224 bp size indicated successful amplification.

[0067] D. pigrum phylogenetic and core genome analysis

[0068] We first analyzed the genetic diversity of 21 D. pigrum whole genome sequences available (n = 7 from NCBI and n = 14 from in-house D. pigrum genomes, Table 1). We extracted 87,993 SNPs from non-recombined regions of the core genome and examined the genetic diversity based on maximum likelihood phylogeny (FIG. 3). This showed multiple distinct D. pigrum lineages, which indicates both the non-clonal nature of D. pigrum and the robustness of the genome collection. We then generated and analyzed the D. pigrum pangenome to identify 1,291 core and 357 accessory genes. For potential assay targets, we focused on the 1,291 core genes.

[0069] D. pigrum assay design

[0070] In an aspect of the disclosure, FIG. 1 is a schematic of a core genome-based approach for assay design in accordance with embodiments of the disclosure. FIG. 1 represents the approach taken to mine the pan-genome for assay targets. In each succeeding step, the pangenome analysis illustrates how genes were filtered to finally retain a unique core genome for the organism of interest. Assay target genes discovery was a multi-step processAttorney Ref.: 70851.3W001(see: FIG. 1). We removed ribosomal genes (n = 71) and genes with homologs in other genera (n = 345). Manual filtering of randomly-chosen assay target genes from the 843 single-copy core species genes was performed, requiring that the assay target gene: a) must be present in all 21 D. pigrum genomes, b) must have less than 70% similarity identity and coverage against sequences from non-Dolosigranulum taxa by BLAST, and c) contain forward and reverse primer sequences meeting Primer3 design criteria and that have less than 50% similarity identity and cover against sequences from non-Dolosigranulum taxa by BLAST.

[0071] In an aspect of the disclosure, FIG. 2 illustrates an aspect of D. pigrum mur] phylogeny and sequence alignment, as a Neighbor joining tree constructed using full length mur] gene sequences from 21 D. pigrum isolates using Jalview 2.11 and ordered by branch lengths, highlighting that mur] is part of the conserved core genome but is also phylogenetically informative, in accordance with embodiments of the disclosure. FIG. 2B illustrates an aspect of D. pigrum mur] phylogeny and sequence alignment, as multiple sequence alignment of mur] amplicon region, where the forward primer is located at 1234-1255 bp and the reverse primer is located at 1436-1457 bp, in accordance with embodiments of the disclosure.

[0072] The first single -copy core species genes (SCSG) that met our selection criteria as a target gene candidate with conserved regions for primer design was murJ (UniProt ID: 034674), a gene with a length of 1665 bp encoding a Lipid II flippase protein. The average uncorrected distance between the isolates for the mur} alignment was 35.84 bp (SD=13.67 bp) (FIG. 2A). After iterations of primer design and in silico analysis, we identified a pair of forward and reverse PCR primers (Table 3) targeting the mur J gene that produces a 224 bp PCR product. On average the amplicon varied by 2.14 bp (SD=1.69 bp) between the isolates (FIG. 2B).Table 3. D. pigrum mur] forward and reverse primer sequencesAssay Primer Size(bp) Tin (annealing GC Sequence (5 '-3')Target temp)murJ murJJP 21 54 °C 50% CAACAGCGTCCAGCAATCTA (SEQ ID NO: 3)mz / rJ_R 21 54 °C 47.5% ATCGCTGTAATCCCGATRAG (SEQ ID NO: 4)

[0073] D. pigrum PCR sensitivity and specificity against clinical isolates and human nasal swabsAttorney Ref.: 70851.3W001

[0074] The murJ assay was highly sensitive and specific in laboratory analysis of DNA from bacterial isolates and from human nasal swabs. We first evaluated the assay using well-characterized I), pigrum isolates (N = 12) and against five common nasal bacterial species namely Moraxella catarrhalis, S. aureus, S. epidermidis, Corynebacterium pseudodiphtheriticum, Corynebacterium propinquum, Corynebacterium accolens, which showed 100% sensitivity and specificity.

[0075] We further evaluated the assay using DNA extracted from human nasal swabs (n = 110) characterized using 16S rRNA V3-V4 gene-based sequencing, including 54 samples that were positive for I), pigrum and 56 samples that were negative for D. pigrum. This showed that the murJ assay was not able to detect D. pigrum in samples (n = 9) with fewer than ten D. pigrum 16S rRNA gene copies per uL of swab eluent, or 1.0 x 104D. pigrum 16S rRNA gene copies per swab. However, among the 45 D. pzgram-positive samples with more than 1.0 x 104D. pigrum 16S rRNA gene copies per swab, the murJ PCR assay was able to detect D. pigrum in 41 (91%) samples (Table 4). There were no false positives in the 56 D. pigrum -negative samples.Table 4. Detection of D. pigrum in nasal samples by PCR in relation to D. pigrum absolute abundanceD. pigrum absolute abundance (16S rRNA gene copies / swab) positive negative<lxl040 58IxlO4- <1 x 1050 7lx 105-<1 x 10617 31 x 106or greater 24 1

[0076] An aspect of the disclosure is directed to, by identifying potential assay targets using the D. pigrum core genome, designing a PCR assay that is both sensitive and specific for D. pigrum. In contrast to other commonly used methods for species confirmation, such as biochemical testing, DNA sequencing, or MALDI-TOF, PCR-based assays are rapid and cost-effective and do not require expensive equipment. This method provides a simpler option for D. pigrum detection and avoids the restriction digestion and analysis challenges of T-RFLP (see, for e.g. Prakash et al., (2014) Indian J. Microbiol.) that has been used previously forAttorney Ref.: 70851.3W001detecting microbial communities in anterior nares (see, for e.g. Camarinha-Silva et al. (2012) FEMS Microbiol. EcoE). We demonstrated the utility of the core genome mining techniques to develop species confirmation assays. The resultant mur] assay was able to identify I), pigrum and diverse bacterial isolates with a 100% sensitivity and specificity. Our assay was also highly sensitive and specific for detecting D. pigrum in clinical samples.

[0077] The above-described utilization of the mur} assay can be expanded to any core genome-based candidates identified for D. pigrum, as outlined in Table 5 below.Table 5. Core Genome-Based Candidate Genes for D. pigrumGene # Gene Annotation1 group 1000 hypothetical proteinn thiE Thiannne-phosphate synthase3 ppaC Manganese-dependent inorganic pyrophosphatase4 plsY Glycerol-3-phosphate acyltransferase5 group 1007 hypothetical protein6 spsB Signal peptidase IB7 group 1010 hypothetical protein8 murC UDP-N-acetylmuramate— L-alanine ligase9 pepA Glutamyl aminopeptidase10 group 1017 hypothetical protein11 hit Protein liit12 group 102 hypothetical protein13 group 1020 hypothetical protein14 xseB Exodeoxyribonuclease 7 small subunit15 group 1023 hypothetical protein16 murJ Lipid II flippase MurJ17 group 1029 hypothetical protein18 group 103 hypothetical protein19 group 1030 hypothetical protein20 ponA Penicillin-binding protein 1A / 1B21 perR Peroxide operon regulator22 mrpC Na( ) / H( ) antiporter subunit C23 group 1034 hypothetical protein24 panT Pantothenate transporter PanT25 group 1036 hypothetical protein26 group 1038 hypothetical protein27 zapA Cell division protein ZapA28 group 104 hypothetical protein29 group 1042 hypothetical proteinAttorney Ref.: 70851.3W00130 ribU Riboflavin transporter RibU31 scpA Segregation and condensation protein A32 fur Ferric uptake regulation protein33 pfkA ATP-dependent 6-phosphofnictokina se34 ifcA Fumarate reductase flavoprotein subunit35 rnfC 1 Electron transport complex subunit RnfC36 group 1050 hypothetical protein37 group 1052 hypothetical protein38 nadD putative nicotinate -nucleotide adenylyltransferase39 group 1054 hypothetical protein40 pat A Putative N-acetyl-LL-diaminopimelate aminotransferase41 ezrA Septation ring formation regulator EzrA42 group 1057 hypothetical protein43 group 1060 hypothetical protein44 yutF Acid sugar phosphatase45 group 1063 Acetyltransferase46 mecA Adapter protein MecA47 group 1066 hypothetical protein48 group 1067 hypothetical protein49 mreC Cell shape-determining protein MreC50 group 107 hypothetical protein51 gcvH 2 Glycine cleavage system H protein52 cobB 2 NAD-dependent protein deacetylase53 nagA N-acetylglucosamine-6-phosphate deacetylase54 fni Isopentenyl-diphosphate delta-isomerase55 group 1075 hypothetical proteinSpermidine / putrescine transport system permease56 potB protein PotB57 rbn Ribonuclease BN58 group 1081 hypothetical protein59 ybbA putative protein YbbA60 group 1083 hypothetical proteinN-formyl-4-amino-5-aminomethyl-2-methylpyrimidine61 ylmB deformylase62 yqjG Glutathionyl-hydroquinone reductase YqjG63 tenA 1 Aminopyrimidine aminohydrolase64 group 1088 hypothetical protein65 group 1089 hypothetical protein66 yfbR 5'-deoxynucleotidase YfbR67 ftsE Cell division ATP-binding protein FtsE68 prdB D-proline reductase subunit gammaputative ABC transporter phosphite binding protein69 phnDl PhnDlAttorney Ref.: 70851.3W001Glyciiie / sarcosine / betaine reductase complex component70 grdC C subunit beta71 trxA 2 Thioredoxin72 trxB 2 Thioredoxin reductaseGlycine / sarcosine / betaine reductase complex component73 grdAl 1 Al74 copZ Copper chaperone CopZPutative N-acetylmannosamine-6-phosphate 2- 75 nanE epimerase76 group 1107 hypothetical protein77 nanM N-acetylneuraminate epimerase78 cycB Cyclodextrin-binding protein79 group 1112 Alpha-monoglucosyldiacylglycerol synthase80 anr Transcriptional activator protein AnrHigh-affinity zinc uptake system ATP-binding protein81 znuC 2 ZnuC82 znuA High-affinity zinc uptake system binding-protein ZnuA83 group 1117 Putative Nrdl-like protein84 group 1118 hypothetical protein85 dus putative tRNA-dihydrouridine synthase86 group 112 hypothetical protein87 hslO 33 kDa chaperonin88 ideR Iron-dependent repressor IdeR89 noxE NADH oxidase90 IrgA 2 Antihohn-like protein LrgA91 group 1129 hypothetical protein92 group 1130 hypothetical proteinTwo-component system WalR / WalK regulatory protein93 yycl Yycl94 yycJ Putative metallo-hydrolase YycJ95 folD Bifunctional protein FolD protein96 group 1138 hypothetical protein97 mscL Large-conductance mechanosensitive channel98 sbcD Nuclease SbcCD subunit D99 group 1143 hypothetical protein100 group 1145 hypothetical protein101 pep Pyrrolidone-carboxylate peptidase102 group 1147 Methionine import system permease protein MetP103 btuD 9 Vitamin B12 import ATP-binding protein BtuD104 group 1150 hypothetical protein105 group 1155 hypothetical protein106 fabll 3-oxoacyl-[acyl-carrier-proteinl synthase 3107 fabF 3-oxoacyl-Iacyl-carrier-protein] synthase 2Acetyl-coenzyme A carboxylase carboxyl transferase108 accD subunit betaAttorney Ref.: 70851.3W001Acetyl-coenzyme A carboxylase carboxyl tr ansferase109 accA subunit alpha110 group 1164 Glycine cleavage system H-like protein111 metQ 3 D-methionine-binding lipoprotein MetQ112 tenA 2 Aminopyrimidine aminohydrolase113 ssuB Aliphatic sulfonates import ATP- binding protein SsuB114 ribX Riboflavin transport system permease protein RibX115 nosF putative ABC transporter ATP-binding protein NosF116 ywpJ Phosphatase YwpJ117 yqliD Alcohol dehydrogenase Y qhD118 atpD 1 ATP synthase subunit delta119 ybaB Nucleoid-associated protein YbaBGlutathione-regulated potassium-effhix system ancillary120 kefG protein KefG121 group 1184 hypothetical protein122 ytrA 1 HTH-type transcriptional repressor YtrA123 ytrB ABC transporter ATP-binding protein YtrB124 group 1187 hypothetical protein125 veg Protein VegHigh-affinity zinc uptake system ATP-binding protein126 znuC 1 ZnuCUndecaprenyl-phosphate 4-deoxy-4-formamido-L- 127 arnC 2 arabinose transferase128 group 1230 hypothetical protein129 group 1249 Putative peptidyl-prolyl cis-trans isomerase130 group 125 hypothetical protein131 gsliAB Glutathione biosynthesis bifunctional protein GshAB132 mniC Mini-ribonuclease 3133 group 129 hypothetical protein134 msrC Free methionine-R-sulfoxide reductase135 group 134 hypothetical protein136 group 136 hypothetical protein137 gtfl Glycosyltransferase Gtfl138 copY 2 Transcriptional repressor CopY139 cadA 2 Cadmium-transporting ATPase140 tils tRNA(Ile)-lysidine synthase141 mapP Maltose 6'-phosphate phosphatase142 fumC Fumarate hydratase class II143 maa Maltose O- acetyltransferase144 group 146 Proteiii-ADP-ribose hydrolase145 fadA 3-ketoacyl-CoA thiolase146 yidD Putative membrane protein insertion efficiency factorAttorney Ref.: 70851.3W001147 group 15 hypothetical protein148 group 150 hypothetical protein149 rihC Non-specific ribonucleoside hydrolase RihC150 pepQ Xaa-Pro dipeptidase151 birA Bifunctional ligase / repressor BirAUndecaprenyl-phosphate 4-deoxy-4-formainido-L- 152 group 160 arabinose transferase153 group 162 hypothetical protein154 group 167 hypothetical protein155 umuC Protein UmuC156 rluD 2 hypothetical protein157 group 172 hypothetical protein158 group 176 hypothetical proteinN-acetylglucosaniinyldiphosphoundecaprenol N-acetyl- 159 tarA beta-D-mannosaminyltransferasetRNA threonylcarbamoyladenosine biosynthesis protein160 tsaB TsaB161 group 181 hypothetical protein162 group 185 hypothetical protein163 rpoNl 2 RNA polymerase sigma-54 factor 1164 group 188 Protein SprT-like protein165 group 192 hypothetical protein166 group 193 hypothetical protein167 fabD Malonyl CoA-acyl carrier protein transacylase168 group 195 hypothetical protein169 group 197 hypothetical protein5-amino-6-(5-phospho-D-ribitylamino)uracil170 ybji phosphatase Ybji171 acpS Holo-1 acyl-carrier-protein] synthase172 prmC Release factor glutamine methyltransferase173 group 211 hypothetical protein174 dagK 2 Di acylglycerol kinase175 folT Folate transporter FolT176 group 214 hypothetical protein177 group 220 hypothetical protein178 yccX Acylphosphatase179 group 223 GTP cyclohydrolase 1 type 2180 group 225 hypothetical protein181 group 227 hypothetical protein182 est 2 Carboxylesterase183 yjbK Putative triphosphatase YjbK184 lacC 1 Tagatose-6-phosphate kinase185 rpiA Ribose-5-phosphate isomerase A186 pyrC Dihydroorotase187 pvrE Orotate phosphonbosvltransferaseAttorney Ref.: 70851.3W001188 group 238 hypothetical protein189 sasA 1 Adaptive-response sensory-kinase SasA190 group 245 hypothetical protein191 fchA Methenyltetrahydrofol ate cyclohydrolasePhosphoribosylformylglycinamidine synthase subunit192 purL PurL193 group 250 hypothetical protein194 truA tRNA pseudouridine synthase A195 group 253 hypothetical protein196 oppF 1 Oligopeptide transport ATP-binding protein OppF197 fadN putative 3-hydroxyacyl-CoA dehydrogenase198 ytnP putative quorum-quenching lactonase YtnP199 group 257 hypothetical protein200 apbE 1 FAD protein FMN transferase201 yfkN Tnfunctional nucleotide phosphoesterase protein YfkN202 group 2672 hypothetical protein203 group 2675 hypothetical protein204 group 2676 hypothetical protein205 acpP 2 Acyl carrier protein206 group 2678 hypothetical protein207 group 2680 putative protein208 ntpK V-type sodium ATPase subunit K209 yugl General stress protein 13210 ndoA Endoribonuclease EndoA211 group 2695 hypothetical protein212 group 2697 hypothetical protein213 cinA Putative competence-damage inducible protein214 group 2704 hypothetical protein215 group 2708 hypothetical protein216 atpE 2 ATP synthase subunit c217 est 1 Carboxylesterase218 group 2710 hypothetical protein219 ntpD V-type sodium ATPase subunit D220 group 2715 hypothetical protein221 mutL DNA mismatch repair protein MutL222 group 2720 hypothetical protein223 ypiQ putative protein YpjQ224 secG putative protein-export membrane protein SecG225 ntpG V-type sodium ATPase subunit G226 rcsC Sensor histidme kinase RcsC227 group 2733 hypothetical protein228 group_2736 Putative gluconeogenesis factorAttorney Ref.: 70851.3W001229 group 2737 hypothetical protein230 lemA Protein LemA231 kti-B Ktr system potassium uptake protein B232 group 274 hypothetical proteinGlycine / sarcosine / betaine reductase complex component233 grdAl 2 Al234 group 2742 hypothetical protein235 dppE 2 hypothetical protein236 group 2744 hypothetical protein237 scpB Segregation and condensation protein B238 group 2748 hypothetical protein239 group 2751 hypothetical proteinAspartyl / glutamyl-tRNA(Asn / Gln) amidotransferase240 gatC 2 subunit C241 ftsA Cell division protein FtsA242 pem Protein-L -isoaspartate O-methyltransferase243 metN2 Methionine import ATP-binding protein MetN 2244 rnj2 Ribonuclease J 2245 yabA Initiation-control protein YabA246 xerC Tyrosine recombinase XerC247 macB Macrolide export ATP-binding / permease protein MacB248 mntB 2 Manganese transport system membrane protein MntBNAD(P)H-quinone oxidoreductase subunit 2,249 ndhB chloroplastic250 deoD Purine nucleoside phosphorylase DeoD-type251 group 2771 putative ABC transporter ATP-binding protein252 atpC ATP synthase epsilon chain253 ribF Riboflavin biosynthesis protein RibF254 group 2781 RNA-binding protein255 group 2782 hypothetical protein256 group 2783 hypothetical protein257 lipA Lipoyl synthase258 sugC Trehalose import ATP-binding protein SugC259 rnc Ribonuclease 3260 group 2793 L-methionine gamma-lyaseEnergy-coupling factor transporter ATP-binding protein261 ecfA2 EcfA2262 nanA 1 N-acctylncuraminatc lyase263 htrA Serine protease Do-like HtrA264 Idh L-lactate dehydrogenase265 in ail Y PTS system mannose-specific EIIC component266 sglT Sodium / glucosc cotransportcr267 lac A Galactose-6-phosphate isomerase subunit LacAAttorney Ref.: 70851.3W001268 ssbA 1 Single-stranded DNA-binding protein A269 group 2809 hypothetical protein270 folA Dihydrofolatc reductase271 yodB HTH-type transcriptional regulator YodB272 penA Penicillin-binding protein 2B273 metQ 1 Methionine-binding lipoprotein MetQ274 rsxD 2 Election transport complex subunit RsxD275 manX 1 PTS system mannose-specific EIIAB component276 rodA Peptidoglycan glycosyltransferase RodA277 gsiC 1 Glutathione transport system permease protein GsiC278 mgtE 1 Magnesium transporter MgtE279 trkA Trk system potassium uptake protein TrkA280 yrrK Putative pre-16S rRNA nuclease281 group 2827 hypothetical protein282 group 2828 hypothetical protein283 dual Primosomal protein Dnal284 gpsB Cell cycle protein GpsB285 mgtE 2 Magnesium transporter MgtE286 gmk 2 Guanylate kinase287 group 2840 Deoxyguanosine kinase288 metQ 2 Methionine -binding lipoprotein MetQ289 yaaQ putative protein YaaQ290 mutS2 Endonuclease MutS2291 group 2849 hypothetical protein292 dnaN Beta sliding clamp293 rasP Regulator of sigma- W protease RasP294 iscU Iron-sulfur cluster assembly scaffold protein IscU295 group 2855 hypothetical protein296 group 2856 hypothetical protein297 nrdR Transcriptional repressor NrdR298 xpt Xanthine phosphoribosyltransferase299 fabG 2 3-oxoacyl-[acyl-carrier-protein] reductase FabG300 whiA Putative sporulation transcription regulator WhiA301 group 2863 hypothetical protein302 atpG ATP synthase gamma chain303 group 2865 hypothetical protein304 manX 2 PTS system mannose-specific EIIAB component305 misCA Membrane protein insertase MisCA306 group 2869 hypothetical protein307 Ipl.T 1 Lipoate-protein ligase LplJAttorney Ref.: 70851.3W001308 hprK HPr kinase / phosphorylase309 gmk 1 Guanylate kinase310 group 2876 hypothetical protein311 oppF 2 Oligopeptide transport ATP-binding protein OppF312 group 2883 hypothetical protein313 ypdA Sensor histidine kinase YpdA314 group 289 hypothetical protein315 trkG Trk system potassium uptake protein TrkG316 pnp 2 Polyribonucleotide nucleotidyltransferase317 ribT Protein RibT318 group 2896 putative protein319 group 2897 hypothetical protein320 fabZ 3-hydroxyacyl-Iacyl-carner-protein] dehydratase FabZ321 group 2900 hypothetical protein322 codY GTP-sensing transcriptional pleiotropic repressor CodY323 pepV Beta-Ala-Xaa dipeptidasesn-glycerol-3-phosphate-binding periplasmic protein324 ugpB UgpB325 group 291 hypothetical protein326 yibN putative protein YibN327 group 2911 putative ABC transporter ATP-binding protein328 gsiC 2 Glutathione transport system permease protein GsiC329 adk Adenylate kinase330 mraZ Transcriptional regulator MraZ331 guaB 1 Inosine-5'-monophosphate dehydrogenase332 rex Redox-sensing transcriptional repressor Rex333 rnfC 2 Electron transport complex subunit RnfC334 oppB Oligopeptide transport system permease protein OppBtRNA threonylcarbamoyladenosine biosynthesis protein335 tsaE TsaEBifunctional aspartate aminotransferase and L-aspartate336 asD beta-decarboxylase337 dnaJ Chaperone protein DnaJ338 yiiP Inner membrane protein YjjP339 pyrP Uracil permease340 secY Protein translocase subunit SecY341 group 2928 hypothetical protein342 alrl Alanine racemase 1343 group 2930 hypothetical protein344 apt Adenine phosphoribosyltransferase345 srrA 1 Transcriptional regulatory protein SrrA346 group 2934 hypothetical proteinAttorney Ref.: 70851.3W001347 group 2935 hypothetical protein348 pepFl 1 Oligoendopeptidase F, plasmidBranched-chain amino acid transport system 2 carrier349 brnQ 1 protein350 malP PTS system maltose-specific EIICB component351 sulA Dihydropteroate synthase352 htpX Protease HtpX353 ctpC putative manganese / zinc-exporting P-type ATPaseUDP-N-acetylglucosamine— N-acety Imur amyl- (pentapeptide) pyrophosphoryl-undecaprenol N- 354 murG acetylglucosamine transferase355 group 2949 hypothetical protein356 group 295 hypothetical protein357 metP 2 Methionine import system permease protein MetP358 abgT p-aminobenzoyl-glutamate transport protein359 group 2952 hypothetical protein360 pbuO Guanine / hypoxanthine permease PbuO361 rpoZ DNA-directed RNA polymerase subunit omega362 yloA putative protein YloA363 hpt Hypoxanthine-guanme phosphoribosyltransferase364 group 2961 putative ABC transporter ATP-binding protein365 glgD Glycogen biosynthesis protein GlgD366 rpiR HTH-type transcriptional regulator RpiR367 hrcA Heat-inducible transcription repressor HrcA368 zwf Glucose-6-phosphate 1 -dehydrogenase369 mraY Phospho-N-acctylmuramoyl-pcntapcptidc -transferase370 plsC l-acyl-sn-glycerol-3-phosphate acyltransferase371 group 2969 hypothetical protein372 pheT 2 Phenylalanine— tRNA ligase beta subunit373 group 2970 hypothetical protein374 group 2973 hypothetical protein375 sepF Cell division protein SepF376 group 2975 Nitronate monooxygenase377 odd Cytidine deaminaseputative glycme dehydrogenase (decarboxylating)378 gcvPA subunit 1Ferric-anguibactin transport system permease protein379 fatD FatD380 npr NADH peroxidase381 group 298 hypothetical protein382 sodA Superoxide dismutase [Mill383 pflA Pyruvate formate-lyase-activating enzyme384 group 2982 hypothetical proteinAttorney Ref.: 70851.3W001385 group 2983 hypothetical protein386 natA ABC transporter ATP-binding protein NatA387 dnaE DNA polymerase III subunit alpha388 uppP Undecaprenyl-diphosphatase389 uvrC UvrABC system protein C390 mntB 1 Manganese transport system membrane protein MntB391 group 2993 hypothetical protein392 group 2994 hypothetical protein393 oat A O- acetyltransferase OatA394 bglK 1 Beta-glucoside kinase395 group 2997 hypothetical protein396 group 2998 hypothetical protein397 sasA 2 Adaptive-response sensory-kinase SasA398 group 30 hypothetical protein399 group 3001 hypothetical protein400 yoliK 2 Inner membrane protein YohK401 group 3003 hypothetical protein402 ybeY Endoribonuclease YbeY403 yumC Ferredoxin— NADP reductase 2404 group 3007 hypothetical protein405 clpQ ATP-dependent protease subunit ClpQ406 group 301 hypothetical protein407 group 3012 hypothetical protein408 group 3013 hypothetical protein409 spoOJ Stage 0 sporulation protein J410 smc 2 Chromosome partition protein Smc411 group 3016 hypothetical protein412 group 3017 hypothetical protein413 group 3018 hypothetical protein414 mngB Mannosylglyccratc hydrolase415 araQ 2 L-arabinose transport system permease protein AraQ416 zur Zinc-specific metallo-regulatory protein417 minC Septum site-determining protein MinC418 group 3026 hypothetical protein419 group 3028 hypothetical protein420 group 3030 hypothetical protein421 ktrA Ktr system potassium uptake protein A422 group 3032 hypothetical protein423 rpoE DNA-directed RNA polymerase subunit delta424 group 3038 putative protein425 group 3039 hypothetical protein426 cutC Copper homeostasis protein CutCAttorney Ref.: 70851.3W001427 gmuE Putative fructokinase428 upp Uracil phosphoribosyltransferase429 ffh Signal recognition particle protein430 group 3044 hypothetical protein431 group 3046 hypothetical protein432 murl Glutamate racemase433 group 3048 hypothetical protein434 group 3049 hypothetical protein435 group 305 putative oxidoreductase / MSMEI 2347436 det2 Peptide deformylase 2437 group 3052 hypothetical protein438 sstT Serine / threonine transporter SstT439 mltG Endolytic murein transglycosylase440 metN 2 Methionine import ATP-binding protein MetN441 group 3058 hypothetical protein442 pgcA Phosphoglucomutase443 group 306 Acetyltransferase444 group 3062 hypothetical protein445 acpP 1 Acyl carrier proteinputative D,D-dipeptide transport system permease446 ddpC protein DdpC447 dnaD DNA replication protein DnaD448 folE GTP cyclohydrolase 1449 ftsW putative peptidoglycan glycosyltransferase FtsW450 ald2 Alanine dehydrogenase 2451 glcR HTH-type transcriptional repressor GlcR452 lipL Lipoyl-[GcvH] protein N-lipoyltransferase453 ntpC V-type sodium ATPase subunit C454 clsA Major cardiolipin synthase ClsA455 dps DNA protection during starvation protein456 group 3076 hypothetical protein457 spxA 2 Regulatory protein Spx458 group 3081 hypothetical proteinGlycine reductase complex component B subunit459 grdB 2 gamma460 group 3083 hypothetical protein461 group 3087 DegV domain-containing protein462 nadE NH(3)-dependent NAD( ) synthetase463 divIB Cell division protein DivIB464 group 3091 hypothetical protein465 group 3092 hypothetical protein466 pheT 1 Phenylalanine— tRNA ligase beta subunit467 lytR_2 Sensory transduction protein LytRAttorney Ref.: 70851.3W001468 group 3095 hypothetical protein469 cggR Central glycolytic genes regulatorDihydroorotate dehydrogenase B (NAD( )), catalytic470 pyrD subunit471 nusG Transcription termination / antitermination protein NusG472 atpB ATP synthase subunit a473 group 3102 Putative aminotransferase / MSMEI 6121474 group 3104 hypothetical protein475 group 3106 hypothetical protein476 prkC Serine / threonine -protein kinase PrkC477 secE Protein translocase subunit SecE478 bioY Biotin transporter BioY479 group 3110 hypothetical protein480 group 3111 hypothetical protein481 glcK Glucokinase482 group 3116 hypothetical protein483 group 3117 Putative glycerol transporter484 group 3118 hypothetical protein485 tadA tRNA-specific adenosine deaminase486 group 3120 Putative transport protein487 cdaR CdaA regulatory protein CdaR488 tdcF Putative reactive intermediate deaminase TdcF489 thiM Hydroxyethylthiazole kinase490 group 3125 hypothetical protein491 group 3126 hypothetical protein492 purK 1 N5-carboxyaminoimidazole ribonucleotide synthase493 group 3128 dITP / XTP pyrophosphatase494 gcvH 1 Glycine cleavage system H protein495 yqeN putative protein YqeN496 yidC Membrane protein insertase YidC497 galK 2 Galactokinase498 group 3133 hypothetical protein499 yoliK 1 Inner membrane protein YohK500 group 3135 hypothetical protein501 group 3137 hypothetical protein502 graR Response regulator protein GraR503 thlA Acetyl-CoA acetyltransferase504 ccpA 2 Catabolite control protein AUbiquinone / menaquinone biosynthesis C- 505 ubiE methyltransferase UbiE506 group 3144 hypothetical protein5-methyltetrahydropteroyltrighitani ate— homocysteine507 metE methyltransferaseAttorney Ref.: 70851.3W001508 atpE 1 V-type proton ATPase subunit Eputative pyridine nucleotide-disulfide oxidoreductase509 rclA RclA510 group 3149 hypothetical protein511 group 3150 hypothetical protein512 group 3154 hypothetical protein513 dacA 1 Diadenylate cyclase514 lacR Lactose phosphotransferase system repressor515 recU Holliday junction resolvase RecU516 oppD 2 Oligopeptide transport ATP-binding protein OppD517 group 3159 hypothetical proteinputative bifunctional oligoribonuclease and PAP518 nrnA phosphatase NmA519 plsX Phosphate acyltransferase520 group 3164 hypothetical protein521 yqeH putative protein YqeH522 group 3166 hypothetical protein523 group 3167 hypothetical protein524 divIVA Cell division protein DivIVA525 murAB UDP-N-acetylglucosamine 1-carboxyvinyltiansferase 2526 group 317 hypothetical protein527 group 3170 Putative glycerol transporter528 group 3171 hypothetical protein529 group 3172 hypothetical protein530 dsdA 1 D- serine dehy dratase531 trmL tRNA (cytidine(34)-2'-O)-methyltransferase532 group 3177 hypothetical protein533 group 3179 hypothetical proteinputative undecaprenyl-phosphate N-acetylglucosaminyl534 tagO 1 -phosphate transferase5-amino-6-(5-phospho-D-ribitylamino)uracil535 yitU phosphatase YitU536 yxdL ABC transporter ATP-binding protein YxdL537 group 3183 hypothetical protein538 thiY Formylaminopyrimidine-binding protein539 rho Transcription termination factor Rho540 recO DNA repair protein RecO541 oxyR hypothetical proteinUDP-N-acetylmuramoyl-L-alanyl-D-glutamate— L- 542 murE lysine ligase543 group 3191 putative transcriptional regulatory protein / MSMEI 2866544 group 3193 hypothetical protein545 group 3195 hypothetical proteinAttorney Ref.: 70851.3W001546 metQ 4 D-methionine-binding lipoprotein MetQ547 ftsL Cell division protein FtsLBranched-chain amino acid transport system 2 carrier548 group 3198 protein549 group 3199 hypothetical protein550 ghrA Glyoxylate / hydroxypyruvate reductase A551 axe2 Acetylxylan esterase552 trxB 1 Thioredoxin reductase553 bepA Beta-barrel assembly-enhancing protease554 yxeO putative ABC transporter ATP-binding protein555 gloA Lactoylglutathione lyase556 fpgS 1 Folylpolyglutamate synthase557 group 324 hypothetical protein558 group 325 Protein ADP-ribosyltransferase559 dppC Dipeptide transport system permease protein DppC560 group 329 hypothetical protein561 pglF UDP-N-acetyl-alpha-D-glucosamuie C6 dehydratase562 dacA 2 D-alanyLD-alanine carboxypeptidase DacA563 argS 1 Arginine— tRNA ligase564 appA hypothetical protein565 apu Amylopullulanase566 yvyl Putative mannose-6-phosphate isomerase Yvyl567 group 342 hypothetical protein568 group 357 hypothetical proteinPTS-dependent dihydroxyacetone kinase, ADP-binding569 dhaL subunit DhaL570 opuCB Carnitine transport permease protein OpuCB571 mutM Formamidopynmidine-DNA glycosylase572 group 368 hypothetical protein573 truB tRNA pseudouridine synthase B574 group 371 Putative phosphatase575 group 374 hypothetical protein576 group 375 hypothetical protein577 glys Glycine— tRNA ligase beta subunit578 group 379 hypothetical protein579 group 380 hypothetical protein580 group 384 hypothetical protein581 thiQ Thiamine import ATP-binding protein ThiQ582 group 387 hypothetical protein583 group 388 hypothetical protein584 glgA Glycogen synthase585 group 39 hypothetical protein586 group 390 SulfurtransferaseAttorney Ref.: 70851.3W001587 group 393 hypothetical protein588 group 394 hypothetical protein589 tlyA Hemolysin A590 fabG 1 3-oxoacyl-facyl-carrier-protein] reductase FabG591 group 399 hypothetical protein592 group 40 hypothetical protein593 group 402 hypothetical proteinFerric-anguibactin transport system permease protein594 fatC FatC595 group 406 hypothetical protein596 group 407 hypothetical protein597 nagB Glucosamine-6-phosphate deaminase598 galK 1 Galactokinase599 Irp Leucine-rich proteinAlkaline phosphatase synthesis transcriptional600 sphR regulatory protein SphR601 group 417 hypothetical protein602 naiiS hypothetical proteinGlycine reductase complex component B subunit603 grdB 1 gamma604 tag DNA-3-methyladenine glycosylase 1605 group 422 hypothetical protein606 bglK 2 Beta-glucoside kinase607 group 426 hypothetical protein608 group 427 hypothetical protein609 folB Dihydroneopterin aldolase610 iolU scyllo-inositol 2-dehydrogenase (NADP( )) IolU611 oppD 1 Oligopeptide transport ATP-binding protein OppD612 group 436 hypothetical protein613 hclrA Protein / nucleic acid deglycase HcliA614 group 439 hypothetical protein615 group 440 hypothetical protein616 rbsK Ribokinase617 group 45 hypothetical protein618 group 463 hypothetical protein619 group 464 hypothetical protein620 group 466 hypothetical protein621 group 468 hypothetical protein622 group 469 hypothetical protein623 group 475 hypothetical protein624 group 476 hypothetical protein625 group 477 hypothetical protein626 grpE Protein GrpEAttorney Ref.: 70851.3W001627 rnhA 147 kDa ribonuclease H-like protein628 rnhB Ribonuclease HII629 group 488 hypothetical protein630 prsA Foldase protein PrsA3',5'-cyclic adenosine monophosphate phosphodiesterase631 cpdA CpdA632 lytG Exo-glucosaniinidase LytG633 groS 10 kDa chaperonin634 group 503 hypothetical protein635 group 508 hypothetical protein636 cmk Cytidylate kinase637 cvfB Conserved virulence factor B638 group 512 hypothetical protein639 group 52 hypothetical protein640 group 525 hypothetical protein641 degV Protein DegV642 yheS putative ABC transporter ATP-binding protein YheS643 tsaD tRNA N6-adenosine threonylcarbamoyltransferase644 sppA Putative signal peptide peptidase SppA645 group 534 hypothetical protein646 selA L-seryl-tRNA(Sec) selenium transferase647 yacP putative protein YacP648 yxeP putative hydrolase YxeP649 group 54 hypothetical protein650 group 541 hypothetical protein651 group 544 hypothetical protein652 troA Periplasmic zinc-binding protein TroA653 nrdH Glutaredoxin-like protein NrdH654 group 549 hypothetical protein655 IrgA 1 Antiholin-like protein LrgA656 group 554 Nudix hydrolase657 yqgN putative protein YqgN658 group 557 hypothetical protein659 group 559 hypothetical protein660 group 56 hypothetical protein661 group 561 Aspartate racemase662 group 562 hypothetical proteinBiotin carboxyl carrier protein of acetyl-CoA663 accB carboxylase664 nylA 2 6-aminohexanoate -cyclic-dimer hydrolase665 group 566 Lipoate— protein ligase 2666 group 570 hypothetical protein667 group 571 hypothetical proteinAttorney Ref.: 70851.3W001668 group 574 hypothetical protein669 group 575 hypothetical protein670 rpc Ribulosc-phosphatc 3-cpimcrasc2-amino-4-hydroxy-6-hydroxymethyldihydropteridine671 folK pyrophosphokinase672 group 595 putative metallo-hydrolase673 group 60 Nitronate monooxygenase674 mro Aldose 1 -epimerase675 agaS D-galactosamine-6-phosphate deaminase AgaS676 group 635 hypothetical protein677 group 641 hypothetical protein678 IplJ 2 Lipoate-protein ligase LplJPEP-dependent dihydroxyacetone kinase 2, phosphoryl679 dhaM-2 donor subunit DhaMHydro xymethylpyrimidine / phosphomethylpyrimidine680 thiD 2 kinase681 group 647 hypothetical protein682 niaR putative transcription repressor NiaR683 group 65 hypothetical protein684 gpsA Glycerol-3-phosphate dehydrogenase INAD(P) 1685 group 661 hypothetical protein686 nth Endonuclease III687 group 664 hypothetical protein688 group 665 putative AAA domain-containing protein689 map Methionine aminopeptidase 1690 InpD UDP-glucose 4-epimerase691 group 67 hypothetical proteinPolyisoprenyl-teichoic acid— peptidoglycan teichoic acid692 tagU transferase TagU693 group 673 hypothetical protein694 group 677 hypothetical protein695 rhaR 1 HTH-type transcriptional activator RhaR696 group 682 hypothetical protein697 group 683 hypothetical protein698 atsA hypothetical protein699 yqeY putative protein YqeY700 dgkA Undecaprenol kinase701 IspA Lipoprotein signal peptidase702 ctpA Carboxy-terminal processing protease CtpA703 group 697 hypothetical protein704 group 701 hypothetical protein705 ypdF Aminopeptidase YpdF706 group 705 hypothetical proteinAttorney Ref.: 70851.3W001707 group 706 hypothetical protein708 group 711 Diacylglycerol kinase709 group 717 Putative mctallophosphocstcrasc MG207710 group 718 Putative TrmH family tRNA / rRNA methyltransferase711 ydbM Putative acyl-CoA dehydrogenase Y dbM712 engB putative GTP-binding protein EngB713 coaD Phosphopantetheine adenylyltransferase714 mshA D-inositol-3-phosphate glycosyltransferase715 group 735 hypothetical protein716 group 737 hypothetical protein717 yodJ Putative carboxypeptidase YodJ718 sufS 1 Cysteine desulfurase SufS719 group 746 hypothetical protein720 group 747 hypothetical protein721 VgaZ hypothetical protein722 group 751 hypothetical protein723 ohrA Organic hydroperoxide resistance protein OhrA724 group 757 hypothetical proteinMethylated-DNA— protein-cysteine methyltransferase,725 adaB inducible726 group 760 hypothetical protein727 group 762 hypothetical protein728 group 765 hypothetical protein729 ftsX Cell division protein FtsX730 group 770 hypothetical protein731 group 772 hypothetical proteinGlycine / sarcosine / betaine reductase complex component732 grdD C subunit alpha733 apbE 2 FAD:protein FMN transferase734 group 785 hypothetical protein735 atzC N-isopropylammelide isopropyl amidohydrolase736 yedJ putative protein YedJtRNA 5-methylaminomethyl-2-thiouridine biosynthesis737 mnmC bifunctional protein MnmC738 group 799 hypothetical protein739 yhaP putative protein YhaP740 group 800 hypothetical protein741 group 801 hypothetical protein742 dut Deoxyuridine 5 '-triphosphate nucleotidohydrolase743 group 808 hypothetical protein744 strH Bcta-N-acctylhcxosaminidasc745 group 810 hypothetical protein746 rpoNl 1 RNA polymerase sigma-54 factor 1Attorney Ref.: 70851.3W001747 group 813 hypothetical protein748 group 817 putative oxidoreductase749 group 82 hypothetical protein750 group 820 hypothetical protein751 atpF ATP synthase subunit b752 tmk Thymidylate kinase753 group 827 hypothetical protein754 group 830 hypothetical protein755 group 833 hypothetical protein756 group 834 hypothetical protein757 recX Regulatory protein RecX758 ebgA Evolved beta-galactosidase subunit alpha759 group 844 hypothetical protein760 lytR 1 Transcriptional regulator LytR761 zosA Zinc -transporting ATPase762 niaX Niacin transporter NiaX763 group 89 hypothetical protein764 group 90 hypothetical protein765 group 904 hypothetical protein766 thiN Thiamine pyrophosphokinase767 gpxl Hydroperoxy fatty acid reductase gpxl768 luxS 2 S-ribosylhomocysteine lyase769 group 94 hypothetical protein770 group 940 hypothetical protein771 ssbA 2 Single-stranded DNA-binding protein A772 rnmV Ribonuclease M5773 chrA hypothetical protein774 group 945 hypothetical protein775 fabG 3 3-oxoacyl-[acyl-carrier-protein] reductase FabG776 group 948 hypothetical protein777 arsD Arsenical resistance operon trans-acting repressor ArsD778 coaE Dephospho-CoA kinase779 group 950 hypothetical protein780 err PTS system glucose-specific ETTA component781 PgpB Phosphatidylglycerophosphatase B782 group 954 hypothetical proteinUDP-N-acetylmuramoyl-tripeptide— D-alanyl-D-alanine783 murF 1 ligase784 group 958 hypothetical protein785 group 959 hypothetical protein786 dsdA 2 D- serine dehydratase787 ruvA Holliday junction ATP-dependent DNA helicase RuvAAttorney Ref.: 70851.3W001788 group 962 hypothetical protein789 group 964 hypothetical protein790 group 965 hypothetical protein791 group 967 hypothetical protein792 group 968 hypothetical protein793 graS Sensor histidine kinase GraS794 group 97 hypothetical protein795 pepFl 2 Oligoendopeptidase F, plasmid796 group 972 hypothetical proteinMonofunctional biosynthetic peptidoglycan797 mtgA transglycosylase798 group 979 hypothetical protein799 group 980 Putative universal stress protein800 diiaB Replication initiation and membrane attachment protein801 deoC Deoxyribose -phosphate aldolase802 dtd D-aminoacyl-tRNA deacylase803 gcvT Annnomethyltransferase804 clcA H( ) / Cl(-) exchange transporter ClcA805 cspA Cold shock protein CspA806 hflX GTPase HflX807 group 996 hypothetical protein808 group 998 hypothetical protein809 group 101 hypothetical protein810 group 1209 hypothetical protein811 group 130 hypothetical proteinAlkaline phosphatase synthesis transcriptional812 phoP regulatory protein PhoP813 group 165 hypothetical proteinputative multidrug ABC transporter ATP-binding814 ybhF protein YbhF815 group 694 hypothetical protein816 group 779 hypothetical proteinExample 2: P. bivia Experimental Protocols

[0078] P. bivia core genome analysis

[0079] A local P. bivia genome database was curated by downloading publicly available genomes from NCBI RefSeq and adding in-house sequenced and assembled P. bivia genomes, as has been described generally in respect of D. pigrum.

[0080] P. bivia assay target identificationAttorney Ref.: 70851.3W001

[0081] The core genome was filtered and only SCSG were retained. Thereafter, following protocols as have been described in respect of D. pigrum, a candidate pool for targets to design P. bivia specific assay was established.

[0082] P. bivia assay design

[0083] The P. bivia assay design was established using the protocols as have been described in respect of D. pigrum.

[0084] P. bivia assay validation

[0085] To assess the sensitivity of the primers, we tested the assays as has been described previously for D. pigrum.

[0086] Human subject research

[0087] The parent study of this project, in which the clinical samples were collected, was granted IRB approval by Uganda Virus Research Institute (UVRI) to the Rakai Health Sciences Program (RHSP) with which GWU has an IRB Authorization Agreement (IAA) in place. This project is NHSR as determined by the GWU IRB and does not have any human subject protection issues.

[0088] Human penile microbiome collection

[0089] Penile coronal sulcus samples for microbiome analysis were collected using premoistened swabs by swabbing around the circumference of the coronal sulcus, after retracting the foreskin in uncircumcised study participants.

[0090] DNA Isolation and purification

[0091] DNA isolation and purification was carried out as described previously in respect of D. pigrum.

[0092] g!021 PCR amplification

[0093] Each g!021 PCR was performed in a 10 pl reaction volume containing 1 pl of template DNA added to 9 pl of master mix containing 3.39 pl Molecular grade water, 5 pl PerfeCTa qPCR ToughMix, 0.1 pl DMSO, 0.23 pl of each primer at 40 pM concentration, and 0.06 pl of each probe at 40 pM concentration. Amplification and fluorescence detections was performed on a Roche 480 II light cycler under the following conditions: initial denaturing of 95°C for 3 minutes, followed by 40 cycles 95°C for 15s for denaturing and 60°C for 1 min for annealing and extension. Amplified DNA was run on a 2% agarose E-gel (ThermoFisher) to assess amplification of D. pigrum DNA. Gels were imaged using a ChemiDoc-It2 (AnalytikAttorney Ref.: 70851.3W001Jena US, Upland, CA). Presence of a visible band near the 172 bp size indicated successful amplification.

[0094] P. bivia phylogenetic and core genome analysis

[0095] P. bivia phylogenetic and core genome analysis was carried out as previously described for D. pigrum.

[0096] P. bivia assay design

[0097] Assay development was performed by Antibiotic Resistance Action Center researchers based on available sequencing data for the taxa of interest. Once a target gene region was identified, using blastn v.2.9.0, primers were developed using the Primer3 program. After iterations of primer design and in silico analysis, we identified a pair of forward and reverse PCR primers (Table 6).Table 6. P. bivia gl021 forward, reverse, and probe primer sequencesAssay Primer Size(bp) Tin GC Sequence (5 -3')Target (annealingtemp)glO21_F 27 58.0 44.4% AGTACTCTTCATCGTGATAGGAGGACT (SEQ ID NO: 5)g!021_R 23 56.4 39.1% TCTTCCTTTTCGTCTTTCAGCTT (SEQ ID NO: 6)g 1021 Probe 20 59.0 40.0% TGTAATCCAAAGAGAGTGTG (SEQ ID NO:7)

[0098] P. bivia PCR sensitivity and specificity against clinical isolates

[0099] The gl021 assay was highly sensitive and specific in laboratory analysis of DNA from bacterial isolates. We first evaluated the assay using well-characterized P. bivia isolates (N = 45) identified as containing P. bivia during previous 16S sequencing, which showed 100% sensitivity. Analysis using extracted genomic DNA resulted in Linear dynamic range of 1E7 -1E2 dsDNA copies / pl, with a linearity of R2>0.99. Reaction efficiency was 82.4%. Crossreactivity evaluation of the gl021 assay demonstrated 100% specificity against a panel of 30 near neighbor, and common co-inhabiting species namely Lactobacillus iners, Dialister micraerophilus, and Peptostresptococcus anaerobius.

[0100] We further evaluated the assay using DNA in experiments to estimate intra-run (i.e., within each PCR plate) and inter-run (i.e.g, between PCR plates) variations using 3 standardAttorney Ref.: 70851.3W001curves per plate across 3 plates, where each standard curve comprises 10-fold dilutions range from 10 copies to 10A7 copies per uL reaction. All reactions are performed in triplicate. Variations were estimated using coefficient of variation for Cp values and interpolated copy numbers. Limit of quantification is determined based on target copy number with CoV < 20%. Limit of detection is determined based on experiments that dilute target down to extinction to estimate the target copy number that will amplify 100% of the time.

[0101] Referring to FIG.5A, in an aspect of P. bivia, P. anaerobius, and D. micraerophilus quantitative validation, quantitative validation was established by calculating the Limit of Detection (LOD), Limit of Quantification (LOQ), Linear Dynamic Range, reaction efficiency, R2, and intra-assay coefficient of variation (CoV). The LOD was the lowest concentration of target DNA at which 95% of the positive samples were detected by the assay. The linear dynamic range of the assay was determined by performing a calibration curve with the target DNA in a 10-fold serial dilution. Performing the serial dilution in triplicate, the threshold cycle values were plotted on a base 10-semi logarithmic graph (referring to FIG. 5B, showing I), micraerophilus data). The linear range of the plot matched the linear dynamic range of the assay. The limit of quantification (LOQ) was the value at the lowest end of this linear relationship. The linearity of the plot, R2, was evaluated based on a result > 0.980. The reaction efficiency was calculated using the following equation: Efficiency=-l+(10A(-l / slope)) in accordance with aspects of the disclosure.

[0102] An aspect of the disclosure is directed to, by identifying potential assay targets using the P. bivia core genome, designing a PCR assay that is both sensitive and specific for P. bivia. In contrast to other commonly used methods for species confirmation, such as biochemical testing, DNA sequencing, or MALDI-TOF, PCR-based assays are rapid and cost-effective and do not require expensive equipment. This method provides a simpler option for P. bivia detection and avoids the restriction digestion and analysis challenges of T-RFLP (see, for e.g.: Prakash et al., (2014) Indian J. Microbiol.') that has been used previously for detecting microbial communities (see, for e.g.-. Camarinha-Silva et al. (2012) FEMS Microbiol. Ecol.). We demonstrated the utility of the core genome mining techniques to develop species confirmation assays. The resultant g!021 assay was able to identify P. bivia and diverse bacterial isolates with al 00% sensitivity and specificity. Our assay was also highly sensitive and specific for detecting P. bivia in clinical samples.Attorney Ref.: 70851.3W001

[0103] The above-described utilization of the g!021 assay can be expanded to any core genome-based candidates identified for P. bivia.Example 3: P. anaerobius Experimental Protocols

[0104] P. anaerobius core genome analysis

[0105] A local P. anaerobius genome database was curated by downloading publicly available genomes from NCBI RefSeq and adding in-house sequenced and assembled P. anaerobius genomes, as has been described generally in respect of D. pigrum.

[0106] P. anaerobius assay target identification

[0107] The core genome was filtered and only SCSG were retained. Thereafter, following protocols as have been described in respect of D. pigrum, a candidate pool for targets to design P. anaerobius specific assay was established.

[0108] P. anaerobius assay design

[0109] The P. anaerobius assay design was established using the protocols as have been described in respect of D. pigrum.

[0110] P. anaerobius assay validation

[0111] To assess the sensitivity of our primers, we tested the assays as has been described previously for D. pigrum.

[0112] Human subject research

[0113] The parent study of this project, in which the clinical samples were collected, was granted IRB approval by Uganda Virus Research Institute (UVRI) to the Rakai Health Sciences Program (RHSP) with which GWU has an IRB Authorization Agreement (IAA) in place. This project is NHSR as determined by the GWU IRB and does not have any human subject protection issues.

[0114] Human penile microbiome collection

[0115] Human penile microbiome collection was carried out as described previously in respect of P. bivia.

[0116] DNA Isolation and purification

[0117] DNA isolation and purification was carried out as described previously in respect of P. bivia.Attorney Ref.: 70851.3W001

[0118] gH2 PCR amplification

[0119] Each g!12 PCR was performed in a 10 pl reaction volume containing 1 pl of template DNA added to 9 pl of master mix containing 3.39 pl Molecular grade water, 5 pl PerfeCTa qPCR ToughMix, 0.1 pl DMSO, 0.23 pl of each primer at 40 pM concentration, and 0.06 pl of each probe at 40 pM concentration. Amplification and fluorescence detections was performed on a Roche 480 II light cycler under the following conditions: initial denaturing of 95°C for 3 minutes, followed by 40 cycles 95°C for 15s for denaturing and 60°C for 1 min for annealing and extension. Amplified DNA was ran on a 2% agarose E-gel (ThermoFisher) to assess amplification of D. pigrum DNA. Gels were imaged using a ChemiDoc-It2 (Analytik Jena US, Upland, CA). Presence of a visible band near the 214 bp size indicated successful amplification.

[0120] P. anaerobius phylogenetic and core genome analysis

[0121] P. anaerobius phylogenetic and core genome analysis was earned out as previously described for / ). pigrum.

[0122] P. anaerobius assay design

[0123] Assay development was performed by Antibiotic Resistance Action Center researchers based on available sequencing data for the taxa of interest. Once a target gene region was identified, using blastn v.2.9.0, primers were developed using the Primer3 program. After iterations of primer design and in silico analysis, we identified a pair of forward and reverse PCR primers (Table 8).Table 8. P. anaerobius g!12 forward, reverse, and probe primer sequencesAssay Primer Size(bp) Tm (annealing GC Sequence (5 '-3')Target temp)g772 g772_F 20 55.3 45.0% ATGTCGAGGTTTAGGGCMAA (SEQ ID NO: 8)gH2_R 22 57.0 54.5% CCCTTTCCAGTGTAGAACCTCC (SEQ ID NO: 9)g772_Probe 19 58.0 42.1% AGGGATTATGGAGGACTAT (SEQ ID NO: 10)

[0124] P. anaerobius PCR sensitivity and specificity against clinical isolates

[0125] The g!12 assay was highly sensitive and specific in laboratory analysis of DNA from bacterial isolates. We first evaluated the assay using well-characterized P. anaerobiusAttorney Ref.: 70851.3W001isolates (N = 45) identified as containing P. anaerobius during previous 16S sequencing, which showed 97.7% sensitivity. Analysis using extracted genomic DNA resulted in Linear dynamic range of 1E7 - 1E2 dsDNA copies / pl, with a linearity of R2>0.98. The reaction efficiency was 107.59%. Cross-reactivity evaluation of the g!12 assay demonstrated 100% specificity against a panel of 30 common co-inhabiting species namely Lactobacillus iners, Dialister micraerophilus, and Prevotella bivia.

[0126] We further evaluated the assay using DNA as previously described for P. bivia.

[0127] Again, and referring to FIG. 5A, in an aspect of P. bivia, P. anaerobius, and I), micraerophilus quantitative validation, quantitative validation was established by calculating the Limit of Detection (LOD), Limit of Quantification (LOQ), Linear Dynamic Range, reaction efficiency, R2, and intra-assay coefficient of variation (CoV). The LOD was the lowest concentration of target DNA at which 95% of the positive samples were detected by the assay. The linear dynamic range of the assay was determined by performing a calibration curve with the target DNA in a 10-fold serial dilution. Performing the serial dilution in triplicate, the threshold cycle values were plotted on a base 10-semi logarithmic graph (referring to FIG.5B, showing D. micraerophilus data). The linear range of the plot matched the linear dynamic range of the assay. The limit of quantification (LOQ) was the value at the lowest end of this linear relationship. The linearity of the plot, R2, was evaluated based on a result > 0.980. The reaction efficiency was calculated using the following equation: Efficiency=-l+(10A(-l / slope)) in accordance with aspects of the disclosure.

[0128] An aspect of the disclosure is directed to, by identifying potential assay targets using the P. anaerobius core genome, designing a PCR assay that is both sensitive and specific for P. anaerobius. In contrast to other commonly used methods for species confirmation, such as biochemical testing, DNA sequencing, or MALDI-TOF, PCR-based assays are rapid and cost-effective and do not require expensive equipment. This method provides a simpler option for P. anaerobius detection and avoids the restriction digestion and analysis challenges of T-RFLP (see, for e.g.'. Prakash el al., (2014) Indian J. Microbiol.) that has been used previously for detecting microbial communities (see, for e.g. : Camarinha-Silva et al. (2012) FEMS Microbiol. Ecol.). We demonstrated the utility of the core genome mining techniques to develop species confirmation assays. The resultant g!12 assay was able to identify P. anaerobius and diverse bacterial isolates with a 100% sensitivity and specificity. Our assay was also highly sensitive and specific for detecting P. anaerobius in clinical samples.Attorney Ref.: 70851.3W001

[0129] The above-described utilization of the gll2 assay can be expanded to any core genome-based candidates identified for P. anaerobius.Example 4: D. micraerophilus Experimental Protocols

[0130] D. micraerophilus core genome analysis

[0131] A local D. micraerophilus genome database was curated by downloading publicly available genomes from NCBI RefSeq and adding in-house sequenced and assembled I), micraerophilus, as has been described generally in respect of D. pigrum.

[0132] D. micraerophilus assay target identification

[0133] The core genome was filtered and only SCSG were retained. Thereafter, following protocols as have been described in respect of D. pigrum, a candidate pool for targets to design a D. micraerophilus specific assay was established.

[0134] £>. micraerophilus assay design

[0135] The D. micraerophilus assay design was established using the protocols as have been described in respect of D. pigrum.

[0136] D. micraerophilus assay validation

[0137] To assess the sensitivity of our primers, we tested the assays as has been described previously for D. pigrum.

[0138] Human subject research

[0139] The parent study of this project, in which the clinical samples were collected, was granted IRB approval by Uganda Virus Research Institute (UVRI) to the Rakai Health Sciences Program (RHSP) with which GWU has an IRB Authorization Agreement (IAA) in place. This project is NHSR as determined by the GWU IRB and does not have any human subject protection issues.

[0140] Human penile microbiome collection

[0141] DNA isolation and purification was carried out as described previously in respect of P. bivia.

[0142] DNA Isolation and purification

[0143] DNA isolation and purification was carried out as described previously in respect of P. bivia.Attorney Ref.: 70851.3W001

[0144] acpP PCR amplification

[0145] Each acpP PCR was performed in a 10 pl reaction volume containing 1 pl of templateDNA added to 9 pl of master mix containing 3.39 pl Molecular grade water, 5 pl PerfeCTa qPCR ToughMix, 0.1 pl DMSO, 0.23 pl of each primer at 40 pM concentration, and 0.06 pl of each probe at 40 pM concentration. Amplification and fluorescence detections was performed on a Roche 480 II light cycler under the following conditions: initial denaturing of 95°C for 3 minutes, followed by 40 cycles 95°C for 15s for denaturing and 60°C for 1 min for annealing and extension. Amplified DNA was run on a 2% agarose E-gel (ThermoFisher) to assess amplification of D. pigrum DNA. Gels were imaged using a ChemiDoc-It2 (Analytik Jena US, Upland, CA). Presence of a visible band near the 214 bp size indicated successful amplification.

[0146] D. micraerophilus phylogenetic and core genome analysis

[0147] D. micraerophilus phylogenetic and core genome analysis was carried out as previously described for D. pigrum.

[0148] D. micraerophilus assay design

[0149] Assay development was performed by Antibiotic Resistance Action Center researchers based on available sequencing data for the taxa of interest. Once a target gene region was identified, using blastn v.2.9.0, primers were developed using the Primer3 program. After iterations of primer design and in ilico analysis, we identified a pair of forward and reverse PCR primers (Table 9).Table 9. D. micraerophilus acpP forward, reverse, and probe primer sequences Assay Primer Size(bp) Tm GC Sequence (5 '-3')Target (annealingtemp)acpP acpP P 22 56.9 45.5% ATGAGCGCATTTGACAGAGTGA (SEQ ID NO: SEQ ID NO: 11)acpP R 26 58.0 46.2% TGTAGTCTACTGCATCACGAACTGTC (SEQ ID NO: SEQ ID NO: 12)acpP Probe 17 51.0 35.3% CTTCAAATGCCATAATC (SEQ ID NO: 13)

[0150] D. micraerophilus PCR sensitivity and specificity against clinical isolatesAttorney Ref.: 70851.3W001The acpP assay was highly sensitive and specific in laboratory analysis of DNA from bacterial isolates. We first evaluated the assay using well-characterized D. micraerophilus isolates (N = 45) identified as containing I), micraerophilus during previous 16S sequencing, which showed 100% sensitivity. Analysis using extracted genomic DNA resulted in Linear dynamic range of 1E6 - 1E2 dsDNA copies / pl, with a linearity of R2>0.99. The reaction efficiency was 105.35%. Cross-reactivity evaluation of the acpP assay demonstrated 100% specificity against a panel of 30 near neighbor, and common co-inhabiting species namely Lactobacillus iners, P. bivia, and Peptostresptococcus anaerobius.

[0151] We further evaluated the assay using DNA as previously described for P. bivia.

[0152] Again, and referring to FIG. 5A, in an aspect of P. bivia, P. anaerobius, and D. micraerophilus quantitative validation, quantitative validation was established by calculating the Limit of Detection (LOD), Limit of Quantification (LOQ), Linear Dynamic Range, reaction efficiency, R2, and intra-assay coefficient of variation (CoV). The LOD was the lowest concentration of target DNA at which 95% of the positive samples were detected by the assay. The linear dynamic range of the assay was determined by performing a calibration curve with the target DNA in a 10-fold serial dilution. Performing the serial dilution in triplicate, the threshold cycle values were plotted on a base 10-semi logarithmic graph (referring to FIG. 5B, showing D. micraerophilus data). The linear range of the plot matched the linear dynamic range of the assay. The limit of quantification (LOQ) was the value at the lowest end of this linear relationship. The linearity of the plot, R2, was evaluated based on a result > 0.980. The reaction efficiency was calculated using the following equation: Efficiency=-l+(10A(-l / slope)) in accordance with aspects of the disclosure.

[0153] An aspect of the disclosure is directed to, by identifying potential assay targets using the D. micraerophilus core genome, designing a PCR assay that is both sensitive and specific for D. micraerophilus. In contrast to other commonly used methods for species confirmation, such as biochemical testing, DNA sequencing, or MALDI-TOF, PCR-based assays are rapid and cost-effective and do not require expensive equipment. This method provides a simpler option for D. micraerophilus detection and avoids the restriction digestion and analysis challenges of T-RFLP (see, for e.g. : Prakash et al., (2014) Indian J. Microbiol.) that has been used previously for detecting microbial communities (see, for e.g.: Camarinha-Silva et al. (2012) FEMS Microbiol. Ecol.). We demonstrated the utility of the core genome mining techniques to develop species confirmation assays. The resultant acpP assay was able toAttorney Ref.: 70851.3W001identify D. micraerophilus and diverse bacterial isolates with a 100% sensitivity and specificity. Our assay was also highly sensitive and specific for detecting D. micraerophilus in clinical samples.

[0154] The above-described utilization of the acpP assay can be expanded to any core genome-based candidates identified for I). micraerophilus.

[0155] Of note, the exemplar embodiments of the disclosure described herein do not limit the scope of the invention since these embodiments are merely examples of the embodiments of the invention. Any equivalent embodiments are intended to be within the scope of this invention. Indeed, various modifications of the disclosure, in addition to those shown and described herein, such as alternative useful combinations of the elements described, may become apparent to those skilled in the art from the description. Such modifications and embodiments are also intended to fall within the scope of the appended claims.

Claims

Attorney Ref.: 70851.3W001CLAIMSWHAT IS CLAIMED IS:

1. A method of determining a genomic biomarker for detecting an organism of interest, comprising:(a) obtaining comprehensive genomic strain data from the organism of interest;(b) removing at least one subset of genomic strain data from the comprehensive genomic strain data; and(c) filtering the data obtained from step (b) by requiring the resultant genomic strain data to have at least one pre-selected genomic feature that is common to the identified genomic biomarker.

2. The method of claim 1, wherein the at least one subset of genomic strain data comprises ribosomal genomic data or homolog genomic data.

3. The method of claim 1, further comprising removing at least two subsets of genomic strain data.

4. The method of claim 3, wherein the at least two subsets of genomic strain data comprise ribosomal genomic data and homolog genomic data.

5. The method of claim 1, wherein the at least one pre-selected genomic feature comprises a requirement that the identified genomic biomarker: (i) is present in all genomes of the organism of interest; (ii) has at least 70% sequence identity with taxa members associated with the organism of interest; or (iii) contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.

6. The method of claim 1 , further comprising at least two pre-selected genomic features.

7. The method of claim 6, wherein the at least two pre-selected genomic features are selected from a requirement that the identified genomic biomarker: (i) is present in all genomes of the organism of interest; (ii) has at least 70% sequence identity with taxa members associated with the organism of interest; and (iii) contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.Attorney Ref.: 70851.3W0018. The method of claim 1, further comprising at least three pre-selected genomic features.

9. The method of claim 8, wherein the at least three pre-selected genomic features comprise a requirement that the identified genomic biomarker: (i) is present in all genomes of the organism of interest; (ii) has at least 70% sequence identity with taxa members associated with the organism of interest; and (iii) contains forward and reverse primer sequences that meets Primer3 design criteria and have less than 50% sequence identity and cover against sequences from taxa members associated with the organism of interest.

10. The method of claim 1, wherein the comprehensive genomic strain data is associated with a pathogenic organism.

11. The method of claim 1, wherein the comprehensive genomic strain data is associated with a non-pathogenic organism.

12. The method of claim 1, wherein the comprehensive genomic strain data is associated with a bacterium.

13. The method of claim 1, wherein the comprehensive genomic strain data is associated with a fungus.

14. The method of claim 1, wherein the comprehensive genomic strain data is associated with a virus.

15. The method of claim 1, wherein the biomarker is used as an epitope in an experimental development, selection, and production of a custom immunoassay reagent of polyclonal antibodies, monoclonal antibodies, recombinant antibodies, camelid nanobodies, and single chain variable fragments.

16. The method of claim 1 , wherein a portion of the biomarker is used as an epitope in an experimental development, selection, and production of a custom immunoassay reagent of polyclonal antibodies, monoclonal antibodies, recombinant antibodies, camelid nanobodies, and single chain variable fragments.

17. The method of claim 1 , wherein the biomarker is used as a protein target for in silico design of an immunoassay reagent of an antibody fragment with at least 70% affinity with a region of the biomarker.Attorney Ref.: 70851.3W00118. The method of claim 1 , wherein a portion of the biomarker is used as a protein target for in silico design of an immunoassay reagent of an antibody fragment with at least 70% affinity with a region of the biomarker.

19. The method of claim 1, wherein the biomarker is used as a polypeptide target for in silico design of an immunoassay reagent of a single chain variable fragment with at least 70% affinity with a region of the biomarker.

20. The method of claim 1, wherein a portion of the biomarker is used as a polypeptide target for in silico design of an immunoassay reagent of a single chain variable fragment with at least 70% affinity with a region of the biomarker.

21. The method of claim 12, wherein the bacterium comprises Dolosigranulum pigrum or variants thereof.

22. The method of claim 12, wherein the bacterium comprises Prevotella bivia or variants thereof.

23. The method of claim 12, wherein the bacterium comprises Peptostreptococcus anaerobius or variants thereof.

24. The method of claim 12, wherein the bacterium comprises Dialister micraerophilus or variants thereof.