Probes and probe sequences for the detection, identification and differentiation of bacteria, pathogenicity elements, and antimicrobial resistance (AMR) genes, and methods of designing, making and using

EP4689192A1Pending Publication Date: 2026-02-11THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024782000
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2024-03-29
Publication Date
2026-02-11

AI Technical Summary

Technical Problem

Current methods for diagnosing bacterial infections are slow, lack insight into antimicrobial resistance (AMR) genes, and fail to provide early and precise antibiotic treatment, leading to morbidity, mortality, and economic burdens.

Method used

A database of probe sequences and probes designed to detect, identify, and differentiate bacteria, pathogenicity elements, and AMR genes using species-specific or clade-specific gene sequences, 16S ribosomal RNA, virulence factor sequences, and AMR genes, enabling high-throughput sequencing and rapid identification.

Benefits of technology

The solution enhances the sensitivity and speed of bacterial detection and identification, allowing for early and precise antibiotic treatment, reducing morbidity, mortality, and healthcare costs by providing rapid insights into bacterial infections and antimicrobial resistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024022175_03102024_PF_FP_ABST
    Figure US2024022175_03102024_PF_FP_ABST
Patent Text Reader

Abstract

Described herein is a database of probe sequences and a set of probes that enable the detection, identification and differentiation of bacteria, and one or more of 16S ribosomal RNA pathogenicity elements, and / or antimicrobial resistance (AMR) genes. These sequences or probes have many uses including but not limited to use in a sequence capture platform and other diagnostic assays. The sequences or probes described herein increase the sensitivity of high- throughput sequencing for detection, identification, and differentiation of bacteria, and one or more of 16S ribosomal RNA, pathogenicity elements, and AMR genes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PROBES AND PROBE SEQUENCES FOR THE DETECTION, IDENTIFICATION AND DIFFERENTIATION OF BACTERIA, PATHOGENICITY ELEMENTS, AND ANTIMICROBIAL RESISTANCE (AMR) GENES, AND METHODS OF DESIGNING, MAKING AND USING

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 455,774, filed March 30, 2023, the content of which is hereby incorporated by reference.

[0003] Throughout this application, various publications are referenced, including referenced in parenthesis. The disclosures of all publications mentioned in this application in their entireties are hereby incorporated by reference into this application in order to provide additional description of the art to which this invention pertains and of the features in the art which can be employed with this invention.

[0004] BACKGROUND OF THE INVENTION

[0005] Early, accurate differential diagnosis of bacterial infections is critical to reducing morbidity, mortality, and health care costs. It can also reduce the inappropriate use of antibiotics. Multiplex PCR methods in common use for differential diagnosis of bacterial infections can identify potential pathogens but do not provide insights into the presence or expression of antimicrobial resistance (AMR) genes. Moreover, culture-based methods require two to several days to identify pathogens and even longer to provide antibiotic susceptibility profiles (Rhee et al., 2017). Accordingly, physicians typically administer broad-spectrum antibiotics pending acquisition of more specific information (Howell and Davis, 2017).

[0006] Antibiotic resistance is the ability of bacteria to resist the effects of antibiotics. This occurs when bacteria evolve mechanisms to neutralize the drugs designed to kill them. Antibiotic resistance is a growing public health concern as it can lead to the spread of antibiotic-resistant infections, which are difficult to treat and can be deadly.

[0007] No platform currently permits rapid and simultaneous insights into phylogeny and pathogenicity markers needed to enable the early and precise antibiotic treatment that could reduce morbidity, mortality and economic burden. Moreover, there is currently no method to quickly and accurately identify if a bacterial infection is resistant to one or more antibiotics. Thus, there is a need for a sensitive cost-effective assay for the detection of bacteria, especially in a clinical setting, as well as features associated with pathogenicity and antibiotic resistance.

[0008] BRIEF SUMMARY OF THE INVENTION

[0009] Described herein is a database of probe sequences and a set of probes that enable the detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or antimicrobial resistance (AMR) genes and / or 16S ribosomal RNA (rRNA). These sequences or probes have many uses including but not limited to use in a sequence capture platform and other diagnostic assays. The sequences or probes described herein increase the sensitivity of high- throughput sequencing for detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA. The current database of probe sequences or set of probes comprises less than one million oligonucleotides.

[0010] To enable efficient detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or antimicrobial resistance and / or 16S ribosomal RNA, the database of probe sequences and set of probes was designed to target species-specific or clade-specific gene sequences; and / or 16S ribosomal RNA sequences; and / or virulence factor sequences; and / or AMR genes.

[0011] Accordingly, disclosed herein is a method of designing and / or making or constructing a database of probe sequences or a set of probes comprising the following steps.

[0012] The first step is to obtain sequence information. In some embodiments, sequence information is obtained for:

[0013] (i) one or more species-specific or clade-specific marker gene sequences; or

[0014] (ii) one or more 16S ribosomal RNA sequences; or

[0015] (iii) one or more virulence factor sequences; or

[0016] (iv) one or more AMR gene sequences; or

[0017] (v) any combination of (i), (ii), (iii) and (iv)

[0018] Sequence information is obtained from any public or private database of sequence information of bacteria and / or 16S ribosomal RNA and / or AMR genes and / or virulence factors, including, but not limited to, Metaphlan4, SILVA, CARD (The Comprehensive Antibiotic Resistance Database) and VFDB (Virulence Factor Database). For example, versions of each of these databases are provided in Table 2, however, additional versions, releases, and updates to these or other databases may be used.

[0019] In some embodiments, the combined target sequence dataset can contain over 101,000 genetic targets.

[0020] The next step of the method is to break the target sequences into fragments to be the basis of the oligonucleotide probes. The probes are designed to be of a length, and spaced at a distance across the target sequences, such that the total number of probe sequences in the database or probes in the probe set corresponds to a desired range or number. For example, the length and spacing of the probes may be configured to result in less than one million probes. In other embodiments, the length and spacing of the probes may be configured to result in about one million probes. In further embodiments, the length and spacing of the probes may be configured to result in over one million probes.

[0021] In some embodiments, the probe length is about 5 nucleotides (“nt”) to about 300 nt. In some embodiments, the probe length is about 10 nt to about 280 nt. In some embodiments, the probe length is about 20 nt to about 260 nt. In some embodiments, the probe length is about 30 nt to about 240 nt. In some embodiments, the probe length is about 40 nt to about 220 nt. In some embodiments, the probe length is about 50 nt to about 200 nt. In some embodiments, the probe length is about 60 nt to about 190 nt. In some embodiments, the probe length is about 70 nt to about 180 nt. In some embodiments, the probe length is about 80 nt to about 170 nt. In some embodiments, the probe length is about 90 nt to about 160 nt. In some embodiments, the probe length is about 100 nt to about 150 nt. In some embodiments, the probe length is about 110 nt to about 140 nt. In some embodiments, the probe length is about 115 nt to about 130 nt. In some embodiments, the probe length is about 120 nt.

[0022] In some embodiments, the inter-probe spacing is about 20 nt to about 100 nt tiled across the target sequences. In some embodiments, the inter-probe spacing is about 30 nt to about 90 nt tiled across the target sequences. In some embodiments, the inter-probe spacing is about 40 nt to about 80 nt tiled across the target sequences. In some embodiments, the inter-probe spacing is about 50 nt to about 70 nt tiled across the target sequences.

[0023] The generated probes can be further clustered for sequence identity to obtain a certain number of probe sequences or probes. In some embodiments, the generated probes are clustered at about 90% to about 99% sequence identity. In some embodiments, the generated probes are clustered at about 92% to about 98% sequence identity. Tn some embodiments, the generated probes are clustered at about 94% to about 97% sequence identity. In some embodiments, the generated probes are clustered at about 95% to about 97% sequence identity. In some embodiments, the generated probes are clustered at about 96% sequence identity to obtain less than 1 million probes.

[0024] Embodiments of the present disclosure also provide automated systems and methods for designing and / or constructing the database of probe sequences and / or set of probes.

[0025] In some embodiments, systems, apparatuses, methods, and computer readable media are provided that use bacterial and sequence information along with analytical tools in a design model for designing and / or constructing the database of probe sequences and / or set of probes. For example, in some embodiments, a first analytical tool using the information from speciesspecific or clade-specific marker genes sequences and / or from 16S ribosomal RNA sequences and / or virulence factor sequences and / or AMR genes and a second analytical tool to fragment the sequences into oligonucleotides with the desired or advantageous parameters for the probes including but not limited to probe length, spacing distance between the probes on the target sequences, and percentage sequence identity.

[0026] A further embodiment of the present disclosure is a database of probe sequences and / or a set of probes designed and / or made or constructed using the methods described herein. In one embodiment, the database of probe sequences and / or set of probes comprises less than one million probes. In another embodiment, the dataset of probe sequences and / or set of probes comprises about one million probes. In a further embodiment, the dataset of probe sequences and / or set of probes comprises more than one million probes.

[0027] In one embodiment, the probes are oligonucleotide probes. In a further embodiment, the oligonucleotide probes are synthetic. In one embodiment, the set of probes is in the form of an oligonucleotide probe library. In one embodiment, the oligonucleotides can comprise DNA, RNA, linked nucleic acids (LNA), bridged nucleic acids (BNA) and / or peptide nucleic acids (PNA) as well as any nucleic acids that can be derived naturally or synthesized now or in the future. In one embodiment, the set of probes is in the form of a solution. In a further embodiment, the set of probes is in a solid-state form such as a microarray or bead. In a further embodiment, the oligonucleotides are modified by a composition to facilitate binding to a solid state. A further embodiment is a database comprising information on the probes including but not limited to the length, nucleotide sequence, and / or origin of each oligonucleotide probe. A further embodiment is a computer-readable storage medium with program code comprising information, e.g., a database, comprising information regarding the probes including but not limited to the length, nucleotide sequence, and / or origin of each oligonucleotide probe.

[0028] Additionally, the present disclosure provides a method for constructing a sequencing library for the detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes using the disclosed set of probes.

[0029] The present disclosure also provides systems and methods using the database of probe sequences and / or the set of probes for detecting, identifying and / or differentiating bacteria and / or pathogenicity elements and / or AMR genes in a single sample.

[0030] The present disclosure also provides for kits.

[0031] The present disclosure also provides a bacterial sequence capture platform for the detection, identification, and / or differentiation of bacterially-derived sequences in a sample.

[0032] In some embodiments, the platform comprises a plurality of oligonucleotide probes, wherein the plurality comprises at least one oligonucleotide probe which comprises a hybridization portion partially or fully complementary to a portion of a bacterially-derived sequence selected from the group consisting of a bacterial gene sequence, a 16S ribosomal RNA sequence, a pathogenicity element sequence, a virulence factor sequence, and an antimicrobial resistance (AMR) gene sequence.

[0033] In some embodiments, the sequences of the hybridization portions of the oligonucleotide probes cluster at about 90-100% sequence identity.

[0034] In some embodiments, each hybridization portion of an oligonucleotide probe is about 5- 300 nucleotides in length,

[0035] In some embodiments, different hybridization portions that each bind a different portion of the same bacterially-derived sequence are tiled across said bacterially-derived sequence and have an inter-probe spacing of about 20-100 nucleotides.

[0036] In some embodiments, the plurality of oligonucleotide probes of the platform comprises 100,000 to 1,000,000 oligonucleotide probes, preferably less than about 1,000,000 oligonucleotide probes. The present disclosure also provides for methods of using the platform and kits comprising the platform.

[0037] BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figs. 1A-1B show identification of bacterial species (Fig. 1A) and resistance genes (Fig. IB) in contrived plasma samples using a bacterial sequence capture platform as described herein. The K. pnemoiiiae strain has AMR genes for carbapenem (KPC), beta-lactamase (0xa9, SHV), trimethoprim (dfrA). and efflux pumps (LptD, Kpne-KpnG).

[0039] DETAILED DESCRIPTION OF THE INVENTION

[0040] Molecular biology

[0041] In accordance with the present disclosure, there may be numerous tools and techniques within the skill of the art, such as those commonly used in molecular immunology, cellular immunology, pharmacology, and microbiology. See, e.g., Sambrook et al. (2001) Molecular Cloning: A Laboratory Manual. 3rd ed. Cold Spring Harbor Laboratory Press: Cold Spring Harbor, N.Y.; Ausubel et al. eds. (2005) Current Protocols in Molecular Biology. John Wiley and Sons, Inc.: Hoboken, N.J.; Bonifacino et al. eds. (2005) Current Protocols in Cell Biology. John Wiley and Sons, Inc.: Hoboken, N.J.; Coligan et al. eds. (2005) Current Protocols in Immunology, John Wiley and Sons, Inc.: Hoboken, N.J.; Coico et al. eds. (2005) Current Protocols in Microbiology, John Wiley and Sons, Inc.: Hoboken, N.J.; Coligan et al. eds. (2005) Current Protocols in Protein Science, John Wiley and Sons, Inc.: Hoboken, N.J.; and Enna et al. eds. (2005) Current Protocols in Pharmacology, John Wiley and Sons, Inc.: Hoboken, N.J.

[0042] Definitions

[0043] The terms used in this specification generally have their ordinary meanings in the art, within the context of this disclosure and the specific context where each term is used. Certain terms are discussed below, or elsewhere in the specification, to provide additional guidance to the practitioner in describing the disclosed methods and how to use them. Moreover, it will be appreciated that the same thing can be said in more than one way. Consequently, alternative language and synonyms may be used for any one or more of the terms discussed herein, nor is any special significance to be placed upon whether or not a term is elaborated or discussed herein. Synonyms for certain terms are provided. A recital of one or more synonyms does not exclude the use of the other synonyms. The use of examples anywhere in the specification, including examples of any terms discussed herein, is illustrative only, and in no way limits the scope and meaning of the invention or any exemplified term. Likewise, the invention is not limited to its preferred embodiments.

[0044] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting.

[0045] In the discussion unless otherwise stated, adjectives such as “substantially” and “about” modifying a condition or relationship characteristic of a feature or features of an embodiment of the invention, are understood to mean that the condition or characteristic is defined to within tolerances that are acceptable for operation of the embodiment for an application for which it is intended. In embodiments, about means within a standard deviation using measurements generally acceptable in the art. In embodiments, about means a range extending to + / - 10% of the specified value. In embodiments, about includes the specified value. Unless otherwise indicated, the word “or” in the specification and claims is considered to be the inclusive “or” rather than the exclusive or, and indicates at least one of and any combination of items it conjoins.

[0046] As used herein and in the claims, the singular forms “a,” “an,” and “the” include the singular and the plural reference unless the context clearly indicates otherwise. Thus, for example, a reference to “an agent” includes a single agent and a plurality of such agents. Accordingly, it should be understood that the terms “a” and “an” as used above and elsewhere herein refer to “one or more” of the enumerated components. It will be clear to one of ordinary skill in the art that the use of the singular includes the plural unless specifically stated otherwise. Therefore, the terms “a,” “an” and “at least one” are used interchangeably in this application.

[0047] For purposes of better understanding the present teachings and in no way limiting the scope of the teachings, unless otherwise indicated, all numbers expressing quantities, percentages or proportions, and other numerical values used in the specification and claims, are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the following specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained. At the very least, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0048] In the description and claims of the present application, each of the verbs, “comprise,” “include” and “have” and conjugates thereof, are used to indicate that the object or objects of the verb are not necessarily a complete listing of components, elements or parts of the subject or subjects of the verb. Other terms as used herein are meant to be defined by their well-known meanings in the art.

[0049] Where a numerical range is provided herein, it is understood that all numerical subsets of that range, and all the individual integers contained therein, are provided as part of the invention. For example, an oligonucleotide probe which is from 100 to 150 nucleotides in length includes the subset of oligonucleotide probes which are 100 to 140 nucleotides in length, the subset of oligonucleotide probes which are 130 to 150 nucleotides in length etc. as well as an oligonucleotide probe which is 100 nucleotides in length, an oligonucleotide probe which is 101 nucleotides in length, an oligonucleotide probe which is 102 nucleotides in length, etc. up to and including an oligonucleotide probe which is 150 nucleotides in length.

[0050] As used herein the terms “database of probe sequences” or “database of sequences” and refers to a database comprising information on the probes disclosed herein for the detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA and possibly including the length, nucleotide sequence, and / or origin of each oligonucleotide probe, and computer-readable storage mediums with program code comprising information on the probes disclosed herein for the detection, identification, and / or differentiation of bacteria, and and / or pathogenicity elements, and / or AMR genes and / or 16S ribosomal RNA and possibly including the length, nucleotide sequence, and / or origin of each oligonucleotide probe.

[0051] As used herein, the terms “set of probes” or “set of oligonucleotide probes” will be used interchangeably and can refer to the set of probes disclosed herein for the detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA in the form of a collection of synthetic oligonucleotides either in solution or attached to a solid support. As used herein, the term "oligonucleotide" or “oli onucleotide probe” refers to a nucleic acid that is hybridizable to a genomic DNA molecule, a cDNA molecule, or an mRNA molecule encoding a gene, mRNA, cDNA, or other nucleic acid of interest. The nucleic acids comprised in the oligonucleotides include but are not limited to DNA, RNA, linked nucleic acids (LNA), bridged nucleic acids (BNA) and peptide nucleic acids (PNA). Oligonucleotides can be labeled, e.g., with32P-nucleotides or nucleotides to which a label, such as biotin, has been covalently conjugated.

[0052] The term “synthetic oligonucleotide” refers to single-stranded DNA or RNA molecules which can be synthesized. In general, these synthetic molecules are designed to have a unique or desired nucleotide sequence, although it is possible to synthesize families of molecules having related sequences and which have different nucleotide compositions at specific positions within the nucleotide sequence. The term synthetic oligonucleotide will be used to refer to DNA or RNA molecules having a designed or desired nucleotide sequence.

[0053] The term “subject” as used in this application can mean an animal with an immune system such as avians and mammals. Mammals include canines, felines, rodents, bovine, equines, porcines, ovines, and primates. Avians include, but are not limited to, fowls, songbirds, and raptors. Thus, the methods can be used in veterinary medicine, e.g., to treat companion / domestic animals, farm animals, laboratory animals in zoological parks, and animals in the wild, such as bats and rodents. The subject may also be an invertebrate, such as a tick, mosquito or sand fly. The methods are particularly desirable for human medical applications.

[0054] The term “patient” as used in this application means a human subject.

[0055] The term “detection”, “detect”, “detecting” and the like as used herein means as used herein means to discover the presence or existence of.

[0056] The terms “identification”, “identify”, “identifying” and the like as used herein means to recognize a specific bacterium or bacteria and / or gene or genes and / or nucleic acid or nucleic acids in a sample from a subject.

[0057] As used herein, the term “isolated” and the like means that the referenced material is free of components found in the natural environment in which the material is normally found. In particular, isolated biological material is free of cellular components. In the case of nucleic acid molecules, an isolated nucleic acid includes a PCR product, an isolated mRNA, a cDNA, an isolated genomic DNA, or a restriction fragment. In another embodiment, an isolated nucleic acid is preferably excised from the chromosome in which it may be found. Isolated nucleic acid molecules can be inserted into plasmids, cosmids, artificial chromosomes, and the like. Thus, in a specific embodiment, a recombinant nucleic acid is an isolated nucleic acid. An isolated protein may be associated with other proteins or nucleic acids, or both, with which it associates in the cell, or with cellular membranes if it is a membrane-associated protein. An isolated material may be, but need not be, purified.

[0058] As used herein, a “nucleic acid”, and “polynucleotide” and “nucleic acid sequence” and “nucleotide sequence” includes a nucleic acid, an oligonucleotide, a nucleotide, a polynucleotide, and any fragment, variant, or derivative thereof. The nucleic acid or polynucleotide may be double-stranded, single-stranded, or triple-stranded DNA or RNA (including cDNA), or a DNA- RNA hybrid of genetic or synthetic origin, wherein the nucleic acid contains any combination of deoxyribonucleotides and ribonucleotides and any combination of bases, including, but not limited to, adenine, thymine, cytosine, guanine, uracil, inosine, and xanthine hypoxanthine. As further used herein, the term “cDNA” refers to an isolated DNA polynucleotide or nucleic acid molecule, or any fragment, derivative, or complement thereof. It may be double-stranded, singlestranded, or triple-stranded, it may have originated recombinantly or synthetically, and it may represent coding and / or noncoding 5’ and / or 3’ sequences.

[0059] The term “fragment” when used in reference to a nucleotide sequence refers to portions of that nucleotide sequence. The fragments may range in size from 5 nucleotide residues to the entire nucleotide sequence minus one nucleic acid residue.

[0060] The term “genome” as used herein, refers to the entirety of an organism’s hereditary information that is encoded in its primary DNA or RNA or nucleotide sequence (DNA or RNA as applicable). The genome includes both the genes and the non-coding sequences. For example, the genome may represent a viral genome, a microbial genome or a mammalian genome.

[0061] A “coding sequence” or a sequence “encoding” an expression product, such as a RNA, polypeptide, protein, or enzyme, is a nucleotide sequence that, when expressed, results in the production of that RNA, polypeptide, protein, or enzyme, i.e., the nucleotide sequence encodes an amino acid sequence for that polypeptide, protein or enzyme. A coding sequence for a protein may include a start codon (usually ATG) and a stop codon.

[0062] As used herein, the terms “complementary” or “complementarity” are used in reference to “polynucleotides” and “oligonucleotides” (which are interchangeable terms that refer to a sequence of nucleotides) related by the base-pairing rules. It may also include mimics of or artificial bases that may not faithfully adhere to the base-pairing rules. For example, the sequence “C-A-G-T,” is complementary to the sequence “G-T-C-A ” In another example, a nucleotide sequence of 5’-CAGT-3’ is complementary to, and is capable of hybridizing to, a nucleotide sequence of 3’-GTCA-5’. Complementarity can be “partial” or “total.” “Partial” complementarity is where one or more nucleic acid bases are not matched according to the base pairing rules. “Total” or “complete” complementarity between nucleic acids is where each and every nucleic acid base is matched with another base under the base pairing rules. The degree of complementarity between nucleic acid strands has significant effects on the efficiency and strength of hybridization between nucleic acid strands. This is of particular importance in amplification reactions, as well as detection methods which depend upon binding between nucleic acids.

[0063] The term “nucleic acid hybridization” or “hybridization” refers to anti-parallel hydrogen bonding between two single-stranded nucleic acids, in which A pairs with T (or U if an RNA nucleic acid) and C pairs with G. Nucleic acid molecules are “hybridizable” to each other when at least one strand of one nucleic acid molecule can form hydrogen bonds with the complementary bases of another nucleic acid molecule under defined stringency conditions. Stringency of hybridization is determined, e.g., by (i) the temperature at which hybridization and / or washing is performed, and (ii) the ionic strength and (iii) concentration of denaturants such as formamide of the hybridization and washing solutions, as well as other parameters. Hybridization requires that the two strands contain substantially complementary sequences. Depending on the stringency of hybridization, however, some degree of mismatches may be tolerated. Under “low stringency” conditions, a greater percentage of mismatches are tolerable (i.e., will not prevent formation of an anti-parallel hybrid).

[0064] As used herein the term “hybridization product” refers to a complex formed between two nucleic acid sequences by virtue of the formation of hydrogen bounds between complementary G and C bases and between complementary A and T bases; these hydrogen bonds may be further stabilized by base stacking interactions. The two complementary nucleic acid sequences hydrogen bond in an antiparallel configuration. A hybridization product may be formed in solution or between one nucleic acid sequence present in solution and another nucleic acid sequence immobilized to a solid support. As used herein the term “stringency” is used in reference to the conditions of temperature, ionic strength, and the presence of other compounds such as organic solvents, under which nucleic acid hybridizations are conducted. “Stringency” typically occurs in a range from about Tmto about 20°C to 25°C below Tm. A “stringent hybridization” can be used to identify or detect identical polynucleotide sequences or to identify or detect similar or related polynucleotide sequences. For example, when fragments are employed in hybridization reactions under stringent conditions the hybridization of fragments which contain unique sequences (i.e., regions which are either non-homologous to or which contain less than about 50% homology or complementarity) are favored. Alternatively, when conditions of “weak” or “low” stringency are used hybridization may occur with nucleic acids that are derived from organisms that are genetically diverse (i.e., for example, the frequency of complementary sequences is usually low between such organisms).

[0065] The terms “percent (%) sequence similarity”, “percent (%) sequence identity”, and the like, generally refer to the degree of identity or correspondence between different nucleotide sequences of nucleic acid molecules or amino acid sequences of proteins that may or may not share a common evolutionary origin. Sequence identity can be determined using any of a number of publicly available sequence comparison algorithms, such as BLAST, FASTA, DNA Strider, and GCG (Genetics Computer Group, Program Manual for the GCG Package, Version 7, Madison, Wisconsin).

[0066] To determine the percent identity between two amino acid sequences or two nucleic acid molecules, the sequences are aligned for optimal comparison purposes. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences i.e., percent identity = number of identical positions / total number of positions (e.g., overlapping positions) x 100). In one embodiment, the two sequences are, or are about, of the same length. The percent identity between two sequences can be determined using techniques similar to those described below, with or without allowing gaps. In calculating percent sequence identity, typically exact matches are counted.

[0067] “Amplification” is defined as the production of additional copies of a nucleic acid sequence and is generally carried out either in vivo, or in vitro, i.e. for example using polymerase chain reaction. As used herein, the term “polymerase chain reaction” (“PCR”) refers to the method disclosed in U.S. Patent Nos. 4,683,195 and 4,683,202, herein incorporated by reference, which describe a method for increasing the concentration of a segment of a target sequence in a mixture of genomic DNA without cloning or purification. The length of the amplified segment of the desired target sequence is determined by the relative positions of two oligonucleotide primers with respect to each other, and therefore, this length is a controllable parameter. By virtue of the repeating aspect of the process, the method is referred to as the “polymerase chain reaction” (hereinafter “PCR”). Because the desired amplified segments of the target sequence become the predominant sequences (in terms of concentration) in the mixture, they are said to be “PCR amplified”. With PCR, it is possible to amplify a single copy of a specific target sequence in genomic DNA to a level detectable by several different methodologies (e. ., hybridization with a labeled probe; incorporation of biotinylated primers followed by avidin-enzyme conjugate detection; incorporation of 32P -labeled deoxynucleotide triphosphates, such as dCTP or dATP, into the amplified segment). In addition to genomic DNA, any oligonucleotide sequence can be amplified with the appropriate set of primer molecules. In particular, the amplified segments created by the PCR process itself are, themselves, efficient templates for subsequent PCR amplifications. With PCR, it is also possible to amplify a complex mixture (library) of linear DNA molecules, provided they carry suitable universal sequences on either end such that universal PCR primers bind outside of the DNA molecules that are to be amplified.

[0068] The terms “next-generation sequencing platform” and “high-throughput sequencing” and “HTS” as used herein, refer to any nucleic acid sequencing device that utilizes massively parallel technology. For example, such a platform may include, but is not limited to, Illumina sequencing platforms.

[0069] The term “sequencing library”, as used herein refers to a library of nucleic acids that are compatible with next-generation high throughput sequencers.

[0070] The term “bacterially-derived sequence” as used herein refers to a sequence which is typically associated with bacteria. For example, the sequence may be a sequence present in a bacterial genome, or a sequence from a plasmid, virus, or bacteriophage known to be harbored by one or more bacterial species.

[0071] The term “hybridization portion” as used herein in the context of an oligonucleotide probe of a bacterial sequence capture platform refers to a portion of a oligonucleotide probe that is partially or fully complementary to a bacterially-derived sequence. For example, the hybridization portion of an oligonucleotide probe may hybridize to a target bacterially-derived sequence on a tested nucleotide molecule when the oligonucleotide probe is exposed to a sample containing the tested nucleotide molecule.

[0072] The term "pathogenicity element sequence” is a nucleotide sequence associated with increasing the pathogenicity (i.e., the capacity to cause disease) of an organism.

[0073] The term “virulence factor sequence” refers to a nucleotide sequence which encodes a product that enables a microorganism to establish itself on or within a host of a particular species and enhance its potential to cause disease. For example, virulence factors include, but are not limited to, bacterial toxins, cell surface proteins that mediate bacterial attachment, cell surface carbohydrates, proteins that protect a bacterium, and hydrolytic enzymes that may contribute to bacterial pathogenicity.

[0074] The term “environmental sample” as used herein refers to a sample obtained from any non-biological media or material(s), including but not limited to, air, soil, water, and swabs of inanimate surfaces. Environmental samples contrast with biological samples, which typically derive from an organism. Examples of biological samples include, but are not limited to, bodily fluids, cells, tissue samples, and swabs of a surface or cavity of a biological organism.

[0075] The following embodiments and examples (including details thereof) are set forth to aid in an understanding of the subject matter of this disclosure but are not intended to, and should not be construed to, limit in any way the invention that is claimed.

[0076] Database of Probe Sequences and Set of Probes

[0077] Described herein is a database of probe sequences and a set of probes that enable the detection, identification and / or differentiation of bacteria, as well as pathogenicity elements, and / or antimicrobial resistance (AMR) genes and / or 16S ribosomal RNA. These sequences or probes have many uses including but not limited to use in a sequence capture platform and other diagnostic assays. The sequences or probes described herein increase the sensitivity of high- throughput sequencing for detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA.

[0078] The database of probe sequences or set of probes is comprised of oligonucleotides that are distributed across informative regions of bacteria. For example, the database of probe sequences or set of probes may comprise about one million or fewer oligonucleotides. To enable efficient detection, identification, and / or differentiation of bacteria, and / or virulence elements and / or antimicrobial resistance and / or 16S ribosomal RNA, the database of probe sequences and set of probes can be designed to target four major components: 1. Sequence-specific or cladespecific marker genes sequences extracted, for example one or more such sequences from the Metaphlan4 database; or 2. 16S ribosomal RNA sequences, for example one or more such sequences extracted from SILVA database for a total of 1333 bacterial species (see Table 1); or 3. Virulence factors genetic sequences in bacterial pathogens, for example one or more such sequences extracted from the VFDB (Virulence Factor Database); or 4. Antibiotic resistance determinants genes, for example one or more such sequences extracted from CARD (The Comprehensive Antibiotic Resistance Database) or any combination of the four. In one embodiment, oligonucleotide probes were designed to bind to regions distributed across the combined target sequence dataset (101,185 genetic fragments = 90,776 for 894 species from Metaphlan4 + 1325 species from SILVA 16S + 4750 AMR + 4334 VFDB) (Table 2). The generated probes were further clustered for sequence identity, which resulted in 988,786 probes.

[0079] The database of probe sequences and set of probes disclosed and described herein are more targeted than prior known databases and sets of probes and can identify the bacteria in any given sample by targeting species-specific or clade-specific marker sequences in bacterial genomes, rather than the entire genome of bacteria.

[0080] Other differences from prior known databases and probe sets are a longer uniform probe size and smaller number of probes (e.g., one million or less). There is also no adjustment of length for Tm of the probes. Additionally, the probe set may include 16S ribosomal RNA sequences and / or AMR genes and / or virulence factor genes. After all of the sequences were obtained, they were clustered for sequence identity to reduce or eliminate redundancy. This resulted in a database of probe sequences and set of probes that was less redundant than previous sets. Additionally, over 1,300 different bacteria can be identified using the disclosed database of probe sequences or set of probes (Table 1). The disclosed database of probe sequences or set of probes also leads to more straightforward analysis. For example, the platform of oligonucleotide probes described herein enables detection of bacterially-derived sequences in environmental samples, for example, to determine the prevalence of medically relevant bacteria, pathogenesis elements, virulence factors, and / or AMR sequences in a sample. The disclosed platform or probe set enables a faster, more cost-effective approach to detecting medically relevant bacterially-derived sequences in environmental or clinical samples without sacrificing coverage or accuracy.

[0081] The current disclosure includes a method of designing and / or making or constructing a database of probe sequences or set of probes and methods of using the set of probes to construct sequencing libraries suitable for sequencing in any high throughput sequencing technology. The disclosure also includes methods and systems for detecting, identifying and / or differentiating bacteria and / or pathogenic elements and / or AMR genes and / or 16S ribosomal RNA in a single sample, of any origin, using the database of probe sequences or set of probes. The database of probe sequences or set of probes enables detection of bacterial sequences in any complex sample background, including those found in clinical specimens and the presence of features associated with pathogenicity and / or antimicrobial resistance.

[0082] The present disclosure includes a method of designing and / or constructing a database of probe sequences or set of probes for the detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA. Accordingly, the method may include the following steps.

[0083] The first step is to obtain sequence information including species-specific or cladespecific marker gene of bacteria, or 16S ribosomal RNA sequences, or AMR genes, or virulence factors, or a combination of any of the four.

[0084] Sequence information is obtained from any public or private database of sequence information of bacteria, 16S ribosomal RNA sequences, AMR genes and / or virulence factors, including, but not limited, to Metaphlan4, SILVA, CARD and VFDB. Any version of these databases, including but not limited to those exemplified in Table 2, as well as future updates, may be used.

[0085] The next step of the method is to break the sequences into fragments to be the basis of the oligonucleotide probes. In the current embodiment, the probes are spaced at a distance across the target sequences, such that the total number of probe sequences in the database or probes in the probe set is about one million or less and cover all target sequences.

[0086] In some embodiments, the probe length is about 5 nt to about 300 nt. In some embodiments, the probe length is about 10 nt to about 280 nt. In some embodiments, the probe length is about 20 nt to about 260 nt. In some embodiments, the probe length is about 30 nt to about 240 nt. In some embodiments, the probe length is about 40 nt to about 220 nt. In some embodiments, the probe length is about 50 nt to about 200 nt. In some embodiments, the probe length is about 60 nt to about 190 nt. In some embodiments, the probe length is about 70 nt to about 180 nt. In some embodiments, the probe length is about 80 nt to about 170 nt. In some embodiments, the probe length is about 90 nt to about 160 nt. In some embodiments, the probe length is about 100 nt to about 150 nt. In some embodiments, the probe length is about 110 nt to about 140 nt. In some embodiments, the probe length is about 115 nt to about 130 nt. In some embodiments, the probe length is about 120 nt.

[0087] In some embodiments, the inter-probe spacing is about 20 nt to about 100 nt tiled across the target sequences. In some embodiments, the inter-probe spacing is about 30 nt to about 90 nt tiled across the target sequences. In some embodiments, the inter-probe spacing is about 40 nt to about 80 nt tiled across the target sequences. In some embodiments, the inter-probe spacing is about 50 nt to about 70 nt tiled across the target sequences. In some embodiments, the interprobe spacing is about 60 nt tiled across the target sequences.

[0088] The generated probes can be further clustered for sequence identity to obtain a certain number of probe sequences or probes. In some embodiments, the generated probes are clustered at about 90% to about 99% sequence identity. In some embodiments, the generated probes are clustered at about 92% to about 98% sequence identity. In some embodiments, the generated probes are clustered at about 94% to about 97% sequence identity. In some embodiments, the generated probes are clustered at about 95% to about 97% sequence identity.

[0089] In some embodiments, the generated probes are clustered at about 96% sequence identity which resulted in less than one million (988,786) probes.

[0090] Specifically, oligonucleotides are selected to bind to regions distributed across the combined target sequence dataset, which in the current embodiment was 101,185 genetic targets, corresponding to 90,776 genes in 894 species from Metaphlan4, 1325 rRNA sequences from SILVA 16S, 4750 AMR genes from CARD, and 4334 virulence factor sequences from VFDB.

[0091] Any bacterially-derived sequences desired to be targeted, preferably sequences which are relevant to pathogenesis and / or virulence or are otherwise medically relevant, may be used to generate oligonucleotides probes for use in any one of the probe sets or bacterial sequence capture platforms described herein, or used to generate a database of probe sequences. For example, sequence information of desired targets may be obtained from any public or private database of sequence information of bacteria and / or 16S ribosomal RNA and / or AMR genes and / or virulence factors, including, but not limited to, Metaphlan4, SILVA, CARD, and VFDB. For example, versions of each of these databases are provided in Table 2, however, additional versions, releases, and updates to these or other databases may be used.

[0092] Metaphlan4 (Metagenomic Phylogenetic Analysis 4) is a computational tool for specieslevel microbial profiling. See huttenhower.sph.harvard.edu / metaphlan and Aitor Blanco-Miguez et al. (2022) “Extending and improving metagenomic taxonomic profiling with uncharacterized species with MetaPhlAn 4”, bioRxiv preprint doi.org / 10.1101 / 2022.08.22.504593, the contents of both of which are incorporated herein by reference.

[0093] SILVA is a high-quality ribosomal RNA database. Release information of the SILVA SSU and LSU databases 138.1 as of August 27, 2020 is available at www.arb- silva.de / documentation / release-1381 / , the content of which is incorporated herein by reference.

[0094] CARD (The Comprehensive Antibiotic Resistance Database) is a bioinformatic database of resistance genes, their products and associated phenotypes. See card.mcmaster.ca / home and Alcock BP et al. “CARD 2023: expanded curation, support for machine learning, and resistome prediction at the Comprehensive Antibiotic Resistance Database.” Nucleic Acids Res. 2023 Jan 6;51(Dl):D690-D699, the contents of both of which are incorporated herein by reference.

[0095] VFDB (Virulence Factor Database) is an integrated and comprehensive online resource for curating information about virulence factors of bacterial pathogens. See mgc.ac.cn / VFs / main.htm and Liu B et al. “VFDB 2022: a general classification scheme for bacterial virulence factors.” Nucleic Acids Res. 2022 Jan 7; 5O(D1): D912-D917, the contents of both of which are incorporated herein by reference.

[0096] Table 1: Medically Important Bacterial Species

[0097] Abiotrophia defectiva Leptospira alexanderi

[0098] Acetobacter nitrogenifigens Leptospira alstonii

[0099] Achromobacter denitrificans Leptospira biflexa

[0100] Achromobacter insolitus Leptospira borgpetersenii

[0101] Achromobacter piechaudii Leptospira broomii

[0102] Achromobacter ruhlandii Leptospira fainei

[0103] Achromobacter xylosoxidans Leptospira inadai

[0104] Acidaminococcus fermentans Leptospira interrogans Acidaminococcus intestini Leptospira kirschneri

[0105] Acidovorax citrulli Leptospira kmetyi

[0106] Acinetobacter baumannii Leptospira licerasiae

[0107] Acinetobacter bereziniae Leptospira mayottensis

[0108] Acinetobacter calcoaceticus Leptospira meyeri

[0109] Acinetobacter haemolyticus Leptospira noguchii

[0110] Acinetobacter j ohnsonii Leptospira santarosai

[0111] Acinetobacter j unii Leptospira terpstrae

[0112] Acinetobacter Iwoffii Leptospira vanthielii

[0113] Acinetobacter parvus Leptospira weilii

[0114] Acinetobacter pittii Leptospira wolbachii

[0115] Acinetobacter radioresistens Leptospira yanagawae

[0116] Acinetobacter schindleri Leptotrichia buccalis

[0117] Acinetobacter seifertii Leptotrichia goodfell owii

[0118] Acinetobacter soli Leptotrichia shahii

[0119] Acinetobacter ursingii Leptotrichia trevisanii

[0120] Actinobacillus hominis Leptotrichia wadei

[0121] Actinobacillus suis Leuconostoc carnosum

[0122] Actinobacillus ureae Leuconostoc citreum

[0123] Actinobaculum massiliense Leuconostoc lactis

[0124] Actinomadura madurae Leuconostoc mesenteroides

[0125] Actinomadura pelletieri Leuconostoc pseudomesenteroides

[0126] Actinomyces Cardiff ensis Levilactobacillus brevis

[0127] Actinomyces georgiae Ligilactobacillus salivarius

[0128] Actinomyces gerencseriae Limosilactobacillus fermentum

[0129] Actinomyces graevenitzii Listeria grayi

[0130] Actinomyces hongkongensis Listeria innocua

[0131] Actinomyces israelii Listeria ivanovii

[0132] Actinomyces massiliensis Listeria monocytogenes

[0133] Actinomyces meyeri Listeria seeligeri

[0134] Actinomyces naeslundii Listeria welshimeri Actinomyces neuii Luteococcus peritonei

[0135] Actinomyces neuii anitratus Luteococcus sanguinis

[0136] Actinomyces neuii neuii Lysinibacillus sphaericus

[0137] Actinomyces oris Mannheimia haemolytica

[0138] Actinomyces radicidentis Massilia timonae

[0139] Actinomyces radingae Megasphaera elsdenii

[0140] Actinomyces timonensis Megasphaera micronuciformis

[0141] Actinomyces turicensis Methylobacterium mesophilicum

[0142] Actinomyces urogenitalis Microbacterium

[0143] Actinomyces viscosus Microbacterium arborescens

[0144] Advenella incenata Microbacterium foliorum

[0145] Aerococcus christensenii Microbacterium maritypicum

[0146] Aerococcus sanguinicola Microbacterium oxydans

[0147] Aerococcus urinae Microbacterium paraoxydans

[0148] Aerococcus urinaeequi Microbacterium resistens

[0149] Aerococcus urinaehominis Microbacterium testaceum

[0150] Aerococcus viridans Micrococcus luteus

[0151] Aeromonas bestiarum Micrococcus luteus ATCC 49442

[0152] Aeromonas caviae Micrococcus lylae

[0153] Aeromonas enteropelogenes Mitsuokella multacida

[0154] Aeromonas hydrophila Mobiluncus curtisii

[0155] Aeromonas salmonicida Mobiluncus curtisii curtisii

[0156] Aeromonas schubertii Mobiluncus curtisii holmesii

[0157] Aeromonas veronii Mobiluncus mulieris

[0158] Afipia birgiae Moellerella wisconsensis

[0159] Afipia broomeae Mogibacterium diversum

[0160] Afipia clevelandensis Mogibacterium neglectum

[0161] Afipia felis Mogibacterium timidum

[0162] Aggregatibacter actinomycetemcomitans Moraxella atlantae

[0163] Aggregatibacter aphrophilus Moraxella catarrhalis

[0164] Aggregatibacter segnis Moraxella lacunata Agrobacterium tumefaciens Moraxella lincolnii

[0165] Alcaligenes faecalis Moraxella nonliquefaciens

[0166] Alistipes finegoldii Moraxella osloensis

[0167] Alistipes onderdonkii Morganella morganii

[0168] Alistipes putredinis Morganella morganii morganii

[0169] Alistipes shahii Morganella morganii sibonii

[0170] Alloiococcus otitis Morococcus cerebrosus

[0171] Alloprevotella tannerae Moryella indoligenes

[0172] Alloscardovia omnicolens Mycobacterium abscessus

[0173] Alysiella crassa Mycobacterium africanum

[0174] Amycolatopsis palatopharyngis Mycobacterium alvei

[0175] Anaerobiospirillum succiniciproducens Mycobacterium arupense

[0176] Anaerococcus hydrogenalis Mycobacterium asiaticum

[0177] Anaerococcus lactolyticus Mycobacterium aurum

[0178] Anaerococcus octavius Mycobacterium avium

[0179] Anaerococcus prevotii Mycobacterium barrassiae

[0180] Anaerococcus tetradius Mycobacterium bohemicum

[0181] Anaerococcus vaginalis Mycobacterium bolletii

[0182] Anaeroglobus geminatus Mycobacterium bovis

[0183] Anaerostipes caccae Mycobacterium branded

[0184] Anaplasma phagocytophilum Mycobacterium brisbanense

[0185] Arcanobacterium haemolyticum Mycobacterium canariasense

[0186] Arcobacter butzleri Mycobacterium celatum

[0187] Arcobacter cryaerophilus Mycobacterium chelonae

[0188] Arcobacter skirrowii Mycobacterium chimaera

[0189] Arthrobacter oxydans Mycobacterium chubuense

[0190] Arthrobacter scleromae Mycobacterium colombiense

[0191] Arthrobacter woluwensis Mycobacterium conceptionense

[0192] Atopobium parvulum Mycobacterium conspicuum

[0193] Atopobium rimae Mycobacterium cosmeticum

[0194] Atopobium vaginae Mycobacterium diemhoferi Aureimonas altamirensis Mycobacterium doricum

[0195] Bacillus anthracis Mycobacterium elephantis

[0196] Bacillus cereus Mycobacterium flavescens

[0197] Bacillus circulans Mycobacterium florentinum

[0198] Bacillus coagulans Mycobacterium fortuitum

[0199] Bacillus glycinifermentans Mycobacterium franklinii

[0200] Bacillus licheniformis Mycobacterium gastri

[0201] Bacillus megaterium Mycobacterium genavense

[0202] Bacillus mycoides Mycobacterium goodii

[0203] Bacillus paralicheniformis Mycobacterium gordonae

[0204] Bacillus paucivorans Mycobacterium grossiae

[0205] Bacillus pumilus Mycobacterium haemophilum

[0206] Bacillus safensis Mycobacterium hassiacum

[0207] Bacillus sphaericus Mycobacterium heckeshomense

[0208] Bacillus subtilis Mycobacterium heidelbergense

[0209] Bacillus thuringiensis Mycobacterium heraklionense

[0210] Bacteroides caccae Mycobacterium hodleri

[0211] Bacteroides distasonis Mycobacterium holsaticum

[0212] Bacteroides eggerthii Mycobacterium houstonense

[0213] Bacteroides faecis Mycobacterium immunogenum

[0214] Bacteroides finegoldii Mycobacterium interj ectum

[0215] Bacteroides fragilis Mycobacterium intermedium

[0216] Bacteroides massiliensis Mycobacterium intracellulare

[0217] Bacteroides merdae Mycobacterium iranicum

[0218] Bacteroides nordii Mycobacterium kansasii

[0219] Bacteroides ovatus Mycobacterium koreense

[0220] Bacteroides pyogenes Mycobacterium kumamotonense

[0221] Bacteroides stercoris Mycobacterium kyorinense

[0222] Bacteroides thetaiotaomicron Mycobacterium lentiflavum

[0223] Bacteroides uniformis Mycobacterium leprae

[0224] Bacteroides vulgatus Mycobacterium lepromatosis Balneatrix alpica Mycobacterium llatzerense

[0225] Bartonella alsatica Mycobacterium mageritense

[0226] Bartonella ancashensis Mycobacterium malmoense

[0227] Bartonella bacilliformis Mycobacterium marinum

[0228] Bartonella birtlesii Mycobacterium massiliense

[0229] Bartonella bovis Mycobacterium microti

[0230] Bartonella clarridgeiae Mycobacterium monacense

[0231] Bartonella doshiae Mycobacterium mucogenicum

[0232] Bartonella elizabethae Mycobacterium nebraskense

[0233] Bartonella grahamii Mycobacterium neoaurum

[0234] Bartonella henselae Mycobacterium nonchromogenicum

[0235] Bartonella koehlerae Mycobacterium novocastrense

[0236] Bartonella quintana Mycobacterium obuense

[0237] Bartonella rattaustraliani Mycobacterium palustre

[0238] Bartonella rochalimae Mycobacterium paraffinicum

[0239] Bartonella schoenbuchensis Mycobacterium parascrofulaceum

[0240] Bartonella taylorii Mycobacterium peregrinum

[0241] Bartonella tribocorum Mycobacterium phlei

[0242] Bartonella vinsonii Mycobacterium phocaicum

[0243] Bartonella vinsonii subsp. Vinsonii_,f Mycobacterium porcinum

[0244] Bergey ella zoohelcum Mycobacterium saopaulense

[0245] Bifidobacterium adolescentis Mycobacterium scrofulaceum

[0246] Bifidobacterium angulatum Mycobacterium septicum

[0247] Bifidobacterium animalis Mycobacterium setense

[0248] Bifidobacterium bifidum Mycobacterium sherrisii

[0249] Bifidobacterium breve Mycobacterium shigaense

[0250] Bifidobacterium dentium Mycobacterium shimoidei

[0251] Bifidobacterium infantis Mycobacterium simiae

[0252] Bifidobacterium longum Mycobacterium smegmatis

[0253] B ifi dob acterium p seudocatenul atum Mycobacterium szulgai

[0254] Bifidobacterium psychraerophilum Mycobacterium talmoniae Bifidobacterium scardovii Mycobacterium terrae

[0255] Bilophila wadsworthia Mycobacterium thermoresistibile

[0256] Bordetella avium Mycobacterium triplex

[0257] Bordetella bronchialis Mycobacterium triviale

[0258] Bordetella bronchiseptica Mycobacterium tuberculosis

[0259] Bordetella flabilis Mycobacterium tusciae

[0260] Bordetella hinzii Mycobacterium ulcerans

[0261] Bordetella holmesii Mycobacterium wolinskyi

[0262] Bordetella parapertussis Mycobacterium xenopi

[0263] Bordetella pertussis Mycolicibacterium aurum

[0264] Bordetella petrii My coli cib acterium chi orophenoli cum

[0265] Bordetella trematum Mycolicibacterium hassiacum

[0266] Borrelia afzelii Mycolicibacterium vaccae

[0267] Borrelia crocidurae Mycolicibacterium wolinskyi

[0268] Borrelia duttonii Mycoplasma amphoriforme

[0269] Borrelia garinii Mycoplasma capricolum

[0270] Borrelia hermsii Mycoplasma faucium

[0271] Borrelia hispanica Mycoplasma fermentans

[0272] Borrelia mayonii Mycoplasma genitalium

[0273] Borrelia miyamotoi Mycoplasma hominis

[0274] Borrelia parkeri Mycoplasma hyopneumoniae

[0275] Borrelia persica Mycoplasma orale

[0276] Borrelia recurrentis Mycoplasma penetrans

[0277] Borrelia sinica Mycoplasma pirum

[0278] Borrelia spielmanii Mycoplasma pneumoniae

[0279] Borrelia turicatae Mycoplasma primatum

[0280] Borrelia valaisiana Mycoplasma salivarium

[0281] Borreliella burgdorferi Mycoplasma spermatophilum

[0282] Bosea massiliensis Mycoplasmopsis arginini

[0283] Brachyspira aalborgi Mycoplasmopsis cynos

[0284] Brachyspira pilosicoli Mycoplasmopsis fermentans Brevibacillus brevis My coplasmopsis pulmonis

[0285] Brevibacillus centrosporus Myroides marinus

[0286] Brevibacillus laterosporus Myroides odoratimimus

[0287] Brevibacillus parabrevis Myroides odoratus

[0288] Brevibacterium casei Neisseria animaloris

[0289] Brevundimonas diminuta Neisseria bacilliformis

[0290] Brevundimonas vesicularis Neisseria canis

[0291] Brucella abortus Neisseria cinerea

[0292] Brucella canis Neisseria elongata

[0293] Brucella inopinata Neisseria elongata nitroreductens

[0294] Brucella melitensis Neisseria flavescens

[0295] Brucella suis Neisseria gonorrhoeae

[0296] Budvicia aquatica Neisseria lactamica

[0297] Bulleidia extructa Neisseria meningitidis

[0298] Burkholderia ambifaria Neisseria mucosa

[0299] Burkholderia anthina Neisseria polysaccharea

[0300] Burkholderia cenocepacia Neisseria sicca

[0301] Burkholderia cepacia Neisseria subflava

[0302] Burkholderia dolosa Neisseria wadsworthii

[0303] Burkholderia fungorum Neisseria weaveri

[0304] Burkholderia gladioli Neisseria zoodegmatis

[0305] Burkholderia glumae Neorickettsia helminthoeca

[0306] Burkholderia mallei Neorickettsia sennetsu

[0307] Burkholderia multivorans Nocardia abscessus

[0308] Burkholderia oklahomensis Nocardia acidivorans

[0309] Burkholderia pseudomallei Nocardia africana

[0310] Burkholderia pyrrocinia Nocardia alba

[0311] Burkholderia stabilis Nocardia amamiensis

[0312] Burkholderia thailandensis Nocardia anaemiae

[0313] Burkholderia vietnamiensis Nocardia aobensis

[0314] Burkholderiales bacterium Nocardia araoensis Burkholderiales bacterium 8X Nocardia arizonensis

[0315] Burkholderiales bacterium C2 Nocardia arthritidis

[0316] Burkholderiales bacterium GJ E10 Nocardia asiatica

[0317] Burkholderiales bacterium JOSHI 001 Nocardia asteroides

[0318] Burkholderiales bacterium LSUCC0115 Nocardia beijingensis

[0319] Buttiauxella agrestis Nocardia brasiliensis

[0320] Buttiauxella brennerae Nocardia brevicatena

[0321] Buttiauxella ferragutiae Nocardia caishijiensis

[0322] Buttiauxella gaviniae Nocardia carnea

[0323] Butyrivibrio fibrisolvens Nocardia cerradoensis

[0324] Campylobacter coli Nocardia concava

[0325] Campylobacter concisus Nocardia coubleae

[0326] Campylobacter corcagiensis Nocardia crassostreae

[0327] Campylobacter cuniculorum Nocardia cummidelens

[0328] Campylobacter curvus Nocardia cyriacigeorgica

[0329] Campylobacter fetus Nocardia elegans

[0330] Campylobacter gracilis Nocardia exalbida

[0331] Campylobacter hominis Nocardia farcinica

[0332] Campylobacter hyointestinalis Nocardia flavorosea

[0333] Campylobacter iguaniorum Nocardia fusca

[0334] Campylobacter j ejuni Nocardia gamkensis

[0335] Campylobacter jejuni doylei Nocardia grenadensis

[0336] Campylobacter jejuni jejuni Nocardia harenae

[0337] Campylobacter lari Nocardia higoensis

[0338] Campylobacter mucosalis Nocardia ignorata

[0339] Campylobacter rectus Nocardia inohanensis

[0340] Campylobacter showae Nocardia j ej uensi s

[0341] Campylobacter sputorum Nocardia jiangxiensis

[0342] Campylobacter upsaliensis Nocardia kruczakiae

[0343] Campylobacter ureolyticus Nocardia lijiangensis

[0344] Candidatus Bartonella Nocardia mexicana Capnocytophaga canimorsus Nocardia mikamii

[0345] Capnocytophaga cynodegmi Nocardia miyunensis

[0346] Capnocytophaga gingivalis Nocardia niigatensis

[0347] Capnocytophaga granulosa Nocardia ninae

[0348] Capnocytophaga ochracea Nocardia niwae

[0349] Capnocytophaga sputigena Nocardia nova

[0350] Cardiobacterium hominis Nocardia otitidiscaviarum

[0351] Cardiobacterium valvarum Nocardia paucivorans

[0352] Catabacter hongkongensis Nocardia pneumoniae

[0353] Catonella morbi Nocardia pseudobrasiliensis

[0354] Cedecea davisae Nocardia pseudovaccinii

[0355] Cedecea lapagei Nocardia puris

[0356] Cedecea neteri Nocardia rhamnosiphila

[0357] Cellulomonas flavigena Nocardia salmonicida

[0358] Cellulomonas hominis Nocardia seriolae

[0359] Cellulosimicrobium cellulans Nocardia shimofusensis

[0360] Cellulosimicrobium funkei Nocardia sienata

[0361] Centipeda periodontii Nocardia soli

[0362] Chlamydia pneumonia Nocardia speluncae

[0363] Chlamydia pneumoniae Nocardia takedensis

[0364] Chlamydia psittaci Nocardia tenerifensis

[0365] Chlamydia trachomatis Nocardia terpenica

[0366] Chromobacterium haemolyticum Nocardia testacea

[0367] Chromobacterium violaceum Nocardia thailandica

[0368] Chryseobacterium Nocardia transvalensis

[0369] Chryseobacterium gleum Nocardia uniformis

[0370] Chryseobacterium indologenes Nocardia vaccinii

[0371] Citrobacter amalonaticus Nocardia vermiculata

[0372] Citrobacter braakii Nocardia veterana

[0373] Citrobacter farmeri Nocardia vinacea

[0374] Citrobacter freundii Nocardia vulneris Citrobacter koseri Nocardia xishanensis

[0375] Citrobacter murliniae Nocardia yamanashiensis

[0376] Citrobacter rodentium Nocardiopsis dassonvillei

[0377] Citrobacter sedlakii Ochrobactrum anthropi

[0378] Citrobacter werkmanii Ochrobactrum intermedium

[0379] Citrobacter youngae Ochrobactrum oryzae

[0380] Clostridium argentinense Odoribacter laneus

[0381] Clostridium baratii Odoribacter splanchnicus

[0382] Clostridium beijerinckii Oerskovia turbata

[0383] Clostridium bifermentans Oligella ureolytica

[0384] Clostridium bolteae Oligella urethralis

[0385] Clostridium botulinum Olsenella uli

[0386] Clostridium butyricum Oribacterium sinus

[0387] Clostridium cadaveris Orientia tsutsugamushi

[0388] Clostridium camis Oscillibacter ruminantium

[0389] Clostridium celatum Paenalcaligenes hominis

[0390] Clostridium cochlearium Paenibacillus alvei

[0391] Clostridium cocleatum Paenibacillus macerans

[0392] Clostridium difficile Paenibacillus mucilaginosus

[0393] Clostridium fallax Paenibacillus polymyxa

[0394] Clostridium ghonii Paenibacillus popilliae

[0395] Clostridium haemolyticum Paeniclostridium sordellii

[0396] Clostridium hylemonae Pandoraea apista

[0397] Clostridium indolis Pandoraea pulmonicola

[0398] Clostridium innocuum Pandoraea sputorum

[0399] Clostridium leptum Pannonibacter phragmitetus

[0400] Clostridium neonatale Pantoea agglomerans

[0401] Clostridium novyi Pantoea ananatis

[0402] Clostridium paraputrificum Pantoea dispersa

[0403] Clostridium perfringens Parabacteroides distasonis

[0404] Clostridium piliforme Parabacteroides faecis Clostridium ramosum Parabacteroides goldsteinii

[0405] Clostridium septicum Parabacteroides gordonii

[0406] Clostridium sordellii Parabacteroides j ohnsonii

[0407] Clostridium sphenoides Parabacteroides massiliensis

[0408] Clostridium spiroforme Parabacteroides merdae

[0409] Clostridium sporogenes Paraburkholderia fungorum

[0410] Clostridium subterminale Parachlamydia acanthamoebae

[0411] Clostridium symbiosum Paraclostridium bifermentans

[0412] Clostridium tertium Paracoccus sanguinis

[0413] Clostridium tetani Paracoccus yeei

[0414] Collinsella aerofaciens Paraeggerthella hongkongensis

[0415] Comamonas kerstersii Parascardovia denti colens

[0416] Comamonas terrigena Parvimonas micra

[0417] Comamonas testosteroni Pasteurella aerogenes

[0418] Corynebacterium accolens Pasteurella bettyae

[0419] Corynebacterium afermentans Pasteurella canis

[0420] Corynebacterium amycolatum Pasteurella dagmatis

[0421] Corynebacterium argentoratense Pasteurella gallinarum

[0422] Corynebacterium aurimucosum Pasteurella haemolytica

[0423] Corynebacterium auris Pasteurella multocida

[0424] Corynebacterium bovis Pasteurella multocida multocida

[0425] Corynebacterium confusum Pasteurella multocida septica

[0426] Corynebacterium coyleae Pediococcus acidilactici

[0427] Corynebacterium diphtheriae Pediococcus pentosaceus

[0428] Corynebacterium durum Pelobacter propionicus

[0429] Corynebacterium falsenii Peptococcus niger

[0430] Corynebacterium freiburgense Peptoniphilus asaccharolyticus

[0431] Corynebacterium freneyi Peptoniphilus coxii

[0432] Corynebacterium glucuronolyticum Peptoniphilus duerdenii

[0433] Corynebacterium halotolerans Peptoniphilus harei

[0434] Corynebacterium imitans Peptoniphilus indolicus Corynebacterium jeikeium Peptoniphilus lacrimalis

[0435] Corynebacterium kroppenstedtii Peptostreptococcus anaerobius

[0436] Corynebacterium kutscheri Peptostreptococcus canis

[0437] Corynebacterium lipophiloflavum Peptostreptococcus stomatis

[0438] Corynebacterium macginleyi Photobacterium damselae

[0439] Corynebacterium massiliense Photorhabdus asymbiotica

[0440] Corynebacterium matruchotii Photorhabdus luminescens

[0441] Corynebacterium minutissimum Plesiomonas shigelloides

[0442] Corynebacterium mucifaciens Pluralibacter gergoviae

[0443] Corynebacterium mycetoides Porphyromonas asaccharolytica

[0444] Corynebacterium pilosum Porphyromonas catoniae

[0445] Corynebacterium propinquum Porphyromonas endodontalis

[0446] Corynebacterium pseudodiphtheriticum Porphyromonas gingivalis

[0447] Corynebacterium pseudotuberculosis Porphyromonas gingivicanis

[0448] Corynebacterium renale Porphyromonas somerae

[0449] Corynebacterium resistens Porphyromonas uenonis

[0450] Corynebacterium riegelii Prevotella bergensis

[0451] Corynebacterium sanguinis Prevotella bivia

[0452] Corynebacterium simulans Prevotella buccae

[0453] Corynebacterium singulare Prevotella buccalis

[0454] Corynebacterium stationis Prevotella corporis

[0455] Corynebacterium striatum Prevotella dentalis

[0456] Corynebacterium sundsvallense Prevotella denticola

[0457] Corynebacterium thomssenii Prevotella disiens

[0458] Corynebacterium timonense Prevotella intermedia

[0459] Corynebacterium tuberculostearicum Prevotella loescheii

[0460] Corynebacterium tuscaniense Prevotella melaninogenica

[0461] Corynebacterium ulcerans Prevotella multiformis

[0462] Corynebacterium urealyticum Prevotella multi saccharivorax

[0463] Corynebacterium ureicelerivorans Prevotella nigrescens

[0464] Corynebacterium vitaeruminis Prevotella oralis Corynebacterium xerosis Prevotella oris

[0465] Coxiella burnetii Prevotella tannerae

[0466] Cronobacter condimenti Prevotella timonensis

[0467] Cronobacter dublinensis Propionib acterium acidifaciens

[0468] Cronobacter malonaticus Propionib acterium propionicum

[0469] Cronobacter sakazakii Propionimicrobium lymphophilum

[0470] Cronobacter turicensis Proteus mirabilis

[0471] Cronobacter universalis Proteus penneri

[0472] Cryptobacterium curtum Proteus vulgaris

[0473] Cupriavidus gilardii Providencia alcalifaciens

[0474] Cupriavidus metallidurans Providencia rettgeri

[0475] Cupriavidus pauculus Providencia rustigianii

[0476] Cupriavidus taiwanensis Providencia stuartii

[0477] Delftia acidovorans Pseudomonas aeruginosa

[0478] Dermabacter hominis Pseudomonas alcaligenes

[0479] Dermacoccus abyssi Pseudomonas cannabina

[0480] Dermacoccus nishinomiyaensis Pseudomonas citronellolis

[0481] Dermatophilus congolensis Pseudomonas fluorescens

[0482] Desulfomicrobium orale Pseudomonas fulva

[0483] Desulfovibrio desulfuricans Pseudomonas luteola

[0484] Desulfovibrio fairfieldensis Pseudomonas mendocina

[0485] Desulfovibrio vulgaris Pseudomonas monteilii

[0486] Dialister invisus Pseudomonas mosselii

[0487] Dialister micraerophilus Pseudomonas oryzihabitans

[0488] Dialister pneumosintes Pseudomonas otitidis

[0489] Dialister propionicifaciens Pseudomonas poae

[0490] Dichelobacter nodosus Pseudomonas protegens

[0491] Dielma fastidiosa Pseudomonas pseudoalcaligenes

[0492] Dietzia maris Pseudomonas putida

[0493] Dolosicoccus paucivorans Pseudomonas stutzeri

[0494] Dolosigranulum pigrum Pseudomonas veronii Dysgonomonas capnocytophagoides Pseudopropionibacterium propionicum

[0495] Dysgonomonas gadei Pseudoramibacter

[0496] Dysgonomonas hofstadii Pseudoramibacter alactolyticus

[0497] Dysgonomonas mossii Psychrobacter cryohalolentis

[0498] Edwardsiella hoshinae Psychrobacter immobilis

[0499] Edwardsiella ictaluri Psychrobacter phenylpyruvicus

[0500] Edwardsiella tarda Rahnella aquatilis

[0501] Eggerthella hongkongensis Ralstonia insi diosa

[0502] Eggerthella lenta Ralstonia mannitolilytica

[0503] Eggerthella sinensis Ralstonia pickettii

[0504] Ehrlichia canis Ralstonia solanacearum

[0505] Ehrlichia chaffeensis Raoultella ornithinolytica

[0506] Ehrlichia muris Raoultella planticola

[0507] Eikenella corrodens Raoultella terrigena

[0508] Elizabethkingia anophelis Rhodococcus equi

[0509] Elizabethkingia meningoseptica Rhodococcus erythropolis

[0510] Elizabethkingia miricola Rhodococcus fascians

[0511] Empedobacter brevis Rhodococcus rhodochrous

[0512] Empedobacter falsenii Rickettsia africae

[0513] Enterobacter aerogenes Rickettsia akari

[0514] Enterobacter cancerogenus Rickettsia amblyommatis

[0515] Enterobacter cloacae Rickettsia australis

[0516] Enterobacter gergoviae Rickettsia canadensis

[0517] Enterobacter hormaechei Rickettsia conorii

[0518] Enterobacter kobei Rickettsia felis

[0519] Enterobacter ludwigii Rickettsia japonica

[0520] Enterobacter mori Rickettsia massiliae

[0521] Enterobacter sakazakii Rickettsia monacensis

[0522] Enterococcus asini Rickettsia parkeri

[0523] Enterococcus avium Rickettsia prowazekii

[0524] Enterococcus casseliflavus Rickettsia raoultii Enterococcus cecorum Rickettsia rickettsii

[0525] Enterococcus columbae Rickettsia sibirica

[0526] Enterococcus di spar Rickettsia slovaca

[0527] Enterococcus durans Rickettsia typhi

[0528] Enterococcus faecalis Riemerella anatipestifer

[0529] Enterococcus faecium Robinsoniella peoriensis

[0530] Enterococcus flavescens Roseobacter denitrificans

[0531] Enterococcus gallinarum Roseomonas cervicalis

[0532] Enterococcus gilvus Roseomonas gilardii

[0533] Enterococcus haemoperoxidus Roseomonas mucosa

[0534] Enterococcus hirae Rothia aeria

[0535] Enterococcus italicus Rothia dentocariosa

[0536] Enterococcus malodoratus Rothia mucilaginosa

[0537] Enterococcus mundtii Rouxiella chamberiensis

[0538] Enterococcus pallens Ruminococcus flavefaciens

[0539] Enterococcus phoeniculicola Salmonella bongori

[0540] Enterococcus pseudoavium Salmonella enterica

[0541] Enterococcus raffinosus Salmonella enterica ssp. Arizonae

[0542] Enterococcus saccharolyticus Salmonella enterica ssp. Diarizonae

[0543] Enterococcus sulfureus Salmonella enterica ssp. Enterica

[0544] Enterococcus thailandicus Salmonella enteritidis

[0545] Erwinia billingiae Salmonella paratyphi

[0546] Erwinia gerundensis Salmonella typhi

[0547] Erysipelatoclostridium ramosum Salmonella typhimurium

[0548] Erysipelothrix rhusiopathiae Sanguibacteroides justesenii

[0549] Escherichia albertii Scardovia inopinata

[0550] Escherichia coli Scardovia wiggsiae

[0551] Escherichia fergusonii Selenomonas artemidis

[0552] Eubacterium brachy Selenomonas flueggei

[0553] Eubacterium infirmum Selenomonas infelix

[0554] Eubacterium limosum Selenomonas noxia Eubacterium minutum Selenomonas sputigena Eubacterium nodatum Serratia ficaria Eubacterium rectale Serratia fonticola Eubacterium saphenum Serratia grimesii Eubacterium sulci Serratia liquefaciens Eubacterium tenue Serratia marcescens Eubacterium ventriosum Serratia odorifera Eubacterium yurii Serratia plymuthica Eubacterium yurii mararetiae Serratia proteamaculans Eubacterium yurii schtitka Serratia quinivorans Eubacterium yurii yurii Serratia rubidaea Ewingella americana Serratia ureilytica Exiguobacterium acetylicum Shewanella algae Exiguobacterium aurantiacum Shewanella putrefaciens Facklamia hominis Shigella boydii Facklamia ignava Shigella dysenteriae Facklamia languida Shigella flexneri Facklamia sourekii Shigella sonnei Faecalicoccus pleomorphus Shimwellia blattae Fenollaria massiliensis Siccibacter turicensis Filifactor alocis Simkania negevensis Finegoldia magna Slackia exigua Franci sella hispaniensis Sneathia sanguinegens Francisella noatunensis Sphingobacterium multivorum Francisella philomiragia Sphingobacterium spiritivorum Francisella tularensis Sphingobium yanoikuyae Franconibacter helveticus Sphingomonas paucimobilis Fusobacterium gonidiaformans Staphylococcus agnetis Fusobacterium mortiferum Staphylococcus argenteus Fusobacterium naviforme Staphylococcus arlettae Fusobacterium necrogenes Staphylococcus aureus Fusobacterium necrophorum Staphylococcus auricularis

[0555] Fusobacterium nucleatum Staphylococcus capitis

[0556] Fusobacterium nucleatum fusiforme Staphylococcus capitis capitis

[0557] Fusobacterium nucleatum nucleatum Staphylococcus capitis ureolyticus

[0558] Fusobacterium nucleatum polymorphum Staphylococcus caprae

[0559] Fusobacterium nucleatum vincentii Staphylococcus carnosus

[0560] Fusobacterium periodonticum Staphylococcus chromogenes

[0561] Fusobacterium russii Staphylococcus cohnii

[0562] Fusobacterium ulcerans Staphylococcus cohnii cohnii

[0563] Fusobacterium varium Staphylococcus cohnii urealyticus

[0564] Gardnerella vaginalis Staphylococcus condimenti

[0565] Gemella bergeri Staphylococcus delphini

[0566] Gemella haemolysans Staphylococcus epidermidis

[0567] Gemella morbillorum Staphylococcus equorum

[0568] Gemella sanguinis Staphylococcus gallinarum

[0569] G1 obi cat ell a sanguinis Staphylococcus haemolyticus

[0570] Gordonia araii Staphylococcus hominis

[0571] Gordonia bronchialis Staphylococcus hominis hominis

[0572] Gordonia otitidis Staphylococcus hominis novobiosepticius

[0573] Gordonia polyisoprenivorans Staphylococcus hyicus

[0574] Gordonia rubripertincta Staphylococcus intermedins

[0575] Gordonia sputi Staphylococcus lugdunensis

[0576] Gordonia terrae Staphylococcus massiliensis

[0577] Gordonibacter pamelaeae Staphylococcus pasteuri

[0578] Granulibacter bethesdensis Staphylococcus pettenkoferi

[0579] Granulicatella adiacens Staphylococcus pseudintennedius

[0580] Granulicatella elegans Staphylococcus saccharolyticus

[0581] Grimontia hollisae Staphylococcus saprophyticus

[0582] Haemophilus aegyptius Staphylococcus schleiferi

[0583] Haemophilus ducreyi Staphylococcus schleiferi coagulans

[0584] Haemophilus haemolyticus Staphylococcus schleiferi schleiferi Haemophilus influenzae Staphylococcus sciuri

[0585] Haemophilus parahaemolyticus Staphylococcus simiae

[0586] Haemophilus parainfluenzae Staphylococcus simulans

[0587] Haemophilus paraphrohaemolyticus Staphylococcus succinus

[0588] Haemophilus pittmaniae Staphylococcus vitulinus

[0589] Haemophilus quentini Staphylococcus wameri

[0590] Haemophilus sputorum Staphylococcus xylosus

[0591] Hafnia alvei Stenotrophomonas acidaminiphila

[0592] Hafnia paralvei Stenotrophomonas maltophilia

[0593] Helcococcus kunzii Streptobacillus moniliformis

[0594] Helcococcus sueciensis Streptococcus acidominimus

[0595] Helicobacter bilis Streptococcus agalactiae

[0596] Helicobacter canadensis Streptococcus anginosus

[0597] Helicobacter canis Streptococcus canis

[0598] Helicobacter cinaedi Streptococcus constellatus

[0599] Helicobacter felis Streptococcus constellatus constellatus

[0600] Helicobacter fennelliae Streptococcus constellatus pharyngis

[0601] Helicobacter heilmannii Streptococcus criceti

[0602] Helicobacter magdeburgensis Streptococcus cri status

[0603] Helicobacter pullorum Streptococcus dentisani

[0604] Helicobacter pylori Streptococcus dysgalactiae

[0605] Helicobacter winghamensis Streptococcus dysgalactiae dysgalactiae

[0606] Holdemania filiformis Streptococcus dysgalactiae equisimilis

[0607] Ignatzschineria larvae Streptococcus equi

[0608] Ignavigranum ruoffiae Streptococcus equi equi

[0609] Inquilinus limosus Streptococcus equi zooepidemicus

[0610] Isoptericola variabilis Streptococcus equinus

[0611] Janibacter indicus Streptococcus ferus

[0612] Janibacter melonis Streptococcus gallolyticus

[0613] Johnsonella ignava Streptococcus gallolyticus ssp. Gallolyticus

[0614] Jonesia denitrificans Streptococcus gallolyticus ssp. Pateurianus Kerstersia gyiorum Streptococcus gordonii

[0615] Kingella denitrificans Streptococcus hyovaginalis

[0616] Kingella kingae Streptococcus infantarius

[0617] Kingella oralis Streptococcus infantis

[0618] Kingella potus Streptococcus iniae

[0619] Klebsiella granulomatis Streptococcus intermedius

[0620] Klebsiella michiganensis Streptococcus lutetiensis

[0621] Klebsiella oxytoca Streptococcus macacae

[0622] Klebsiella pneumoniae Streptococcus macedonicus

[0623] Klebsiella pneumoniae ssp. Ozaenae Streptococcus massiliensis

[0624] Klebsiella pneumoniae ssp. Pneumoniae Streptococcus mitis

[0625] Klebsiella quasipneumoniae Streptococcus mutans

[0626] Klebsiella variicola Streptococcus oralis

[0627] Kluyvera ascorbata Streptococcus parasanguinis

[0628] Kluyvera cryocrescens Streptococcus pasteurianus

[0629] Kluyvera intermedia Streptococcus peroris

[0630] Kocuria kristinae Streptococcus pneumoniae

[0631] Kocuria palustris Streptococcus porcinus

[0632] Kocuria rhizophila Streptococcus pseudopneumoniae

[0633] Kocuria rosea Streptococcus pseudoporcinus

[0634] Kocuria varians Streptococcus pyogenes

[0635] Kurthia gibsonii Streptococcus ratti

[0636] Kurthia huakuii Streptococcus salivarius

[0637] Kurthia massiliensis Streptococcus sanguinis

[0638] Kytococcus schroeteri Streptococcus sinensis

[0639] Kytococcus sedentarius Streptococcus sobrinus

[0640] Lactobacillus acidophilus Streptococcus suis

[0641] Lactobacillus antri Streptococcus thermophilus

[0642] Lactobacillus brevis Streptococcus tigurinus

[0643] Lactobacillus casei Streptococcus uberis

[0644] Lactobacillus coleohominis Streptococcus urinalis Lactobacillus crispatus Streptococcus vestibularis

[0645] Lactobacillus fermentum Streptomyces bikini ensis

[0646] Lactobacillus gasseri Streptomyces cattleya

[0647] Lactobacillus iners Streptomyces griseus

[0648] Lactobacillus j ensenii Streptomyces somaliensis

[0649] Lactobacillus paracasei Succinivibrio dextrinosolvens

[0650] Lactobacillus paraplantarum Sutterella wadsworthensis

[0651] Lactobacillus plantarum Suttonella indologenes

[0652] Lactobacillus pontis Tannerella forsythia

[0653] Lactobacillus rhamnosus Tatumella ptyseos

[0654] Lactobacillus saerimneri Taylorella asinigenitalis

[0655] Lactobacillus sakei Taylorella equigenitalis

[0656] Lactobacillus salivarius Tissierella praeacuta

[0657] Lactobacillus ultunensis Treponema amylovorum

[0658] Lactobacillus vaginalis Treponema denticola

[0659] Lactococcus garvieae Treponema lecithinolyticum

[0660] Lactococcus lactis Treponema maltophilum

[0661] Laribacter hongkongensis Treponema medium

[0662] Latilactobacillus sakei Treponema pallidum

[0663] Lautropia mirabilis Treponema parvum

[0664] Lawsonella clevelandensis Treponema pectinovorum

[0665] Lawsonia intracellularis Treponema pertenue

[0666] Leclercia adecarboxylata Treponema putidum

[0667] Legionella adelaidensis Treponema socranskii

[0668] Legionella anisa Treponema vincentii

[0669] Legionella birmingham ensis Tropheryma whipplei

[0670] Legionella brunensis Trueperella pyogenes

[0671] Legionella cherrii Tsukamurella paurometabola

[0672] Legionella Cincinnati ensis Tsukamurella pulmonis

[0673] Legionella clemsonensis Tsukamurella tyrosinosolvens

[0674] Legionella drancourtii Turicella otitidis Legionella dumoffii Ureaplasma parvum

[0675] Legionella erythra Ureaplasma urealyticum

[0676] Legionella fairfieldensis Vagococcus fluvialis

[0677] Legionella fallonii Veillonella dispar

[0678] Legionella feeleii Veillonella montpellierensis

[0679] Legionella geestiana Veillonella parvula

[0680] Legionella gormanii Veillonella seminalis

[0681] Legionella hackeliae Vibrio alginolyticus

[0682] Legionella israelensis Vibrio cholerae

[0683] Legionella jamestowniensis Vibrio cincinnatiensis

[0684] Legionella jordanis Vibrio fluvialis

[0685] Legionella lansingensis Vibrio fumissii

[0686] Legionella londiniensis Vibrio harveyi

[0687] Legionella longbeachae Vibrio metschnikovii

[0688] Legionella maceachemii Vibrio mimicus

[0689] Legionella massiliensis Vibrio navarrensis

[0690] Legionella nautarum Vibrio parahaemolyticus

[0691] Legionella norrlandica Vibrio vulnificus

[0692] Legionella oakridgensis Waddlia chondrophila

[0693] Legionella parisiensis Wautersiella falsenii

[0694] Legionella pneumophila Weeksella virosa

[0695] Legionella quateirensis Weissella confusa

[0696] Legionella quinlivanii Weissella paramesenteroides

[0697] Legionella rubrilucens Weissella viridescens

[0698] Legionella sainthelensi Williamsia muralis

[0699] Legionella santicrucis Wohlfahrtiimonas chitiniclastica

[0700] Legionella shakespearei Wolbachia pipientis

[0701] Legionella spiritensis Xanthomonas axonopodis

[0702] Legionella steelei Xanthomonas campestris

[0703] Legionella tucsonensis Xylanimonas cellulosilytica

[0704] Legionella tunisiensis Yersinia bercovieri Legionella wadsworthii Yersinia enterocolitica

[0705] Legionella waltersii Yersinia frederiksenii

[0706] Legionella worsleiensis Yersinia intermedia

[0707] Leifsonia aquatica Yersinia kristensenii

[0708] Leifsonia xyli Yersinia pestis

[0709] Leminorella grimontii Yersinia pseudotuberculosis

[0710] Leminorella richardii Yersinia ruckeri

[0711] Yokenella regensburgei

[0712] The present disclosure also relates to methods and systems that use computer-generated information to design and / or construct a database of probe sequences or set of probes. For example, in some embodiments, a first analytical tool using the information from speciesspecific or clade-specific marker gene sequences and / or 16S ribosomal RNA sequences and / or virulence factor sequences and / or AMR genes and a second analytical tool to fragment the sequences into oligonucleotides with the desired or advantageous parameters for the probes, including but not limited to length, distance spaced between the probes on the target sequences, and percentage sequence identity.

[0713] In a further aspect, analytical tools such as a first module configured to perform the choice of species-specific or clade-specific marker gene sequences and / or 16S ribosomal RNA sequences and / or virulence factor sequences and / or AMR genes, and a second module to perform the fragmentation of the sequences may be provided that determines desired or advantageous features of the oligonucleotides such as the length, distance spaced between the oligonucleotides on the sequences, and / or percentage sequence identity. The results of these tools form a model for use in designing the oligonucleotides for the disclosed database of probe sequences or set of probes.

[0714] An illustrative system for generating a design model includes an analytical tool such as a module configured to include species-specific or clade-specific marker gene sequences extracted from the Metaphlan4 database 16S ribosomal RNA sequences extracted from SILVA database for a total of 1333 bacterial species, virulence factor sequences extracted from the VFDB, and / or AMR extracted from CARD. The analytical tool may include any suitable hardware, software, or combination thereof for determining correlations. A second analytical tool such as module is used to fragment the sequences. This analytical tool may include any suitable hardware, software, or combination for determining the desired or advantageous features of the oligonucleotides including but not limited to length, distance spaced between the probes on the sequences, and percentage sequence identity.

[0715] After the sequence information is obtained for the oligonucleotide probes, the oligonucleotides can be synthesized by any method known in the art including but not limited to solid-phase synthesis using phosphoramidite method and phosphoramidite building blocks derived from protected 2’-deoxynucleosides (dA, dC, dG, and T), ribonucleosides (A, C, G, and U), or chemically modified nucleosides, e.g. linked nucleic acids (LNA), bridged nucleic acids (BNA) or peptide nucleic acids (PNA).

[0716] One embodiment is a library or platform comprising the set of oligonucleotide probes with the sequences in the database that is capable of capturing nucleic acids from at least one bacterium. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than one bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than ten bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than fifty bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than one hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than one hundred and fifty bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than two hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than two hundred and fifty bacteria. In some embodiments, the library or platform comprising the oligonucleotide probes is capable of capturing nucleic acids from more than three hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than four hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than five hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than six hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than seven hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than eight hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than nine hundred bacteria. In some embodiments, the library comprising the oligonucleotide probes is capable of capturing nucleic acids from more than one thousand hundred bacteria.

[0717] In one embodiment, the oligonucleotides are in solution.

[0718] In one embodiment, the oligonucleotides are pre-bound to a solid support or substrate. Preferred solid supports include, but are not limited to, beads (e.g., magnetic beads (i.e., the bead itself is magnetic, or the bead is susceptible to capture by a magnet)) made of metal, glass, plastic, dextran (such as the dextran bead sold under the tradename, Sephadex (Pharmacia)), silica gel, agarose gel (such as those sold under the tradename, Sepharose (Pharmacia)), or cellulose); capillaries; flat supports (e.g., fdters, plates, or membranes made of glass, metal (such as steel, gold, silver, aluminum, copper, or silicon), or plastic (such as polyethylene, polypropylene, polyamide, or polyvinylidene fluoride)); a chromatographic substrate; a microfluidics substrate; and pins (e.g., arrays of pins suitable for combinatorial synthesis or analysis of beads in pits of flat surfaces (such as wafers), with or without filter plates). Additional examples of suitable solid supports include, without limitation, agarose, cellulose, dextran, polyacrylamide, polystyrene, sepharose, and other insoluble organic polymers. Appropriate binding conditions (e.g., temperature, pH, and salt concentration) may be readily determined by the skilled artisan.

[0719] The oligonucleotides may be either covalently or non-covalently bound to the solid support. Furthermore, the oligonucleotides may be directly bound to the solid support (e.g., the oligonucleotides are in direct van der Waal and / or hydrogen bond and / or salt-bridge contact with the solid support), or indirectly bound to the solid support (e.g., the oligonucleotides are not in direct contact with the solid support themselves). Where the oligonucleotides are indirectly bound to the solid support, the nucleotides of the capture nucleic acid are linked to an intermediate composition that, itself, is in direct contact with the solid support.

[0720] To facilitate binding of the oligonucleotides to the solid support, the oligonucleotides may be modified with one or more molecules suitable for direct binding to a solid support and / or indirect binding to a solid support by way of an intermediate composition or spacer molecule that is bound to the solid support (such as an antibody, a receptor, a binding protein, or an enzyme). Examples of such modifications include, without limitation, a ligand (e.g., a small organic or inorganic molecule, a ligand to a receptor, a ligand to a binding protein or the binding domain thereof (such as biotin and digoxigenin)), an antigen and the binding domain thereof, an aptamer, a peptide tag, an antibody, and a substrate of an enzyme. In a preferred embodiment, the oligonucleotides comprise biotin.

[0721] Linkers or spacer molecules suitable for spacing biological and other molecules, including nucleic acids / polynucleotides, from solid surfaces are well-known in the art, and include, without limitation, polypeptides, saturated or unsaturated bifunctional hydrocarbons, and polymers (e.g., polyethylene glycol). Other useful linkers are commercially available.

[0722] In a further embodiment, the sequences of the oligonucleotides are the complement of (i.e., is complementary to) a sequence of the marker sequences of one or more bacteria as well as AMR genes and / or virulence factors and / or 16S ribosomal RNA. In another embodiment, the oligonucleotides are capable of hybridizing to a sequence of the marker sequences of one or more bacteria as well as AMR genes and / or virulence factors and / or 16S ribosomal RNA under stringent conditions.

[0723] The "complement" of a nucleic acid sequence refers, herein, to a nucleic acid molecule which is completely complementary to another nucleic acid, or which will hybridize to the other nucleic acid under conditions of high stringency. High-stringency conditions are known in the art. See, e.g., Maniatis et al., Molecular Cloning: A Laboratory Manual, 2nd ed. (Cold Spring Harbor: Cold Spring Harbor Laboratory, 1989) and Ausubel et al., eds., Current Protocols in Molecular Biology (New York, N.Y.: John Wiley & Sons, Inc., 2001). Stringent conditions are sequence-dependent, and may vary depending upon the circumstances.

[0724] In one embodiment, the oligonucleotides are synthesized using a cleavable programmable array. The oligonucleotides are cleaved from the array and hybridized with the nucleic acids from the sample in solution.

[0725] The set of probes can be in the form of a collection of oligonucleotides, preferably designed as set forth above, i.e., a probe library. The oligonucleotides can be in solution or attached to a solid state, such as an array or a bead. Additionally, the oligonucleotides can be modified with another molecule. In a preferred embodiment, the oligonucleotides comprise biotin. The database of probe sequences can also be in the form of a database or databases which can include information regarding the sequence and length of each oligonucleotide probe, and the bacterium and / or marker sequence from which the oligonucleotide sequence derived as well as AMR genes and virulence factors and 16S ribosomal RNA. The database can searchable. From the database, one of skill in the art can obtain the information needed to design and synthesis the oligonucleotide probes. The databases can also be recorded on machine-readable storage medium, any medium that can be read and accessed directly by a computer. A machine- readable storage medium can comprise, for example, a data storage material that is encoded with machine-readable data or data arrays. Machine-readable storage medium can include but are not limited to magnetic storage media, optical storage media, electrical storage media, and hybrids. One of skill in the art can easily determine how presently known machine-readable storage medium and future developed machine-readable storage medium can be used to create a manufacture of a recording of any database information. “Recorded” refers to a process for storing information on a machine-readable storage medium using any method known in the art.

[0726] Construction of a Sequencing Library

[0727] A further embodiment of the present disclosure is a method of constructing a sequencing library suitable for sequencing with any high throughput sequencing method utilizing the set of probes.

[0728] Accordingly, the method may include the following steps.

[0729] Nucleic acids from a sample are obtained. The sample used in the present methods may be an environmental sample, a food sample, or a biological sample. The preferred sample is a biological sample or an environmental sample (e.g., a wastewater sample or sewage sample). A biological sample may be obtained from a tissue of a subject or bodily fluid from a subject including, but not limited to, nasopharyngeal aspirate, blood, cerebrospinal fluid, saliva, serum, urine, sputum, bronchial lavage, pericardial fluid, or peritoneal fluid, or a solid such as feces. A biological sample can also be cells, cell culture or cell culture medium. The sample may or may not comprise or contain any bacterial nucleic acids. In one embodiment, the sample is from a vertebrate subject, and in a further embodiment, the sample is from a human subject. In another embodiment, the sample comprises blood. In another preferred embodiment, the sample comprises cells, cell culture, cell culture medium or any other composition being used for developing pharmaceutical and therapeutic agents. In some embodiments, the sample is from food or a food supply.

[0730] The nucleic acids from the sample are subjected to fragmentation, to obtain a nucleic acid fragment. There are no special limitations on the type of the nucleic acid sample which may be used and there are no special limitations on means for performing the fragmentation. Any chemical or physical method which randomly fragments nucleic acid samples may be used. It is preferred that the nucleic acid sample is fragmented to obtain a nucleic acid fragment having a length of about 200 bp to about 300 bp or any other size distribution suitable for the respective sequencing platform.

[0731] After being obtained, the nucleic acid fragments can be ligated to an adaptor. In one embodiment, the adaptor is a linear adaptor. Linear adaptors can be added to the fragments by end-repairing the fragments, to obtain an end-repaired fragment; adding an adenine base to the 3’ ends of the fragment, to obtain a fragment having an adenine at the 3’ end; and ligating an adaptor to the fragment having an adenine at the 3 ’end.

[0732] In some embodiments, the adaptor comprises an identifier sequence. In some embodiments, the adaptor comprises sequences for priming for amplification. In some embodiments, the adaptor comprises both an identified sequence and sequences for priming for amplification.

[0733] After the nucleic acid fragment is ligated to the adaptor, it is contacted with the oligonucleotide probes described herein, under conditions that allow the nucleic acid fragment to hybridize to the oligonucleotide probes if the nucleic acid comprises any sequences from bacteria or genes represented in the database, set of sequences, or oligonucleotide probes described herein. This step may be performed in solution or in a solid phase hybridization method.

[0734] After contact with the oligonucleotides, any hybridization product(s) may be subject to amplification conditions. In one embodiment, primers for amplification are present in the adaptor ligated to the nucleic acid fragment. The resulting amplified product(s) comprise the sequencing library that is suitable to be sequenced using any HTS system now known or later developed.

[0735] Amplification may be carried out by any means known in the art, including polymerase chain reaction (PCR) and isothermal amplification. PCR is a practical system for in vitro amplification of a DNA base sequence. For example, a PCR assay may use a heat-stable polymerase and two primers: one complementary to the (+)-strand at one end of the sequence to be amplified; and the other complementary to the (-)-strand at the other end. Because the newly- synthesized DNA strands can subsequently serve as additional templates for the same primer sequences, successive rounds of primer annealing, strand elongation, and dissociation may produce rapid and highly -specific amplification of the desired sequence. PCR also may be used to detect the existence of a defined sequence in a DNA sample. In one embodiment, the hybridization products are mixed with suitable PCR reagents. A PCR reaction is then performed to amplify the hybridization products.

[0736] In one embodiment, the sequencing library is constructed using the probe set in a cleavable array. Nucleic acids from the sample are extracted and subjected to reverse transcriptase treatment and ligated to an adaptor comprising an identifier and sequences for priming for amplification. The oligonucleotides are synthesized using a cleavable array platform wherein the oligonucleotides are biotinylated. The biotinylated oligonucleotides are then cleaved from the solid matrix into solution with the nucleic acids from the sample to enable hybridization of the oligonucleotides to any bacterial nucleic acids in solution. After hybridization, nucleic acid(s) from the sample bound to the biotinylated oligonucleotides comprising the probe set, i.e., hybridization product(s), is collected by streptavidin magnetic beads, and amplified by PCR using the adaptor sequences as specific priming sites, resulting in an amplified product for sequencing on any known HTS systems (Ion, Illumina, 454) and any HTS system developed in the future.

[0737] In some embodiments, a sample comprising nucleic acids is exposed to the oligonucleotide probes described under hybridization conditions. After hybridization, the probes are captured (e.g., biotinylated probes are captured on streptavidin magnetic beads) and hybridization products are purified. Nucleic acids which bound the probes can be released and subsequently prepared for amplification and / or HTS sequencing, for example, by adding adaptor sequence portions to the released nucleic acids and / or size selecting the released nucleic acids.

[0738] In a further embodiment, the sequencing library can be directly sequenced using any method known in the art. In other words, the nucleic acids captured by the probes can be sequenced without amplification.

[0739] Methods and Systems Using the Disclosed Database of Sequences and Set of Probes The present disclosure includes methods and systems for the detection, identification and / or differentiation of bacteria and / or pathogenicity elements, and / or AMR genes, and / or 16S ribosomal RNA, in any sample, utilizing the database of probe sequences or set of probes.

[0740] The methods and systems may be used to detect bacteria and / or pathogenicity elements and / or AMR and / or 16S ribosomal RNA genes, in research, clinical, environmental, and food samples. Additional applications include, without limitation, detection of infectious pathogens, the screening of blood products (e.g., screening blood products for infectious agents), biodefense, food safety, environmental contamination, forensics, and genetic-comparability studies. The present disclosure also provides methods and systems for detecting bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA in cells, cell culture, cell culture medium and other compositions used for the development of pharmaceutical and therapeutic agents. Accordingly, the present disclosure provides methods and systems for a myriad of specific applications, including, without limitation, a method for determining the presence of bacteria and / or pathogenicity elements and / or AMR genes, and / or 16S ribosomal RNA, in a sample, a method for screening blood products, a method for assaying a food product for contamination, a method for assaying a sample for environmental contamination, and a method for detecting genetically-modified organisms. The present disclosure further provides use of the system in such general applications as biodefense against bioterrorism, forensics, and genetic-comparability studies.

[0741] The subject may be any animal, particularly a vertebrate and more particularly a mammal or avian, including, without limitation, a cow, dog, human, monkey, mouse, pig, rat, chicken or wildlife species such as a bat or a rodent. The subject may also be an invertebrate such as tick, mosquito or sand fly. In some embodiments, the subject is a human. The subject may be known to have a pathogen infection, suspected of having a pathogen infection, or believed not to have a pathogen infection.

[0742] The systems and methods described herein support the multiplex detection of multiple bacteria and bacterial transcripts in any sample.

[0743] Thus, one embodiment provides a system for the detection, identification and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA, in any sample. The system includes at least one subsystem wherein the subsystem includes the database of probe sequences or set of oligonucleotide probes as described herein. The system can also include additional subsystems for the purpose of preparation of oligonucleotides from the database of probe sequences; isolation and preparation of the nucleic acid from the sample; hybridization of the nucleic acid from the sample with the oligonucleotides to form hybridization product(s); amplification of the hybridization product(s); sequencing the hybridization product(s); amplification of the nucleic acid(s) from the sample which do not form hybridization product(s); sequencing the nucleic acid(s) from the sample which do not form hybridization product(s); and identification and characterization of the bacteria, and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA by the comparison between the sequences of the hybridization product(s) and / or nucleic acids, and known bacteria and / or pathogenicity elements and / or AMR genes and / or 16S ribosomal RNA.

[0744] Additionally, the present disclosure provides a method for the detection, identification, and / or differentiation of bacteria and / or pathogenicity elements and / or AMR genes, and / or 16S ribosomal RNA, in any sample, including the steps of: obtaining the sample; isolating and preparing the nucleic acid from the sample; contacting the nucleic acid or derivatives thereof from the sample with the oligonucleotides generated from the disclosed database of probe sequences or set of oligonucleotide probes as described herein under conditions sufficient for the nucleic acid fragments and the oligonucleotides to hybridize; and detecting any hybridization products formed between the nucleic acid and the oligonucleotides.

[0745] These methods can also include additional steps to: amplify hybridization product(s); sequence the hybridization product(s); amplify nucleic acid(s) from the sample which do not form hybridization product(s); sequence nucleic acid(s) or derivatives thereof from the sample which do not form hybridization product(s); and comparison of hybridization product(s) and / or nucleic acid(s) from the sample which do not form hybridization product(s) with sequences of known bacteria, 16S ribosomal RNA, AMR genes and / or pathogenicity elements.

[0746] As disclosed above, the methods can be performed on any sample, including but not limited to biological samples, environmental samples, or food samples. One such sample is a biological sample. A biological sample may be obtained from a tissue of a subject or bodily fluid from a subject including but not limited to nasopharyngeal aspirate, blood, cerebrospinal fluid, saliva, serum, urine, sputum, bronchial lavage, pericardial fluid, or peritoneal fluid, or a solid such as feces. A biological sample can also be cells, cell culture or cell culture medium. The sample may or may not comprise or contain any bacterial nucleic acids. In one embodiment, the sample is from a vertebrate subject, and in a further embodiment, the sample is from a human subject. In another embodiment, the sample is from an invertebrate subject.

[0747] In another embodiment, the sample comprises cells, cell culture, cell culture medium or any other composition being used for developing pharmaceutical and therapeutic agents.

[0748] In some embodiments, the nucleic acids from the sample are further processed by shearing, adaptor, etc., forming derivatives of the isolated nucleic acid.

[0749] Kits

[0750] The disclosure also includes reagents and kits for practicing the disclosed methods. These reagents and kits may vary.

[0751] One reagent would be the disclosed set of probes, which can be in the form of a collection of oligonucleotide probes which comprise sequences derived from the disclosed database of probe sequences. This collection of oligonucleotide probes can be in solution or attached to a solid state. Additionally, the oligonucleotide probes can be modified for use in a reaction. A preferred modification is the addition of biotin to the probes.

[0752] A further reagent is a searchable database with information regarding the oligonucleotides including at least sequence information, length, and the origin.

[0753] Other reagents in the kit could include reagents for isolating and preparing nucleic acids from a sample, hybridizing the nucleic acid fragments from the sample with the oligonucleotides of the probe set, amplifying the hybridization products, and obtaining sequence information.

[0754] Kits may include any of the above-mentioned reagents, as well as reference / control sequences that can be used to compare the test sequence information obtained, by for example, suitable computing means based upon an input of sequence information.

[0755] In addition, kits would also further include instructions.

[0756] A further embodiment is a kit for designing and / or constructing the database of probe sequences comprising analytical tools to choose sequence information and break the sequences into fragments for oligonucleotides with the proper parameters including proper length, distance spaced between the oligonucleotides on the target sequences, and percentage sequence identity. This kit could also include instructions as to database and target sequence choice. Additional Embodiments

[0757] According to embodiments of the present invention, there is provided a bacterial sequence capture platform for the detection, identification, and / or differentiation of bacterially- derived sequences in a sample, the platform comprising a plurality of oligonucleotide probes, wherein the plurality comprises at least one oligonucleotide probe which comprises a hybridization portion partially or fully complementary to a portion of a bacterially-derived sequence selected from the group consisting of a bacterial gene sequence, a 16S ribosomal RNA sequence, a pathogenicity element sequence, a virulence factor sequence, and an antimicrobial resistance (AMR) gene sequence, wherein the sequences of the hybridization portions of the oligonucleotide probes cluster at about 90-100% sequence identity, wherein each hybridization portion of an oligonucleotide probe is about 5-300 nucleotides in length, wherein different hybridization portions that each bind a different portion of the same bacterially-derived sequence are tiled across said bacterially-derived sequence and have an interprobe spacing of about 20-100 nucleotides, and wherein the plurality of oligonucleotide probes of the platform comprises 100,000 to 1,000,000 oligonucleotide probes, preferably less than about 1,000,000 oligonucleotide probes.

[0758] In some embodiments, each hybridization portion of an oligonucleotide probe is about 50-200 nucleotides in length, preferably about 100-150 nucleotides in length, more preferably about 120 nucleotides in length.

[0759] In some embodiments, the average length of the plurality of hybridization portions of oligonucleotide probes is about 120 nucleotides.

[0760] In some embodiments, different hybridization portions that each bind a different portion of the same bacterially-derived sequence are tiled across said bacterially-derived sequence and have an inter-probe spacing of about 60 nucleotides.

[0761] In some embodiments, the sequences of the hybridization portions of the oligonucleotide probes cluster at about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity.

[0762] In some embodiments, the plurality of oligonucleotide probes comprises hybridization portions partially or fully complementary to portions of bacterially-derived sequences comprising one or more bacterial gene sequences, one or more 16S ribosomal RNA sequences, one or more pathogenicity element sequences, one or more virulence factor sequences, and / or one or more antimicrobial resistance (AMR) gene sequences.

[0763] In some embodiments, the bacterial gene sequence is a species-specific or clade-specific gene sequence.

[0764] In some embodiments, the species-specific or clade-specific gene sequences are obtained from Metaphlan4 database.

[0765] In some embodiments, the 16S ribosomal RNA sequences are obtained from the SILVA database.

[0766] In some embodiments, the virulence factor sequences are obtained from the Virulence Factor Database (VFDB).

[0767] In some embodiments, the AMR genes are obtained from the Comprehensive Antibiotic Resistance Database (CARD).

[0768] In some embodiments, each bacterially-derived sequence comprises a portion that is about 50-300 nucleotides in length and is partially or fully complementary to a hybridization portion of an oligonucleotide probe.

[0769] In some embodiments, each hybridization portion is at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% complementary to a portion of a bacterially-derived sequence.

[0770] In some embodiments, the plurality of oligonucleotide probes comprises at least one oligonucleotide probe which comprises a hybridization portion partially or fully complementary to a portion of a bacterially-derived sequence from a bacterial species listed in Table 1.

[0771] In some embodiments, every bacterial species listed in Table 1 comprises a sequence, preferably a unique sequence relative to any other bacterial species listed in Table 1, that is partially or fully complementary to a hybridization portion of a oligonucleotide probe of the plurality of the platform.

[0772] In some embodiments, each oligonucleotide probe comprises a capture portion.

[0773] In some embodiments, the capture portion is selected from the group consisting of biotin, digoxygenin, a ligand, a small organic molecule, a small inorganic molecule, an aptamer, an antigen, an antibody, and a substrate.

[0774] In some embodiments, each oligonucleotide probe is biotinylated. In some embodiments, and means for capturing, isolating, and / or purifying the plurality of oligonucleotide probes from a mixture of other nucleic acid molecules.

[0775] In some embodiments, the oligonucleotide probes comprise DNA, RNA, bridged nucleic acids, locked nucleic acids, and / or peptide nucleic acids. In some embodiments, the hybridization portion of an oligonucleotide probe comprises DNA, RNA, bridged nucleic acids, locked nucleic acids, and / or peptide nucleic acids. In some embodiments, the oligonucleotide probes are capable of hybridizing DNA, cDNA, RNA, and / or mRNA molecules.

[0776] In some embodiments, the oligonucleotide probes of the platform may be in solution or attached to a solid support. In some embodiments, the platform comprises oligonucleotide probes generated in an array format, e.g., a cleavable array format. In some embodiments, the platform comprises oligonucleotide probes generated from semiconductor-based synthetic DNA manufacturing.

[0777] In some embodiments, the sample is a biological sample or an environmental sample.

[0778] In some embodiments, the sample is selected from the group consisting of saliva, mucus, a nasopharyngeal swab, serum, plasma, blood, urine, feces, cerebrospinal fluid, a bodily fluid, cultured cells, an organ tissue, and biopsied tissue.

[0779] In some embodiments, the sample is selected from the group consisting of an aqueous sample, a liquid sample, water, wastewater, sewage, greywater, blackwater, freshwater, liquid waste, seawater, drinking water, air, a gaseous sample, soil, a food sample, culture medium, and a swab of an inanimate surface or object.

[0780] In some embodiments, the sample is obtained from a sewage system, a drainage system, a plumbing system, or a water treatment facility.

[0781] In some embodiments, the sample is obtained from a human subject.

[0782] According to embodiments of the present invention, there is provided a method of screening a sample for bacterially-derived sequences, the method comprising: a) exposing the sample, or nucleic acids isolated, amplified, and / or enriched from the sample, to any one of the bacterial sequence capture platforms described herein to form one or more hybridization products, wherein each hybridization product comprises a nucleic acid of the sample and an oligonucleotide probe of the platform; b) capturing the one or more hybridization products; and c) identifying the presence of one or more bacterially-derived sequences in the sample based on the sequences of the one or more captured hybridization products; thereby screening the sample for bacterially-derived sequences.

[0783] In some embodiments, nucleic acids in the sample are isolated and / or enriched prior to the exposing in step (a).

[0784] In some embodiments, the sample is a biological sample or an environmental sample.

[0785] In some embodiments, the sample is selected from the group consisting of saliva, mucus, a nasopharyngeal swab, serum, plasma, blood, urine, feces, cerebrospinal fluid, a bodily fluid, cultured cells, an organ tissue, and biopsied tissue.

[0786] In some embodiments, the sample is selected from the group consisting of an aqueous sample, a liquid sample, water, wastewater, sewage, greywater, blackwater, freshwater, liquid waste, seawater, drinking water, air, a gaseous sample, soil, a food sample, culture medium, and a swab of an inanimate surface or object.

[0787] In some embodiments, the sample is obtained from a sewage system, a drainage system, a plumbing system, or a water treatment facility.

[0788] In some embodiments, the sample is obtained from a human subject.

[0789] In some embodiments, the method further comprises: sequencing one or more detected hybridization products; comparing the nucleotide sequence of the one or more hybridization products to nucleotide sequences of known bacterially-derived sequences; and identifying and / or differentiating one or more bacterially-derived sequences in the sample based on sequence identity of the hybridization product to the nucleotide sequences of known bacterially-derived sequences.

[0790] According to embodiments of the present invention, there is provided a kit comprising any one of the bacterial sequence capture platforms described herein and instructions for using the platform.

[0791] In some embodiments, the kit further comprises a sample, wherein the platform is used for the detection, identification, and / or differentiation of bacterially-derived sequences in the sample.

[0792] In some embodiments, the sample is a biological sample or an environmental sample. In some embodiments, the sample is a liquid sample or an aqueous sample.

[0793] In some embodiments, the sample is selected from the group consisting of a water sample, wastewater, sewage, greywater, blackwater, freshwater, liquid waste, seawater, drinking water, air, a gaseous sample, soil, a food sample, culture medium, and a swab of an inanimate surface or object.

[0794] In some embodiments, the sample is a wastewater sample or a sewage sample.

[0795] In some embodiments, the sample is a wastewater sample.

[0796] In some embodiments, the sample is a sewage sample.

[0797] In some embodiments, the sample is obtained from a sewage system, a drainage system, a plumbing system, or a water treatment facility.

[0798] In some embodiments, the sample is selected from the group consisting of saliva, mucus, a nasopharyngeal swab, serum, plasma, blood, urine, feces, cerebrospinal fluid, a bodily fluid, cultured cells, an organ tissue, and biopsied tissue.

[0799] In some embodiments, the sample comprises nucleic acids. In some embodiments, the nucleic acids in the sample are purified, enriched, and / or isolated. The platform of the kit may then be applied to the nucleic acids derived from the sample for the detection, identification, and / or characterization of vertebrate-infecting viruses in the sample.

[0800] According to embodiments of the present invention, there is provided a method for designing and / or constructing a database of probe sequences or a probe set comprising oligonucleotide probes for the detection, identification, and / or differentiation of bacteria and / or one or more of 16S ribosomal RNA, pathogenicity elements and / or AMR genes, comprising: a) obtaining i) one or more species-specific or clade-specific marker gene sequences; or ii) one or more 16S ribosomal RNA sequences; or iii) one or more virulence factor sequences; or iv) one or more AMR gene sequences; or v) any combination of (i), (ii), (iii), and (iv); and b) breaking the sequences obtained in step a. into fragments, wherein the fragments are the basis of the probes and are designed to be of a length, and spaced at a distance across the target sequences, such that the total number of probe sequences in the database or probes in the probe set corresponds to a desired range or number. In some embodiments, the species-specific or clade-specific gene sequences are obtained from Metaphlan4 database.

[0801] In some embodiments, the 16S ribosomal RNA sequences are obtained from the SILVA database.

[0802] In some embodiments, the virulence factor sequences are obtained from the Virulence Factor Database (VFDB).

[0803] In some embodiments, the AMR genes are obtained from the Comprehensive Antibiotic Resistance Database (CARD).

[0804] In some embodiments, the desired range or number is less than one million.

[0805] In some embodiments, the method comprises a further step of synthesizing one or more of the oligonucleotide probes for which the sequence information was obtained in step b.

[0806] In some embodiments, the oligonucleotide probes are chosen from the group consisting of DNA, RNA, Bridged Nucleic Acids, Locked Nucleic Acids, and Peptide Nucleic Acids.

[0807] In some embodiments, the one or more oligonucleotide probes are synthesized on a cleavable microarray.

[0808] In some embodiments, the oligonucleotides are modified to comprise a composition for binding to a solid support, chosen from the group consisting of biotin, digoxygenin, ligands, small organic molecules, small inorganic molecules, aptamers, antigens, antibodies, and substrates.

[0809] According to embodiments of the present invention, there is provided a database of probe sequences for the detection, identification, and / or differentiation of bacteria and / or one or more of 16S ribosomal RNA, pathogenicity elements and AMR genes constructed by the method of constructing described herein and comprising one or more of sequence information, length, and origin of each oligonucleotide probe for which sequence information was obtained from the fragments in step b.

[0810] According to embodiments of the present invention, there is provided a probe set comprising oligonucleotides for the detection, identification, and / or differentiation of bacteria and / or one or both of pathogenicity elements and / or AMR genes, constructed by the method of constructing described herein.

[0811] In some embodiments, the probe set comprises approximately less than one million oligonucleotides. According to embodiments of the present invention, there is provided a method for the detection, identification, and / or differentiation of bacteria and / or one or more of 16S ribosomal RNA, pathogenicity elements and / or AMR genes in a sample, comprising: a) isolating nucleic acid from the sample; b) contacting the nucleic acid or derivatives thereof with oligonucleotide probes of any one of the probe sets described herein to form hybridization products; and c) detecting hybridization products between the nucleic acids from the sample and the oligonucleotide probes.

[0812] In some embodiments, the sample is chosen from the group consisting of a biological sample, an environmental sample, and a food sample.

[0813] In some embodiments, the sample is from a human.

[0814] In some embodiments, the subject is selected from the group consisting of domestic vertebrate animals, wild vertebrate animal and invertebrate animals.

[0815] In some embodiments, the method further comprises amplifying and sequencing one or more of the hybridization products from step (c).

[0816] In some embodiments, the method further comprises comparing one or more sample- derived sequences from the hybridization products from step (c) to one or more sequences of known bacteria, AMR genes and / or pathogenicity elements.

[0817] According to embodiments of the present invention, there is provided a kit for the detection, identification, and / or differentiation of bacteria, and / or one or more of 16S ribosomal RNA, pathogenicity elements and / or AMR genes, comprising any one of the databases or probe sets described herein.

[0818] For the foregoing embodiments, each embodiment disclosed herein is contemplated as being applicable to each of the other disclosed embodiments.

[0819] As used herein, all headings are simply for organization and are not intended to limit the disclosure in any manner. The content of any individual section may be equally applicable to all sections. All combinations of the various elements disclosed herein are within the scope of the invention.

[0820] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.

[0821] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment of the invention. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.

[0822] All publications discussed and / or referenced herein are incorporated herein in their entirety.

[0823] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is solely for the purpose of providing a context for the present invention. It is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present invention as it existed before the priority date of each claim of this application.

[0824] Examples are provided below to facilitate a more complete understanding of the invention. The following examples illustrate the exemplary modes of making and practicing the invention. However, the scope of the invention is not limited to specific embodiments disclosed in these Examples, which are for purposes of illustration only.

[0825] EXAMPLES

[0826] Example 1 - Design of probes from sequence databases for detection and differentiation of bacteria, pathogenicity elements and antibiotic resistance

[0827] To identify bacteria and associated virulence and resistance markers by capture sequencing, 120 bp oligonucleotide probes matching species-specific genomic or plasmid- encoded regions of bacteria, AMR genes / elements, and virulence factors were generated. These regions included species-specific genomic marker sequences, 16s rRNA genes, and AMR and virulence-associated genes from genomic and plasmid sequences. The marker sequences are the unique interspersed regions within genomes of a particular bacterial species within its core genomic sequence. These are termed as clade-specific marker genes in Metaphlan4. In the initial design 1333 bacterial species that are reported to be medically important (Table 1) were included. The design also included AMR genes and virulence associated factors from CARD and VFDB databases. The 120-mer oligonucleotides probes were spaced with a 60 nt distance along the target sequences. The resulting probe sets were clustered at 96% to obtain a final set of 988,786 probes. See Table 2.

[0828] Table 2: Databases used in probe design

[0829] Example 2 - In silico validation of marker sequences for bacterial identification - Marker Sequence Validation

[0830] As an example, to show the use of the selected species-specific marker sequences for identifying bacterial species, bacterial species belonging to the same genus were taken and BLAST analysis was performed.

[0831] GenBank Refseq sequences for all bacterial species in Table 3 were downloaded and used for BLASTN analysis (-max_target_seqs 3 -max_hsps 3 -evalue 0.1) against the selected marker sequences for all Helicobacter species; for example, Helicobacter pylori (155 specific marker sequences), Helicobacter heilmannii (200 specific marker sequences), Helicobacter felis (200 specific marker sequences). All the species in Table 3 were evaluated for uniqueness.

[0832] Table 4 shows the number and percentage of our marker sequences that gave a BLAST hit with each of the tested species. For example, of the 155 marker sequences for H. pylori all hit H. pylori strain MT5135'. only one hit in addition to tested Helicobacter species, H. felis (Table 4). In all instances, marker sequences (98-100%) hit the RefSeq genome for the respective Helicobacter species to which they are assigned. The only exception was H. cineadi, which belongs to the H. cinaedi / caniola / magdeburgensis complex of closely related species. In this case 99% of markers showed a BLAST hit with H. magdebur gensis. Accordingly, positive signal can represent multiple species within the complex; thus, further downstream analysis will be required for species designation. Table 3: Bacterial species selected for validation of targeted marker regions

[0833] Table 4: Results of BLASTN analysis for Species-specific regions of Helicobacter species a Only species with one or more BLAST hit are listed. b Helicobacter cinaedi is a member of larger Helicobacter cinaedi / caniola / magdeburgenesis complex

[0834] 5 Example 3 - In silico validation of marker sequences for bacterial identification - Probe set validation

[0835] To validate the selected probes, multiple and single sequence alignments were performed and the number of probes that aligned to Refseq genomes of the genus Helicobacter as well as the specific contributions of each marker region were recorded. Of the total -0.99M probes, 10 11,196 probes mapped to species of the genus Helicobacter. These probes ranged in specificity from 89-100% for their designated species. In the 750-probe set designed for the H. cinaedi / caniola / magdeburgensis complex, 168 and 561 probes mapped to H. cinaedi and H. magdeburgensis, respectively. For H. pylori, additional probes from a virulence factor database (VFDB) that target specific virulence markers of this bacterium were designed. In summary, the

[0836] 15 analysis revealed discrete discriminatory alignment of probes which leads to efficient species level identification even within closely related genomes of same genus such as helicobacter .

[0837] Table 5: Results of clustering analysis from multiple and single sequence analysis

[0838] c additional probes mapped are majorly related to specific virulence factors of H. pylori from

[0839] VFDB

[0840] REFERENCES

[0841] 5 Howell and Davis. 2017. Management of sepsis and septic shock. JAMA 317:847- 848.

[0842] Rhee et al. 2017. Incidence and trends of sepsis in US hospitals using clinical vs claims data, 2009-2014. JAMA 318: 1241-1249.

[0843] Aitor Blanco-Miguez et al. (2022) “Extending and improving metagenomic taxonomic profiling with uncharacterized species with MetaPhlAn 4”, bioRxiv preprint 10 doi.org / 10.1101 / 2022.08.22.504593.

[0844] Alcock BP et al. “CARD 2023: expanded curation, support for machine learning, and resistome prediction at the Comprehensive Antibiotic Resistance Database.” Nucleic Acids Res. 2023 Jan 6;51(Dl):D690-D699.

[0845] Liu B et al. “VFDB 2022: a general classification scheme for bacterial virulence factors.” 15 Nucleic Acids Res. 2022 Jan 7; 50(Dl): D912-D917.

Claims

CLAIMS1. A bacterial sequence capture platform for the detection, identification, and / or differentiation of bacterially-derived sequences in a sample, the platform comprising a plurality of oligonucleotide probes, wherein the plurality comprises at least one oligonucleotide probe which comprises a hybridization portion partially or fully complementary to a portion of a bacterially-derived sequence selected from the group consisting of a bacterial gene sequence, a 16S ribosomal RNA sequence, a pathogenicity element sequence, a virulence factor sequence, and an antimicrobial resistance (AMR) gene sequence, wherein the sequences of the hybridization portions of the oligonucleotide probes cluster at about 90-100% sequence identity, wherein each hybridization portion of an oligonucleotide probe is about 5-300 nucleotides in length, wherein different hybridization portions that each bind a different portion of the same bacterially-derived sequence are tiled across said bacterially-derived sequence and have an inter-probe spacing of about 20-100 nucleotides, and wherein the plurality of oligonucleotide probes of the platform comprises 100,000 to 1,000,000 oligonucleotide probes, preferably less than about 1,000,000 oligonucleotide probes.

2. The platform of claim 1, wherein each hybridization portion of an oligonucleotide probe is about 50-200 nucleotides in length, preferably about 100-150 nucleotides in length, more preferably about 120 nucleotides in length.

3. The platform of claims 1 or 2 wherein the average length of the plurality of hybridization portions of oligonucleotide probes is about 120 nucleotides.

4. The platform of any one of claims 1 -3, wherein different hybridization portions that each bind a different portion of the same bacterially-derived sequence are tiled across said bacterially-derived sequence and have an inter-probe spacing of about 60 nucleotides.

5. The platform of any one of claims 1-4, wherein the sequences of the hybridization portions of the oligonucleotide probes cluster at about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity.

6. The platform of any one of claims 1-5, wherein the plurality of oligonucleotide probes comprises hybridization portions partially or fully complementary to portions of bacterially-derived sequences comprising one or more bacterial gene sequences, one or more 16S ribosomal RNA sequences, one or more pathogenicity element sequences, one or more virulence factor sequences, and / or one or more antimicrobial resistance (AMR) gene sequences.

7. The platform of any one of claims 1-6, wherein the bacterial gene sequence is a speciesspecific or clade-specific gene sequence.

8. The platform of claim 7, wherein the species-specific or clade-specific gene sequences are obtained from Metaphlan4 database.

9. The platform of any one of claims 1-8, wherein the 16S ribosomal RNA sequences are obtained from the SILVA database.

10. The platform of any one of claims 1-9, wherein the virulence factor sequences are obtained from the Virulence Factor Database (VFDB).

11. The platform of any one of claims 1-10, wherein the AMR genes are obtained from the Comprehensive Antibiotic Resistance Database (CARD).

12. The platform of any one of claims 1-11, wherein each bacterially-derived sequence comprises a portion that is about 50-300 nucleotides in length and is partially or fully complementary to a hybridization portion of an oligonucleotide probe.

13. The platform of any one of claims 1 -12, wherein each hybridization portion is at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% complementary to a portion of a bacterially-derived sequence.

14. The platform of any one of claims 1-13, wherein the plurality of oligonucleotide probes comprises at least one oligonucleotide probe which comprises a hybridization portion partially or fully complementary to a portion of a bacterially-derived sequence from a bacterial species listed in Table 1.

15. The platform of any one of claims 1-14, wherein every bacterial species listed in Table 1 comprises a sequence, preferably a unique sequence relative to any other bacterial species listed in Table 1, that is partially or fully complementary to a hybridization portion of a oligonucleotide probe of the plurality of the platform.

16. The platform of any one of claims 1-15, wherein each oligonucleotide probe comprises a capture portion.

17. The platform of claim 16, wherein the capture portion is selected from the group consisting of biotin, digoxygenin, a ligand, a small organic molecule, a small inorganic molecule, an aptamer, an antigen, an antibody, and a substrate.

18. The platform of any one of claims 1-17, wherein each oligonucleotide probe is biotinylated.

19. The platform of any one of claims 1-18, and means for capturing, isolating, and / or purifying the plurality of oligonucleotide probes from a mixture of other nucleic acid molecules.

20. The platform of any one of claims 1-19, wherein the oligonucleotide probes comprise DNA, RNA, bridged nucleic acids, locked nucleic acids, and / or peptide nucleic acids.

21. The platform of any one of claims 1-20, wherein the sample is a biological sample or an environmental sample.

22. The platform of any one of claims 1-21, wherein the sample is selected from the group consisting of saliva, mucus, a nasopharyngeal swab, serum, plasma, blood, urine, feces, cerebrospinal fluid, a bodily fluid, cultured cells, an organ tissue, and biopsied tissue.

23. The platform of any one of claims 1-22, wherein the sample is selected from the group consisting of an aqueous sample, a liquid sample, water, wastewater, sewage, grey water, blackwater, freshwater, liquid waste, seawater, drinking water, air, a gaseous sample, soil, a food sample, culture medium, and a swab of an inanimate surface or object.

24. The platform of any one of claims 1-23, wherein the sample is obtained from a sewage system, a drainage system, a plumbing system, or a water treatment facility.

25. The platform of any one of claims 1-23, wherein the sample is obtained from a human subject.

26. A method of screening a sample for bacterially-derived sequences, the method comprising: a) exposing the sample, or nucleic acids isolated, amplified, and / or enriched from the sample, to the bacterial sequence capture platform of any one of claims 1-25 to form one or more hybridization products, wherein each hybridization product comprises a nucleic acid of the sample and an oligonucleotide probe of the platform; b) capturing the one or more hybridization products; and c) identifying the presence of one or more bacterially-derived sequences in the sample based on the sequences of the one or more captured hybridization products; thereby screening the sample for bacterially-derived sequences.

27. The method of claim 26, wherein nucleic acids in the sample are isolated and / or enriched prior to the exposing in step (a).

28. The method of claim 26 or 27, wherein the sample is a biological sample or an environmental sample.

29. The method of any one of claims 26-28, wherein the sample is selected from the group consisting of saliva, mucus, a nasopharyngeal swab, serum, plasma, blood, urine, feces, cerebrospinal fluid, a bodily fluid, cultured cells, an organ tissue, and biopsied tissue.

30. The method of any one of claims 26-29, wherein the sample is selected from the group consisting of an aqueous sample, a liquid sample, water, wastewater, sewage, grey water, blackwater, freshwater, liquid waste, seawater, drinking water, air, a gaseous sample, soil, a food sample, culture medium, and a swab of an inanimate surface or object.

31. The method of any one of claims 26-30, wherein the sample is obtained from a sewage system, a drainage system, a plumbing system, or a water treatment facility.

32. The method of any one of claims 26-29, wherein the sample is obtained from a human subject.

33. The method of any one of claims 26-32, the method further comprising: sequencing one or more detected hybridization products; comparing the nucleotide sequence of the one or more hybridization products to nucleotide sequences of known bacterially-derived sequences; and identifying and / or differentiating one or more bacterially-derived sequences in the sample based on sequence identity of the hybridization product to the nucleotide sequences of known bacterially-derived sequences.

34. A kit comprising the bacterial sequence capture platform of any one of claims 1-25 and instructions for using the platform.

35. The kit of claim 34, further comprising a sample, wherein the platform is used for the detection, identification, and / or differentiation of bacterially-derived sequences in the sample.

36. The kit of claim 35, wherein the sample is a biological sample or an environmental sample.

37. The kit of claim 35 or 36, wherein the sample is a liquid sample or an aqueous sample.

38. The kit of any one of claims 35-37, wherein the sample is selected from the group consisting of a water sample, wastewater, sewage, grey water, blackwater, freshwater,liquid waste, seawater, drinking water, air, a gaseous sample, soil, a food sample, culture medium, and a swab of an inanimate surface or object.

39. The kit of any one of claims 35-38, wherein the sample is a wastewater sample or a sewage sample.

40. The kit of any one of claims 35-39, wherein the sample is a wastewater sample.

41. The kit of any one of claims 35-39, wherein the sample is a sewage sample.

42. The kit of any one of claims 35-41, wherein the sample is obtained from a sewage system, a drainage system, a plumbing system, or a water treatment facility.

43. The kit of any one of claims 35-37, wherein the sample is selected from the group consisting of saliva, mucus, a nasopharyngeal swab, serum, plasma, blood, urine, feces, cerebrospinal fluid, a bodily fluid, cultured cells, an organ tissue, and biopsied tissue.

44. The kit of any one of claims 35-43, wherein the sample comprises nucleic acids.

45. A method for designing and / or constructing a database of probe sequences or a probe set comprising oligonucleotide probes for the detection, identification, and / or differentiation of bacteria and / or one or more of 16S ribosomal RNA, pathogenicity elements and / or AMR genes, comprising: a) obtaining i) one or more species-specific or clade-specific marker gene sequences; or ii) one or more 16S ribosomal RNA sequences; or iii) one or more virulence factor sequences; or iv) one or more AMR gene sequences; or v) any combination of (i), (ii), (iii), and (iv); and b) breaking the sequences obtained in step a. into fragments, wherein the fragments are the basis of the probes and are designed to be of a length, and spaced at a distance across the target sequences, such that the total number of probe sequences in the database or probes in the probe set corresponds to a desired range or number.

46. The method of claim 45, wherein the species-specific or clade-specific gene sequences are obtained from Metaphlan4 database.

47. The method of claim 45, wherein the 16S ribosomal RNA sequences are obtained from the SILVA database.

48. The method of claim 45, wherein the virulence factor sequences are obtained from the Virulence Factor Database (VFDB).

49. The method of claim 45, wherein the AMR genes are obtained from the Comprehensive Antibiotic Resistance Database (CARD).

50. The method of claim 45, wherein the desired range or number is less than one million.

51. The method of claim 45, comprising a further step of synthesizing one or more of the oligonucleotide probes for which the sequence information was obtained in step b.

52. The method of claim 51, wherein the oligonucleotide probes are chosen from the group consisting of DNA, RNA, Bridged Nucleic Acids, Locked Nucleic Acids, and Peptide Nucleic Acids.

53. The method of claim 51, wherein the one or more oligonucleotide probes are synthesized on a cleavable microarray.

54. The method of claim 51, wherein the oligonucleotides are modified to comprise a composition for binding to a solid support, chosen from the group consisting of biotin, digoxygenin, ligands, small organic molecules, small inorganic molecules, aptamers, antigens, antibodies, and substrates.

55. A database of probe sequences for the detection, identification, and / or differentiation of bacteria and / or one or more of 16S ribosomal RNA, pathogenicity elements and AMR genes constructed by the method of claim 45 and comprising one or more of sequence information, length, and origin of each oligonucleotide probe for which sequence information was obtained from the fragments in step b.

56. A probe set comprising oligonucleotides for the detection, identification, and / or differentiation of bacteria and / or one or both of pathogenicity elements and / or AMR genes, constructed by the method of claim 45.

57. The probe set of claim 56, comprising approximately less than one million oligonucleotides.

58. A method for the detection, identification, and / or differentiation of bacteria and / or one or more of 16S ribosomal RNA, pathogenicity elements and / or AMR genes in a sample, comprising: a) isolating nucleic acid from the sample; b) contacting the nucleic acid or derivatives thereof with oligonucleotide probes of the probe set of claim 56 to form hybridization products; and c) detecting hybridization products between the nucleic acids from the sample and the oligonucleotide probes.

59. The method of claim 58, wherein the sample is chosen from the group consisting of a biological sample, an environmental sample, and a food sample.

60. The method of claim 58, wherein the sample is from a human.

61. The method of claim 58, wherein the subject is selected from the group consisting of domestic vertebrate animals, wild vertebrate animal and invertebrate animals.

62. The method of claim 58, further comprising amplifying and sequencing one or more of the hybridization products from step (c).

63. The method of claim 62, further comprising comparing one or more sample-derived sequences from the hybridization products from step (c) to one or more sequences of known bacteria, AMR genes and / or pathogenicity elements.

64. The method of claim 58, further comprising amplifying and sequencing one or more nucleic acids or derivatives thereof from the sample which do not form hybridization products with any of the probes in the probe set.

65. The method of claim 64, further comprising comparing one or more sequences of nucleic acids from the sample which do not form hybridization products to one or more sequences of known bacteria, 16S ribosomal RNA, AMR genes and / or pathogenicity elements.

66. A kit for the detection, identification, and / or differentiation of bacteria, and / or one or more of 16S ribosomal RNA, pathogenicity elements and / or AMR genes, comprising the database of claim 55 or the probe set of claim 56.