Simultaneous pathogen detection methods and systems

Metagenomic sequencing is used to derive pathogen profiles from complex microbial communities, addressing the challenge of distinguishing between conditions with similar symptoms, thereby enhancing clinical decision-making and reducing unnecessary procedures.

WO2025179352A1PCT designated stage Publication Date: 2025-09-04MICROBA IP PTY LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/AU2025/050187
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing methods struggle to distinguish between conditions with indistinguishable symptoms, leading to prolonged clinical journeys for patients with chronic diarrhea, often requiring unnecessary procedures like endoscopy and colonoscopy, and there is a need for methods that can inform clinical decisions and diagnoses based on complex microbial communities.

Method used

A method involving metagenomic sequencing to derive a pathogen profile, including the presence or absence of pathogens, AMR genes, and SNPs, to determine indicators for clinical decisions and diagnoses, using metagenomic sequencing information to analyze complex microbial communities, particularly from fecal samples, to identify pathogens and inform treatment or diagnosis.

Benefits of technology

Enables accurate differentiation between conditions with similar symptoms, facilitating informed clinical decisions and diagnoses, reducing unnecessary procedures and improving patient management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050187_04092025_PF_FP_ABST
    Figure AU2025050187_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates generally to methods and systems of determining an indicator to inform diagnosis and / or clinical decisions with respect to a subject. More specifically, the methods related to the diagnosis and / or clinical decision is generally with respect to a gastrointestinal infection or symptoms.
Need to check novelty before this filing date? Find Prior Art

Description

JAWS Ref: 750704PCT TITLE OF THE INVENTION “SIMULTANEOUS PATHOGEN DETECTION METHODS AND SYSTEMS” FIELD OF THE INVENTION

[0001] This invention relates generally to the field of pathogen detection. More particularly, the present invention relates to methods of detecting and / or identifying the presence, absence or abundance of a disease-causing infectious marker in a clinical sample, and systems for performing said methods. BACKGROUND OF THE INVENTION

[0002] Samples may be analysed for various purposes, including detecting the presence, absence, or amount of target (such as a microorganism) in a sample. Analysis of a sample that contains one or more nucleic acid molecules may involve sequencing the nucleic acid molecules, or portions or derivatives thereof.

[0003] Nucleic acid sequencing generally to facilitate the identification of contaminants and / or species of potential interest within a sample. For example, sequencing may be used to identify a microorganism or pathogen within a symbol.

[0004] It is estimated between 6-7% of the western population suffer from chronic diarrhea. Generally, initial screening of these patients includes a faecal pathogen test that involves culture, microscopy and / or PCR of between 6 – 12 common pathogens to rule out infection. If initial screening tests are negative, the clinical journey for these patients can be long. Eventually, many of these patients will be referred to a gastroenterologist for an endoscopy and colonoscopy.

[0005] There is a requirement in the art for new methods of distinguishing between conditions with indistinguishable symptoms. Often these symptoms are broadly related to the presence or absence of a known disease or condition. There is also a requirement for these methods to facilitate or provide information to make clinical decisions. SUMMARY OF THE INVENTION

[0006] The present invention is predicated in part on the inventors identifying a method of distinguishing between conditions known to have similar symptoms and issues to the host.

[0007] Accordingly, in one aspect the present invention provides methods of determining an indicator to inform a clinical decision, the method comprising: receiving metagenomic sequencing information of a complex microbial community; analysing the metagenomic sequencing information to derive a pathogen profile;JAWS Ref: 750704PCT wherein the pathogen profile comprises: a determination of the presence or absence of pathogens from a plurality of classes; a determination of the presence or absence of AMR genes; and single nucleotide polymorphisms. and determining an indicator on the basis of the pathogen profile, wherein the indicator may be used to inform a clinical decision.

[0008] In some embodiments, the complex microbial community is derived from a sample that comprises over 100 microbial species.

[0009] In some embodiments, the complex microbial community is derived from a fecal sample. In embodiments of this type, the microbial community derived from the fecal sample is representative of the gut microbiome. In some embodiments, the plurality of pathogen classes include bacteria, virus, fungus, helminth, protozoan, microsporidia, and / or invertebrate.

[0010] In some embodiments, the determination of the presence or absence of a pathogen is performed by (i) aligning sequence reads from the metagenomic data to sequences present to a reference genome; and (ii) for genomes that are considered present by the alignment analysis of step (i), determining whether a diagnostic loci is present in the aligned sequence reads.

[0011] In some embodiments, the clinical decision is the initiating of a treatment regimen for a particular infection or disease condition. In other embodiments, the clinical decision is the modification of a treatment regimen (e.g., modification of the dose of treatment, or changing the treatment).

[0012] In some embodiments, the clinical decision is the cessation of a treatment regimen.

[0013] Accordingly, another aspect the present invention provides methods of determining an indicator to inform a diagnosis, the method comprising: receiving metagenomic sequencing information of a complex microbial community; analysing the metagenomic sequencing information to derive a pathogen profile; wherein the pathogen profile comprises: a determination of the presence or absence of pathogens from a plurality of classes; a determination of the presence or absence of AMR genes; and single nucleotide polymorphisms (SNPs). and determining an indicator on the basis of the pathogen profile, wherein the indicator may be used to inform a diagnosis.JAWS Ref: 750704PCT

[0014] Suitably, the SNPs are selected from virulence factors, antibiotic resistance traits, metabolic traits.

[0015] In some embodiments, the condition being diagnosed is inflammatory bowel disease (IBD). In some of the same embodiments and some other embodiments, the condition being diagnosed is a pathogenic infection.

[0016] In some embodiments of this type, the pathogenic infection is selected from bacteria, archaea, fungi, viruses, phages, invertebrates, protozoa, microsporidia, helminths, parasites. In some embodiments of this type, the pathogenic infection comprises Aeromonas caviae, Aeromonas veronii, Campylobacter concisus, Enteropathogenic Escherichia coli (EPEC), Giardia intestinalis, H. pylori and Tropheryma whipplei.

[0017] In some embodiments, the subject is suffering with chronic diarrhea.

[0018] In some embodiments, the pathogen profile includes the presence of one or more virulence factor genes (for example, eae and / or bfpA). BRIEF DESCRIPTION OF THE FIGURES

[0019] The following figures form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The disclosure may be better understood by reference to one or more of these figures in combination with the detailed description of specific embodiments presented herein.

[0020] Figure 1 provides a high-level overview of the bioinformatics components of the present invention.

[0021] Figure 1 provides a graphical workflow for building the prokaryotic genome reference database. Constructing the database consists of seven steps: (i) downloading human gut associated genomes; (ii) assigning genomes to species delineated by genomic similarity; (iii) grouping novel genomes into de novo species clusters; (iv) assigning taxonomic information to species; (v) assigning specific taxonomic information to novel species; (vi) selecting MProkDB representative genomes for each species; and (vii) constructing read mapping indices used to map sequencing reads to the MProkDB genomes.

[0022] Figure 2 provides a graphical representation of masking putative contamination in a eukaryotic genome. Overlapping pseudoreads are generated across each of the contigs in the eukaryotic genome being decontaminated. Pseudoreads are classified using Kraken 2 against a custom reference database consisting of prokaryotic, viral, vector and common host genomes. Classified reads represent putative contamination and are masked from the contig if they cover a sufficiently large region. Sufficiently short unmasked regions are bridged when forming putative regions to mask.

[0023] Figure 3 provides a workflow for building the viral genomes reference database. Constructing the database consists of five steps: (i) downloading viral genomes from large metagenomic studies, (ii) using CheckV to calculate genome quality statistics, (iii)JAWS Ref: 750704PCT performing quality-control of the viral genomes, (iv) assigning genomes passing QC to species- level vOTUs, and (v) constructing a BWA index used to map sequencing reads to MViralDB genomes.

[0024] Figure 5 provides a workflow for estimating the relative abundance of species with the Microba Community Profiler (MCP). Criteria are for identification of prokaryotic and eukaryotic species.

[0025] Figure 6 provides a workflow for creating databases and mock communities used for parameter tuning and validating the performance of MCP on viral species. The MViralDB genomes were divided into a “testing” and “holdout” set in order to facilitate the generation of mock communities with genomes not in the reference DB. Initial prokaryotic mock communities were generated by considering Microba’s extensive proprietary metagenomic dataset (i.e., MCP species profiles of >11K human gut samples) in order to generate mock communities with similar numbers of species and relative abundance distributions. Strain diversity was simulated by including multiple genomes per selected species. Prophage were then identified in these prokaryotic-exclusive mock communities using a BLAST-based procedure (Appendix B). These mock communities were then extended with viral species in the “holdout” set of MViralDB ensuring that species identified as putative prophage in a mock were not selected. Viral strain diversity was simulated by selecting multiple genomes per viral species. Parameter tuning of MCP was performed by mapping the in silico read pairs to the MViralDB “testing database” and inferring MCP species profiles under a range of parameter settings.

[0026] Figure 7 provides the recall rate of MCP with vOTUs below a given depth of coverage filtered from consideration. The recall rate increases with an increasing depth of coverage requirement as lower abundance vOTUs are harder to identify.

[0027] Figure 8 provides the mapping and MCP workflow for producing viral species profiles. The viral profile is determined by applying MCP to viral mappings in isolation to ensure prophage can be robustly identified. All other profiles are produced by considering the mappings to all reference DBs, including the viral DB, as this has been rigorously evaluated and reads assigned to viruses (prophage or otherwise) reduce the number of unclassified reads which in turn produce more accurate and interpretable profiles.

[0028] Figure 9 provides a workflow to construct a database of diagnostic regions for a MetaPanel target species.

[0029] Figure 10 provides a workflow for determining the presence of low- abundance pathogens in MCPdx.

[0030] Figure 11 provides a comparison of pathogens, virulence factors, and antimicrobial resistance (AMR) genes detected using the methods described herein in patients with Crohn’s Disease (CD) and ulcerative colitis (UC), stratified by disease state (remission vs. active disease). Each row represents a detected factor, with values indicating the number and percentage of patients in whom it was identified. Columns denote remission (left) and activeJAWS Ref: 750704PCT disease (right) for each condition. The colour gradient reflects detection frequency, with darker shades indicating higher prevalence.

[0031] Figure 12 shows Pathogen Frequency Data. (A) The frequency (i.e. number of positive tests) of detected pathogens in the study population. (B) The distribution of the number of detected pathogens per test.

[0032] Figure 13 shows the number of co-infections detected. Upset plot describing the frequency of pathogen combinations in all individuals where >1 pathogen was detected.

[0033] Figure 14. AMR Gene Family Frequency. (A) The frequency (i.e. number of positive tests) of detected AMR genes in the study population. (B) The distribution of the number of detected AMR genes per test. DETAILED DESCRIPTION OF THE INVENTION 1. Definitions

[0034] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, preferred methods and materials are described. For the purposes of the present invention, the following terms are defined below.

[0035] The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.

[0036] The term “about” as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field. Reference to “about” a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se.

[0037] As used herein, the term “administering,” refers to the placement of an agent (e.g., bacteria) as disclosed herein into a subject by a method or route which results in at least partial delivery of the agent at the desired site. Compositions comprising the compounds disclosed herein can be administered by any appropriate route which results in an effective biological activity or therapeutic effect in the subject. In some embodiments, administration comprises physical human activity (e.g., an injection, act of ingestion, an act of application, and / or manipulation of a delivery device or machine). Such activity can be performed (e.g., by a medical professional and / or the subject being treated).

[0038] Specifically, as used herein “administer” and “administration” encompasses embodiments in which one person directs another to consume live bacteria, dead bacteria, spent mediums derived from bacteria, cell pellets of bacteria, purified metabolitesJAWS Ref: 750704PCT produced by bacteria, purified proteins produced by bacteria, prebiotics, small molecules, or combinations thereof in a certain manner and / or for a certain purpose independently of or in variance to any instructions received from a second person. Non-limiting examples of embodiments include the situation in which one person directs another to consume live bacteria, dead bacteria, spent mediums derived from bacteria, cell pellets of bacteria, purified metabolites produced by bacteria, purified proteins produced by bacteria, prebiotics, small molecules, or combinations thereof in a certain manner and / or for a certain purpose include when a physician prescribes a course of conduct and / or treatment to a patient, when a parent commands a minor user (such as a child) to consume such a product, when a trainer advises a user (such as an athlete) to follow a particular course of conduct and / or treatment, or when a manufacturer, distributer, or marketer recommends conditions of use to an end user, for example through advertisements or labeling on packing or on other materials provided in association with the sale or marketing of a product. In some embodiments, the disclosed compositions can be administered orally, intravenously, intramuscularly, intrathecally, subcutaneously, sublingually, buccally, rectally, vaginally, by the ocular route, by the optic route, nasally, via inhalation, by nebulization, cutaneously, transdermally, or combinations thereof, and formulated for delivery with a pharmaceutically acceptable excipient, carrier or diluent. Of note, although the disclosed compositions encompass multiple formulations and modes of delivery for treatments to ameliorate dysbiosis and its sequelae, it should be noted that live biotherapeutic products such as probiotics are not typically administered intravenously, intramuscularly, or intraperitoneally. These modes of delivery would likely be reserved for small-molecule products of bacterial metabolism.

[0039] The “amount” or “level” of a biomarker is a detectable level in a sample. These can be measured by methods known to one skilled in the art and also disclosed herein. The expression level or amount of biomarker assessed can be used to determine the response to treatment.

[0040] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (or).

[0041] As used herein, the terms “antimicrobial resistance marker” or “AMR marker” refers to a measurable and / or detectable marker indicating that a respective microorganism has antimicrobial resistance. As used herein, the term “antimicrobial resistance” refers to a property of or exhibited by a respective microorganism, such that the respective microorganism is resistant to one or more antimicrobial interventions (e.g., where an effect of an antimicrobial intervention is attenuated, obstructed, or negated). As used herein, the term “antimicrobial susceptibility” refers to a property of or exhibited by a respective microorganism, such that the respective microorganism is susceptible to one or more antimicrobial interventions (e.g., where an effect of an antimicrobial intervention serves to kill, diminish, slow or prevent growth in one or a population of microorganisms).JAWS Ref: 750704PCT

[0042] In some embodiments, antimicrobial resistance is conferred by a genetic sequence (e.g., an antimicrobial resistance gene). In some embodiments, the antimicrobial resistance marker is a genetic marker (e.g., a nucleic acid sequence for the antimicrobial resistance gene indicating that the gene comprises a mutation that confers resistance). In some embodiments, the antimicrobial resistance marker is a restriction fragment length polymorphism (RFLP), a random amplified polymorphic DNA (RAPD), an amplified fragment length polymorphism (AFLP), a variable number tandem repeat (VNTR), an oligonucleotide polymorphism (OP), a single nucleotide polymorphism (SNP), an allele specific associated primer (ASAP), an inverse sequence-tagged repeat (ISTR), an inter-retrotransposon amplified polymorphism (IRAP), and / or a simple sequence repeat (SSR. or microsatellite). In some embodiments, an antimicrobial resistance marker is detected based on a mapping (e.g., an alignment) of one or more nucleotide sequences to a reference sequence (e.g., a reference genome). In some embodiments, an antimicrobial resistance marker is an amino acid sequence and / or an amino acid residue. In some embodiments, an antimicrobial resistance marker is a biochemical marker.

[0043] In some embodiments, an antimicrobial resistance marker indicates that a respective microorganism is resistant to one or more interventions for a corresponding type of microorganism (e.g., antibacterial resistance, antiprotozoal resistance, antifungal resistance, antihelminthic resistance, and / or antiviral resistance). For example, in some embodiments, an antimicrobial intervention is a drug that targets a specific gene in a respective microorganism, and a mutation in the gene confers resistance to the microorganism. In some such embodiments, an antimicrobial resistance marker can be a genetic marker for the target gene that indicates a resistance to the antimicrobial drug.

[0044] As used herein, the term “antimicrobial resistance status” refers to an indication of a presence or absence of an antimicrobial resistance marker. For example, the term antimicrobial resistance status or AMR status will be understood to include an indication that a respective biological and / or non-biological sample and / or a microorganism detected in a sample has either antimicrobial resistance or antimicrobial susceptibility. In some embodiments, an antimicrobial resistance status includes an indication that an antimicrobial resistance marker is present (e.g., has been detected) in the respective sample and / or microorganism. In some embodiments, an antimicrobial resistance status includes an indication of any one or more features for the respective antimicrobial resistance marker (e.g., gene identifier, gene name, intervention (drug) information, intervention (drug) classes, associated organisms, gene families, and / or resistance mechanisms).

[0045] In some embodiments, an antimicrobial resistance marker is associated with one or more microorganisms in a plurality of microorganisms (e.g., where the respective microorganism has been reported or annotated as expressing the respective antimicrobial resistance marker). In some embodiments, a first antimicrobial resistance marker is associated with a first respective microorganism in a plurality of microorganisms, and a second antimicrobialJAWS Ref: 750704PCT resistance marker is associated with a second respective microorganism, other than the first microorganism, in the plurality of microorganisms.

[0046] Examples of antimicrobial resistance markers (e.g., genes and / or amino acid residues) include, but are not limited to, the antimicrobial resistance markers CTX-MG1, CTX-MG2, CTX-M G5, CTX-M G8, CTX-M G9, armA, blaACC, blaACT, blaAIM, blaCMY G1, blaCMY G2 blaDHA, blaDIM, blaFOX, blaGES, blaGIM, blaIMI, blaIMP, blaKPC, blaMIR, blaNDM, blaOXA-48 like, blaSHV, blaSIM, blaTEM, blaVEB, blaVIM, fosA5 fam, fosA7 fam, fosA8 fam, fosA PA1129, fosA gen, fos A3 A4, fos A A2, mcr1, mcr10, mcr2, mcr3, mcr4, mcr5, mcr6, mcr7, mcr8, mecA, mecC, qnrA, qnrB, rmtA, rmtB, rmtC, rmtD, rmtE, rmtF, rmtG, vanA, and vanB.

[0047] Further examples of antimicrobial resistance markers can be found, for example, in Capela et al., 2019, “An Overview of Drug Resistance in Protozoal Diseases.” Int J. Mol. Sci. 2022: 5748; Beech et al., 2011, “Anthelmintic resistance: markers for resistance, or susceptibility?” Parasitology 138(2): 160-174; and Toledu-Rueda et al., 2018, “Antiviral resistance markers in influenza virus sequences in Mexico, 2000-2017,” Infect. Drug Resist.11: 1751-1756; each of which is hereby incorporated herein by reference in its entirety.

[0048] In some embodiments, the term “antimicrobial resistance marker” will be understood to include any one or more genes, amino acid sequences amino acid residues, genetic markers, and / or biochemical markers selected from a database. In some embodiments, an antimicrobial resistance marker is selected from a database that is one or more of locally maintained, proprietary, and / or open access. In some embodiments, an antimicrobial resistance marker is selected from a national and / or international database.

[0049] Examples of such databases include, but are not limited to, the National Database of Antibiotic Resistant Organisms (NDARO), the Comprehensive Antibiotic Resistance Database (CARD), ResFinder, PointFinder, ARG-ANNOT, ARGs-OSP, PlasmoDB, the nMycology Antifungal Resistance Database (MARDy), DBDiaSNP, the HIV Drug Resistance Database, the Virus Pathogen Resource (ViPR), and / or any of the databases used for selecting one or more microorganisms, as disclosed above (see, for example, McArthur et al., 2013, “The Comprehensive Antibiotic Resistance Database,” Antimicrob. Ag. Chemother., 57(7) 3348-3357; Zankari et al., 2017, “PointFinder: a novel web tool for WGS-based detection of antimicrobial resistance associated with chromosomal point mutations in bacterial pathogens,” Antimicrob. Chemother., 72 (10) 2764-2768; Gupta et al., 2013, “ARG-ANNOT, a New Bioinfonnatic Tool To Discover Antibiotic Resistance Genes in Bacterial Genomes.” Antimicrob. Ag. Chemother.58(1) 212-220; Zhang et al., “ARGs-OSP: online searching platform for antibiotic resistance genes distribution in metagenomic database and bacterial whole genome database,” bioRxiv 337675; Nash et al., 2018, “MARDy: Mycology Antifungal Resistance Database,” 34 (18) 3233-3234; and Mehla and Ramana, 2015, “DBDiaSNP: An Open-Source Knowledgebase of Genetic Polymorphisms and Resistance Genes Related to Diarrheal Pathogens,” OMICS 19 (6) 354-360; each of which is hereby incorporated herein by reference in its entirety.JAWS Ref: 750704PCT

[0050] Throughout this specification, unless the context requires otherwise, the words “comprise”, “comprises” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements. Thus, use of the term “comprising” and the like indicates that the listed elements are required or mandatory, but that other elements are optional and may or may not be present. By “consisting of” is meant including, and limited to, whatever follows the phrase “consisting of”. Thus, the phrase “consisting of” indicates that the listed elements are required or mandatory, and that no other elements may be present. By “consisting essentially of” is meant including any elements listed after the phrase, and limited to other elements that do not interfere with or contribute to the activity or action specified in the disclosure for the listed elements. Thus, the phrase “consisting essentially of” indicates that the listed elements are required or mandatory, but that other elements are optional and may or may not be present depending upon whether or not they affect the activity or action of the listed elements.

[0051] The terms “decrease”, “reduced”, “reduction”, “inhibit”, “suppress”, “attenuate” and the like are all used herein to mean a decrease by a statistically significant amount. In some embodiments, these terms typically mean a decrease by at least 10% as compared to a reference level (e.g., the absence of a given treatment or agent) and can include, for example, a decrease by at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or more. As used herein “reduction”, “suppression”, and “inhibition” does not necessitate a complete inhibition or reduction as compared to a reference level. “Complete inhibition” and the like is a 100% inhibition as compared to a reference level. A decrease can be preferably down to a level accepted as within the range of normal (e.g., for an individual without a given disorder).

[0052] The terms “increased”, “increase”, enhance”, or “activate” are all used herein to mean an increase by a statistically significant amount. In some embodiments, the terms “increased”, “increase”, “enhance”, or “activate” can mean an increase of at least 10% as compared to a reference level (e.g., the absence of a given treatment or agent) and can include, for example, of at least about 10% as compared to a reference level, for example an increase of at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or up to and including a 100% increase or any increase between 10-100% as compared to a reference level or at least about a 2-fold, or at least about a 3-fold, or at least about a 3-fold, or at least about a 4-fold, or at least about a 5-fold, or at least about a 10-fold increase, or any increase between 2-fold and 10- fold or greater as compared to a reference level. In the context of a marker or symptom, an “increase” is a statistically significant increase in such level.JAWS Ref: 750704PCT

[0053] As used herein, the terms “genome” or “reference genome” refer to any particular known, sequenced or characterized genome, whether partial or complete, of any organism or virus that may be used to reference identified sequences from a subject. Example reference genomes used for human subjects as well as many other organisms are provided in the online genome browser hosted by the National Center for Biotechnology Information (“NCBI”) or the University of California, Santa Cruz (UCSC). As used herein, a reference sequence or reference genome often is an assembled or partially assembled genomic sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genomic sequence from one or more human individuals. In some embodiments, a reference genome is an assembled or partially assembled genomic sequence from one or more microorganisms of the same species. The reference genome can be viewed as a representative example of a species’ set of genes. In some embodiments, a reference genome comprises sequences assigned to chromosomes. Exemplary human reference genomes include but are not limited to NCBI build 34 (UCSC equivalent: hgl6), NCBI build 35 (UCSC equivalent: hgl7), NCBI Build 6.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hgl9), and GRCh38 (UCSC equivalent: hg38). Exemplar bacterial reference genomes include but are not limited to the NCBI reference genome for Escherichia coli str. K-12 substr. MG1655 (GCF_000005845.2), the NCBI representative genome for Giardia intestinalis WB C6 (GCF_000002435.2), and the NCBI genome for human gammaherpesvirus 4 B95-8 (Epstein-Barr virus; GCF_002402265.1).

[0054] In some embodiments, a genome is a complete genome. In some embodiments, a genome is an incomplete genome. For example, in some embodiments, an incomplete genome is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the complete genome.

[0055] In some embodiments, a complete or incomplete genome is less than 1 megabase pairs (Mb), less than 0.5 Mb, less than 0.4 Mb, less than 0.3 Mb, less than 0.2 Mb, or less than 0.1 Mb. In some embodiments, a complete or incomplete genome is at least 1 Mb, at least 2 Mb, at least 3 Mb, at least 4 Mb, at least 5 Mb, at least 6 Mb, at least 7 Mb, at least 8 Mb. at least 9 Mb, at least 10 Mb. at least 15 Mb, at least 20 Mb, at least 25 Mb, at least 30 Mb, at least 35 Mb, at least 40 Mb, at least 45 Mb, at least 50 Mb, at least 100 Mb, at least 200 Mb, at least 500 Mb, at least 1,000 Mb, at least 2,000 Mb, at least 3,000 Mb, at least 4,000 Mb, at least 5,000 Mb, at least 10 gigabase pairs (Gb), at least 20 Gb, or at least 50 Gb.

[0056] In some embodiments, a complete or incomplete genome spans a region of a reference genome comprising at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at leastJAWS Ref: 750704PCT 35, at least 40, at least 45, at least 50, at least 100, at least 200, at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 10,000, or at least 50,000 genes. In some embodiments, a complete or incomplete genome spans a region of a reference genome comprising between 1 and 10, between 10 and 50, between 50 and 100, between 100 and 500, between 500 and 1000, between 1000 and 2000, between 2000 and 5000, between 5000 and 10,000, between 10,000 and 50,000, or more than 50,000 genes.

[0057] In some embodiments, a complete or incomplete genome spans a region of a reference genome comprising at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 100, at least 200, or at least 500 antimicrobial resistance markers. In some embodiments, a complete or incomplete genome spans a region of a reference genome comprising between 1 and 10, between 10 and 50, between 50 and 100, or more than 100 antimicrobial resistance markers.

[0058] In some embodiments, a complete or incomplete genome is obtained from one or more nucleotide sequence databases and / or microorganism databases, including but not limited to GTDB, NCBI, SRA, BLAST, EMBL-EBI, GenBank, Ensembl, EuPathDB, The Human Microbiome Project, Pathogen Portal, RDP, SILVA, GREENGENES, EBI Metagenomics, EcoCyc, PATRIC, TBDB, PlasmoDB, the Microbial Genome Database (MBGD), and / or the Microbial Rosetta Stone Database. See, for example, Zhulin, 2015, “Databases for Microbiologists,” J. Bacteriol. 197:2458-2467; Uchiyama et al., 2019, “MBGD update 2018: microbial genome database based on hierarchical orthology relations covering closely related and distantly related comparisons,” Nuc. Acids Res., 47 (DI), D382—D389; and Ecker et al., 2005, “The Microbial Rosetta Stone Database: A compilation of global and emerging infectious microorganisms and bioterrorist threat agents,” BMC Microbiology 5, 19; each of which is hereby incorporated by reference herein in its entirety.

[0059] As used herein, the term “gut” is understood to refer to the human gastrointestinal tract, also known as the alimentary canal. The gut includes the mouth, pharynx, oesophagus, stomach, small intestine (duodenum, jejenum, ileum), large intestines (cecum and colon) and rectum. While the entire alimentary canal can be colonized by varying species of microbes, the majority of the gut microbiome, in terms of both numbers of species and biomass, resides in the intestines (small and large).

[0060] As used herein the term “k-mer” refers to a subsequence of a given length k within a longer sequence, where k is a positive integer of two or greater. In some embodiments, k is between 3 and 100. In some embodiments, k is between 4 and 50. In some embodiments, k is between 5 and 40. In one example, the sequence “AGCTCT” is divided into the three-nucleotide subsequences “AGC”, “GCT”, “CTC”, and “TCT”. In this example, each of these subsequences is a k-mer, where k = 3. K-mers may be overlapping or non-overlapping. in some embodiments, k-mers overlap each other by one residue. K-mers and their use in sequence alignment and mapping are further described in Stokes and Glick, 2006, “MICA: desktop softwareJAWS Ref: 750704PCT for comprehensive searching of DNA databases”, BMC Bioinformatics 7:427; Kalafus, 2004, “Pash: Efficient Genome-Scale Sequence Anchoring by Positional Hashing”, Genome Research 14:672-678; and Mann and Noble, “Efficient identification of DNA hybridization partners in a sequence database”, Bioinformatics 12(22), e350-e358, each of which is hereby incorporated by reference.

[0061] The terms “marker”, “biomarker” and the like, refer to any compound that can be measured as an indicator of the physiological status of a biological system. The marker may be a biomarker that comprises an amino acid sequence, a nucleic acid sequence and fragments thereof. Exemplary biomarkers include, but are not limited to cytokines, chemokines, growth and angiogenic factors, metastasis related molecules, cancer antigens, apoptosis related proteins, enzymes, proteases, adhesion molecules, cell signalling molecules and hormones. The marker may also be a sugar that, in some embodiments, may not be significantly metabolized in the biological system. The sugar may be, for example, mannitol, lactulose, sucrose, sucralose and combinations of any of the forgoing.

[0062] “Measuring" or “measurement” means assessing the presence, absence, quantity or amount (which can be an effective amount) of a given substance within a sample, including the derivation of qualitative or quantitative concentration levels of such substances, or otherwise evaluating the values or categorization of a subject's clinical parameters. Alternatively, the term “assaying,” “detecting" or “detection” may be used to refer to all measuring or measurement as described in this specification.

[0063] As used herein, the term “microorganism,” or “microbe,” refers to a microscopic organism. In some embodiments, the term “microorganism” will be understood to include bacteria, fungi, protozoa (e.g., protozoan parasites), viruses (e.g., DNA viruses and / or RNA viruses), microsporidia, algae, archaea, phages, and / or helminths (e.g., multicellular eukaryotic parasites).

[0064] In some embodiments, a microorganism is a single-celled organism and / or a colony of single-celled organisms. In some embodiments, a microorganism is eukaryotic or prokaryotic. In some embodiments, a microorganism is a pathogen (e.g., disease-causing), such as a human, animal, or plant-infective pathogen.

[0065] Examples of bacteria include, but are not limited to, disease-causing agents such as Acinetobacter sp. (such as Acinetobacter baumannii, Acinetobacter calcoaceticus, Acinetobacter haemolyticus, Acinetobacter johnsonii, Acinetobacter lwoffii and Acinetobacter schindleri) Actinobacillus sp., Actinomycetes, Actinomyes sp. (such as Actinomyces israelii and Actinomyces naeslundii), Aeromonas sp. (such as Aeromonas hydrophiha, Aeromonas veronii biovar sobria (Aeromonas sobria), Aeromonas caviae, Aeromonas allosaccharophila, and Aeromonas dhakensis), Anaplasma phagocytophilum, Anaplasma marginale, Alcaligenes xylosoxidans, Actinobacillus actinomycetemcomitans, Arcobacter sp. (such as Arcobacter butzleri and Arcobacter cryaerophilus), Bacillus sp. (such as Bacillus anthracis, Bacillus cereus, Bacillus subtilis, Bacillus thuringiensis, and BacillusJAWS Ref: 750704PCT stearothermophaus), Bacteroides sp. (such as Bacteroides fragelis), Bartonella sp. (such as Bartonella bacilliformis and Bartonella henselae), Bifidobacterium sp., Bordetella sp. (such as Bordetella pertussis, Bordetella parapertussis, and Bordetella bronchiseptica), Borrelia sp. (such as Borrelia recurrentis, and Borrelia burgdorferi), Brucella sp. (such as Brucella abortus, Brucella canis, Brucella melitensis and Brucella suis), Burkholderia sp. (such as Burkholderia pseudomallei and Burkholderia cepacia), Campylobacter sp. (such as Campylobacter jejuni, Campylobacter coli, Campylobacter lari, Campylobacter fetus, Campylobacter upsaliensis, and Campylobacter concisus), Capnocytophaga sp., Cardiobacterium hominis, Chlamydia trachomatis, Chlamydophila pneumoniae, Chlamydia psittaci, Citrobacter sp. Coxiella burnetii, Corynebacterium sp. (such as, Corynebacterium diphtheriae, Corynebacterium jeikeum and Colynebacterium), Clostridioides sp. (such as Clostridioides perfringens, Clostridioides difficile, Clostridioides botulinum and Clostridioides tetani), Edwardsiella tarda, Eikenella corrodens, Enterobacter sp. (such as Enterobacter aerogenes, Enterobacter agglomerans, Enterobacter cloacae and Escherichia coli, including opportunistic Escherichia coli, such as enterotoxigenic E. coli, enteroinvasive E. coli, enteropathogenic E. coli, enterohemorrhagic E. coli (such as EHEC O157:H7), enteroaggregative E. coli, and uropathogenic E. coli, Enterococcus sp. (such as Enterococcus faecalis and Enterococcus faecium), Ehrlichia sp. (such as Ehrlichia chafeensia and Ehrlichia canis), Epidermophyton floccosum, Erysipelothrix rhusiopathiae, Eubacterium sp., Francisella tularensis, Fusobacterium nucleatum, Gardnerella vaginalis, Gemella morbillorum, Grimontia hollisae, Haemophilus sp. (such as Haemophilus influenzae, Haemophilus ducreyi, Haemophilus aegyptius, Haemophilus parainfluenzae, Haemophilus haemolyticus and Haemophilus parahaemolyticus), Helicobacter sp. (such as Helicobacter pylori, Helicobacter cinaedi, Helicobacter fennelliae, Helicobacter canadensis, and Helicobacter pullorum), Kingella kingae, Klebsiella sp. (such as Klebsiella pneumoniae, Klebsiella granulomatis, Klebsiella oxytoca, Klebsiella quasipneumoniae, and Klebsiella variicola), Kluyvera sp. (such as Kluyvera cryocrescens, Klebsiella variicola, Kluyvera ascorbate, Kluyvera cryocrescens, Kluyvera georgiana, Kluyvera intermedia), Lactobacillus sp., Listeria monocytogenes, Leptospira interrogans, Legionella pneumophila, Leptospira interrogans, Peptostreptococcus sp., Mannheimia haemolytica, Microsporum canis, Mordrelle catarrhalis, Morganella sp. (such as Morganella morganii), Mobiluncus sp., Micrococcus sp., Mycobacterium sp. (such as Mycobacterium leprae, Mycobacterium tuberculosis, Mycobacterium paratuberculosis, Mycobacterium intracellulare, Mycobacterium avium, Mycobacterium bovis, and Mycobacterium marinum), Mycoplasma sp. (such as Mycoplasma pneumoniae, Mycoplasma hominis; and Mycoplasma genitalium), Nocardia sp. (such as Nocardia asteroides, Nocardia cyriacigeorgica and Nocardia brasiliensis), Neisseria sp. (such as Neisseria gonorrhoeae and Neisseria meningitidis), Orientia tsutsugamushi, Pasteurella multocida, Pityrosporum orbiculare (Malassezia furfur), Plesiomonas shigelloides, Prevotella sp., Porphyromonas sp., Prevotella melaninogenica, Proteus sp. (such as Proteus vulgaris and Proteus mirabilis), Providencia sp. (such as Providencia alcalifaciens, Providencia rettgeri and Providencia stuartii), Pseudomonas aeruginosa, Propionibacterium acnes, Rhodococcus equi, Rickettsia sp. (such as RickettsiaJAWS Ref: 750704PCT rickettsii, Rickettsia akari, and Rickettsia prowazekii, and Rickettsia typhi), Rhodococcus sp., Stenotrophomonas maltophilia, Salmonella sp. (such as Salmonella bongori, Salmonella enterica, Salmonella typhi, Salmonella paratyphi, Salmonella enteritidis, Salmonella choleraesuis and Salmonella typhimurium), Serratia sp. (such as Serratia fonticola, Serratia marcescens and Serratia liquefaciens), Shewanella sp., Shigella sp. (such as Shigella dysenteriae, Shigella flexneri, Shigella boydii, and Shigella sonnei), Staphylococcus sp. (such as Staphylococcus aureus, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus saprophyticus), Streptococcus sp. (such as Streptococcus pneumoniae (for example chloramphenicol-resistant serotype 4 Streptococcus pneumoniae, spectinomycin-resistant serotype 6B Streptococcus pneumoniae, streptomycin-resistant serotype 9V Streptococcus pneumoniae, erythromycin-resistant serotype 14 Streptococcus pneumoniae, optochin-resistant serotype 14 Streptococcus pneumoniae, rifampicin-resistant serotype 18C Streptococcus pneumoniae, tetracycline-resistant serotype 19F Streptococcus pneumoniae, penicillin-resistant serotype 19F Streptococcus pneumoniae, and trimethoprim-resistant serotype 23F Streptococcus pneumoniae, chloramphenicol-resistant serotype 4 Streptococcus pneumoniae, spectinomycin-resistant serotype 6B Streptococcus pneumoniae, streptomycin-resistant serotype 9V Streptococcus pneumoniae, optochin-resistant serotype 14 Streptococcus pneumoniae, rifampicin-resistant serotype 18C Streptococcus pneumoniae, penicillin-resistant serotype 19F Streptococcus pneumoniae, or trimethoprim-resistant serotype 23F Streptococcus pneumoniae), Streptococcus agalactiae, Streptococcus mutans, Streptococcus pyogenes, Group A streptococci, Group B streptococci, Group C streptococci, Streptococcus anginosus, Streptococcus equisimilis, Group D streptococci, Streptococcus bovis, Group F streptococci, and Group G streptococci), Spirillum minus, Streptobacillus mondiformi, Treponema sp. (such as Treponema carateum, Treponema pertenue, Treponema pallidum and Treponema endemicum), Trichophyton rubrum, Trichophyton mentagrophytes, Tropheryma whipplei, Ureaplasma urealyticum. Veillonella sp., Vibrio sp. (such as Vibrio cholerae, Vibro parahaemolyticus, Vibrio vulnificus, Vibrio alginolyticus, Vibrio mimicus, Vibrio fluvialis, Vibrio metschnikovii, Vibrio damsela and Vibrio furnissii), Yersinia sp. (such as Yersinia enterocolitica, Yersinia pestis, Yersinia pseudotuberculosis, and Yersinia frederiksenii) and Xanthomonas maltophilia.

[0066] Examples of fungi include, but are not limited to, Aspergillus sp., Candida auris, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida kruseli, Candida lusitaniae, Candida orthopsilosis, Candida parapsilosis, Candida sake, Candida tropicalis, Cryptococcus gattii, Cryptococcus neoformans, Fusarium sp., Malassezia furfur, Rhodotorula sp., Trichosporon sp., Histoplasma capsulatum, Coccidioides immitis, and Pneumocystis carinii, as well as the causative agents of Apergillosis, Blastomycosis, Candidiasis, Coccidioidomycosis, fungal eye infections, fungal nail infections, histoplasmosis, mucormycosis, mycetoma, Pneumocystis pneumonia, ringworm, sporotrichosis, crypococcosis, and Talaromycosis.

[0067] Examples of protozoan parasites include, but are not limited to, Blastocystis hominis, Blastocystis sp. subtype 1, Blastocystis sp. subtype 2, Blastocystis sp.JAWS Ref: 750704PCT subtype 3, Blastocystis sp. subtype 4, Blastocystis sp. subtype 5, Blastocystis sp. subtype 6, Blastocystis sp. subtype 7, Blastocystis sp. subtype 8, Blastocystis sp. subtype 9, Cryptosporidium andersoni, Cryptosporidium hominis / parvum / cuniculus / tyzzeri complex, Cryptosporidium felis, Cryptosporidium meleagridis, Cryptosporidium muris, Cryptosporidium ubiquitum, Cryptosporidium viatorum, Cyclospora cayetanensis, Dientamoeba fragilis, Entamoeba histolytica, Entamoeba moshkovskii, Giardia intestinalis, G. muris, G. lamblia, Plasmodium falciparum, P. vivax, P. ovale, P. malariae, P. berghei, Leishmania donovani, L. infantum, L. chagasi, L. mexicana, L. amazonensis, L. venezuelensis, L. tropica, L. major, L. minor, L. aethiopica, L. braziliensis, L. (V) guyanensis, L. (V) panamensis, L. (V) peruviana, Trypanosoma brucei rhodesiense, T. brucei gambiense, T. cruzi, Toxoplasma gondii, Entamoeba histolytica, Trichomonas vaginalis, and Pneumocystis carinii.

[0068] Examples of helminths include, but are not limited to, Filarioidea sp., Wuchereria sp. (such as Wuchereria bancrofti), Brugia sp. (such as Brugia malayi and Brugia timori), Loa sp. (such as Loa loa), Mansonella sp. (such as Mansonella streptocerca, Mansonella perstans, and Monsonella ozzardi), Onchocerca sp. (such as Onchocerca volvulus), Enterobius vermicularis, Ascaris sp. (such as Ascaris lumbricoides), Dracunculus (such as Dracunculus medinensis), Ancylostoma sp. (such as Ancylostoma duodenale, Ancylostoma braziliense, Ancylostoma tubaeforme, and Ancylostoma caninum), Necator sp. (such as Necator americanus), Trichuris sp. (such as Trichuris trichiura, Trichuris vulpis, Trichuris campanula, Trichuris suis, and Trichuris muris), Strongyloides sp. (such as Strongyloides stercoralis, Strongyloides canis, Strongyloides fuelleborni, Strongyloides cebus, and Strongyloides kellyi), Nematodirus sp., Moniezia sp., Oesophagostomum sp. (such as Oesophagostomum bifurcum, Oesophagostomum aculeatum, Oesophagostomum brumpti, Oesophagostomum stephanostomum, and Oesophagostomum stephanostomum var thomasi), Cooperia sp. (such as Cooperia ostertagi and Cooperia oncophora), Haemonchus sp., Ostertagia sp. (such as Ostertagia ostertagi), Trichostrongylus sp. (such as Trichostrongylus axei), Dirofilaria sp. (such as Dirofilaria immitis, Dirofilaria tenuis and Dirofilaria repens), and Schistosoma sp. (such as Schistosoma incognitum, Schistosoma ovuncatum, Schistosoma sinensium. Schistosoma indicum, Schistosoma nasale, Schistosoma spindale, Schistosoma japonicam, Schistosoma malayensis, Schistosoma mekongi, Schistosoma haematobium, Schistosoma bovis, Schistosoma curassoni, Schistosoma guineensis, Schistosoma haematobium, Schistosoma intercalatum, Schistosoma leiperi, Schistosoma margrebowiei, Schistosoma mattheei, Schistosoma mansoni, Schistosoma edwardiense, Schistosoma hippotami, and Schistosoma rodhaini).

[0069] Examples of viruses include, but are not limited to, disease-causing agents such as Adeno-associated virus, Aichivirus A, Australian bat lyssavirus, BK polyomavirus, Banana virus, Barmah forest virus, Bunyamwera virus, Bunyavirus La Crosse, Bunyavirus snowshoe hare, Cercopithecine herpesvirus, Chandipura virus, Chikungunya virus, Coronavirus, Cosavirus A, Cowpox virus, Coxsackievirus, Crimean-Congo hemorrhagic fever virus, Dengue virus, Dhori virus, Dugbe virus, Duvenhage virus, Eastern equine encephalitis virus, Ebolavirus,JAWS Ref: 750704PCT Echovirus, Encephalomyocarditis virus, Epstein-Barr virus, European bat lyssavirus, GB virus C, Hantavirus, Hendra virus, Hepatitis A virus, Hepatitis B virus, Hepatitis C virus, Hepatitis E virus, Hepatitis delta virus, Horsepox virus, Human adenovirus, Human astrovirus, Human coronavirus, Human cytomegalovirus, Human enterovirus D68, Human enterovirus 70, Human herpesvirus 1, Human herpesvirus 2, Human herpesvirus 6, Human herpesvirus 7, Human herpesvirus 8, Human immunodeficiency virus, Human papillomavirus 1, Human papillomavirus 2, Human papillomavirus 16, Human papillomavirus 18, Human parainfluenza, Human parvovirus B19, Human respiratory syncytial virus, Human rhinovirus, Human SARS coronavirus. Human spumaretrovirus, Human T-lymphotropic virus, Human torovirus, Influenza A virus, Influenza B virus, Influenza C virus, Isfahan virus, JC polyomavirus, Japanese encephalitis virus, Junin virus, ICI Polyomavirus, Kunjin virus, Lagos bat virus, Lake Victoria Marburgvirus, Langat virus, Lassa virus, Lordsdale virus, Louping ill virus, Lymphocytic choriomeningitis virus, Machupo virus, Mayaro virus, MERS coronavirus, Measles virus, Mengo encephalomyocarditis virus, Merkel cell polyomavirus, Mokola virus, Molluscum contagiosum virus, Monkeypox virus, Mumps virus, Murray valley encephalitis virus, New York virus, Nipah virus, Norwalk virus, Norovirus, O’nyong’nyong virus, Orf virus, Oropouche virus, Pichinde virus, Poliovirus, Punta Toro phlebovirus. Puumala virus, Rabies virus, Rift valley fever virus, Rotavirus A, Ross river virus, Rotavirus A, Rotavirus B, Rotavirus C, Rubella virus, Sagiyama virus, Salivirus A, Sandfly fever Sicilian virus, Sapporo virus, Semliki forest virus, Seoul virus, Severe acute respiratory syndrome coronavirus 2, Simian foamy virus, Simian virus 5, Sindbis virus, Southampton virus, St. Louis encephalitis virus, Tick-borne powassan virus, Torque teno virus, Toscana virus, Uukuniemi virus, Vaccinia virus, Varicella-zoster virus, Variola virus, Venezuelan equine encephalitis virus, Vesicular stomatitis virus, Western equine encephalitis virus, WU polyomavirus, West Nile virus, Yaba monkey tumor virus, Yaba-like disease virus, Yellow fever virus, and Zika virus.

[0070] Examples of microsporidia pathogens include, but are not limited to, Anncaliia algerae, Encephalitozoon cuniculi, Encephalitozoon hellem, Encephalitozoon intestinalis, and Encephalitozoon bieneusi.

[0071] Examples of invertebrate pathogens include, but are not limited to, Ancylostoma duodenale, Ascaris lumbricoides, Clonorchis sinensis, Dibothriocephalus latus, Dipylidium caninum, Enterobius vermicularis, Fasciola hepatica, Fasciolopsis buski, Hymenolepis diminuta, Hymenolepis nana, Necator americanus, Opisthorchis felineus, Opisthorchis viverrine, Schistosoma japonicum, Schistosoma mansoni, Strongyloides stercoralis, Taenia saginata, Taenia solium, and Trichuris trichiura.

[0072] In some embodiments, the term “microorganism” will be understood to include any one or more bacteria, fungi, protozoa, viruses, algae, archaea, phages, and / or helminths selected from a database (e.g., a microbial genome database, a transcriptomic database, a proteomic database, a metabolomics database, a taxonomic database, and / or a clinical database). In some embodiments, the database comprises one or more entries corresponding to and / or identifying a microorganism (e.g., an annotation, for a respectiveJAWS Ref: 750704PCT microorganism, to a genome, transcriptome, nucleic acid sequence, protein sequence, metabolite, taxonomic record and / or clinical record). In some embodiments, a microorganism is selected from a database that is locally maintained, proprietary, and / or open access. In some embodiments, a microorganism is selected from a national and / or international database. Examples of such databases include, but are not limited to, GTDB, NCBI, BLAST, EMBL-EBI, GenBank, Ensembl, EuPathDB, The Human Microbiome Project, Pathogen Portal, RDP, SILVA, GREENGENES, EBI Metagenomics, EcoCyc, PATRIC, TBDB, PiasmoDB, the Microbial Genome Database (MBGD), and / or the Microbial Rosetta Stone Database. For example, MBGD comprises all complete genome sequences of bacteria, archaea, and unicellular eukaryotes, including fungi and protozoa, available at the NCBI genomes site. The Microbial Rosetta Stone is a database that provides information on disease-causing organisms (e.g., bacteria, fungi, protozoa, DNA viruses, RNA viruses, plants, and animals) and the toxins produced therefrom (see, Zhulin, 2015, “Databases for Microbiologists,” J. Bacteriol. 197:2458-2467; Uchiyama et al., 2019, "MBGD update 2018: microbial genome database based on hierarchical orthology relations covering closely related and distantly related comparisons,” Nuc. Acids Res., 47 (D1), D382-D389; and Ecker et al., 2005, “The Microbial Rosetta Stone Database: A compilation of global and emerging infectious microorganisms and bioterrorist threat agents,” BMC Microbiology 5:19; each of which is hereby incorporated by reference herein in their entirety).

[0073] As used herein, the terms “nucleic acid” and “nucleic acid molecule” are used interchangeably. The terms refer to nucleic acids of any composition form, such as ribonucleic acid (RNA), deoxyribonucleic acid (DNA, e.g., complementary DNA (cDNA), genomic DNA (gDNA) and the like), and / or DNA or RNA analogs (e.g., containing base analogues, sugar analogues and / or a non-native backbone and the like). In some embodiments, nucleic acids are in single- or double-stranded form. Unless otherwise limited, a nucleic acid can comprise known analogs of natural nucleotides, some of which can function in a similar manner as naturally occurring nucleotides. A nucleic acid can be in any form useful for conducting processes herein (e.g., linear, circular, supercoiled, single-stranded, double-stranded and the like). In some embodiments a nucleic acid can be from a single chromosome or fragment thereof (e.g., a nucleic acid sample may be from one chromosome of a sample obtained from a diploid organism). In certain embodiments nucleic acids comprise nucleosomes, fragments or parts of nucleosomes or nucleosome-like structures. Nucleic acids sometimes comprise protein (e.g., histones, DNA binding proteins, and the like). Nucleic acids analyzed by processes described herein sometimes are substantially isolated and are not substantially associated with protein or other molecules. Nucleic acids also include derivatives, variants and analogues of DNA synthesized, replicated or amplified from single-stranded (“sense" or "antisense”, “plus” strand or “minus” strand, “forward” reading frame or “reverse” reading frame) and double-stranded polynucleotides. Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine and deoxythymidine. A nucleic acid may be prepared using a nucleic acid obtained from a subject as a template.JAWS Ref: 750704PCT

[0074] As used herein, the term “reference sequence” refers to a sequence of nucleotide bases. In some embodiments, a reference sequence is a reference genome. In some embodiments, a reference sequence is a complete or incomplete genome. In some embodiments, a reference sequence is less than 1 megabase pairs (Mb), less than 0.5 Mb, less than 0.4 Mb, less than 0.3 Mb, less than 0.2 Mb, or less than 0.1 Mb in length. In some embodiments, a reference sequence is at least 1 Mb, at least 2 Mb, at least 3 Mb, at least 4 Mb. at least 5 Mb, at least 6 Mb, at least 7 Mb, at least 8 Mb, at least 9 Mb, at least 10 Mb, at least 15 Mb, at least 20 Mb, at least 25 Mb, at least 30 Mb, at least 35 Mb, at least 40 Mb, at least 45 Mb, at least 50 Mb, at least 100 Mb, at least 200 Mb, at least 500 Mb, at least 1,000 Mb, at least 2,000 Mb, at least 3,000 Mb, at least 4,000 Mb, at least 5,000 Mb, at least 10 gigabase pairs (Gb), at least 20 Gb, or at least 50 Gb in length.

[0075] In some embodiments, a reference sequence spans a region of a reference genome comprising at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 100, at least 200, at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 10,000, or at least 50,000 genes. In some embodiments, a reference sequence spans a region of a reference genome comprising between 1 and 10, between 10 and 50, between 50 and 100, between 100 and 500, between 500 and 1000, between 1000 and 2000, between 2000 and 5000, between 5000 and 10,000, between 10,000 and 50,000, or more than 50,000 genes.

[0076] In some embodiments, a reference sequence comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 100, at least 200, or at least 500 antimicrobial resistance markers. In some embodiments, a reference sequence comprises between 1 and 10, between 10 and 50, between 50 and 100, or more than 100 antimicrobial resistance markers. The term “species” is defined as a collection of closely related organisms with greater than 97% 16S ribosomal RNA (rRNA) sequence homology and greater than 70% genomic hybridization and sufficiently different from all other organisms so as to be recognized as a distinct unit. Species and other phylogenic identifications are according to the classification known to a person skilled in the art of microbiology.

[0077] As used herein, a “host” or “subject” means a human or animal. Usually the animal is a vertebrate such as a primate, rodent, domestic animal or game animal. Primates include chimpanzees, cynomolgus monkeys, spider monkeys, and macaques (e.g., Rhesus). Rodents include mice, rats, woodchucks, ferrets, rabbits, and hamsters. Domestic and game animals include cows, horses, pigs, deer, bison, buffalo, feline species (e.g., domestic cat), canine species (e.g., dog, fox, wolf), avian species (e.g., chicken, emu, ostrich), and fish (e.g., trout, catfish, and salmon). In some embodiments the subject is a mammal (e.g., a primate (e.g., a human)). The terms “host”, “individual”, “patient” and “subject” are used interchangeably herein. Preferably the subject is a mammal. The mammal can be a human, non-human primate, mouse,JAWS Ref: 750704PCT rat, dog, cat, horse or cow, but is not limited to these examples. Mammals other than humans can be advantageously used as subjects that represent animal models of inflammatory and autoimmune disorders (e.g., models of gut barrier function). A subject can be male or female.

[0078] As used herein, the terms “treat”, “treatment”, “treating” and the like, refer to therapeutic treatments, wherein the object is to reverse, alleviate, ameliorate, inhibit, slow down or stop the progression or severity of a condition associated with a disease or disorder (e.g., an inflammatory or autoimmune disorder). The term “treating” includes reducing or alleviating at least one adverse effect or symptom of a condition, disease or disorder associated with an inflammatory or autoimmune disorder. Treatment is generally “effective” if one or more symptoms or clinical markers are reduced. Alternatively, treatment is “effective” if the progression of a disease is reduced or halted. That is, “treatment” includes not just the improvement of symptoms or markers, but also a cessation of, or at least slowing of, progress or worsening of symptoms compared to what would be expected in the absence of treatment. Beneficial or desired clinical results include, but are not limited to, alleviation of one or more symptom(s), diminishment of extent of disease, stabilized (i.e., not worsening) state of disease, delay or slowing of disease progression, amelioration, or palliation of the disease state, remission (whether partial or total), and / or decreased mortality, whether detectable or undetectable. The term “treatment” of a disease also includes providing relief from the symptoms or side-effects of the disease (including palliative treatment). A treatment need not cure a disorder (i.e., complete reversal or absence of disease) to be considered effective.

[0079] Each embodiment described herein is to be applied mutatis mutandis to each and every embodiment unless specifically stated otherwise. 2. Samples

[0080] In some embodiments, a biological sample is collected, prepared, sequenced (e.g., by next-generation sequencing), and mapped (e.g., aligned) to one or more reference sequences (e.g., complete and / or incomplete genomes, parts of genomes, genes or other nucleotide or protein sequences) prior to the analysis of the presence of microorganisms. In some embodiments, sample processing is performed using any of the methods as disclosed in International PCT Patent Application No. PCT / AU2019 / 050618, entitled “Methods for sample preparation and microbiome characterisation” filed 15 June 2019, which is hereby incorporated by reference herein in its entirety. In some embodiments, sample processing is performed using the method described in Example 1 and Figure 1 (see Examples, below).

[0081] In some embodiments, the biological sample is obtained from a subject (e.g., a biological subject). For example, in some embodiments, the subject is a human (e.g., a patient). In some embodiments, the subject is not a human (e.g. a cow, horse, pig, chicken, or another animal). In some embodiments, the biological sample is obtained from any microbiome from the subject (e.g., gut microbiome, skin microbiome, vaginal microbiome, etc.). In some embodiments, a plurality of biological samples is obtained from the subject (e.g., a plurality of replicates and / or a plurality of samples including a healthy sample and a diseased sample). InJAWS Ref: 750704PCT some embodiments, the biological or sample is obtained from a human with a disease condition. In some embodiments, the disease condition presents as an undifferentiated gastrointestinal issue. By way of an example, the undifferentiated gastrointestinal issue may be selected from one or more of chronic or acute diarrhea, inflammatory bowel disease (IBD), irritable bowel syndrome (IBS). In some embodiments, the biological or non-biological sample is obtained from a human with a gastrointestinal disease. In some embodiments, the biological sample is obtained from a human with a bacterial gastrointestinal tract infection. In some embodiments, the sample is obtained from a human with a C. difficile infection. In some embodiments, the disease condition present as a urinary tract infection, respiratory infection. In some embodiments, the subject may be immunocompromised or immunosuppressed. In some embodiments, the sample is taken from a human suspected to have a pathogen infection. In some embodiments, the human is a returned traveler, visitor or immigrant with gastrointestinal symptoms

[0082] In some embodiments, the biological or non-biological sample is obtained from a human with a gastrointestinal disease. In some embodiments, the biological sample is obtained from a human with a bacterial gastrointestinal tract infection. In some embodiments, the sample is obtained from a human with a C. difficile infection.

[0083] In some embodiments, the biological sample is an analysis (e.g., test) sample or a control sample (e.g., a positive control, negative control, and / or blank control).

[0084] In some embodiments, the biological or non-biological sample comprises nucleic acid material (e.g., RNA or DNA). In some embodiments, the nucleic acids included in the biological sample comprise any of the embodiments described above and / or elsewhere herein. 3. Sequencing

[0085] As described above and elsewhere herein, in some implementations, the sequencing generates a plurality of nucleotide sequences that can be mapped against a plurality of reference sequences. In some embodiments, the sequencing is performed on a sample or portion thereof that has undergone a nucleic acid amplification process. Alternatively, in some embodiments, the sequencing is performed on a sample or portion thereof that has not undergone a nucleic acid amplification process. In some embodiments, nucleic acid molecules within a sample or portion thereof are fragmented prior to undergoing sequencing.

[0086] Alternatively, in some embodiments, nucleic acid molecules are not fragmented prior to undergoing sequencing. Multiple different schemes may be applied to identify nucleic acid sequences within a sample.

[0087] Different types of nucleic acid molecules may undergo the same or different processing and sequencing. For example, in some embodiments, DNA molecules undergo a first sequencing process and RNA molecules undergo a second sequencing process, where the first and second sequencing processes include at least one process difference. In an example, genomic DNA is processed according to a first sequencing method (e.g., standard Sanger sequencing or sequencing by synthesis of DNA) while RNA molecules are processedJAWS Ref: 750704PCT according to a second sequencing method (e.g., a sequencing method that targets RNA molecules that include a polyA sequence, such as messenger RNA (mRNA) molecules). In some embodiments, different sequencing procedures are performed on the same or different samples. For example, in some embodiments, a first sequencing method to analyze a first type of nucleic acid molecule and a second sequencing method to analyze a second type of nucleic acid molecule, where the first and second sequencing methods are different and the first and second types of nucleic acid molecules are different, are performed on a same sample (e.g., at the same or different times). Alternatively or in addition, in some embodiments, a first sequencing method to analyze a first type of nucleic acid molecule is performed using a first sample and a second sequencing method to analyze a second type of nucleic acid molecule may be performed using a second sample, where the first and second sequencing methods are different, the first and second types of nucleic acid molecules are different, and the first and second samples are different. In some embodiments, the first and second samples are aliquots of the same sample.

[0088] In some embodiments, the sequencing is quantitative or approximately quantitative. Alternatively, in some embodiments, nucleic acid sequencing is qualitative and does not provide significant insight into the relative or absolute amounts of different nucleic acid molecules included within a sample. In some embodiments, the sequencing is semi-quantitative, or represents the relative abundance of the nucleic acid molecules.

[0089] Various sequencing schemes can be employed. For example, in some embodiments, the sequencing is sequencing by synthesis, sequencing by hybridization, sequencing by ligation, nanopore sequencing, sequencing using nucleic acid nanoballs, pyrosequencing, single molecule sequencing (e.g., single molecule real time sequencing), single cell / entity sequencing, massively parallel signature sequencing, polony sequencing, combinatorial probe anchor synthesis, SOLiD sequencing, chain termination (e.g., Sanger sequencing), ion semiconductor sequencing, tunneling currents sequencing, heliscope single molecule sequencing, sequencing with mass spectrometry, transmission electron microscopy sequencing, RNA polymerase-based sequencing, or any other method, or a combination thereof. In some embodiments, the sequencing is a sequencing technology like Heliscope (Helices), SMRT technology (Pacific Biosciences), or nanopore sequencing (Oxford Nanopore) that allows direct sequencing of single molecules without prior clonal amplification. In some embodiments, the sequencing is performed with or without target enrichment. In some embodiments, the sequencing whole genome sequencing. In some other embodiments, the sequencing is Helicos True Single Molecule Sequencing (tSMS) (e.g., as described in Harris T. D. et al., Science 320:106-109

[2008] ). In some other embodiments, the sequencing is 454 sequencing (Roche) (e.g., as described in Margulies, M. et al., Nature 437:376-380 (2005)). In some embodiments, the sequencing is SOLiD™ technology (Applied Biosystems). In some embodiments, the sequencing is single molecule, real-time (SMRT™) sequencing technology of Pacific Biosciences.JAWS Ref: 750704PCT

[0090] In some embodiments, the systems and methods described herein are used with any sequencing platform, including, but not limited to, Illumina NGS platforms, Ion Torrent (Thermo) platforms, and GeneReader (Qiagen) platforms.

[0091] In some embodiments, the sequencing is performed as described in PCT Application No. PCT / US2019 / 060915, entitled “Directional Targeted Sequencing,” filed November 12, 2019, which is hereby incorporated by reference herein in its entirety.

[0092] In some embodiments, the sequencing reaction is a whole genome sequencing reaction (e.g., shotgun workflow). In some instances, the sequencing is digital polymerase chain reaction (PCR) sequencing. In some embodiments, the sequencing reaction is a whole transcriptome sequencing reaction (e.g., RNASeq). In some embodiments, the sequencing reaction is a panel enriched sequencing reaction. In some embodiments, the panel is pathogen-specific and / or disease condition-specific. For example, in some embodiments, the panel is a respiratory virus oligo panel (RVOP). 4. Detection and Identification

[0093] In some embodiments, the plurality of nucleotide sequences (e.g., in a nucleotide sequence data store (130)) includes a first subset of nucleotide sequences that map (e.g., align) to a first plurality of reference sequences (e.g., a first genomes) and a second subset of nucleotide sequences that map (e.g., align) to a second reference sequence (e.g., a second genome) (e.g., where the first genome is a reference genome of a host organism, parts of a genome or one or more genes and the second genome is a reference genome, parts of a genome or one or more genes of microorganism). In some embodiments, the plurality of nucleotide sequences includes a plurality of subsets of nucleotide sequences, each respective subset of nucleotide sequences mapping to a corresponding reference sequence in a plurality of reference sequences (e.g., in reference sequence data store (132). In some such embodiments, the plurality of subsets of nucleotide sequences includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 x 106, at least 1 x 107or at least 1 x 108subsets of nucleotide sequences that map to a corresponding reference sequence.

[0094] In some embodiments, the mapping of the plurality of nucleotide sequences against the plurality of reference sequences corresponding to a set of microorganisms (e.g., genomes), where the set of microorganisms comprises at least 1, at least 3, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 microorganisms, collectively maps against at least 0.5, at least 1, at least 2, at least 3, at least 4, at least 5, or at least 6 megabases of the respective reference sequences (e.g., genomes). In some embodiments, the mapping of the plurality of nucleotide sequencesJAWS Ref: 750704PCT against the plurality of reference sequences corresponding to the set of microorganisms collectively maps against at least 0.0001, at least 0.001, at least 0.01, at least 0.1, at least 0.5, at least 0.8, at least 1, at least 1.5, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 100, at least 200, at least 500, or at least 1000 megabases of the respective reference sequences. In some embodiments, the mapping of the plurality of nucleotide sequences against the plurality of reference sequences corresponding to the set of microorganisms collectively maps against no more than 107, no more than 106, no more than 105, no more than 10000, no more than 5000, no more than 2000, no more than 1000, no more than 500, no more than 100, no more than 80, no more than 60, no more than 40, no more than 20, no more than 10, no more than 5, no more than 3, no more than 2, or no more than 1, no more than 0.5, no more than 0.1, no more than 0.01, no more than 0.001 or no more than 0.0001 megabases of the respective reference sequences. In some embodiments, the mapping of the plurality of nucleotide sequences against the plurality of reference sequences corresponding to the set of microorganisms collectively maps against from 0.001 to 0.05, 0.001 to 1, 0.5 to 10, from 1 to 6, from 2 to 5, from 4 to 15, from 8 to 20, from 12 to 30, from 10 to 60, from 20 to 100, from 75 to 500, from 100 to 1000, from 300 to 800, from 500 to 2000, from 800 to 104, from 103to 105, or from 104to 106, or from 105to 107megabases of the respective reference sequences.

[0095] In some embodiments, the mapping of the plurality of nucleotide sequences against the plurality of reference sequences corresponding to the set of microorganisms collectively maps against at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 50, or at least 100 sets of subsequent reference sequences.

[0096] In some embodiments, the result set further includes a plurality of nucleotide sequences mapped (e.g., aligned) to a human reference genome. Accordingly, in some embodiments, the mapping of the plurality of nucleotide sequences against a plurality of reference sequences, where the plurality of reference sequences includes a set of reference sequences corresponding to a set of microorganisms (e.g., at least 1, at least 3, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 microorganisms) and a host reference genome, collectively maps against at least 0.0001, at least 0.001, at least 0.01, at least 0.1, at least 1, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1000, or at least 3000 megabases. In some embodiments, the mapping of the plurality of nucleotide sequences against a plurality of reference sequences, where the plurality of reference sequences includes a set of reference sequences corresponding to a set of microorganisms and a human reference genome, collectively maps against no more than 5000, no more than 3000, no more than 1000, no more than 500, no more than 100, no more than 50, no more than 10, or no more than 5 megabases. In some embodiments, the mapping of the plurality of nucleotide sequences against a plurality of reference sequences, where the plurality of reference sequences includes a set of reference sequences corresponding to a set ofJAWS Ref: 750704PCT microorganisms and a human reference genome, collectively maps against from 1 to 10, from 2 to 20, from 15 to 60, from 40 to 200, from 150 to 800, or from 500 to 5000 megabases. In some embodiments, the mapping of the plurality of nucleotide sequences against a plurality of reference sequences, where the plurality of reference sequences includes a set of reference sequences corresponding to a set of microorganisms and a human reference genome, collectively maps against another range of megabases starting no lower than 1 megabase and ending no higher than 5000 megabases.

[0097] In some embodiments the reference database includes 250,000 genomes and over 850 gigabases. In some of the same embodiments and some other embodiments the reference database includes over 4,250 funghi genomes, and over 175 gigabases.

[0098] In some embodiments, the analysis comprises preprocessing and / or pre- sorting of the sequencing data. In some embodiments, pre-sorting includes sorting each nucleotide sequence obtained from the sequencing of the biological or non-biological sample into one or more bins, where each bin corresponds to a different microorganism, family, genus, species, sub-species, genome, gene, gene family, gene group or single nucleotide variant, depending on the likelihood that the nucleotide sequence originated from the respective microorganism, trait or group. Each nucleotide sequence is then mapped (e.g., using a k-mer alignment and / or a full alignment) to one or more reference sequences (e.g., genomes, parts of genomes, genes or other sequences) corresponding to or derived from different microorganisms. In some embodiments, nucleotide sequences are represented as k-mers, signatures or motifs before mapping to the reference sequences. In some embodiments, nucleotide sequences are compared to one or more reference sequences using artificial intelligence or machine learning models (e.g. hidden Markov models, large language models, or decision trees).

[0099] In some embodiments, the analysis is performed using an analysis pipeline. Methods of mapping nucleotide sequences obtained from sequencing nucleic acids are provided in, for example Muller A., et al., “MetaCache: context-aware classification of metagenomic reads using minhashing, 2017, Bioinformatics, 33:23, 3740-3748; and Li and Durbin, Fast and accurate short read alignment with Burrows-Wheeler transform, 2009, Bioinformatics 25:14, 1754-1760; the contents which are hereby incorporated by in their entirety. Other methods of mapping nucleotide sequences to a reference sequence are possible, as will be apparent to one skilled in the art. See, for example, Parks et al., 2021, “Evaluation of the Microba Community Profiler for Taxonomic Profiling of Metagenomic Datasets From the Human Gut Microbiome,” Front. Microbiol., 12; and Breitwieser et al., 2019, “A review of methods and databases for metagenomic classification and assembly” Brief Bioinform. 20(4): 1125-1136, which is hereby incorporated herein by reference in its entirety.

[0100] In some embodiments, the analysis involves assembling the nucleotide sequences together, such that the plurality of nucleotides are overlapped according to matched nucleotides, and joined, scaffolded, linked or assembled into contiguous nucleotide sequences.JAWS Ref: 750704PCT In some embodiments, the plurality of nucleotides are then mapped (e.g. aligned), to the contiguous sequences in order to determine the abundance of those contiguous sequences. In some embodiments, the assembled contiguous sequences are mapped (e.g. aligned) to the reference sequences). The contiguous sequences can be placed into bins, based on the abundance of the contiguous sequences, nucleotide sequence composition, translated protein composition and protein coding profile, and other measures. These bins represent metagenome assembled genomes (MAGs) (Kang et al., “MetaBAT 2: an adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies, PeerJ, 2019, 7:e7359; and Yong et al., A review of computational tools for generating metagenome-assembled genome from metagenomic sequencing data, Comp. Struct. Biotech. J., 2021, 19: 6301-6314), and a taxonomic description applied to the MAG, by comparing the genome and gene sequences it contains to reference databases and known phylogenies. The completeness of the MAGs relative to the endogenous entire genome of the organism can be estimated using bioinformatics tools such as Parks et al., CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes, Genome Res., 2015, 25: 7, 1043-1055).

[0101] The methods described above and elsewhere herein are able to accurately identify microorganisms and define their abundance within a microbial community. In some embodiments, nucleic acid sequence information can be extracted from a sample and their sequence read into a computing system. The methods described herein can provide a metagenomic sequencing platform that is faster, less computationally intensive, and / or more sensitive and specific than existing technology.

[0102] In some embodiments, the reference sequences are derived from bacteria, archaea, fungi, viruses, phages, invertebrates, protozoa, microsporidia, helminths, parasites, AMR genes, and / or virulence factors.

[0103] These methods comprise the characterisation of one or more bacteria included in a reference genome database (e.g., dereplicated phylogenomic database) as being present or absent in a biological sample. Some selected examples are Bacillus clausii, Bifidobacterium animalis, Pediococcus acidilactici, Acinetobacter indicus, Lactobacillus salivarius, Acinetobacter, Bacillus amyloliquefaciens, Lactobacillus helveticus, Bacillus subtilis, Lactobacillus plantarum, Bifidobacterium longum subsp infantis, Enterococcus hirae, Lactobacillus delbrueckii subsp bulgaricus, Enterococcus, Lactobacillus rhamnosus, Lactococcus lactis, Pseudomonas stutzeri, Lactobacillus acidophilus, Klebsiella and Enterobacter cloacae strain. The methods may alternatively or additionally comprise the characterisation and / or analysis of one or more phage or viruses contained within our database. Some selected examples are Bacillus phage phi29, Enterobacteria phage HK022, Lactobacillus phage A2, Escherichia phage HK639, Phage cdtI, Sclerotinia sclerotiorum partitivirus S segment 2, Burkholderia phage BcepMu, Lactococcus prophage bIL311, Enterococcus phage phiFL4A and Streptococcus phage SM1.JAWS Ref: 750704PCT

[0104] The methods of the present invention provide methods of estimating the relative abundance of prokaryotic, eukaryotic and / or virus species within a biological sample.

[0105] In some embodiments, the system may comprise a processing unit 100. The processing unit 100 may include a processor 104 and a memory 106. The processing unit may be configured to perform the characterization of biological material in a sample. Alternatively, the system may comprise units in the form of hardware and / or software each configured to perform one or more portions of the characterization of biological material. Further each of the units may comprise its own processor and memory, or each of the units may share a processor and / or memory with one or more of the other units. The processor can be, for example, an application-specific integrated circuit designed to achieve one or more specific functions or enable one or more specific devices or applications. The processor can control all of the other functional elements of the sequencing device. For example, the processor can send / receive the DNA sequence data to be stored in a data store (memory). The data store can also include any suitable types of forms of memory for storing data in a form retrievable by the processor.

[0106] In some embodiments, the system may utilize sequence information. The sequence information may be derived from a biological sample. In some embodiments, the biological sample may contain genetic material from a plurality of organisms. In an example, the sample may contain a plurality of microbial organisms, including bacteria, viruses, parasites, fungi, plasmids and other exogenous DNA or RNA fragments available in the biological sample type.

[0107] In one embodiment, the sequence information may be produced by collecting a biological sample containing genetic material, extracting fragments (e.g., nucleic acid and / or protein and / or metabolites) and sequencing the fragments. In some embodiments, the sample is a metagenomic sample, and the extracted and sequenced fragments are metagenomic fragments.

[0108] The sequencing device can further include a communication component to which the processor can send data retrieved from the data store. The communication component can include any suitable technology for communicating with the communication network, such as wired, wireless, satellite, etc.

[0109] In some embodiments, the sequence information may be produced by collecting a biological sample containing genetic material, extracting fragments (e.g., nucleic acid) and sequencing the fragments. In some embodiments, the biological sample is a metagenomic sample, and the extracted and sequenced fragments are metagenomic fragments. Typically, the biological sample may be a subject sample which includes microorganisms that are present within the subject. In some embodiments the sample is selected from a fecal sample, a skin sample, a biopsy, a swab, and a respiratory sample. In some particularly preferred embodiments, the sample is a fecal sample.JAWS Ref: 750704PCT

[0110] In some embodiments, the sequence information may include or be in the form of nucleotide fragment reads (i.e., nucleic acid read sequences). In some of the same embodiments and other embodiments, the sequence information may be unassembled sequence information (i.e., sequence information that has not been assembled into larger contigs or full genomes). For example, in a non-limiting embodiment, the sequence information utilized by the processing unit may include unassembled nucleotide fragment reads.

[0111] System 100 may utilize sequence information including hundreds, thousands, millions or billions of short fragment reads (e.g., unassembled fragment reads). The sequence information may be in the form of a sequence information file 108 produced from the fragment reads.

[0112] Although fragment reads included in the sequence information and utilized by the processing unit 102 may be greater than 100 base pairs in length, the fragment reads included in the sequence information and utilized by the processing unit 102 typically have lengths of approximately 50 to 300 base pairs. For example, for DNA, the fragment reads may have read lengths of less than 100 base pairs, and the sequence information file 108 produced therefrom may contain millions of DNA fragment reads. In some preferred embodiments, the fragment reads may have read lengths of around 150 base pairs. In some other embodiments, the fragment reads may have read lengths of around 250 base pairs. In some other embodiments, fragment reads may have several hundred, or several thousand base pairs (e.g.10-100kb or 100- 300kb). In some embodiments, fragment reads are produced as single-end reads, in other embodiments as paired-end reads.

[0113] A range from 1000 or greater reads of sequencing for short insert methods can be used for this method. Large insert methods such as Pac Bio™, Nanopore™ (and other next gene sequencing methods) can use <1000 sequencing reads. Bioinformatic quality filtering is generally performed before taxonomy assignment. Quality trimming of raw sequencing files may include one of more of removal of sequencing adaptors or indexes; trimming 3' or 5' end of reads based on quality scores, base pairs of end, or signal intensity; removal of reads based on quality scores, GC content, homopolymer nucleotides, DUST filtering, low complexity, PCR duplications, overly abundant, insufficient read length, or non-aligned base pairs; removal of overlapping reads at set number of base pairs. In some embodiments, quality filtering may include removal of host genome sequence reads (e.g. sequence reads from the human genome).

[0114] In the embodiment illustrated in Figure 1, system 100 may receive a sequence information file 108 as an input. Typically, the sequence information file 108 comprises DNA fragment sequences (“read sequences”) that correlate to the nucleic acid material from the sample taken from an individual.

[0115] In some other embodiments, such as that illustrated in Figure 2, system 100 may additionally comprise an extraction unit 210 and a sequencing unit 230 and be capable of receiving a sample or isolate as an input and producing a sequence information file 108 therefrom. In some embodiments, the sequencing unit 230 may perform sequencing based onJAWS Ref: 750704PCT any suitable sequencing method, including but not limited to, sequencing-by-synthesis, sequencing-by-ligation, single-molecule- sequencing, and pyrosequencing.

[0116] In some embodiments, the sequencing unit 230 may be interchangeable and removably coupled to the system 100. In some alternative embodiments, system 100 may be coupled to an external sequencer and may receive a sequence information file 108 directly from the external sequencer, but this is not required. System 100 may also receive the sequences information file 108 indirectly from one or more external sequencers that are not coupled to system 100. For example, system 100 may receive a sequence information file 108 over a communication network from a sequencer, which may be located remotely. Alternatively, a sequence information file 108 which has previously been stored on a storage medium, such as a hard disk drive or optical storage medium, may be input into system 100.

[0117] In addition, system 100 may receive a sequence information file 108 or fragments reads in real-time, immediately following sequencing by a sequencer or in parallel with sequencing by a sequencer, but this also is not required. System 100 may also receive the sequence information file 108 indirectly from one or more external sequencers that are not coupled to system 100. For example, system 100 may receive a sequence information 108 over a communication network from a sequencer, which may be located remotely. Alternatively, a sequence information 108 which has previously been stored on a storage medium, such as a hard disk drive or optical storage medium, may be input into system 100.

[0118] System 100 may also receive a sequence information file 108 or fragments at a later time. In other words, the profiling of a microbial community in a sample performed by system 100 may be performed in-line with sample collection, fragment extraction, and fragment sequencing, but all of the steps may be handled separately and / or in a stepwise fashion.

[0119] System 100 may operate under the control of a sequencer that sequences the fragments extracted from a biological sample, but no connected processing or even direct communication between system 100 and sequencer is required. Instead the profiling of the microbial community in a biological sample performed by system 100 and a sequencer is required. In embodiments of this type, profiling of a microbial community in a biological sample performed by system 100 may be performed separately from sample collection, fragment extraction and / or fragment sequencing.

[0120] Figure 3 illustrates an exemplary method 300 that may be performed to profile microorganisms in a biological sample. In some embodiments, the steps of process 300 are performed by processing unit 102. In some embodiments, the processing unit 302 receives sequence information 301. The sequence information may be in the form of a sequence information file, such as that described above and elsewhere herein. The sequence information is derived from a biological sample containing genetic material from one or more microorganisms. In some preferred embodiments, the sequence information may include fragment reads. In non- limiting embodiments, the fragment reads may be unassembled fragment reads (e.g.,JAWS Ref: 750704PCT unassembled nucleotide fragment reads). In some embodiments, the sequence information may have been derived from genetic material contained in a sample (e.g., fragment reads produced by extracting fragments of the genetic material from the sample and sequencing the extracted fragments). In some embodiments, the genetic material may be from one or more microorganisms present in a microbiome (for example, a gut microbiome).

[0121] The methods described above and elsewhere herein may utilize a reference database (e.g., reference database 304) containing the genomic identities of microorganisms (i.e., a reference genome database). In some non-limiting embodiments, the reference database may be a microbial genome database (e.g., a bacterial genome database, or a database of genomes from a plurality of microbes, like bacteria, archaea, eukaryotes, and viruses). In some of the same embodiments and other embodiments, the reference database includes sequences present in the publicly accessible GenBank database. The reference database typically contains reference sequence information. In some embodiments, the reference sequence information may be, for example, assembled or partially assembled sequence information. In some embodiments the reference information may be complete or partial genome sequence information. In some preferred embodiments, the reference database is dereplicated. In some of the same embodiments and other embodiments, the reference database comprises phylogenomic information. In some preferred embodiments, the reference genomes are clustered into species clusters. In some embodiments, the reference genomes are clustered into genus, sub-species, or strain clusters. In some embodiments, the reference genome information is represented as k-mers, signatures or motifs.

[0122] In some embodiments, the methods may include comparing fragment reads (e.g., unassembled nucleotide fragment reads) included in the received sequence information (e.g., sequence information file 108) with reference sequence information contained in a dereplicated reference genome database. These steps are typically performed by processor 303. In some examples, the methods are performed as probabilistic methods and may include, but are not limited to, perfect matching, subsequence uniqueness, pattern matching, multiple sub- sequence matching within n length, inexact matching, seed and extend, distance measurements, phylogenetic tree mapping, decision trees, large language models or hidden Markov models. In one non-limiting embodiment, the probabilistic methods performed in step may include probabilistic matching. In some embodiments, fragment reads are compared against reference sequences using pairwise sequence alignments, multiple sequence alignments, global alignments, local alignments, heuristic approaches (e.g., BLAST or BWA), hidden Markov models, alignment-free methods (e.g., k-mer based approaches or spectral methods such as Fourier or wavelet transform of sequence data), or machine learning approaches (e.g., decision trees or large language models).

[0123] In some embodiments, a plurality of the DNA fragment sequences is each aligned to genomic sequences included in the dereplicated reference genome database. Typically, at least 1%, 5%, 10%, 15%, 20%, 25%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%,JAWS Ref: 750704PCT 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the DNA fragment sequences are aligned to genomic sequences in the dereplicated reference genome database. In some preferred embodiments, substantially all of the DNA fragment sequences are aligned to genomic sequences in the dereplicated reference genome database.

[0124] In some preferred embodiments, the methods include calculating the percent coverage (PC) for at least one gene or genomic sequence in a reference database containing a plurality of genomes, genes or other nucleotide sequences to which one or more read sequences align. For example, in some embodiments the PC is calculated for each gene or genomic sequence in the dereplicated reference database. To profile the microbial community, microorganisms corresponding to those gene or genomic sequences with a PC above a predetermined threshold are considered to be present in the microbial community. Conversely, microorganisms corresponding to those gene or genomic sequences with a PC under a predetermined threshold are considered not to be present in the microbial community.

[0125] The percent coverage (PC) is a determination of percentage of nucleotides of the reference genome or gene sequences that have aligned nucleotides from the plurality of nucleotide sequences from the sample in the assay, divided by the total mappable number of nucleotides in the reference genome or gene sequence. PC is defined as, total nucleotides mapped total mappable nucleotidesembodiments, the methods include calculating the percent coverage error (PCE) for at least one gene or genomic sequence in the reference databases containing a plurality of genomes, genes or other nucleotide sequences to which one or more read sequences align. For example, in some embodiments the PCE is calculated for each gene or genomic sequence in the dereplicated reference database. To profile the microbial community, microorganisms corresponding to those gene or genomic sequences with a PCE under a predetermined threshold are considered to be present in the microbial community. Conversely, microorganisms corresponding to those gene or genomic sequences with a PCE over a predetermined threshold are considered not to be present in the microbial community.

[0127] In some embodiments, the predetermined threshold for PC is selected from about 0.01% to 100%. More particularly, the predetermined threshold for PC is selected from about 1% to 50%. Even more particularly, the predetermined threshold for PC is about 1%, 2%, 3%, 4%, or 5%. For example, the predetermined threshold can be about 5%. The specific threshold PC that is sufficient for the particular application being performed would be readily determined by a skilled practitioner in this field (for example, based upon testing the application with computer simulated data across a range of PC values, calculating the specificity and sensitivity and determining the required sensitivity and specificity comparing the requirements of the intended use of the application).JAWS Ref: 750704PCT

[0128] The percent coverage error (PCE) is a determination of the error in the empirically observed percent coverage (PC) compared to the theoretical percent coverage. ^^^^^^(TPC − PC) × 100PCE =TPC wherein TPC is the theoretical percent coverage (given by 1- e-C); and C is average depth of coverage.

[0129] In some embodiments, the predetermined threshold for PCE is selected from about 20% to about 150%. More particularly, the predetermined threshold for PCE is selected from about 50% to about 100%. Even more particularly, the predetermined threshold for PCE is about 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90%. For example, the predetermined threshold can be about 75%. The specific threshold PCE that is sufficient for the particular application being performed would be readily determined by a skilled practitioner in this field (for example, based upon testing the application with computer simulated data across a range of PCE values, calculating the equirement for specificity and sensitivity and determining the required sensitivity and specificity comparing the requirements of the intended use of the application).

[0130] The calculation of PCE, in combination with a dereplicated reference genome or gene sequence database, allows for the methods described above and elsewhere herein to have sufficient sensitivity to determine the identities of the organisms contained in the sample. Therefore, in some embodiments the profiling methods include determining the identities of the organisms contained in the sample at the species level and / or strain level.

[0131] In some of the same embodiments and some other embodiments, the methods further comprise the step of calculating percent coverage for at least one gene or genomic sequence to which one or more nucleic acid read sequences align. For example, percent coverage may be calculated for every gene or genomic sequence to which one or more read sequences align. The percent coverage can then be used to determine the likelihood of the organism to which the reference sequence relates is present or absent in the microbial community. For example, reference sequences with a percent coverage greater than about 1% are likely to be present in the microbial community. In another example, a percent coverage greater than about 5% indicates that the organism to which the gene or genomic reference sequence relates is even more likely to be present in the microbial community. In more specific examples, a genomic reference sequence with a percent coverage greater than about 10% indicates that the microorganism to which the genomic reference sequence relates is even more likely to be present in the microbial community. In even more specific embodiments, a reference sequence with a percent coverage greater than about 15% indicates that the microorganism to which the gene sequence or genomic reference sequence relates is considered likely to be present in the microbial community. In yet other specific embodiments, a reference sequence with a percent coverage greater than about 20% indicates that the microorganism to which the geneJAWS Ref: 750704PCT sequence, genomic reference sequence relates is considered likely to be present in the microbial community. In yet other specific embodiments, a reference sequence with a percent coverage greater than about 30% indicates that the microorganism or gene to which the gene sequence or genomic reference sequence relates is considered likely to be present in the microbial community. A person of skill in the art would determine the specific threshold of percent coverage that is sufficient for the particular application being performed (i.e., based upon the requirement for specificity and sensitivity). A person of skill in the art would determine the specific threshold of percent coverage that is sufficient for the particular application being performed (i.e., based upon the requirement for specificity and sensitivity).

[0132] Conversely, in some of the same embodiments and other embodiments, a genomic reference sequence with a percent coverage less than about 5% is considered likely not to be present in the microbial community. In some specific embodiments, a genomic reference sequence with a percent coverage less than about 10% is considered likely not present in the microbial community. In other specific embodiments, a genomic reference sequence with a percent coverage less than about 15% is considered likely not present in the microbial community. In other specific embodiments, a genomic reference sequence with a percent coverage less than about 30% is considered likely not present in the microbial community.

[0133] In some embodiments, in order for a species to be determined as likely present in the microbial community the cumulative percent coverage of all genomic sequences relating to that species must be between about 5% and about 50%. In other words, the cumulative total of the percent coverage of a plurality of reference sequences all relating to a single species or gene can be used to determine whether that species or gene is likely to be present or absent in the microbial community. By way of an illustrative example, if the cumulative total is at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% then the species is determined as likely present in the microbial community. In some embodiments, in order for a species to be determined as likely present in the microbial community the cumulative percent coverage of all genomic sequences relating to that specific species must be between about 10% and about 40%. In more specific embodiments, in order for a species to be determined as being likely present in the microbial community the cumulative percent coverage of all genomic sequences that relate to that specific species must be between about 15% and about 35%. In even more specific embodiments, in order for a species to be determined as being likely present in the microbial community the cumulative percent coverage of all genomic sequences that relating to that specific species must be between at least about 20% and about 30%. For example, the cumulative percent coverage of all genomic sequences that relate to the same species must be between about 15% and about 35%, in order for that species to be determined to be likely present in the microbial community, a person of skill in the art would determine the specific threshold of cumulative percent coverage that is sufficient for the particular application being performed (i.e., based upon the requirement for specificity and sensitivity).JAWS Ref: 750704PCT

[0134] In some of the same embodiments and other embodiments, the depth of coverage is calculated for at least one genomic sequence to which one or more read sequence aligns. The average depth of coverage (C) is calculated by the read length (in base pairs) multiplied by the number or reads mapped to a genomic sequence, divided by the length of the genomic sequence (in base pairs). In other words, the depth of coverage is the total number of mapped nucleotides divided by the length of the genome. Depth of coverage ranges from 0 to n, and is generally expressed as a “fold coverage,” for example 5.3× (i.e., the average coverage by sequences of each nucleotide in the reference sequence is 5.3). The depth of coverage can be used to accurately determine the relative abundance of the organism in the microbial community. This can be done by comparing the relative abundance of two or more organisms, in order to assess the relative abundance of each in the microbial community. A person of skill in the art would determine the specific threshold of depth of coverage that is sufficient for the particular application being performed (i.e., based upon the requirement for specificity and sensitivity). For example, in some embodiments, a microorganism is determined as being present in the sample if the depth of coverage is at least about 1 x 10-9, 1 x 10-8, 1 x 10-7, 1 x 10-6, 1 x 10-5, 1 x 10-4, 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, or greater than 5.0 times.

[0135] Upon broadly profiling the microbial community as described herein, the presence or absence of specific genomic profiling mapping and alignment methods defined above, are retrieved. The presence of specific genomic loci are then mapped to these reads, in order to confirm the presence of the target pathogen.

[0136] For an organism to be confirmed as present in sample, at least one specific genomic locus must be found.

[0137] In some of the same embodiments and other embodiments, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 80, at least 90, at least 100, specific genomic loci are required to be found. A person of skill in the art would determine the specific threshold for number of specific genomic loci that are sufficient for the particular application being performed (i.e., based upon the requirement for specificity and sensitivity).

[0138] In some preferred embodiments the methods described above and elsewhere herein may identify such specific genomic loci within a reference sequence by comparing subsequences within the reference sequence to subsequences within a different reference sequence. The percentage identity (PI) is the percentage of the nucleotides in the query subsequence that match the nucleotides in the subject subsequence. The PI of the subsequences between different sub-species, species, genera, gene or different gene groups can be used to determine how specific that subsequence is to the sub-species, species, genus, gene group or gene. If the subsequence is very similar to other subsequences from the same species, but different to subsequences from a different species, then it can be considered to be a genomicJAWS Ref: 750704PCT locus of interest. A person of skill in the art would be able to determine the PI of the two subsequences to determine whether they were considered sufficiently similar or different between different reference groups (e.g. different species), in order to determine whether these subsequences are a genomic loci of interest. In some of the same embodiments and other embodiments, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or exactly 100% PI may be used to determine the required level of similarity.

[0139] In some preferred embodiments the methods described above and elsewhere herein may utilize a trait-specific database catalogue. The trait-specific database catalogue may contain trait-specific reference sequence information (i.e., sequence information contained in the trait-specific database catalogue may be associated with one or more particular organism traits). The trait-specific reference sequence information may be, for example, closed- genomes, draft genomes, contigs, and / or short-reads may be associated with a particular organism trait. Particular organism traits with which the sequence information contained in the trait-specific database catalogue include, but are not limited to, virulence (i.e., fitness) factors, antibiotic resistance traits, metabolic traits etc..

[0140] In some embodiments, if the sample contains genetic material from one or more microorganisms identified in the reference database (i.e., known organisms), the methods may include determining the identities of the one or more organisms contained in the sample and identified in the reference database. In some embodiments, if the sample contains genetic material from one or more organisms not identified in the reference database (i.e., unknown organisms), an additional step that includes determining the identities of the organisms in the reference database that are nearest neighbours to the one or more organisms contained in the sample and not identified in the reference database allows resolving the taxonomic identity of the organisms.

[0141] In some preferred embodiments, substantially all of the read sequences from the biological sample are used in the reference database alignment analysis. These embodiments are particularly beneficial in the calculation of relative abundance of each microorganism in the microbial community. In some of the same embodiments and other embodiments, at least about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 100% of the fragment reads are used in the reference sequence alignment analysis.

[0142] In some embodiments, the system may perform a quality check of the received sequence information. If the quality of the received sequence information is determined to be good, the method may proceed to the alignment analysis. However, if the quality of the received sequence information is determined to be poor, the received sequence information may go through methods of correction before proceeding to the next step of the method. In someJAWS Ref: 750704PCT embodiments, the quality check is performed after the DNA fragments are sequenced, because the quality of data may be important for various downstream analyses, such as sequence assembly, single nucleotide polymorphisms identification, gene expression studies as well as microbial identification. Several sequence artefacts, including read errors (base calling errors and small insertions / deletions), poor quality reads and / or primer / adaptor contaminations are quite common in the NGS data and can impose a significant reduction on the accuracy of the downstream sequence processing / analysis. The quality check and subsequent correction removes these sequence artefacts before downstream analyses to reduce erroneous conclusions. In some embodiments, the quality check may be performed using a quality score assigned to the received sequence information software integrated into sequencing platforms(s) (e.g., sequencing unit). In some non-limiting embodiments, reads with quality scores of at least Q5, of at least Q10 or of at least Q15 are included, and software to trim the ends of sequences is applied.

[0143] In some embodiments, the present invention may enable precise determination of microbial populations in a given sample with respect to the specific taxa (e.g., genus, species, sub-species, and / or strain) of archaea, bacteria, viruses, phage, prophage, eukaryotes including helminths, worms, animals, protozoa, invertebrates, microsporidia and or fungi, or nucleic acid fragments including plasmids, transposons, retrotransposons and mobile genetic elements. Some particular embodiments of the present invention may enable simultaneous identification of a plurality of organisms in a given sample with a single test without having any prior knowledge of organisms present in the sample. Some particular embodiments of the present invention may distinguish between very similar or interrelated species, sub-species and strains for medical, agricultural, food, wastewater surveillance, epidemiological and industrial applications.

[0144] Some particular embodiments of the present invention may rapidly determine background bacterial population or microbiomes, mycobiomes (fungi) and viromes (viruses), at the species and / or subspecies or strain levels.

[0145] In some embodiments, in cases when a microbial sequence does not exist in the reference database, the nearest neighbours of the microbial sequence may be identified to enable resolving the taxonomic classification of the unknown sequence. Some embodiments of the present invention may query one or more specific database catalogues of nucleic acid “signature” sequences or genomes to confirm the presence of particular traits of phenotypes of interest, including, but not limited to, antibiotic resistance traits, and pathogenicity traits.

[0146] In some embodiments, in cases when a microbial sequence does not exist in the reference database, the sequences may be assembled together, and a metagenome assembled genome (MAG) recovered from the assembled contiguous sequences. The MAG can then be taxonomically identified to its’ most recent ancestor. Some embodiments of the present invention may search this MAG and the traits encoded against one or more specific databaseJAWS Ref: 750704PCT catalogues of nucleic acid “signature” sequences or genomes to confirm the presence of particular traits of phenotypes of interest, including, but not limited to, antibiotic resistance traits, and pathogenicity traits.

[0147] In some embodiments, single nucleotide variants (SNVs) comprising single nucleotide polymorphisms (SNPs) or small insertions of deletions (INDELs), are identified within the plurality of nucleotide sequences by analysis methods such as that described in Anyansi et al., Computational Methods for Strain-Level Microbial Detection in Colony and Metagenome Sequencing Data, 2020, Front. Microbiol, 11, for example. These SNVs occur within genes conferring specific traits such as antimicrobial resistance. 5. Specific genomic loci.

[0148] Upon broadly profiling the microbial community as described in section 4, the presence or absence specific genomic loci may be used to provide additional sensitivity and / or specificity of the detection of pathogens or microorganisms.

[0149] In embodiments of this type, read alignments assigned to target microorganism species, in the profiling methods defined above, are retrieved. The presence of specific genomic loci are then mapped to these reads, in order to confirm the presence of the target microorganism.

[0150] For a organisms to be confirmed as present in sample, at least one specific genomic locus must be found.

[0151] In some of the same embodiments and other embodiments, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 80, at least 90, at least 100, specific genomic loci are required to be found. A person of skill in the art would determine the specific threshold for number of specific genomic loci that are sufficient for the particular application being performed (i.e., based upon the requirement for specificity and sensitivity). 6. Result combinations and pathogenicity logic

[0152] In some embodiments, the nucleic acid sequencing data (e.g., nucleotide sequence data store prepared for identifying the presence of a subset of microorganisms and / or antimicrobial resistance markers in a biological or non-biological sample (e.g., sample)) comprises output or results data from the sequencing and / or mapping (e.g., result set), which can be performed, as described above, using any sequencing and / or mapping method as will be apparent to one skilled in the art.

[0153] In some embodiments, some or all of the nucleic acid sequencing data is accessed via a system (e.g., in accordance with. the example system 100 embodiments described above) for review and / or visualization. In some embodiments, the review and / or visualization is performed on the display of a computer. In some embodiments, the review and / or visualization is performed using a cloud-based interface such as an online portal.JAWS Ref: 750704PCT

[0154] In some embodiments, some or all of the nucleic acid sequencing data is transmitted from a first system for performing sequencing and / or mapping analysis, to a second system (e.g., in accordance with the example system embodiments described above) for performing review and / or visualization. In some embodiments, some or all of the nucleic acid sequencing data is transmitted from a first system for performing sequencing and / or mapping analysis, to a cloud-based interface, such as an online portal for performing the review and / or visualization. In some embodiments, the sequencing and / or mapping analysis is performed using an analysis pipeline.

[0155] In some embodiments, the method comprises generating an alert when no nucleic acid sequencing data is available to perform the method (e.g., receiving an email notification when data upload fails).

[0156] In some embodiments, the review and / or visualization is performed on the same system as the sequencing and / or mapping analysis, where the sequencing, mapping, review, and / or visualization of some or all of the nucleic acid sequencing data is performed within an analysis workflow. In some embodiments, the sequencing, mapping, review, and / or visualization is performed at a cloud-based interface such as an online portal comprising an analysis pipeline. In some embodiments, the sequencing, mapping, review, and / or visualization is performed using a software program.

[0157] In order that the invention may be readily understood and put into practical effect, particular preferred embodiments will now be described by way of the following non-limiting experimental examples. EXAMPLES EXAMPLE 1 Construction of a Genome Reference Database.

[0158] Microba Genome Databases (MGDB) is a dereplicated set of genomes that aims to provide high genomic coverage across all bacterial, archaeal, microbial eukaryote and viral species. Individual reference genome databases are built for prokaryote (bacterial and archaeal), fungal, protozoan, invertebrate, viruses, and quality control genomes (human, mouse, dog, cat, vector sequences, phiX). These database partitions are used by the Microba Community Profiler (MCP) and MCP (diagnostic) (MCPdx). Prokaryotic Genome Database

[0159] The Microba Prokaryotic Genome Database (MProkDB) is a dereplicated set of bacterial and archaeal genomes covering a wide range of habitats, including extensive coverate of the human gut microbiome.

[0160] MProkDB adopts the Genome Taxonomy Database (GTDB, Parks et al., 2018, 2020, https: / / gtdb.ecogenomic.org / ) as the primary source of genome sequences and species classification. GTDB defines species based on an average nucleotide identity (ANI) andJAWS Ref: 750704PCT alignment fraction (AF) criteria. Under this definition, genomes are grouped into operational species clusters (OSC), based on ANI distances no less than 95% ANI and an AF ≥50%. Several resources, including NCBI, now use ANI to circumscribe or validate species assignments (Ciufo et al., 2018).

[0161] An important aspect of this definition is that each species is defined by a single representative genome within the OSC, which acts as the nomenclatural type (i.e., reference points to which a name for the OSC is affixed). Ideally, this genome is assembled from the type strain of the species (i.e., the strain selected by the scientific community as the nomenclatural type of a species) though this is not always possible. MProkDB spans all species defined by the GTDB along with additional OSCs formed from metagenome-assembled genomes (MAGs) (see below). For target pathogen species, the differences between OSCs and traditional clinical classifications are limited to the separation of single species into multiple species (data not shownError! Reference source not found.). The inventors mapped OSCs to target pathogen species for consistency of nomenclature and ease of clinical interpretation.

[0162] The use of a broad genome database such as MProkDB rather than a specific database of target genomes only, is of critical importance to accurate taxonomic profiling of metagenomic samples, as it provides the greatest likelihood that read alignments are correctly placed. Building the prokaryotic reference database

[0163] MProkDB was built from GTDB version R04-RS89, metagenome- assembled genomes (MAGs) mined by the inventors, and MAGs mined by other researchers (Table 1). The MAGs mined by the inventors were mined from cohort samples obtained from the general Australian population, as well as samples that are available in the public domain. MAGs and isolate genomes were also obtained from five major studies published in 2019 focused on the human gut microbiome (Almeida et al., 2019, 2020; Forster et al., 2019; Nayfach et al., 2019; Pasolli et al., 2019; Zou et al., 2019). TABLE 1 STATISTICS OFGENOMERESOURCES ANDSTUDIESCONSIDERED FORMPROKDB.JAWS Ref: 750704PCT# Number of samples with >1 genome passing QC & Three genomes were removed as they were flagged as low quality

[0164] Constructing the microbial database c of seven steps, as illustrated in Figure 1. Briefly, the present inventors obtained genomes from human gut studies. Human gut associated MAGs and isolate genomes not included in the GTDB were downloaded for inclusion in the MProkDB. Stool samples from a general Australian population and SRA samples were mined for genomes using the Microba MAG-mining pipeline (Microba MAGs). All genomes were collected and screened for quality. Only genomes meeting the following criteria were retained for consideration: (i) completeness ≥80%, (ii) contamination ≤5%, (iii) number of scaffolds <1000, and (iv) N50 ≥5kb.

[0165] The characterized MAGs and isolate genomes from studies not included in the GTDB were assigned to the species clusters defined by the GTDB. Species assignments were performed using the species-specific ANI circumscription criteria specified by the GTDB. This is typically 95% ANI, though can be up to 97% when required to distinguish between closely related species (as described in Parks et al., 2020). MAGs and isolate genomes that could not be assigned to a GTDB species were grouped into operational species clusters (OSCs). This is done using the same species definition and clustering method used by GTDB (Parks et al., 2020). Briefly, genomes were clustered using a greedy approach based on genome quality, as determined by CheckM (Parks et al., 2015). The highest-quality unassigned genome is selected and all unassigned genomes within a species-specific ANI threshold (typically 95%) to this genome form an OSC. This step is repeated until all genomes have been assigned to an OSC. OSC are assigned an automatically generated specific species name of the form MICx (e.g., Escherichia MIC1234).

[0166] The OSCs were then classified according to the GTDB taxonomy. This is accomplished by running the representative genome of each OSC through the GTDB-Tk (Chaumeil et al., 2020), a publicly available toolkit for classifying genomes according to the GTDB

[0167] As OSCs represent new species, the GTDB classifications obtained in the previous step will necessarily be incomplete taxonomy assignments. In some instances, OSCs may be highly novel and lack genus, family, or order classifications within the GTDB. Internally, OSCs are assigned taxonomic classification information at all taxonomic ranks. This is accomplished by propagating the most specific GTDB taxonomic assignment to lower taxonomicJAWS Ref: 750704PCT ranks. For example, an OSC classified by the GTDB-Tk as d__Bacteria; p__Proteobacteria; c__Gammaproteobacteria; o__Enterobacteriales; f__Enterobacteriaceae and with a specific species name of MIC1234 will be given the genus name g__Enterobacteriaceae MIC1234 and the species name Enterobacteriaceae MIC1234, and interpreted as indicating the genome represents a novel genus and species relative to the GTDB taxonomy used to construct the MProkDB.

[0168] In order to reduce the size of the MProkDB, genomes within a species cluster were dereplicated with the goal of covering the genomic diversity of each species. Dereplication was performed in a greedy manner based on genome quality (defined as completeness – 5*contamination + 100I + 100C, where I and C are 1 if a genome is an isolate or complete assembly, respectively, according to metadata in the NCBI Assembly database (Kitts et al., 2016) or are 0 otherwise. This favours selection of high-quality genomes with preference given to isolates and assemblies annotated as being complete. Dereplication was performed using a greedy dereplication strategy where genomes were processed in descending order of genome quality and a genome was selected as a representative if its ANI was ≤98% or alignment fraction (AF) >90% to all previously selected representatives as calculated using FastANI (Jain et al., 2018). If >100 representatives occur in a species, the 100 highest-quality representatives fwere selected. Notably, only 437 of the 28,246 (1.5%) species clusters contain >100 representative genomes.

[0169] BWA (Li, 2013) indices were created in order to map the sequencing reads from metagenomic samples to the genomes in the MProkDB. The MProkDB is divided into multiple BWA indices (shards) of a roughly equal number of genomes. Creating multiple indices allows for distributed mapping of reads while keeping genomes in the same genus together reduces the number of mappings that must be retained. MProkDB is divided into 32 shards as this allows all genomes from even the largest genera, Prevotella at 2,355 genomes, to be kept together. Statistics for prokaryotic reference database

[0170] MProkDB was built from GTDB R04-RS89 and MAGs mined from human stool samples as well as samples available in the public domain (Table 1). Together these sources span 411,415 quality-controlled redundant genomes. The resulting database spans 28,246 prokaryote species of which 3,540 are novel relative to GTDB R04-RS89 and 13,726 are comprised exclusively of environmental genomes (i.e., MAGs or single amplified genomes, SAGs). There are 625 species comprised exclusively of MAGs recovered from collected samples or mined by the inventors from samples in the NCBI Sequence Read Archive (SRA). MProkDB v2 contains 73,646 genomes after selecting representatives within each of the 28,246 species clusters. Eukaryotic Genome Databases

[0171] Microbial eukaryotes were divided into three reference databases covering fungi, protozoa, and invertebrate genomes obtained from the NCBI Assembly database.JAWS Ref: 750704PCT The fungal and protozoan databases cover all species with available reference genomes, whereas the invertebrate database was limited to known human pathogens and species that occur in the human gut. Taxonomic curation of microbial eukaryotic genomes

[0172] There is no systematic genome-based taxonomic framework of eukaryotes comparable to GTDB that the present inventors could rely upon. Therefore, fungal, protozoan, and invertebrate genomes were classified with the NCBI Taxonomy. The present inventors undertook an evaluation of these NCBI classifications for specific eukaryotic genera in order to confirm that they conformed with the systematic operational species cluster (OSC) expectations of MCPdx. Specifically, the present inventors adopted the operational definition for prokaryotic species to guide the taxonomic curation of eukaryotic genomes into OSCs. Three distinct categories of reclassification were undertaken: (i) reclassification of misclassified genomes; (ii) amalgamation of species into a single OSC; and (iii) division of genomes into multiple OSCs. Invertebrate genomes were not included in these reclassifications as reliable ANI values could not be calculated on these large and variable genomes (frequently >100 Mb).

[0173] A number of parameters were applied during manual curation in order to identify genomes requiring reclassification to conform to the OSC expectations of MCPdx. First, genomes with an ANI ≥95% and an AF ≥50% to the representative genomes of a species were considered in the same species cluster. Second, genomes were considered incorrectly classified if the ANI and AF to the nearest representative genome is ≥95% and ≥50%, respectively, and this representative genome had a different species classification (e.g., reclassification of a Blastocystis hominis genomes to B. sp. subtype 3). Third, MCPdx is not designed to distinguish between species containing strains with highly similar ANI values as such species are over- classified in that they are equivalent to subspecies belonging to the same species. These species must be reclassified as a single OSC to avoid misleading results. Species within a 97% ANI cluster with ≥50% AF were merged into a single OSC. Named species meeting these criteria were renamed as a group (e.g., Cryptosporidium hominis, C. parvum, C. cuniculus, and C. tyzzeri were combined into a single OSC named Cryptosporidium group A). In contrast, placeholder species merged with named species were considered nomenclatural synonyms (e.g., Brettanomyces sp. HC-2020a was merged with B. bruxellensis and all genomes in this OSC classified as B. bruxellensis). Finally, genomes from the same species forming distinct ANI clusters at <90% ANI were divided into multiple OSCs. For such species, the nomenclature methodology of the GTDB was applied to label species with an alphabetical suffix with the cluster containing the representative genome retaining the validly published name (e.g., genomes classified as Giardia intestinalis at NCBI were divided into five OSC named G. intestinalis, G. intestinalis_A, G. intestinalis_B, G. intestinalis_C, G. intestinalis_D).

[0174] The present inventors mapped the eukaryotic OSCs to target pathogen species for consistency of nomenclature and ease of clinical interpretation.JAWS Ref: 750704PCT Removal of contamination in eukaryotic genomes

[0175] Genome assemblies can contain DNA from contaminating species. Contamination can be due to co-assemblies when a presumed axenic culture contains other species (e.g., bacterial contamination in a protist culture), human DNA from lab technicians, or to mis-assemblies derived from bioinformatic errors. Genome contamination may lead to reads sourced from one species being erroneously aligned to regions of contaminating DNA in a different species, and in rare cases, may cause MCPdx to report a false positive species. In order to remove potential contamination, screening of microbial eukaryotic genomes was performed. A pipeline based partially on the work of Lu & Salzberg (2018) was developed that employs Kraken 2 (Wood et al., 2019) to rapidly identify prokaryotic, viral, and host contamination in a eukaryotic genome as illustrated in Figure 2. First, the present inventors split the genome into 150 bp pseudoreads which overlap by 50 bp. The pseudoreads were then profiled with Kraken 2 which uses a k-mer approach to assign each pseudoread to taxa in a custom reference database comprised of prokaryotic, viral, vector, and common mammalian hosts genomes, including human. Since the reference database does not contain eukaryote DNA (apart from host genomes), reads that are classified are marked as highly conserved or putative contamination. In either case, these regions may cause errors in mapping, and so both are considered problematic.

[0176] Genomic coordinates of classified pseudoreads are then used to determine regions within the eukaryotic genome that should be masked to designate putative contamination. By default, regions >250 bp covered by classified pseudoreads were masked along with any regions <=150 bp located between masked regions. Building the eukaryotic reference databases

[0177] Eukaryotic genome assemblies were obtained from the NCBI Assembly database on December 8, 2021. These genomes range substantially in quality, largely due to the complex and often repetitive nature of eukaryotic genomes. The following criteria were used to exclude lower-quality eukaryotic genomes from the curated reference databases: (i) assemblies targeting only part of the genome as indicated by the NCBI Genome representation field being set to “partial”; (ii) diploid assemblies as indicated by the NCBI Assembly type field being set to diploid or unresolved-diploid; (iii) genomes with < 1 Mb of ungapped sequence data; and (iv) genomes with >= 20% masked as contamination (see, Removal of contamination in eukaryotic genomes);

[0178] Quality control resulted in the exclusion of 200 fungal and protozoa genomes and a further 18 genomes were excluded after manual review (see, Table 3). TABLE 3 Taxonomic Classification of Target Eukaryotic Species in the MGDB Taxonomic FrameworkJAWS Ref: 750704PCTJAWS Ref: 750704PCT

[0179] In order to provide comprehensive coverage of the genomic diversity within a species, genomes within an OSC passing quality control were dereplicated using an ANI and AF criteria of 99% and 90%, respectively. These criteria are stricter than those used for the prokaryotic database as there are relatively few eukaryotic assemblies available. Statistics for eukaryotic reference databases

[0180] The fungi, protozoa, and invertebrate databases were built from the 10,688 genomes passing quality control and dereplicated to 4,962 genomes organized into 3,418 OSCs (see, Table 4). TABLE 4 STATISTICS FOREUKARYOTICREFERENCEDATABASESViral Genome Database

[0181] The Microba Viral Genome Database (MViralDB) is a dereplicated set of viral genomes obtained for the NCBI Assembly database and several large viral metagenomic studies. Operational definition of a viral species

[0182] Viral genomes are organized into species-level, viral operational taxonomic units (vOTUs) using the Minimum Information about an Uncultivated Virus Genome (MIUViG) recommended criteria of 95% ANI over 85% of the length of the shorter sequence (Roux, et al., 2019). These criteria have been interpreted and implemented in different ways (Roux, et al., 2019; Nayfach, et al., 2021; Roux, et al., 2021), and there is currently no consensus on how best to form vOTUs. Microba had adopted the following methodology that aims to reduce computational requirements and form tightly-clustered vOTUs that are likely to be more informative for tracking viral genomes across samples. Specifically, the MViralDB uses a >=95% ANI and >=85% AF criteria for forming vOTUs as follows: (i) Mash (Ondov, et al., 2016) with a k-mer size of 15, sketch size of 500, and maximum distance of 0.09 is used to get initial ANI estimates for all pairwise combinations of genomes. (ii) Viral genomes are dereplicated using a greedy approach where: (1) genomes are sorted in descending order of genome length, (2) the longest genome is selected as a new cluster representative, and (3) unclustered genomes with a Mash ANI >=99.5% to the representative are assigned to the new cluster. Steps 2 and 3 areJAWS Ref: 750704PCT repeated until all genomes have been assigned to a cluster. This step is purely done to reduce the number of genomes for which FastANI (Jain, Rodriguez-R, Phillippy, Konstantinidis, & Aluru, 2018) values must be calculated. (iii) FastANI with a fragment length of 600 and minimum fraction of 0.5 is used to get ANI and AF values between all pairwise combinations of the cluster representatives. A fragment length of 600 was selected as a trade-off between obtaining values correlated with traditional BLAST- and MUMMER-based ANI statistics and computational requirements. (iv) vOTUs are formed using a greedy approach where cluster representatives are first determined and then genomes are assigned to the closest representative. Cluster representatives are determined by: (1) sorting genomes in descending order of genome length, (2) selecting the longest genome to be a new representative, and (3) removing all genomes with an ANI >=95% and AF >=85% from further consideration. Steps 2 and 3 are repeated until all genomes have been processed. Genomes that have not been selected as a cluster representative are then assigned to the closest representative in terms of ANI with AF being used to break ties.

[0183] The present inventors mapped vOTUs to target pathogen species for consistency of nomenclature and ease of clinical interpretation (table not shown). Building the viral reference database

[0184] The viral reference database (MViralDB) was built from 515,090 genomes obtained from the NCBI Assembly database on 23 June 2022 and several large-scale metagenomic studies focused on recovering viral genomes (Table 5). TABLE 5 STATISTICS OF GENOME RESOURCES AND STUDIES CONSIDERED FOR MVIRALDB.

[0185] Construction of the viral database consists of five steps, as illustrated in Figure 3. The present inventors first sought to gather viral genomes to be considered for the viral reference database. This included obtaining all viral genomes from the NCBI Assembly databaseJAWS Ref: 750704PCT and downloading data from large-scale metagenomic studies being considered. CheckV (Nayfach, et al., 2021) was used to obtain viral quality statistics for all viral genomes being considered except for NCBI genomes (as these are not recovered from metagenomic samples and therefore are assumed to be of reasonable quality). The present inventions then identified viral genomes obtained from metagenomic data passing QC based on meeting all the following CheckV criteria: (i) completeness is 90% or greater; (ii) quality is defined as being either ‘Complete’ or ‘High-quality’; (iii) at least two viral genes (within the viral region for proviruses); (iv) at least 4× more viral genes than host genes (within the viral region for proviruses); (v) warning does not indicate greater than one viral region detected; (vi) warning does not indicate contig greater than 1.5× longer than expected genome length; (vii) warning does not indicate high kmer_freq may indicate large duplication; (viii) 500 bp <= genome size <= 5 Mb; and / or (ix) 100% of duplicate genomes are filtered (also applied to NCBI viral genomes).

[0186] Viral genomes were subsequently organized into species-level clusters (vOTUs) as described above. BWA indices are created in order to map the sequencing reads from metagenomic samples to the genomes in the MViralDB. Taxonomic assignment of vOTUs

[0187] Taxonomic classification of vOTUs is determined using a most common vote across the viral genomes that have a taxonomic assignment. Only the NCBI, MGV, IMG / VR, and CHVD datasets provide taxonomic classifications, and these are often of limited resolution. While this means many vOTUs have limited taxonomic resolution, this is a simple, pragmatic approach. The final taxonomy only allows taxa defined by the NCBI viral taxonomy and follows the NCBI classification hierarchy. It is not uncommon for viral species to be undefined at intermediate ranks as viral classification does not enforce classifications at all ranks. In such cases, the most specific classification available is used to establish placeholder names at higher ranks by appended the rank (e.g., the Pandoravirus genus is not assigned to a family, order, class, or phyla at NCBI and thus these ranks are indicated as Pandoravirus_family, Pandoravirus_order, etc.). Statistics for viral reference database

[0188] The viral database consists of 110,734 vOTUs spanning 266,248 genomes passing QC (Table 5) with the majority of vOTUs having a relatively unresolved taxonomic classification (Table 6). This lack of taxonomic resolution emphasizes the benefits ofJAWS Ref: 750704PCT including metagenomic viral genomes, organizing viral genomes into “species-level” vOTUs, and tracking viruses across samples using “species-level” vOTUs which are taxonomy agnostic. The determination of pathogenic vOTUs against non-pathogenic vOTUs is an area that will require investigation as the MetaPanel product is validated against these targets, and ultimately will provide excellent resolution of viral pathogenicity. TABLE 6 PERCENTAGE OF GENOMES AND VOTU WITH TAXONOMIC ASSIGNMENTS AT DIFFERENT RANKSEXAMPLE 2

[0189] MCP is used to estimate the relative abundance of prokaryotic, eukaryotic, and viral species in a metagenomic sample. Reads passing QC are mapped to all genomes across all domains in Micron Genome Databases (MGDB) including comprehensive prokaryotic, protozoan, fungal, and viral databases, in addition to a database of host and other quality control sequences including the human genome. Many reads will map to multiple genomes as MGDB contains closely related strains and there are regions of high genomic similarity between strains from different species. Mapping to all databases is performed in order to properly determine the most likely origin of each read pair. Only read alignments with high Percent Identity and Alignment length (PIA ≥95%) are considered as true alignments. These strict parameters result in the average percent identity and alignment length being >99.5% across typical human faecal samples. Strict alignment filtering and selection is required to ensure reads are assigned to genomes that are closely related to the organisms in the metagenomic sample. The complete workflow for MCP is shown in Figure 5.

[0190] MCP operates on quality controlled paired-end sequencing data; however, MCP maps reads separately and identifies all mappings for each read in each pair with high percent identity and alignment length. Mappings where read pairs are assigned to genomes from different species are discarded as these represent unreliable mappings. The highest quality remaining mappings (i.e., ≥98% PIA) are used to calculate the initial depth of coverage (DoC) and percent coverage (PC) of each genome. This information is used to establish the single best mapping for each read pair by assuming reads with multiple mapping locations should beJAWS Ref: 750704PCT assigned to the genome with the highest DoC. Best mappings are retained and used to calculate species abundances if they have either (i) percent identity and alignment ≥98% or (ii) percent identity and alignment ≥95% and are assigned to a genome with sufficient coverage that it is likely closely related to a strain present in the sample. Best mappings failing these criteria are classified as unmapped as they map to genomes without sufficient coverage to be confidently declared present in the sample. Prokaryotic and eukaryotic species containing a genome with a PC ≥5% and a PCE ≤75% have sufficient evidence to declare them present in the sample. Viral species are reported at a PC ≥50% and a PCE ≤75%.

[0191] Reads assigned to target pathogen species are further interrogated using MCPdx and in some instances lower PC values are permitted (see Identification of Target Pathogens with MCPdx). Reads assigned to species failing the general PC and PCE criteria are classified as unassigned, and the species is not reported as present. Final species abundances are determined by considering the depth of coverage of all genomes within the species classified as present in the sample along with consideration of the number of unassigned and unmapped reads. Robust identification of prokaryotic and eukaryotic species

[0192] Microbial genomes from different species can share a substantial amount of sequence similarity. This similarity presents challenges in accurately identifying species. MCP has several parameters that determine when a species is identified as present (see Figure 5). These parameters must be set to provide robust identification of species at an acceptable false positive rate. This was accomplished by using a set of in silico communities to evaluate the performance of MCP across a parameter sweep of key parameters.

[0193] Testing metagenomic profilers on simple mock communities consisting of only a handful of species (such as the ATCC MSA-1003 Standard) is insufficient to adequately evaluate profiler performance. Real faecal communities are composed of many hundreds of species with relative abundances >0.01%, as well as novel species and novel strains that vary from those in reference databases. MCP was rigorously tested using a wide range of carefully constructed mock communities designed to stress test its performance against known results. Additionally, the performance of MCP was compared to nine publicly available metagenomic classifiers. These classifiers use a variety of methodologies and have previously been shown to be among the best performing classifiers (Lindgreen et al., 2016; Sczyrba et al., 2017; Seppey et al., 2020; Ye et al., 2019). The results of this comparative analysis indicates that MCP produces less false positive species calls than all other classifiers, with the highest precision, and the lowest limit of detection. The results of this study are presented in Parks et al. (2021). Robust identification of viral species

[0194] The present inventors validated the use of MCP for identifying viral species. Parameter tuning results in the maximum percent coverage error (PCE) and minimum genome coverage (GC) parameters of MCP being set to 75% and 50%, respectively (Figure 5; Appendix B). PCE was set to 75% as the performance of MCP was found to be largely insensitiveJAWS Ref: 750704PCT to this parameter and 75% is used for the identification of prokaryotic and eukaryotic species. GC was set to 50% to ensure false positive predictions were sufficiently rare that the precision of MCP would be high (>94% on validation mocks). This relatively high genome coverage criterion also addresses the issue that some viral MAGs are likely to contain prokaryotic contamination due to inexact identification of start and stop coordinates of prophage. Ultimately, setting GC to 50% is a trade-off between precision and recall as many viral species will be present at <50% genome coverage in a sample and thus effectively below the level of detection of MCP. Modification of viral genome reference database

[0195] A challenge in evaluating MCP performance on viral species is that MViralDB includes all available viral genomes. As such, parameter tuning and validation was performed using a modified version of the MViralDB that is divided into a “testing database” and a set of “holdout” genomes (Figure 6). Genomes in the “testing database” were used to create BWA indices (Li, 2013) for parameter tuning while the “holdout” genomes were used to allow mock communities to be constructed with viral genomes not in the reference DB. Specifically, viral operational taxonomic units (vOTUs) in MViralDB with ≥3 genomes were divided with 2 / 3 of genomes retained for the “testing database” and the other 1 / 3 used as “holdout” genomes for inclusion in mock communities. All genomes in vOTUs with <3 genomes were included in the “testing database”. This procedure was repeated independently to create distinct testing databases and holdout genomes for parameter tuning and validation. Validation results on independent in silico mock communities

[0196] To ensure parameter tuning and validation are independent, a new MViralDB “testing database” and set of “holdout” genomes was created along with a new set of 100 in silico mock communities (Figure 6). In contrast to the “testing database” produced for parameter tuning, the dereplication performed for validation was modified to ensure the representative of each vOTU was retained. This is expected to produce results that will be in better agreement to what would be obtained using the full MViralDB as retaining the vOTU representatives ensures a good sampling of the viral diversity captured by MViralDB.

[0197] Read pairs from the mock communities were mapped to the MViralDB “testing database” and MCP profiling performed with PCE set to 75% and GC set to 25%, 50%, or 75% (Table 7). As expected, results on the parameter tuning and validation mocks are similar. This suggests the performance of MCP is robust to the exact set of genomes comprising the viral “testing” DB and to the composition of the mock communities. TABLE 7 PERFORMANCE OF MCP V3 AVERAGED OVER 100 MOCK COMMUNITIES FOR PCE = 75% AND VARYING GCSETTINGS. RESULTS FOR BOTH PARAMETER TUNING AND VALIDATION ARE SHOWN.JAWS Ref: 750704PCT

[0198] The false positives (FPs) reported by MCP using GC = 50% (selected default) do not show any systematic bias towards specific vOTUs or species. In total, there were 145 FP predictions across the 100 mocks and these spanned 127 distinct vOTUs. The most common vOTUs were only identified as FPs in three samples. Almost all the FP vOTUs were from a novel species with no genus affiliation (124; s_) or specific name (11; e.g., s_Bacteriophage sp.), and no named species was identified as a FP more than once.

[0199] The reported false negatives (FN) predictions assume that a vOTU represented by even a single read pair should be considered present in the sample. As such, FN predictions and the recall rate of MCP were further explored by systematically filtering out vOTUs based on their depth of coverage in the mock community (Figure 7; Table 8). Theoretically, a genome needs a depth of coverage ≥0.7× before it is expected to have a genome coverage >50% (i.e., the selected GC criterion). Here, filtering was performed on vOTUs, which may contain multiple strains (genomes), by removing a vOTU from consideration if the genome with the highest depth of coverage was below the filtering criterion. At a depth of coverage of 0.7×, MCP reported 13.6 FNs on average over the 100 mock communities and had an average recall of 86.8% (Table 8).

[0200] Interestingly, the recall rate of MCP starts to asymptote at a depth of coverage of 0.7× and only reaches a recall rate of 89.6% even at a depth of coverage of 2.0× (Table 8). No systematic issue was identified in the FNs reported at a depth of coverage of 2.0. Across the 100 samples, 769 FNs were reported and these spanned 677 different vOTUs with no vOTU reported as a FN more than three times.

[0201] vOTU25557 was reported as a FN only three times, but this represents 100% of the mock communities in which this vOTU is present. This vOTU is represented by three genomes in MViralDB of which two were retained in the “testing database”. The ANI between the single holdout genome and the closest genome from this species in the “testing database” is only 96.0% ANI with an AF of 87.5%. A single mock (#33) with vOTU25557 reported as a FN was explored in detail to understand why this FN occurs despite having a depth of coverage of 56×. In this mock, only 38 of the 8,827 pairs simulated from this species were mapped to the “testing” DB and all of these were to genome GCA_007311285.1 in vOTU25557. Nearly all pairs being unmapped is presumably a result of the closest genome in the reference DB being relatively divergent at 96.0% ANI. vOTU25557 is not reported as these 38 pairs result in a percent coverage of only 3.5%.JAWS Ref: 750704PCT TABLE 8 PERFORMANCE OF MCP WHEN CONSIDERING VOTUS ABOVE DIFFERENT DEPTH OF COVERAGE THRESHOLDS AVERAGED OVER 100 MOCK COMMUNITIES

[0202] The reported parameter tuning and validation results only consider mappings to the MViralDB “testing database”. It is possible that prokaryotic reads are erroneously mapping to the viral DB causing FP predictions that would not occur if mappings to the MProkDB were also considered (i.e., prokaryotic reads were assigned to prokaryotic genomes instead of erroneously assigned to a viral genome). To evaluate if this is the case, the viral component of each mock community was mapped to the MViralDB “testing database” and profiled with MCP v3. These profiles are identical to those produced when also considering the prokaryotic component of each mock community with the exception of the number of putative prophage identified (Table 9). This demonstrates that the FP predictions (or any of the performance statistics) are not impacted by the presence of prokaryotic reads. TABLE 9 PERFORMANCE OFMCPAVERAGED OVER100MOCK COMMUNITIES FORPCE = 75%ANDGC = 50%ON THE VALIDATION AND VIRAL COMPONENT ONLY MOCKSonly Viral species profiles

[0203] MCP resolves equally-good, best mappings by selecting the mapping to the genome with the highest depth of coverage. This has proven to be a reasonable strategy as it preferentially assigns reads to highly abundant organisms in the community as such organisms are the most likely to be the true source of a read. However, this method of resolving equally-JAWS Ref: 750704PCT good, best mappings is more complicated for viruses due to their relatively small size and the presence of prophage in prokaryotic genomes. In principle, a read from a prophage contained in both the viral and prokaryotic DB has two valid mappings, but MCP cannot currently resolve such cases.

[0204] In order to produce a viral profile without the conflating impact of prophage in the prokaryotic DB, the viral profile is produced by processing mappings to the viral DB independently (Figure 8). The aim of this methodology is to ensure that reads from prophage are assigned to genomes in the viral DB when possible (i.e., the prophage strain is adequately represented in the reference viral DB). A potential issue with this methodology is that non-viral reads might be erroneously assigned to viral genomes resulting in increased numbers of FP predictions. However, the analysis in paragraph

[0191] provides direct evidence that this is not the case and that the FP predictions observed in the mock communities are due to errors in establishing the correct viral species for all viral reads.

[0205] Based on this finding, MCP is configured such that prokaryotic and eukaryotic profiles are obtained in unison but also including consideration of the viral mapping files (BAM format) but that viral profiles are obtained by considering the viral mapping files only. The rationale here is that unclassified reads are effectively treated by MCP as belonging to unidentified prokaryotic species. As such, it improves the accuracy of profiles to have viral reads assigned to the viral DB which reduces the number of reads reported as unclassified. Furthermore, this allows a more accurate assessment of the fraction of a community that is unclassified which is important when reasoning about the quality of species profiles. Handling of reads from prophage is complicated in this case and will depend on the mapping quality to the viral and prokaryotic DBs, and the MCP criteria for resolving equally-good, best mappings. MCP has been tested extensively with reads being mapped to all genome reference DBs and issues surrounding prophage have not been observed to compromise prokaryotic profiles. This is unsurprising given that prophage ultimately comprise only a small fraction of a community (in terms of number of reads) and a small fraction of a prokaryotic genome (i.e., should not compromise coverage statistics).

[0206] MCP profiles estimate the relative abundance of organisms in a community by considering their genome size and assuming unmapped reads are from an unidentified prokaryotic species with an average genome size. This is convenient and gives good estimates for prokaryotic species profiles but is not ideal when considering the viral species profiles as genome sizes vary more substantially and the majority of the community (in terms of DNA) will be prokaryotic. As such, the abundance estimates for viral species is relative to the number of input reads. Pathogen Detection with the diagnostic MCP (MCPdx)

[0207] The Microba Community Profiler (MCP) achieves superior limit of detection, sensitivity and specificity compared to other available metagenomic classifiers (Parks, et al., 2021). However, pathogen detection requires sensitivity and specificity equivalent to theJAWS Ref: 750704PCT current standard of care (e.g., PCR). MCP is run using diagnostic mode (MCPdx) in order to achieve increased sensitivity while maintaining specificity for the identification of target prokaryotic and eukaryotic pathogens. In MCPdx, genomes from every target species in MetaPanel are systematically interrogated to find regions of high diagnostic value. DNA alignments are further inspected to determine if they overlap these predetermined trusted genomic loci, and the combined evidence used to increase the limit of detection of specific pathogens. The present inventors implemented a two-step process: (i) a database of diagnostic genomic fragments is built for each target species, and (ii) read alignments assigned to a target species by MCP that align to diagnostic regions are identified and counted.

[0208] The basic steps in building a MCPdx database of diagnostic regions are as follows (Figure 9): 1. Species are defined as Operational Species Clusters as described in Construction of Genome Reference Databases. 2. A list of trusted genomic loci is collated for every species in each genus that contains a target pathogen. Core genomic regions are defined as regions shared by the majority of genomes in a species, but that are not conserved between species or replicated within a genome. These regions are identified as follows: a. Simulate DNA sequencing reads from the representative genome in each species. Reads are simulated as 100 bp fragments at overlapping 50 bp intervals. b. Map these fragments back to the corresponding MGDB database and classify fragments as core that map to ≥90% of genomes in the species at ≥90% identity. c. Core fragments that occur more than once in any genome are classified as repeat fragments and discarded. d. Core fragments that occur in another species in the genus with an identity ≥98% are classified as conserved fragments and discarded. 3. Coding regions are determined for each genome and each core fragment is assigned a locus based on the coding region it is contained in.

[0209] A small subset of target pathogenic species were identified to contain insufficient numbers of diagnostic regions and so have been merged and treated as a single species in MCPdx. These species are known to contain highly similar strains possibly due to homologous recombination. This limits the identification of diagnostic loci and therefore the resolving power of MCPdx (Table 10).JAWS Ref: 750704PCT TABLE 10 SPECIESMERGED INTO ASINGLESPECIES INMCPDX

[0210] Reads assigned to a target species by MCP are further interrogated by MCPdx to determine the presence or absence of the species as follows (Figure 10):

[0211] 1. Read alignments assigned to target species by MCP are retrieved.

[0212] 2. Diagnostic loci in each target species that have at least one aligned read that overlaps the locus by ≥140 bp are counted.

[0213] 3. In low abundance species (defined as having less than a quarter of all core loci found in a sample), loci with >5 reads mapping are removed as <2 reads are expected to map per loci at this abundance.

[0214] 4. MCPdx reports a species as present in a sample if the following conditions are met:

[0215] Prokaryotic species: If the percent genome coverage (GC) is >5%, and the total percent species coverage (TSC) is >5% and the percent coverage error (PCE) <75%, or, If GC <5% or TSC <5%, and sequencing reads map to ≥4 diagnostic loci of that species, and the species is the most abundant intrageneric species in the sample;

[0216] Assuming a target sequencing depth of 16 million read pairs, this equates to a theoretical limit of detection of ~0.0000615% relative abundanceJAWS Ref: 750704PCT

[0217] Eukaryotic species: If the percent genome coverage (GC) is >1%, and the total percent species coverage (TSC) is >1% and the percent coverage error (PCE) <75%, or, If GC <1% or TSC <1%, and sequencing reads map to ≥4 diagnostic loci of that species, and the species is the most abundant intrageneric species in the sample;

[0218] Assuming a target sequencing depth of 16 million read pairs, this equates to a theoretical limit of detection of ~0.00003125% relative abundance. In silico Validation of MetaPanel Target Pathogen Identification

[0219] In silico testing refers to the computational generation of data from reference sequences that represents as close as possible a true sample, or extreme forms of difficult samples, and then testing the analytical pipelines with these simulated data. This approach is useful in development and validation of genomic-based analytical pipelines because genomic resources are often available for target organisms enabling simulated data to be created which has precise expected outputs. The informatics tools underpinning the present pathogen detection technology was extensively validated using in silico data sets, both generally for MCP (Parks, et al., 2021), and for specific prokaryotic and eukaryotic organisms with MCPdx (see Species Detection with MCP Diagnostic Mode). EXAMPLE 3 IN SILICO VALIDATION OF TARGET PROKARYOTIC AND EUKARYOTIC SPECIES IDENTIFICATION

[0220] In silico data sets were generated and used to validate the performance of MCPdx on target prokaryotic and eukaryotic species. Two types of mock communities were generated to test the sensitivity and specificity of the diagnostic technology on target bacterial, fungal, invertebrate, and protozoan pathogens (Table 3). MCPdx results on sensitivity mock communities provide information on the number of read pairs which must be assigned to a target species for robust identification. In contrast, results on cross-reactivity mock communities evaluate specificity by examining the ability of MCPdx to assert a target species as absent even in a community containing strains from all non-target intrageneric species. Mock communities were created for MGDB species mapping to target pathogens (Table 3) in order to assess the performance of MCPdx at this more resolved level of taxonomic classification. Quality-control of genomes used for creating mock communities

[0221] MProkDB contains large number of metagenome-assembled genomes (MAGs) of varying quality. Use of genomes with contamination adversely impacts interpretation of results on mock communities. As such, only prokaryotic genomes passing the following genome assembly quality criteria were considered for use in mock communities: - CheckM v1 completeness estimate ≥90%; - CheckM v1 contamination estimate ≤2.5%;JAWS Ref: 750704PCT - N50 ≥ 10 kb; and - Number of contigs ≤ 500

[0222] The present inventors screened the eukaryotic genomes for prokaryotic contamination (see, Removal of contamination in eukaryotic genomes) and filtered from consideration if they failed to meet the quality-control criteria used to exclude genomes from eukaryotic reference databases (see Building the eukaryotic reference databases). Evaluation of MCPdx on prokaryotic and eukaryotic sensitivity mocks

[0223] Sensitivity mocks consist of simulated reads from a single target MGDB species. These mocks are used to evaluate the sensitivity or limit of detection of MCPdx by assessing its ability to correctly assert the target species as present in a sample with progressively fewer read pairs. This is assessed by considering the number of diagnostic loci for the target species identified by MCPdx.

[0224] First, the present inventors randomly selected a genome from the target species, wherein the selected genome is not in the corresponding MGDB reference databases when possible.2 × 150 bp read pairs were simulated with an inner distance of 250 ± 25 bp for each selected genome to achieve 3× depth of coverage. The in silico read pairs with then profiled with MCP.

[0225] Next, the present inventors downsampled the set of read pairs mapped to the target species to 10, 20, 50, 100, 500, and 2500 read pairs. This step is repeated in order to create 10 replicates. Each replicate sample is then profiled with MCPdx.

[0226] The above method is reappeared in order to create three sets of downsampled read pairs using different genomes from the target species. If a species has less than three genomes available, less than three replicates were created.

[0227] A total of 32,220 sensitivity mocks were built which cover 13 fungal, 17 invertebrate, 23 protozoa, and 167 prokaryotic target species. This includes mocks for each of the MGDB species merged in MCPdx (Error! Reference source not found.). Sensitivity mocks were not generated for Tropheryma whipplei or Bacillus_A cereus_AK as these species do not contain any genomes passing the required QC criteria (see Quality-control of genomes used for creating mock communities).

[0228] The present inventors demonstrated that MCPdx robustly identifies the majority of target species when present at sufficient abundance (Table 2). As expected, the number of species identified increases with the number of read pairs assigned to the species. At 500 read pairs, all fungi, protozoan, and the vast majority (>92%) of prokaryotic target species are correctly inferred to be present in all sensitivity mocks by MCPdx (i.e., generally, three distinct strains with 10 independent replicates of 500 read pairs). This demonstrates the robustness of MCPdx to different strains and the genomic origin of read pairs. Results on invertebrate species are less robust with four species failing to be identified by MCPdx even at 2,500 read pairs. As itJAWS Ref: 750704PCT stands, this will affect the sensitivity for these species, but does not preclude the species from being identified with good specificity. The invertebrate genomes raise several challenges. In this instance, the lack of informative loci was attributed to difficulty in robustly inferring loci for these genomes as well as a lack of additional genomes from intragenic species to these targets.

[0229] Notably, the 13 MGDB prokaryotic species not identified across all mocks with 500 reads come from only three pathogens as clinically defined (Table 2; Error! Reference source not found.): Bacillus cereus: Bacillus_A cereus, Bacillus_A cereus_P, Bacillus_A cereus_AD, Bacillus_A thuringiensis_J, Bacillus_A thuringiensis_K Campylobacter concisus: Campylobacter_A concisus_A, Campylobacter_A concisus_C, Campylobacter_A concisus_D, Campylobacter_A concisus_F, Campylobacter_A concisus_J, Campylobacter_A concisus_R, Campylobacter_A concisus_T. Acinetobacter lwoffii: Acinetobacter sp002367455 TABLE 11 Prokaryote Pathogens of InterestJAWS Ref: 750704PCT

[0230] The large number of OSCs comprised of strains clinically defined as B. cereus or C. concisus speaks to the diversity of strains within these clinical species. This diversity poses challenges in identifying diagnostic loci which is the rationale for treating clinical C. concisus strains as two effective OSCs in MCPdx (Table 12). Sensitivity for these OSCs can likely be improved through a more judicious selection of OSCs to merge in MCPdx and this is being actively explored.JAWS Ref: 750704PCT

[0231] The performance of MCPdx was also examined across individual sensitivity mocks to better understand the limit of detection of MCPdx (Table 3;

[0232] Table 4). Even at 20 read pairs target species are robustly identified in the majority of mock communities, with many fungal and protozoan species being robustly identified at 10 read pairs. Notably, at 20 read pairs 12 (92.3%) fungal, 13 (76.5%) invertebrate, 23 (100.0%) protozoan, and 110 (65.9%) prokaryote target species were identified by MCPdx in at least one of the mock communities. This demonstrates MCPdx can identify most target species with few read pairs though reliable identification benefits from additional reads.

[0233] With a target sequencing depth of 16 million read pairs, the lower bound on the required abundance (in terms of read pairs) of a species for detection with MCPdx is: - 10 read pairs = 0.0000625%; - 20 read pairs = 0.000125%; - 50 read pairs = 0.0003125%; - 100 read pairs = 0.000625%; - 500 read pairs = 0.003125%; and - 2,500 read pairs = 0.15625%. TABLE 2 NUMBER OF TARGET MGDB SPECIES IDENTIFIED BY MCPDX ACROSS ALL SENSITIVITY MOCK COMMUNITYTABLE 3 NUMBER OF TARGET MGDB SPECIES IDENTIFIED BY MCPDX ACROSS ANY SENSITIVITY MOCK COMMUNITYJAWS Ref: 750704PCTTABLE 4 NUMBER OF SENSITIVITY MOCK COMMUNITIES WHERE THE TARGET SPECIES IS IDENTIFIED BY MCPDXEvaluation of MCPdx on prokaryotic and eukaryotic cross reactivity mocks

[0234] Cross reactivity mocks consist of simulated reads from all species in a genus except from the single, intrageneric species being evaluated. These mocks are used to evaluate the specificity of MCPdx by assessing its ability to correctly assert the absence of a target species in samples under the extreme scenario of a sample containing all closely related species. This is assessed by considering the number of diagnostic loci for the target species identified by MCPdx. Ideally, no diagnostic loci would be identified in the target species though in practice a small number of misclassifications are expected due to regions of high intraspecific genomic similarity due to shared ancestry and homologous recombination. Cross reactivity mocks for MGDB species mapping to a target pathogen were built as (Error! Reference source not found. and Table 3): - Randomly select 1 to 3 genomes from each non-target species in the target genus, selecting genomes not in the corresponding MGDB reference databases when possible. - Simulate 2 × 150 bp read pairs with an inner distance of 250 ± 25 bp for each selected genome to achieve 1x depth of coverage. - Profile reads with MCP followed by MCPdx. - Repeat steps 1 to 3 in order to create three replicates for each target pathogen.

[0235] A total of 627 cross reactivity mocks were built which cover 11 fungal, 10 invertebrate, 22 protozoa, and 135 prokaryotic target species (Error! Reference source not found.). Cross reactivity mocks could not be built for 2 fungal, 7 invertebrate, 1 protozoan, and 3 prokaryotic target species as these lacked any other intrageneric species.JAWS Ref: 750704PCT TABLE 15 Taxonomic Classification of MetaPanel Target Eukaryotic Species in the MGDBJAWS Ref: 750704PCT

[0236] MCPdx performs well on eukaryotic species with 41 of 43 (95.3%) target MGDB species being correctly reported as absent across all cross reactivity replicates at the default eukaryotic reporting criterion of ≥5 diagnostic loci (

[0237]

[0238]

[0239] Table 5). The two false positives are the protozoan species Cryptosporidium muris and Cryptosporidium group A. The three C. muris replicates have 2, 3, and 5 diagnostic loci identified by MCPdx indicating this species is generally correctly reported as absent. All reads erroneously assigned to C. muris originate from C. andersoni which is a closely related species (>95.5% ANI with ~99% AF for all interspecific genome pairs). C. group A had 0, 1, and 5 diagnostic loci identified by MCPdx across the three replicates again indicating it is generally correctly reported as absent. Erroneous assignments originated from C. meleagridis and C. ‘sp chipmunk’ which have <95% ANI to strains in the C. group A species cluster. FurtherJAWS Ref: 750704PCT investigation is required to determine if this is due to legitimate genomic similarity due to shared ancestry or recombination, or due to contamination in one or more genomes.

[0240] Prokaryotic species are reported as present by MCPdx using a default criterion of ≥10 diagnostic loci. With this criterion, 91 of 135 (67.4%) MGDB species are correctly across all three cross reactivity replicates (

[0242]

[0243] Table 5). A number of the false positive predictions are due to MGDB dividing strains from a single clinical species into multiple OSCs (Error! Reference source not found.). For example, strains clinically-defined as belonging to Salmonella enterica are divided into 4 closely related OSCs: S. enterica, S. enterica_C, S. enterica_D, and S. enterica_E. The false positive predictions occurred due to cross reactivity with, for example, reads from S. enterica_E being incorrectly mapped to diagnostic loci in S. enterica. It is important to note that MCPdx and MGDB provide very high-resolution species profiles, and the above example would result in the correct clinical pathogen being reported (S. enterica). Even with these challenges, MCPdx results in a true negative prediction in 378 of 498 (75.9%) prokaryotic cross reactivity mocks.

[0244] Modifying the diagnostic loci criterion of MCPdx from 10 to 50 increases the true negative rate to 467 of 498 (93.8%), at the expensive of minor loss of sensitivity for those species requiring addition identified diagnostic loci. The 10 MGDB species resulting in false positives at 50 loci are (

[0245]

[0246]

[0247] Table 5): Escherichia coli, Klebsiella pneumoniae, Klebsiella variicola, Listeria monocytogenes, Listeria monocytogenes_B, Listeria monocytogenes_C, Salmonella enterica, Serratia marcescens, Serratia marcescens_I, and Yersinia enterocolitica. For Salmonella enterica and the three Listeria species this is the result of cross reactivity between closely related OSCs. This limits the ability of MCPdx to robustly distinguish between these OSCs, but again, still allows for accurate reporting of Salmonella enterica and Listeria monocytogene infections as clinically defined.

[0248] If cross reactivity is examined for species as clinically defined, the number of prokaryotic species correctly reported as absent across all cross reactivity replicates, at the default MCPdx criterion of 10 diagnostic loci, increases to 43 of 61 (70.5%; Table 6 and Table 7). Notably, 431 or 498 (86.5%) cross reactivity mocks result in a true negative prediction by MCPdx for clinical species (Table 7). Increasing the MCPdx criterion from 10 to 50 loci increases the number of species correctly reported as absent to 57 of 61 (93.4%), and the number of mocks with a true negative prediction to 483 of 498 (97.0%).JAWS Ref: 750704PCT

[0249] The cross reactivity mock communities represent an extreme challenge as they contain strains from all non-target, intrageneric species, a scenario never likely to occur in vivo. Despite this, all clinical pathogens and nearly all MGDB species are correctly reported as absent at a diagnostic loci criterion of 100. The number of identified diagnostic loci required to report a pathogen as present can be set independently for each pathogen. Notably, even at 100 required loci, pathogens can be identified at a relative abundance as low as 0.000625% at the target sequencing depth of 16 million read pairs, indicating excellent limit of detection for these targets. TABLE 5 MCPDX RESULTS FOR TARGET MGDB SPECIES ON CROSS REACTIVITY MOCK COMMUNITIESTABLE 6 MCPDX RESULTS FOR TARGET PATHOGENS ON CROSS REACTIVITY MOCK COMMUNITIES.TABLE 7 MCPDXRESULTSFORTARGETPROKARYOTICPATHOGENSHIGHLIGHTINGSPECIES WITH ONE OR MORE FALSE POSITIVE PREDICTIONS AT THE DEFAULT CRITERION OF 10 LOCIJAWS Ref: 750704PCTNumbers in parenthesis indicate the true negative rate as the loci threshold increases up to 100 loci. # Escherichia coli and Shigella spp. Are known to be challenging to distinguish so are effectively treated as a single species by MCPdx. In silico validation of target viral species

[0250] Target viral pathogens are identified with MCP as identification of the diagnostic loci required by MCPdx is challenged by the high mutation rate of viral genomes and a lack of sufficient evolutionary conservation of viral genes. However, the smaller size of viral genomes means cross reactivity simulations can be run across the entire database. Here, the present inventors assess the specificity of MCP on all viral species using extreme cross reactivity mock communities where DNA sequencing reads have been simulated from all vOTUs within a genus other than the target vOTU. Creating cross reactivity mock communities

[0251] MViralDB was divided into a “testing database” and a set of “holdout” genomes as previously described (see Error! Reference source not found.). Cross reactivity mock communities for a target vOTU were then generated as follows:JAWS Ref: 750704PCT (1) Determine all vOTUs in the same genus as the target vOTU according to the MViralDB “most common” vOTU taxonomic assignments (see, Taxonomic assignment of vOTUs) (2) Simulate reads from all intrageneric vOTUs except the target vOTU as follows: a. If a vOTU has genomes in the “holdout” set randomly select at most 10 genomes from this vOTU or all genomes if there are <10; otherwise, select all genomes (there will be at most 2) in the “testing” DB for this vOTU. b. For each genome, simulate 2 × 150 bp read pairs with an inner distance of 250 ± 25 bp across each contig using a step size of 50 bp. That is, the first read pair is generated at base 0 followed by base 50, 100, …, until the end of the contig. Evaluation of MCP on viral cross reactivity viral mocks

[0252] The performance of MCP was assessed across cross reactivity mocks from all 1,208 genera in MViralDB with >1 vOTU. Only 35 (2.9%) of the 1,208 genera contain mocks where the target vOTU was erroneously identified and this FP prediction only occurs for 58 (0.5%) of the 11,654 vOTUs considered (Table 19). Across all genera, the mean recall, precision, and F1 score were 98.7%, 99.9%, and 99.2%, respectively, indicating robust performance on these mock communities.

[0253] Many genera contain only vOTUs comprised exclusively of genomes in the reference “testing DB”. Arguably, these represent artificially easy test cases so should be excluded from consideration. To test this, the present inventors restricted the analysis to the 423 genera with at least 1 genome in the holdout genome. All FPs are contained within this subset of genera. With this filtering, 35 (8.3%) of the 423 genera contain mocks where the target vOTU was erroneously identified and this FP prediction occurs for 58 (0.91%) of 6,383 vOTUs. The mean recall, precision, and F1 scores decrease slightly to 96.4%, 99.8%, and 97.8%, respectively.

[0254] Only four human pathogens were found to be erroneously identified across the entire set of cross reactivity mocks (Table ). The reason for these false positive identifications were as follows: Monkeypox: Monkeypox virus vOTU757 was erroneously identified due to reads simulated from three Vaccinia virus genomes in vOTU1188 being predominately assigned to two genomes, GCA_023535905.1 and GCA_006457925.1, in vOTU757. While the “most common” species assignment for vOTU757 is Monkeypox virus, these two genomes are classified as Vaccinia virus at NCBI. Rotavirus A: Eight vOTUs with this species assignment were erroneously identified. The present inventors only explored vOTU82330, as the cause of the FP identification is likely to be the same for the other vOTUs. Reads assigned to vOTU82330 predominately came from four vOTUs all classified as Rotavirus A. The most common misassignment was reads simulated from GCA_002644435.1 being mapped to GCA_002674275.1 with 100% identityJAWS Ref: 750704PCT and alignment in most cases. These two genomes have an ANI of 98.4% and AF of 84%. This indicates the misclassification is the result of these vOTUs forming a genetic continuum. Notably, the mock does contain the species Rotavirus A so these eight FP predictions are limited to the vOTU level. Human gammaherpesvirus 4: vOTU1668 was erroneously identified due to reads being assigned to genomes in this vOTU from genomes in other vOTUs classified as Human gammaherpesvirus 4. The most common misassignment was reads simulated from GCA_900411595.1 being mapped to GCA_900003785.1 with 100% identity and alignment in most cases. These two genomes have an ANI of 99.98% and AF of 79%. This misclassification is the result of the Human gammaherpesvirus 4 vOTUs forming a genetic continuum. Human immunodeficiency virus 1: vOTU90441 was erroneously identified due to reads being assigned to genomes in this vOTU from genomes in other vOTUs classified as Human immunodeficiency virus. The most common misassignment was reads simulated from GCA_003099015.1 being mapped to GCA_003098995.1 with >98% identity and alignment. These two genomes have an ANI of 98.1% and AF of 100%. This misclassification is the result of Human immunodeficiency virus 1 vOTUs forming a genetic continuum. TABLE8GENERA WHERE ≥1 TARGET VOTU ERRONEOUSLY IDENTIFIEDJAWS Ref: 750704PCTTABLE 20 ERRONEOUSLY IDENTIFIED HUMAN PATHOGENS IN CROSS REACTIVITY MCKSvirus 1 Performance on MetaPanel target viral pathogens

[0255] Several viral human pathogens are being targeted, including adenovirus, cytomegalovirus (CMV), Epstein-Barr virus (EBV), herpes simplex virus (HSV), human herpesvirus-8 (HHV8), and human papillomavirus (HPV). MCP was able to robustly identify the vOTUs in each of the genera containing target MetaPanel viral pathogens, except for Lymphocryptovirus (Table ). Importantly, the target vOTU removed from each cross-reactivity mock was only erroneously identified in 3 (1.2%) of the 254 mocks despite the mock communities containing reads from all other vOTUs in the genera (

[0256]

[0257]

[0258] Table ). The two FP predictions made in the Mastadenovirus genus are discussed above and the single FP prediction made in the Lymphocryptovirus genus is to vOTU1668 which is classified as Human gammaherpesvirus 4 (i.e., EBV). TABLE 21 PERFORMANCE OF MCP ON CROSS REACTIVITY MOCKSJAWS Ref: 750704PCTTABLE 22 PERFORMANCE OF MCP ON CROSS REACTIVITY MOCKS OF TARGET VIRAL PATHOGENSEXAMPLE 4 IDENTIFICATION OF TARGET GENES AND MUTATIONS

[0259] Functional validation of antimicrobial resistance (AMR) and virulence genes is well supported by the scientific community. The present inventors employed NCBI’s curated AMR gene catalogue (Feldgarden, et al., 2021), the Bacterial Antimicrobial Resistance Reference Gene Database (also referred to as AMRFinderPlus database), as the source of AMRJAWS Ref: 750704PCT gene sequences. For virulence genes, the present inventors utilized gene sequences in the Genes catalogue at National Centre for Biotechnology Information (NCBI).

[0260] The presence of 47 AMR gene families were envisaged to be reported. These genes have been selected based on their capacity to confer resistance to clinically relevant antibiotic drug classes. Also reported are 16 amino acid mutations known to expand the resistance profile of four of the identified AMR genes. Allelic sequences for each of the AMR genes and gene families were retrieved from the NCBI AMRFinderPlus database.

[0261] For virulence determinants, the present inventors successfully reported on the presence of greater than 20 clinically relevant virulence factors. These virulence genes are commonly found in eight pathogenic bacterial species, including Escherichia coli (EPEC, EAEC, STEC, EHEC O157:H7, EIEC, ETEC), Shigella spp., Salmonella spp., Staphylococcus aureus, Clostridioides difficile, Listeria monocytogenes, Yersinia enterocolitica, and Vibrio cholerae. The sequences for these genes were obtained from NCBI, and collated into an in-house database referred to as the Microba Reference Gene Catalogue (MRGC), totalling 1650 gene families (totalling 7893 sequence variants). The MRGC database is organised by sorting reference genes into their ARG and VF gene families by both examining their taxonomy description from the source databases, and manually curated with peer reviewed literature. The Microba Gene Profiler (MGP)

[0262] The detection of antimicrobial resistance (AMR) genes and virulence factors is an important component. Because many of these genes are located on mobile elements, may be strain specific, and also these genes contain a large amount of genetic diversity not represented in MGDB, the present inventors could not use MCPdx for this task. To address this, the Microba Gene Profiler (MGP) was developed. MGP operates by aligning reads directly to a database of gene sequences, identifies high-quality alignments, deconvolutes multi-mapped reads, and then reports which genes are present in the sample based on the coverage of gene sequences in the database. MGP can use any database of genes, but here is applied to the MRGC as described above. Mapping and coverage parameters for the reliable detection of genes were determined similarly to the parameters for MCP, to ensure a low false positive rate. In silico validation of MGP

[0263] Similar to the in silico validation of MCP and MCPdx, the present inventors evaluated the performance of MGP using simulated sequencing data sets. To assess the specificity of MGP, reads were simulated from a randomly selected representative gene sequence from each of the 1649 gene families within MRGC. The bioinformatics tool ART-Illumina (Huang, Li, Myers, & Marth, 2011) was used to generate 150 samples of 150 base-pair (bp) paired-end reads and random errors introduced using an error profile simulating that of the HiSeqX TruSeq Illumina platform. Coverage values were randomly generated for each reference gene in each sample with a minimum value of 0.5× depth of coverage (DoC) and a maximum value of 15× DoC. Fifty samples were randomly chosen for each reference gene to contain 0× (zero) DoC asserting that the gene group should not be detected within that sample. Including allJAWS Ref: 750704PCT MRGC gene families within the in silico reads increases the difficulty of correctly calling the clinically relevant target as absent while also allowing the simulation to assess sensitivity.

[0264] The results of the in silico analysis on the target pathogen sequences are shown in Table 9,Table 10

[0265]

[0266] Table 11. Table 9 shows the average in silico performance of MGP on the clinically relevant gene families. Based on these tests MGP achieved good precision, recall and F1 scores for both ARG and VF groups (both being on average above 98.7%). TABLE 9 AVERAGE IN SILICO PERFORMANCE OF MGP ON CLINICALLY RELEVANT ARG AND VF GENE FAMILIES

[0267] Table 10 shows the performance of MGP on each ARG family. MGP performs well on all AMR gene families except blaSHV. This gene family produced false positive results due to sequence similarity with blaOKP and blaLEN. Further investigation is underway to utilize Single Nucleotide Variants (SNVs) to further resolve the blaSHV, blaOKP, and blaLEN family targets. TABLE 10 SUMMARY OF THE IN SILICO PERFORMANCE OF MGP ON EACH CLINICALLY RELEVANT ARG GROUPJAWS Ref: 750704PCT

[0268] Table 25 shows the results of in silico benchmark for each clinically relevant VF family. MGP achieved perfect precision for each gene family and perfect recall for all but the estA, stx1B, stx2B, and yst VF families which had a sensitivity of 95%, 98%, 98%, and 97% respectively. These results indicate that MGP performs well for these targets. TABLE 11 SUMMARY OF THE IN SILICO ANALYSIS OF MGP WHEN DETECTING THE PRESENCE OF VF FAMILIESJAWS Ref: 750704PCTEXAMPLE 4 MICROBA SEQUENCE VARIANT (MSV) PIPELINE

[0269] In addition to identifying target genes, the pathogen assay reports the presence of known mutations in some genes that confer extended antimicrobial resistance. Variant identification is achieved using the Microba Sequence Variant (MSV) pipeline. MSV consists of three main steps, as outlined below. (1) Identifying reads mapped to target genes

[0270] Sequencing reads for each sample are processed with MGP to identify all reads associated with each target gene. (2) Identifying variants within genes of interest:

[0271] Reads identified in Step (1) are processed with the bioinformatics tool Lorikeet (https: / / github.com / rhysnewell / lorikeet), to identify nucleotide sequence variation for specific AMR gene sequences. Lorikeet is a publicly available high-quality variant caller developed for metagenomic communities based on the highly cited and widely used GATKJAWS Ref: 750704PCT HaplotypeCaller (Van der Auwera, 2020). It adopts local reassembly of reference sequences using sample reads to ensure accurate retrieval of all variants found within a sample. (3) Filter for target variants:

[0272] Variants identified in Step (2) are screened for target variants using MSV, and the ‘som.py’ haplotype comparison toolkitwhich incorporates translation of nucleotide variants into amino acid changes to confirm the nucleotide mutation is clinically relevant. In silico evaluation of AMR variant identification

[0273] The present inventors next achieved the in silico evaluation of precision and recall of MSV to detect all required target mutations for AMR genes (

[0274] Table 12). A single alternate gene sequence known to harbour resistance inducing Single Nucleotide Variants (SNVs) was chosen for each target AMR gene. For each reference and alternate gene sequence, 100 replicate data sets were generated with a depth of coverage randomly generated between 1× and 15×. Paired reads were simulated from the gene sequences using these coverage values using ART-Illumina (Huang, Li, Myers, & Marth, 2011). The gene sequences were extended on both ends with a randomly generated 1000bp DNA sequence to allow simulation of reads with only partial overlap with the AMR genes as would occur in real metagenomic sequencing data. The simulated reads were then combined into 100 individual samples. Each sample was guaranteed to have simulated reads from both the reference and alternate gene sequences. However, depending on the random sampling that occurred during the read generation process, it was not guaranteed that the reads would contain evidence of the resistance inducing SNVs, especially at low depth of coverage values. TABLE 12 SUMMARY OF THE REFERENCE AND ALTERNATE GENES USED FOR BENCHMARKING VARIANT CALLING.Performance statistics for AMR variant in silico mock samples

[0275] MSV was run on each simulated sample providing all reference gene sequences as input. The resulting variant list for each reference gene was then compared to the known set of true variants that occur in the alternate allele. The F1 score was above 99% for all discriminatory SNVs for each allele when the coverage was 5× or greater (Table 13). PrecisionJAWS Ref: 750704PCT remained above 84% for all alleles at all DoC values, however, recall dropped to as low as 30% when DoC was less than 2×. The minimum observed number of read pairs required to accurately call all variants within an alternate sequence was 21, 10, 8, and 20 for blaCTX-M, blaGES, blaSHV, and blaTEM, respectively (Error! Reference source not found.) The lower recall for blaCTX-M and blaTEM is likely due to the higher number of SNVs in the alternate gene. There are cases where recall drops at higher Doc (e.g., blaTEM), this is due to stochastic variability in the sampling. Further investigation is underway for these differences. TABLE 13 VARIANT CALLING PERFORMANCE OF MSV ON SIMULATED SAMPLES ACROSS DISCRIMINATORY SNVS.TABLE 28 Variant Calling Performance of MSVJAWS Ref: 750704PCT

[0276] Performance of MSV was further evaluated solely considering the alleles of interest for each gene (

[0277] Table 12). All target variants were identified for each gene in at least 89% of samples (recall = 1.0, Error! Reference source not found.). The minimum observed number of read pairs required to accurately call all variants for each gene was 11, 10, 8, and 15 for blaCTX-M, blaGES, blaSHV, and blaTEM, respectively. A depth of coverage of greater than 5x will allow for accurate calling of variants from a single sample (Table ). TABLE 30 . Variant calling performance of MSV on simulated samples across target alleles.TABLE 30 SUMMARY OF VARIANT CALLING PERFORMANCE OF MSV ON SIMULATED SAMPLES FOR VARIANTS OF INTERESTJAWS Ref: 750704PCTResults are stratified by DoC of variant allele. Theoretical limit of detection of MGP and MSV

[0278] Depth of coverage (DoC) is important when detecting the presence of genes and gene variants with a high degree of confidence. The present inventors conducted a theoretical calculation of required depth for reliable calling genes and variants. Table shows the theoretical limit of detection of MGP and MSV in terms of DoC, number of 2 × 150 bp read pairs, and species relative abundance according to the Lander and Waterman model (percent coverage = 1 – e-depth). As stated in previous sections, additional read pairs will be required to achieve the required DoC due to community strain complexity, low complexity (unmappable) genomic regions, and sequencing biases and error. Plasmid copy number variation may improve depth of coverage for some targets. TABLE 31THEORETICAL LIMIT OF DETECTION FORMGPANDMSVREPRESENTED BYDOCOF A GENE WITHIN THE SPECIFIED LENGTH QUARTILE AND RELATIVE ABUNDANCE OF THEORETICAL GENOME OF 3.5 MBP IN A 16 GBP SAMPLE.JAWS Ref: 750704PCTEXAMPLE 5 Pathogen Detection in Active and Remission States of Inflammatory Bowel Disease

[0279] The inventors then sought to investigate the association between pathogen presence and IBD activity, focusing on patients in flare versus remission.

[0280] Analysis of faecal samples demonstrated a higher prevalence of pathogens in the active IBD group compared to the remission group. This trend was consistent across CD (19.3% difference), UC (18.2% difference) and combined subsets (IBD all; 19% difference; Table 32). A similar difference was not found for AMR and VFs alone, indicating that the presence of pathogens plays a role in triggering or sustaining inflammatory flares in IBD. In CD, several pathogens, virulence factors, and AMR genes were detected more frequently in the active disease group than in remission, notably Campylobacter jejunum and Campylobacter concisus (13.3% vs. 0%; Table 33). A few factors, such as amr blaDHA and amr qnrB, were present in both groups but at low frequencies. For UC detection rates were generally lower, but some factors, including Campylobacter concisus and amr blaDHA, were found exclusively in active disease (9.1%). Overall, microbial and resistance gene presence was more pronounced in active disease compared to remission in both CD and UC. TABLE 32 NUMBER AND PERCENTAGE OF PATIENTS WITH ACTIVE / REMISSION CD, UC, OR COMBINED IBD 91] U I 01] 0 1 11] 0 3 21] + +JAWS Ref: 750704PCT

[0322] Analysis of faecal samples demonstrated a higher prevalence of pathogens in the active IBD group compared to the remission group. This trend was consistent across CD (19.3% difference), UC (18.2% difference) and combined subsets (IBD all; 19% difference; Table 32). A similar difference was not found for AMR and VFs alone, indicating that the presence of pathogens plays a role in triggering or sustaining inflammatory flares in IBD. In CD, several pathogens, virulence factors, and AMR genes were detected more frequently in the active disease group than in remission, notably Campylobacter jejunum and Campylobacter concisus (13.3% vs. 0%; Figure 11). A few factors, such as amr blaDHA and amr qnrB, were present in both groups but at low frequencies. For UC detection rates were generally lower, but some factors, including Campylobacter concisus and amr blaDHA, were found exclusively in active disease (9.1%). Overall, microbial and resistance gene presence was more pronounced in active disease compared to remission in both CD and UC.

[0323] The results of this study support the hypothesis that microbial pathogens may contribute to the pathogenesis of IBD flares. Methods

[0324] An observational study was conducted at the Mater Hospital, Brisbane, Australia.77 patients diagnosed with IBD were enrolled, categorized into UC (11 with active flare, 22 in remission) and CD (15 with active flare, 29 in remission). Faecal samples were collected and sequenced at a depth of 10 Gb per sample. Data analysis was performed as described above. EXAMPLE 5 Applicability of Pathogen, Virulence Factor, and AMR Determination on Clinical Care

[0325] The inventors then sought to confirm the ability of the testing in patient clinical care to assist in the diagnosis of gastrointestinal infection.

[0326] The study population presented herein included patients who undertook testing as part of their clinical care. Eligible participants were those whose treating clinician believed there was a clinical need. To ensure the study accurately captured real-world performance, no additional inclusion or exclusion criteria were applied

[0327] The testing identified a pathogenic target in 137 / 658 (21%) of samples (see, Figure 12 A & B). The total number of unique pathogens detected was 28, including 21 species of bacteria, one virus, five species of parasite, and one microsporidium (Figure 12A). Campylobacter concisus, Enteropathogenic Shigella / Escherichia coli complex (EPEC) and toxigenic Clostridium perfringens were the most frequently detected pathogens, found in 54 / 658 (8.21%), 15 / 658 (2.28%), and 12 / 658 (1.8%) samples, respectively (Figure 1A). While most positive tests 115 / 137 (83.94%) contained a single pathogenic target, two pathogenic targets were detected in 22 / 137 (16.05%) samples (Figure 1B). In the majority of these co-infections (16 / 22, 72.72%), Campylobacter concisus was one of the detected pathogens (Figure 13).JAWS Ref: 750704PCT

[0328] Antimicrobial Resistance (AMR) genes were detected in 176 / 658 (27%) of samples (Figure 14A & B). The total number of unique AMR genes detected was 26 (Figure 14A). β-lactam resistance, quinolone resistance and vancomycin resistance gene families were the most frequently detected, found in 171 / 658 (25.99%), 54 / 658 (8.21%) and 25 / 658 (3.80%) samples, respectively. Of the positive tests, 104 / 176 (59.10%) contained a single pathogen, and the remaining 72 / 176 (40.90%) contained multiple AMR targets (Figure 14B). The maximum number of AMR targets was 6, a number detected in two samples (Figure 14B). REFERENCES Albert E, Walker J, Thiesen A, Churchill T, Madsen K (2010) cis-Urocanic Acid Attenuates Acute Dextran Sodium Sulphate-Induced Intestinal Inflammation. PLoS ONE 5(10): e13676. Almeida, A. (2019). A new genomic blueprint of the human gut microbiota. Nature, 499-504. Armitage , P., & Berry, G. (1994). Statistical Methods in Medical Research (3rd Ed). London: Blackwell. Cantero, D. L. (1997). Difference Between Analytical Sensistivity and Detection Limit. American Journal of Clinical Pathology, 619. Ciufo S, K. S. (2018). Using average nucleotide identity to improve taxonomic assignments in prokaryotic genomes at the NCBI. Int J Syst Evol Microbiol, 2386- 2392. Council, N. P. (2013). Requirements for Medical Pathology Services. NPAAC. David Berendes, J. K. (2019). Gut carriage of antimicrobial resistance genes among young children in urban Maputo, Mozambique: Associations with enteric pathogen carriage and environmental risk factors. PLoS ONE, 11:e0225464. Erdo, S. a. (2016). Alternative Confidence Interval Methods Used in the Diagnostic Accuracy Studies. Computational and Mathematical Methods in Medicine, 1-7. Federhen, S. (2015). Type material in the NCBI Taxonomy Database. Nucleic Acids Res, D1086-98. Forster, S. K. (2019). A human gut bacterial genome and culture collection for improved metagenomic analyses. Nat Biotechnol , 186-192. Gargis. (2016). Assuring the Quality of Next-Generation Sequencing in Clinical Microbiology and Public Health Laboratories. Journal of Clinical Microbiology, 2857- 65. Gargis AS, K. L.-R. (2012). Assuring the quality of next-generation sequencing in clinical laboratory practice. Nature Biotechnology, 1033-6. Hazra, A. (2017). Using the confidence interval confidently. Journal of Thoracic Disease, 4125-30. International Organization for Standardization . (2013). AS ISO15189 Medical Laboratories- Requirements for quality and competence. Geneva: International Standardization for Organization. Jain, C., Rodriguez-R, L. M., Phillippy, A. M., Konstantinidis, K. T., & Aluru, S. (2018). High throughput ANI analysis of 90K prokaryotic genomes reveals clear species boundaries. Nat Commun, 5114. Jennings L1, V. D., & Committee., C. o. (2009). Recommended principles and practices for validating clinical molecular pathology tests. Archives of Pathology and Laboratory Medicine, 743-55.JAWS Ref: 750704PCT Kim, D. (2016). Centrifuge: rapid and sensitive classification of metagenomic sequences. Genome Res, 1721-1729. Kitts, P. (2016). Assembly: a resource for assembled genomes at NCBI. Nucleic Acids Res, D73-80. Lander ES, W. M. (1988). Genomic mapping by fingerprinting random clones: a mathematical analysis. Genomics, 231-9. Mattocks CJ, M. M., & Group., E. V. (2010). A standardized framework for the validation and verification of clinical molecular genetic tests. European Journal of Juman Genetics, 1276-88. Mattocks CJ1, M. M., & Group., E. V. (2010). A standardized framework for the validation and verification of clinical molecular genetic tests. European Journal of Human Genetics, 1276-88. Meric, G. (2019). Correcting index databases improves metagenomic studies. bioRxiv https: / / doi.org / 10.1101 / 712166. Milanese, A. (2019). Microbial abundance, activity and population genomic profiling with mOTUs2. Nat Commun, 1014. National Association of Testing Authorities . (2018). General Accreditation Guidance- Validation and verification of quantitative and qualitative test methods. National Association of Testing Authorities Australia. (2019). General Accreditation Criteria: ISO 15189 Standard Application Document. National Pathology Accreditation Advisory Council. (2013). Requirements for Medical Testing of Microbial Nucleic Acids. National Pathology Accreditation Advisory Council. (2018). Requirments for the Development and Use of In-House In VitroDiagnostic Medical Devices (IVDs). Nayfach, S. S. (2019). New insights from uncultivated genomes of the global human gut microbiome. Nature, 505–510. Nayfach, S., Camargo, A. P., Schulz, F., Eloe-Fadrosh, E., Roux, S., & Kyrpides, N. C. (2021). CheckV assesses the quality and completeness of metagenome-assembled viral genomes. Nature Biotechnology, 578-585. Nayfach, S., Páez-Espino, D., Call, L., Low, S., Sberro, H., & Kyrpides, N. C. (2021). Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome. Nat Microbiol, 960-970. Ondov, B. (2016). Mash: Fast genome and metagenome distance estimation using MinHash. Genome Biol, 132. Ondov, B. D., Treangen, T. J., Melsted, P., Mallonee, A. B., Bergman, N. H., & Phillippy, A. M. (2016). Mash: fast genome and metagenome distance estimation using MinHash. Genome Biology. Ounit, T. (2015). 2015. CLARK: fast and accurate classification of metagenomic and genomic sequences using discriminative k-mers, 236. Parks D, I. M. (2015). CheckM: assessing the quality of microbial genomes. Genome Research, 1043-1055. Parks, D. (2018). A standardized bacterial taxonomy based on genome phylogen substantially revises the tree of life. Nat Biotechnol, 996-1004. Parks, D. H., Rigato, F., Vera-Wolf, P., Krause, L., Hugenholtz, P., Tyson, G. W., & Wood, D. L. (2021). Evaluation of the Microba Community Profiler for Taxonomic Profiling of Metagenomic Datasets From the Human Gut Microbiome. Front. Microbiol. Pasoli E, A. F. (2019). Extensive Unexplored Human Microbiome Diversity Revealed by Over 150,000 Genomes from Metagenomes Spanning Age, Geography, and Lifestyle. Cell, 649-662.JAWS Ref: 750704PCT Piro, V. (2019). Ganon: precise metagenomics classification against large and up-to- date sets of reference sequences. bioRxiv doi:https: / / doi.org / 10.1101 / 406017. Rehm HL1, B. S.-T., & Commitee., W. G. (2013). ACMG clinical laboratory standards for next-generation sequencing. Genetics in Medicine, 733-47. Roux, S., Adriaenssens, E. M., Dutilh, B. E., Koonin, E. V., Kropinski, A. M., & Eloe- Fadrosh, E. A. (2019). Minimum Information about an Uncultivated Virus Genome (MIUViG). Nature Biotechnology volume, 29-37. Roux, S., Páez-Espino, D., Chen, I.-M. A., Palaniappan, K., Ratner, A., & Kyrpides, N. C. (2021). IMG / VR v3: an integrated ecological and evolutionary framework for interrogating genomes of uncultivated viruses. Nucleic Acids Res, D764-D775. Sczyrba, A. (2017). Critical Assessment of Metagenome Interpretation- a benchmark of metagenomics software. Nat Methods, 1063-1071. Seppey, M. (2019). LEMMI: A Live Evaluation of Computational Methods for Metagenome Investigation. bioRxiv https: / / doi.org / 10.1101 / 507731. Steve Miller, S. N. (2019). Laboratory validation of a clinical metagenomic sequencing assay for pathogen detection in cerebrospinal fluid. Genome Research, 831-42. Varvara K. Kozyreva, C.-L. T. (2017). Validation and Implementation of a Clinical Laboratory Improvements Act- Compliant Whole-Genome Seqeuncing in the Public Health Microbiology Laboratory. Journal of Clinical Microbiology, 2502-2520. Wood, D. (2019). Improved metagenomic analysis with Kraken 2. bioRxiv doi:https: / / doi.org / 10.1101 / 762302. Ye, S. (2019). Benchmarking metagenomics tools for taxonomic classification. Cell, 779-794. Zou, Y. X. (2019). 1,520 reference genomes from cultivated human gut bacteria enable functional microbiome analyses. Nat Biotechnol , 179-185.

Claims

JAWS Ref: 750704PCT WHAT IS CLAIMED IS:

1. A method of determining an indicator to inform a clinical decision, the method comprising: receiving metagenomic sequencing information of a complex microbial community; analysing the metagenomic sequencing information to derive a pathogen profile; wherein the pathogen profile comprises: - a determination of the presence or absence of pathogens from a plurality of classes; - a determination of the presence or absence of AMR genes; and - single nucleotide polymorphisms. determining an indicator on the basis of the pathogen profile, wherein the indicator may be used to inform a clinical decision.

2. The method of claim 1, wherein the complex microbial community is derived from a sample that comprises over 100 microbial species.

3. The method of claim 1 or claim 2, wherein the complex microbial community is derived from a fecal sample.

4. The method of any one of claims 1 to 3, wherein the plurality of pathogen classes include bacteria, virus, fungus, helminth, protozoan, microsporidia, and / or invertebrate.

5. The method of any one of claims 1 to 4, wherein the determination of the presence or absence of a pathogen is performed by (i) aligning sequence reads from the metagenomic data to sequences present to a reference genome; and (ii) for genomes that are considered present by the alignment analysis of step (i), determining whether a diagnostic loci is present in the aligned sequence reads.

6. The method of any one of claims 1 to 5, wherein the clinical decision is the initiating of a treatment regimen for a particular infection or disease condition.

7. The method of any one of claims 1 to 5, wherein the clinical decision is the modification of a treatment regiment (e.g., modification of the dose of treatment, or changing the treatment).

8. The method of any one of claims 1 to 5, wherein the clinical decision is the cessation of a treatment regimen.

9. A method of determining an indicator to inform a diagnosis, the method comprising: receiving metagenomic sequencing information of a complex microbial community; analysing the metagenomic sequencing information to derive a pathogen profile; wherein the pathogen profile comprises: a determination of the presence or absence of pathogens from a plurality of classes;JAWS Ref: 750704PCT a determination of the presence or absence of AMR genes; and single nucleotide polymorphisms (SNPs). and determining an indicator on the basis of the pathogen profile, wherein the indicator may be used to inform a diagnosis.

10. The method of claim 9, wherein the SNPs are selected from virulence factors, antibiotic resistance traits, metabolic traits.

11. The method of claim 9 or claim 10, wherein the condition being diagnosed is inflammatory bowel disease (IBD).

12. The method of any one of claims 8 to 10, wherein the condition being diagnosed is a pathogenic infection.

13. The method of any one of claims 1 to 12, wherein the subject is suffering with chronic diarrhea.

14. The method of any one of claims 1 to 13, wherein the pathogenic infection comprises Aeromonas caviae, Aeromonas veronii, Campylobacter concisus, Enteropathogenic Escherichia coli (EPEC), Giardia intestinalis, H. pylori and Tropheryma whipplei.

15. The method of any one of claims 1 to 14, wherein the pathogen profile includes the presence of one or more virulence factor genes (for example eae and / or bfpA).

Citation Information

Patent Citations

  • Method and device for constructing prediction model of helicobacter pylori on multidrug resistance phenotype

    CN117153264A

  • Systems, apparatus, and methods for generating and analyzing resistome profiles

    WO2015184017A1

  • Methods and systems for detection and identification of pathogens and antibiotic resistance genes

    WO2024018485A1