Disease diagnosis and monitoring using microbial ribosomal RNA present in extracellular vesicles

The use of ribosomal RNA from extracellular vesicles in liquid biopsy samples addresses the challenges of mcfDNA degradation and contamination, enabling accurate and cost-effective microbial signature detection for disease diagnostics and monitoring.

WO2026024929A1PCT designated stage Publication Date: 2026-01-29GUSTO GLOBAL LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039021
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-07-24
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Current methods for detecting microbial signatures in liquid biopsy samples are hindered by the rapid degradation of microbial cell-free DNA (mcfDNA) due to its small fragment size and high contamination from human DNA, making deep next-generation sequencing costly and impractical for routine screening.

Method used

A method utilizing ribosomal RNA (rRNA) from extracellular vesicles (EVs) secreted by microbes, which are protected from degradation and less contaminated by DNA, is used to inform microbial signatures through Ribosome Informed Phylogeny (RIP) analysis, involving DNase treatment and nucleic acid sequence-based amplification (NASBA) to convert rRNA into cDNA for sequencing.

Benefits of technology

This approach provides a high-throughput, low-cost method for accurately identifying microbial signatures in liquid biopsy samples, reducing interference from contaminating DNA and enabling effective disease diagnostics and monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025039021_29012026_PF_FP_ABST
    Figure US2025039021_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Methods are provided for detecting disease in liquid biopsy samples by identifying microbial rRNA. The methods include isolating nucleic acids from extracellular vesicles, treating the nucleic acids with DNase to remove DNA including human cfDNA, mcfDNA, and contaminating microbial DNA, and then amplifying a region of a microbial phylogenetic rRNA. To further reduce contaminating DNA, the amplifying step can include reverse transcription to translate the phylogenetic region into cDNA and extend the cDNA ends with a T7 promoter and sequencing adaptors. The translated region is amplified with primers to the extended ends of the cDNA and sequenced. The resulting phylogenetic microbial sequences are useful for diagnosis and monitoring of diseases including cancer. The method can further include amplification of other human or bacterial DNA or RNA sequences to augment disease diagnosis. The method results in a very low contribution of contaminating microbial DNA and RNA to the overall microbial signature.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 673 / 16 PCT DISEASE DIAGNOSIS AND MONITORING USING MICROBIAL RIBOSOMAL RNA PRESENT IN EXTRACELLULAR VESICLES TECHNICAL FIELD

[0001] The presently disclosed subject matter relates to a high-throughput, high- resolution and low-cost method for determining microbial signatures in biological samples. CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No.63 / 675,916, filed on July 26, 2024, which is incorporated by reference herein in its entirety. BACKGROUND

[0003] Liquid biopsy based on circulating cell-free DNA (cfDNA) provides a new prospect for the diagnosis, monitoring, and risk assessment of a range of diseases. The cfDNA molecules circulating in peripheral blood originate from dying human cells, including tumor cells, as well as from viruses, parasites, and colonizing or invasive microbes that release their nucleic acids into the blood as they die and break down. Human-derived cfDNA has evolved into an indispensable biomarker in clinical practice for rapid and noninvasive diagnosis in prenatal screening, organ transplantation, and oncology, where the focus has been on circulating tumor DNA (ctDNA).

[0004] An increasing number of studies have demonstrated that microbial cell free DNA (mcfDNA) detection offers the potential to identify a wide variety of infections, such as invasive fungal infection, tuberculosis, sepsis, cystic fibrosis and chorioamnionitis. In addition to their role in infectious diseases, several studies have shown the presence of distinct cultivable bacteria and fungi in cancer tissues, including lung, prostate, pancreas, colon cancers, breast cancer and brain cancer. The microbial compositions of these tumor microbiomes are cancer specific and often linked to treatment outcomes, thus allowing for the specific detection of the types of cancer and their prognostics based on distinct microbial derived DNA signatures.

[0005] Liquid biopsy samples, especially peripheral blood, represent unique challenges for the analysis of microbial signatures. Due to their rapid degradation by DNases, the majority of mcfDNA fragments in blood were found to be approximately 40 - 100 bp in size. As a result of their small sizes, conventional amplicon-based sequencing approaches that target DNA fragments of several hundred nucleotides (>400) are not suitable for determining the composition of colonizing or invasive microorganisms using mcfDNA from liquid biopsyAttorney Docket No.: 673 / 16 PCT samples. In addition, human cfDNA accounts for the vast majority of cfDNA in plasma (>90% or even >99%), while mcfDNA accounts for only a small fraction with 0.08%-4.85% from bacteria, 0.00%-0.01% from fungi, and 0.00%-0.16% from viruses / phages. Further technical issues that complicate the use of mcfDNA come from the presence of contaminating DNA from bacterial origin in materials used for their processing. As a result, the analysis of mcfDNA requires deep next generation sequencing (NGS) of plasma cfDNA to overcome the limitations of small mcfDNA fragment size, low concentration, and the presence of contaminating bacterial DNA fragments. Unfortunately, despite major breakthroughs in metagenome sequencing technologies to reduce its costs, it is currently still too expensive to be used for routine screening of human associated microbial communities in large population screenings. Another disadvantage of deep microbial metagenome sequencing is the need for relatively large amounts of high-quality microbial DNA. This has hindered its application to study the microbial communities associated with liquid and solid biopsy samples, where only a small fraction of the total DNA is of microbial origin.

[0006] Accordingly, there remains an unmet need for high resolution, high-throughput, and low-cost methods for detection of microbial signatures in liquid biopsy samples. The present disclosure provides such methods for detecting microbial signatures for use in disease diagnostics and monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1A is a graph showing an overview of the relative abundance of bacterial contaminants detected in negative control water sample 1. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNA extracted from two negative control water samples. After sequencing, bacteria were identified on the genus level.

[0008] Figure 1B is a graph showing an overview of the relative abundance of bacterial contaminants detected in negative control water sample 2. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNA extracted from two negative control water samples. After sequencing, bacteria were identified on the genus level.

[0009] Figure 2A is a graph showing an overview of the relative abundance of bacterial contaminants detected in negative control water sample 1. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNAAttorney Docket No.: 673 / 16 PCT extracted from two negative control water samples. After sequencing, bacteria were identified on the family level.

[0010] Figure 2B is a graph showing an overview of the relative abundance of bacterial contaminants detected in negative control water sample 2. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNA extracted from two negative control water samples. After sequencing, bacteria were identified on the family level.

[0011] Figure 3 is a schematic representation of a protocol for Ribosome Informed Phylogeny (RIP) based analysis of microbial signatures in liquid biopsy samples. Extracellular vesicles (EVs) are isolated from liquid biopsy samples, including but not limited to blood plasma. Once isolated, the EVs are heat treated and subsequently total RNA is isolated. The nucleotide sample is treated with DNase to remove any DNA from the sample. Subsequently, nucleic acid sequence-based amplification (NASBA) targeting the V3-V4 region of the 16S rRNA ribosomal subunit is performed using RNA as input. Once completed, the enzymes for the NASBA step are inactivated, and the amplified RNA representing the V3-V4 region of the 16S rRNA ribosomal subunit is converted into cDNA. The cDNA is converted into double stranded DNA and adaptors for RNAseq library sequencing are added. The library is finally amplified using indexing primers and sequenced.

[0012] Figure 4 is a schematic representation of the nucleic acid sequence-based amplification (NASBA) protocol for amplification of the V3-V4 region of the 16S rRNA ribosomal subunit (adapted from Thermo Fisher Scientific).

[0013] Figure 5 shows the microbial signatures obtained after different iterations of the Ribosome Informed Phylogeny (RIP) protocol shown in Figure 3. The abundance of sequencing reads for the 30 most abundant genera observed after analysis of the sequences obtained for samples 1 to 17 are shown. The differences in protocol iterations for the samples is described in Table 3.

[0014] Figure 6 is a schematic representation of a modified protocol for Ribosome Informed Phylogeny-based analysis of microbial signatures in liquid biopsy samples. Extracellular vesicles (EVs) are isolated from liquid biopsy samples, including but not limited to blood plasma. Once isolated, the EVs are heat treated and subsequently total RNA is isolated. The nucleotide sample is treated with DNase to remove any DNA from the sample. Subsequently, a Reverse Transcriptase (RT) step is performed to translate the V3-V4 region of the 16S rRNA ribosomal subunit into cDNA. During this step the 5’ end of the cDNA is extended with sequences for the T7 promoter, the Rd2 sequencing adaptor and a UniqueAttorney Docket No.: 673 / 16 PCT Molecular Identifier (UMI) sequence, and the 3’ end of the cDNA is extended with a sequence for the Rd1 sequencing adaptor. This region is subsequently amplified by nucleic acid sequence- based amplification (NASBA) using primers that target the extended 5’ and 3’ ends of the cDNA. Once completed, the enzymes for the NASBA step are inactivated, and in a second RT step the amplified RNA representing the V3-V4 region of the 16S rRNA ribosomal subunit is converted into cDNA. The cDNA is converted into double stranded DNA, and the library is amplified using indexing primers and sequenced. This protocol is referred to as the RIP2 protocol.

[0015] Figure 7 is a schematic representation of a reverse transcriptase step, part of the RIP2 protocol, for synthesis of cDNA of the V3-V4 region of the 16S rRNA ribosomal subunit. During this step, the T7 promoter sequence, the Rd2 sequencing adaptor sequence and a unique molecular identifier are introduced at the 5’ end of the V3-V4 region. In addition, the Rd1 sequencing adaptor sequence is introduced at the 3’ end of the V3-V4 region.

[0016] Figure 8 is a schematic representation of the nucleic acid sequence-based amplification (NASBA) step, part of the RIP2 protocol, for amplification of the V3-V4 region of the 16S rRNA ribosomal subunit. For this step, primers targeting the T7 promoter plus the Rd2 sequencing adaptor and the Rd1 sequencing adaptor are used.

[0017] Figure 9 is a schematic representation of the workflow that integrates the two alternative versions of the Ribosome Informed Phylogeny (RIP) protocol and the Single Point Amplicon (SPA) fragment sequencing protocol for obtaining microbial signatures from liquid biopsy samples. DETAILED DESCRIPTION

[0018] The presently disclosed subject matter now will be described more fully hereinafter. The presently disclosed subject matter may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Indeed, many modifications and other embodiments of the presently disclosed subject matter set forth herein will come to mind to one skilled in the art to which the presently disclosed subject matter pertains having the benefit of the teachings presented in the descriptions provided herein. Therefore, it is to be understood that the presently disclosed subject matter is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims.Attorney Docket No.: 673 / 16 PCT

[0019] Following long-standing patent law convention, the terms “a,” “an,” and “the” refer to “one or more” when used in this application, including the claims. Thus, for example, reference to “a sample” includes a plurality of samples, unless the context clearly is to the contrary, and so forth.

[0020] Throughout this specification and the claims, the terms “comprise,” “comprises,” and “comprising” are used in a non-exclusive sense, except where the context requires otherwise. Likewise, the terms “having” and “including” and their grammatical variants are intended to be non-limiting, such that recitation of items in a list is not to the exclusion of other like items that can be substituted or added to the listed items.

[0021] For the purposes of this specification and claims, the term “about” when used in connection with one or more numbers or numerical ranges, should be understood to refer to all such numbers, including all numbers in a range and modifies that range by extending the boundaries above and below the numerical values set forth. The recitation of numerical ranges by endpoints includes all numbers, e.g., whole integers, including fractions thereof, subsumed within that range (for example, the recitation of 1 to 5 includes 1, 2, 3, 4, and 5, as well as fractions thereof, e.g., 1.5, 2.25, 3.75, 4.1, and the like) and any range within that range. In addition, as used herein, the term "about", when referring to a value can encompass variations of, in some embodiments + / -20%, in some embodiments + / -10%, in some embodiments + / -5%, in some embodiments + / -1%, in some embodiments + / -0.5%, and in some embodiments + / - 0.1%, from the specified amount, as such variations are appropriate in the disclosed compositions and methods. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0022] Throughout this specification and the claims, the term “subject” as used herein generally refers to humans and animals and can be used interchangeably with the term “human” and the term “patient”.

[0023] The term “cell free DNA (cfDNA)” as used herein generally refers to cfDNA that is extracellular DNA while in a living body of a subject.

[0024] The term “circulating tumor DNA (ctDNA)” as used herein generally refers to extracellular DNA released by tumor cells in a living body of a subject.

[0025] The term “microbial cell free DNA (mcfDNA)” as used herein generally refers to extracellular DNA released by microbial cells in a living body of a subject.Attorney Docket No.: 673 / 16 PCT

[0026] The term “extracellular vesicles (EVs)” as used herein generally refers to extracellular vesicles released by Eukaryotic and microbial cells, including archaeal and bacterial cells, in a living body of a subject.

[0027] The term “bacterial membrane vesicles (BMVs)” as used herein generally refers to extracellular vesicles released by microbial cells in a living body of a subject. The microbial cells are not limited to bacterial cells and the term BMVs is intended to include bacterial and non-bacterial microbes including fungi, parasites, and viruses.

[0028] The term “microbial phylogenetic rRNA” as used herein means any conserved rRNA from any organism, including but not limited to bacteria, fungi, and parasites, that is suitable for phylogenetic identification.

[0029] The term “RIP protocol” as used herein generally refers to either version of the Ribosome Informed Phylogeny protocol as described in Example 2 and Example 3.

[0030] The term “biopsy” as used herein, including in the phrase “liquid biopsy sample,” is intended to be construed in its broadest sense as an examination of tissue, including a fluid tissue such as, but not limited to, blood, spinal fluid, cerebral fluid, urine, saliva, sputum, and lymph fluid, removed from a living body to discover the presence, cause, or extent of a disease or disorder and can include predicting the likelihood of the disease or disorder and / or the stage of the disease or disorder.

[0031] There are a wide range of diseases where microbial community analysis provides important information regarding the disease, its progression, or its treatment options. This includes conditions such as IBD, metabolic diseases, diseases of the central nervous system, including Alzheimer’s disease and Parkinson’s disease, and cancer. For instance, the associations between microbial community dynamics, tumor immunity and immunotherapy efficacy have been highlighted in published results by characterizing microbiome genomes at the species and functional level in over 4,000 metastatic tumor biopsies. These and other published findings show that (1) the tumor microbiome can be used for diagnostics, prognostics and disease monitoring, and (2) the gut microbiome can actively modify responses to chemotherapeutic agents and immunotherapies by influencing host immunosurveillance. Thus, to optimize patient treatment regimens and long-term patient outcomes it is important to monitor, and if possible, modulate both the tumor and gut microbiomes in cancer patients for synergistic beneficial interactions.

[0032] Deep microbial metagenome sequencing is the most informative approach when it comes to microbial community analysis, as it provides detailed information regarding community composition as well as the key functions encoded by the community members.Attorney Docket No.: 673 / 16 PCT Unfortunately, despite major breakthroughs in metagenome sequencing technologies to reduce its costs, it is currently still too expensive to be used for routine screening of human associated microbial communities in large population screenings. Another disadvantage of deep microbial metagenome sequencing is the need for relatively large amounts of high-quality microbial DNA. This has hindered its application to study the microbial communities associated with liquid and solid biopsy samples, where only a small fraction of the total DNA is of microbial origin.

[0033] The amplification and subsequent sequencing of phylogenetic marker genes provides an alternative, cheaper high throughput method for microbial community analysis, but only for samples where there is sufficient DNA having average fragment length of about 1,000 bp or more. The amplification-based sequencing approaches have been successfully applied to identify differences in microbial communities between healthy individuals and patients suffering from a wide range of diseases, and in studies that link disease phenotypes to the presence of specific microbes. Advantages of the amplification and subsequent sequencing method include that it requires significantly less DNA than metagenome sequencing, and because specific DNA primers are used to amplify phylogenetic target genes, there is little contamination of the sequencing libraries with host DNA.

[0034] However, the method of amplification and sequencing of phylogenetic marker genes is ineffective for determining mcfDNA in liquid biopsy samples, due to the short fragment length of this DNA. More than 70% of plasma cfDNA is smaller than 300 bp, with an average size of 170 bp, while the average size of mcfDNA fragments is approximately 40-100 bp. As a result of this size limitation, conventional amplicon-based sequencing approaches including 16S rRNA gene and rpoB gene amplicon sequencing that target DNA fragments of several hundred nucleotides, are not suitable for determining the composition of colonizing or invasive microorganisms using mcfDNA from peripheral blood and other liquid biopsy samples. The small size of mcfDNA makes it nearly impossible to use mcfDNA in amplicon-based sequencing protocols, such as 16S rRNA gene sequencing. It should also be noted that contaminating DNA fragments often have larger fragment sizes than mcfDNA, resulting in their enrichment during amplification.

[0035] Bacteria, Archaea, and cells from Eukaryotic organisms, including fungi, secrete extracellular vesicles (EVs) in their environment that can interact with other organisms, including humans. Released in mass, these EVs overwhelm the host immune system and injure host tissues; however, there is also evidence that EVs may take part in processes that promote host health.Attorney Docket No.: 673 / 16 PCT

[0036] EVs are known to be responsible for shuttling different intracellular components by protecting their cargo from degradation. Regardless of their source, the cargo of bacterial extracellular vesicles (BMVs) displays quite a heterogeneous arrangement, including inner- membrane, periplasmic and cytoplasmic components, genetic material (DNA / RNA), toxins, and factors involved in antibiotic resistance and virulence. It has been discovered that BMVs are essential for different functions such as cell-to-cell communication, the formation of biofilms, bacterial infections, and the transfer of proteins and genetic material. Publications indicate that these membrane vesicles can be found in stool and several human biofluids including urine, saliva and blood and there is evidence that the BMV’s are able to cross the blood – brain barrier into the spinal fluid, where they could play a role in the onset of diseases of the central nervous system, including Parkinson’s disease and Alzheimer’s disease. As such, it has been proposed that DNA present in BMVs that colonize the human body including the gut and the tumor microenvironment can be used as an alternative to mcfDNA to determine which microorganisms contribute to specific disease phenotypes.

[0037] However, the present inventors have discovered that the level of contaminating bacterial DNA in reagents commonly used to quantify bacterial DNA in liquid biopsy samples is often similar to or higher than the amount of the bacterial DNA in the liquid biopsy samples themselves. The high levels of contaminating bacterial DNA interfere with accurate measurement of microbial signatures in liquid biopsy samples for disease diagnostics and monitoring.

[0038] To address this problem, the present inventors provide a method that uses RNA from ribosomes (rRNA) present in EVs secreted by microbes including bacteria to inform the microbial signatures in biopsy samples for disease diagnostic and monitoring purposes. In one embodiment, the 16S ribosomal RNA subunit present in EVs secreted by bacteria is used in the method. This method of targeting microbial rRNA has the following advantages over targeting DNA: 1) due to the ubiquitous presence of RNases, the levels of contaminating RNA are significantly lower than contaminating DNA; 2) a typical bacterial cell may have as many as 15,000 ribosomes compared to an average of 5 copies of the rRNA genes encoding for the various subunits of which the ribosomes are comprised, thus significantly increasing the chances for detection; 3) due to their secondary and tertiary structure, ribosomes and their RNA subunits are relatively stable to degradation by RNases; and 4) while inside EVs, RNA will be protected from RNases present in blood plasma and other bodily fluids, thereby enabling a more accurate quantification. Furthermore, by including a DNase treatment during the RNA isolation step, contaminating DNA will be removed, thus significantly reducing the amount of contaminatingAttorney Docket No.: 673 / 16 PCT DNA and facilitating the accurate interpretation of microbial signatures in biopsy samples for disease diagnostics and monitoring.

[0039] In one published study, analysis of the V3 – V416S rRNA gene region amplified from DNA present in BMVs from blood of biliary tract cancer patients was used as a diagnostic biomarker (Lee H., et al. (2020)). In this study, among the bacteria found to differentiate between healthy controls and biliary track cancer patients, were several bacteria belonging to the families Pseudomonaceae, Corynebacteriaceae and Comomonadaceae. However, DNA from these bacterial families has been frequently observed as a contaminant in DNA isolated from biological samples, including blood samples. Notably, the Lee et al (2020) paper does not report use of experimental controls, including negative water controls that went through the same experimental procedure as blood plasma samples for BMV isolation, DNA extraction, and amplification and sequencing of the 16S rRNA gene region.

[0040] The present inventors provide experimental results in EXAMPLE 1 confirming that a strong signal from contaminating DNA complicates the analysis and interpretation of relatively weak DNA-informed microbial signatures in liquid biopsy samples. More specifically, EXAMPLE 1 shows that bacterial DNA from the Pseudomonaceae, Corynebacteriaceae and Comomonadaceae families was detected in negative control water samples analyzed by amplifying the V3-V4 region of 16S rRNA gene. Figures 1A and B are graphs showing an overview of the relative abundance of bacterial contaminants detected in negative control water samples. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNA extracted from two negative control water samples. After sequencing, bacteria were identified on the genus level. Figures 2A and B are graphs showing an overview of the relative abundance of bacterial contaminants detected in negative control water samples. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNA extracted from two negative control water samples. After sequencing, bacteria were identified on the family level. These results indicate that the Pseudomonaceae, Corynebacteriaceae and Comomonadaceae DNA detected in liquid biopsy samples by amplifying the V3-V4 region of 16S rRNA gene likely represents DNA contamination rather than an accurate indication of a microbial signature for the biopsy sample.

[0041] Ribosomal RNAs (rRNAs) are the most abundant RNA species in the cell. Compared to an average of 5 copies of ribosomal RNA genes, around 15,000 ribosomes can be present in a single bacterial cell. As such the ribosomal RNA should theoretically provide a 3000-fold stronger signal than the DNA coding for the ribosomal RNA subunits. Furthermore,Attorney Docket No.: 673 / 16 PCT published results indicate that mRNA of many other genes can be found in BMVs, and these RNAs can also be interrogated using the method of the present invention.

[0042] To overcome the issue of contaminating DNA, a factor that has major impact on the interpretation of the microbial signatures in liquid biopsy samples, the present inventors provide a method that uses RNA instead of DNA present in EVs secreted by microbes for determining the microbial signatures present in biopsy samples. The advantage of using RNA instead of DNA is the omnipresence of RNases, whose enzymatic activity very significantly reduces levels of contaminating RNA. On the other hand, RNA confined to the EVs is protected from degradation by RNases present in blood plasma.

[0043] To overcome the limitations associated with detecting mcfDNA in blood and other biopsy samples due to the intrinsic properties of mcfDNA (small fragment sizes, low copy number) and the presence of contaminating DNA, the present inventors provide RNA-based methods for analysis of microbial signatures in liquid biopsy samples. The RNA-based methods are exemplified in EXAMPLES 2 and 3. More specifically, in the methods provided herein, RNA is isolated from EVs that are present in stool or liquid biopsy samples including blood, urine, saliva, sputum, and spinal fluid, and subsequent amplification and sequencing of microbial ribosomal RNA (rRNA) is used to phylogenetically identify the microbes that released the EVs. In one example the microbial rRNA is derived from the V3-V4 region of the 16S ribosomal subunit. As exemplified in EXAMPLE 2 herein, by focusing on RNA, the method provided herein results in a very low contribution of contaminating microbial RNA to the overall microbial signature.

[0044] As shown in EXAMPLE 3 herein, by modifying or eliminating amplification steps that use primers that recognize conserved sequences flanking the V3-V4 region of 16S rRNA gene, the contribution of contaminating DNA to the overall microbial signature can be further reduced.

[0045] As described in EXAMPLE 4, the phylogenic identity of the microbe (e.g., bacterium) can be linked to probabilistic models describing the functioning of the microbe, including therapeutic functions such as the synthesis of beneficial secondary metabolites and functions involved in drug resistance or virulence, including functions that elucidate a response by the host’s immune system. This response is of particular interest when considering the effects of bacteria on disease responses in distantly located parts of the body, as secondary metabolites, virulence factors and other functions that elucidate a response by the host’s immune system will also be present in their BMVs. This is in addition to the immunogenic properties of the BMVs’ elicited by their associated lipopolysaccharides (LPS). As such, the present invention enablesAttorney Docket No.: 673 / 16 PCT use of microbial signatures informed by the ribosomes present in EVs for making microbe- specific predictions on disease responses elicited by the microbial EVs’ cargo.

[0046] The utility of the methods of the invention is exemplified in EXAMPLES 1 to 4 of the present disclosure. In EXAMPLE 1 of the present disclosure, results of negative water controls used in PCR amplification of the V3 -V4 region of the bacterial 16S rRNA gene are presented that show the sensitivity of DNA informed methods to the presence of bacterial DNA contamination, justifying the approach described herein to instead target ribosomal and other RNA molecules for the analysis of microbial signatures in liquid biopsy samples. For the negative water controls, DNA extraction was performed on nucleotide-free water samples using the same protocol for extracting DNA from blood plasma. The results show that when DNA fragments of the V3 -V4 region of the 16S rRNA gene are PCR amplified and sequenced, several bacterial species are identified in the negative water control samples that are also commonly found as the dominant species associated with biological samples. Thus, the presence of this strong contaminating DNA signal significantly complicates the analysis and interpretation of relatively weak microbial signatures from biopsy samples.

[0047] In EXAMPLE 2 of the present disclosure, a method is provided (referred to as Ribosome Informed Phylogeny (RIP)) that targets the ribosomes present in microbial EVs for the analysis of microbial signature in liquid biopsy samples, such as blood plasma. The RIP method exemplified in EXAMPLE 2 is designed to target the 16S rRNA ribosomal subunit, but other ribosomal RNAs from bacterial and archaeal origin (the 23S rRNA and the 5S rRNA) or eukaryotic origin (the 5S rRNA, 18S rRNA, 28S rRNA and 5.8S rRNA) can also be targeted, as well as messenger RNA (mRNA) and regulatory RNA (e.g. siRNA) molecules. A schematic description of an exemplary workflow for the RIP method is presented in Figure 3. Based on the results presented in EXAMPLE 2 and compared to methods that target the DNA fragments that code for the 16S rRNA subunit gene, the RIP method is significantly less sensitive to interference by contaminating nucleic acids, making the method superior for the identification of the bacteria that contribute to the microbial signature in liquid biopsy samples, such as blood plasma.

[0048] Figure 3 shows a schematic representation of a protocol for Ribosome Informed Phylogeny (RIP)-based analysis of microbial signatures in liquid biopsy samples. Extracellular vesicles (EVs) are isolated from liquid biopsy samples, including but not limited to blood plasma. Once isolated, the EVs are heat treated and subsequently total RNA is isolated. The nucleotide sample is treated with DNase to remove any DNA from the sample. In this example, nucleic acid sequence-based amplification (NASBA) targeting the V3-V4 region of the 16SAttorney Docket No.: 673 / 16 PCT rRNA ribosomal subunit is performed using RNA as input. Once completed, the enzymes for the NASBA step are inactivated, and the amplified RNA representing the V3-V4 region of the 16S rRNA ribosomal subunit is converted into cDNA. The cDNA is converted into double stranded DNA and adaptors for RNAseq library sequencing are added. The library is finally amplified using indexing primers and sequenced.

[0049] Figure 4 is a schematic representation of a nucleic acid sequence-based amplification (NASBA) protocol for the amplification of the V3-V4 region of the 16S rRNA ribosomal subunit (adapted from Thermo Fisher Scientific). In summary, NASBA can include the following steps: 1. RNA of the 16S rRNA ribosomal subunit is the target for amplification of the V3-V4 region. Total RNA, which includes ribosomal RNA, is added to the reaction mixture, after which a first V4 region targeting primer with the T7 promoter region on its 5' end (T7-16SV4-R785 primer, Table 1) is attached to its complementary site at the 3' end of the template. Alternatively, primer T7plus1-16SV4-R785 with an extended T7 promoter region (Table 1, based on Conrad et al, 2020) can be used. 2. Reverse transcriptase synthesizes the opposite complementary DNA strand extending the 3' end of the primer, moving upstream along the RNA template. 3. RNAse H destroys the RNA template from the DNA-RNA compound (RNAse H only destroys RNA in RNA-DNA hybrids, but not single-stranded RNA). 4. A second primer (16SV3-F350 primer, Table 1) attaches to the 5' end of the (antisense) DNA strand. 5. Reverse transcriptase again synthesizes another DNA strand from the attached primer resulting in double stranded DNA. 6. T7 RNA polymerase binds to the promoter region on the double stranded DNA. Since the T7 RNA polymerase can only transcribe in the 3' to 5' direction, the sense DNA is transcribed and an anti-sense (-)RNA is produced. This is repeated, and the polymerase continuously produces complementary RNA strands of this template which results in amplification. 7. Now a cyclic amplification phase begins, similar to the previous steps. Here, however, the 16SV3-F350 primer first binds to the (-)RNA 8. The reverse transcriptase now produces a (+)cDNA / (-)RNA duplex.Attorney Docket No.: 673 / 16 PCT 9. RNAse H again degrades the RNA and the T7-16SV4-R785 primer binds to the now single stranded +(cDNA) 10. The reverse transcriptase now produces the complementary (-)DNA, creating a dsDNA duplex 11. Exactly like step 6, the T7 polymerase binds to the promoter region to produce (-) RNA, and the cycle is complete.

[0050] EXAMPLE 2 describes performance of seventeen different iterations of the RIP protocol represented by sequencing libraries 1 to 17. Table 4 and Figure 5 summarize the results obtained after sequencing and analysis of the microbial signatures obtained for libraries 1 to 17. To simplify the interpretation of the results, only the 30 most abundant genera found in libraries 1 to 17 are shown. Table 4 shows microbial signatures obtained after different iterations of the Ribosome Informed Phylogeny (RIP) protocol, and Figure 5 is a graph depicting the data from Table 4. Abundance of sequencing reads for the 30 most abundant genera observed after analysis of the sequences obtained for RIP libraries 1 to 17 are reported. The differences in protocol iterations for the samples are described in Table 3.

[0051] In EXAMPLE 3 of the present disclosure, the inventors exemplify a variation of the RIP method which is outlined in Figure 6. Compared to the method described in EXAMPLE 2 and Figure 3, this method has the following modifications.1) Before the NASBA amplification step, a reverse transcriptase step (shown in Figure 7) is used to translate the V3-V4 region of the 16S ribosomal RNA into cDNA. During this process, a T7 promoter sequence, a sequencing adaptor sequence (e.g., Rd2 sequencing adaptor sequence), and a unique molecular identifier (UMI) sequence are introduced at the 5’ end of the cDNA fragment covering the V3- V4 region. In addition, during the synthesis of the complementary cDNA strand, a sequencing adaptor sequence (e.g., Rd1 sequencing adaptor sequence) is introduced at the 3’ end of the V3- V4 region.2) In the subsequent NASBA amplification step (shown in Figure 8), primers targeting the T7 promoter plus the Rd2 sequencing adaptor and the Rd1 sequencing adaptor are used to amplify the V3-V4 region of the 16S rRNA.3) After a second RT step to convert the amplified RNA into cDNA, the library is immediately amplified (I-PCR) using indexing primers without performing the amplification PCR (A-PCR) step. Thus, the key differences between this modified protocol and the RIP protocol described in Example 2 are that: 1) during the NASBA amplification step no primers are being used that target the conserved sequences naturally present upstream and downstream of the V3-V4 region of the 16S ribosomal RNA gene, and 2)Attorney Docket No.: 673 / 16 PCT the A-PCR step that also uses primers targeting the V3-V4 region is omitted from the protocol, further reducing the possibility for this region to be amplified from contaminating DNA.

[0052] The term Unique Molecular Identifier (UMI) is intended to mean a stretch of nucleotides that are added as an identifier barcode to the region of a microbial rRNA before amplification of the region. The objective is to aid in distinguishing between identical copies of the region present in a sample. The uniqueness of a copy of an amplified region can be inferred from a combination of its UMI and its sequence. As a result of this combination of features, the number of UMIs can be smaller than the number of microbial rRNAs that are targeted in the sample. Given that the pool of UMI’s does not have to be in excess of the number of microbial rRNAs that are targeted in the sample, each of the UMI’s present on the region of a microbial rRNA are not required to be unique. Therefore, the term “UMI” is not intended to be understood according to the dictionary definition of “unique”. In one embodiment, the UMIs can range in length from 4 to 6 nucleotides, encoding for 256 or 4096 unique sequences, respectively. In another embodiment, a UMI can include any nucleotide (A, G, C or T), wherein A = adenine, G = guanidine, C = cytosine and T = thymine.

[0053] In EXAMPLE 4, the use of the RIP method is described as a stand-alone method or in combination with methods for targeting human and bacterial DNA or RNA and their use for diagnosis, monitoring, and prognostics of disease including cancer. This includes a workflow, presented in Figure 9, for using the RIP method in combination with the Single Point Amplicon (SPA) fragment sequencing method (PCT / US2023 / 011406). Combining these two methods in a single workflow allows for determination of which microorganisms contribute to the mcfDNA and which contribute to the EVs present in a liquid biopsy sample. EXAMPLE 4 also describes how the phylogenic identity of the bacterial microbe can be linked to probabilistic models describing its functioning, including therapeutic functions such as the synthesis of beneficial secondary metabolites and functions involved in drug resistance or virulence, including functions that elucidate a response by the host’s immune system.Attorney Docket No.: 673 / 16 PCT Primer Sequence Application V3-V416S rRNA NASBA amplification – EXAMPLE 2Table 1. Overview of primers. The sequences of the extended T7 promoter regions in primers T7plus1-16SV4-R785, T7plus2-16SV4-R785, T7plus2-Rd2-UMI-16SV4-R781 and T7plus2-5’-Attorney Docket No.: 673 / 16 PCT Rd2 are based on Conrad et al (2020). The * indicates a phosphorothioated bond between two nucleotides, a modification which renders the nucleotides linkage resistant to nuclease degradation.

[0054] A method provided herein, which is schematically presented in Figure 3 and exemplified in EXAMPLE 2, can involve one or more of the following steps: 1. Purification of extracellular vesicles (EVs) from blood plasma. Currently, the protocol does not differentiate between microbial or bacterial membrane vesicles (BMVs) and EVs released by human cells. If required, human EVs can be lysed by treatment with 1% saponin, after which their DNA and RNA content can be degraded by nuclease activity. 2. The purified plasma containing the EVs is treated with proteinase K to degrade any proteins including RNase present in the plasma and to weaken the integrity of the EVs. Subsequently, EVs are lysed via a denaturation step (heat treatment). 3. Total RNA is isolated and DNA, including contaminating DNA, is degraded via DNase treatment. This step is essential as the presence of DNA, including DNA of human origin, will inhibit the nucleic acid sequence-based amplification (NASBA) step. Furthermore, high amounts of contaminating DNA can also result in aspecific synthesis of non-target RNA during the NASBA step. 4. Total RNA is used in the NASBA step for isothermal amplification of single stranded RNA. In our protocol the V3-V4 region of the 16S rRNA gene is targeted for amplification. However, this method can be used for targeted and non-targeted amplification of any RNA. The principle of NASBA to amplify the V3-V4 region of the 16S rRNA gene and the specific primers used in the NASBA step, the T7-16SV4-R785 primer and the 16SV3-F350 primer (see Table 1), are shown in Figure 4. Alternatively, the T7plus1-16SV4-R785 primer that contains an extended T7 promoter region sequence is used. The outcome of the NASBA reaction is (-)RNA representing the V3-V416S ribosomal RNA region 5. After heat inactivation of the enzymes performing the NASBA step, the amplified (-)RNA representing the V3-V416S ribosomal RNA region is translated into cDNA using the Rd1- 16SV3-F350 primer and amplified by PCR using the Rd1-16SV3-F350 and Rd2-16SV4- R781 primers that will introduce the appropriate adapters for sequencing library construction. It should be noted that compared to the T7-16SV4-R785 or T7plus1-16SV4- R785 primers used during NASBA, the Rd2-16SV4-R781 primer used for the PCR stepAttorney Docket No.: 673 / 16 PCT contains a four-nucleotide extension at its 3’ end, allowing for enrichment PCR to further reduce the risk of converting and amplifying aspecific RNA fragments into cDNA. 6. Sequencing libraries are constructed and amplified using indexing primers, after which the libraries are sequenced followed by data analysis.

[0055] A method provided herein, which is schematically presented in Figure 6 and exemplified in EXAMPLE 3, can involve one or more of the following steps: 1. Purification of extracellular vesicles (EVs) from blood plasma. The protocol does not differentiate between microbial or bacterial membrane vesicles (BMVs) and EVs released by human cells. If required, human EVs can be lysed by treatment with 1% saponin, after which their DNA and RNA content can be degraded by nuclease activity. 2. The purified plasma containing the EVs is treated with proteinase K to degrade any proteins including RNase present in the plasma and to weaken the integrity of the EVs. Subsequently, EVs are lysed via a denaturation step (heat treatment). 3. Total RNA is isolated and DNA, including contaminating DNA, is degraded via DNase treatment. This step is essential as the presence of DNA, including DNA of human origin, will inhibit the Reverse Transcriptase (RT) step and the nucleic acid sequence-based amplification (NASBA) step. Furthermore, high amounts of contaminating DNA can also result in aspecific synthesis of non-target RNA during the NASBA step. 4. A reverse transcriptase step (RT1, described in Figure 7) is used to translate the V3-V4 region of the 16S ribosomal RNA into cDNA. During this process, a T7 promoter sequence, the Rd2 sequencing adaptor sequence, and a unique molecular identifier (UMI) are introduced at the 5’ end of the cDNA fragment covering the V3-V4 region. In addition, during the synthesis of the complementary cDNA strand the Rd1 sequencing adaptor sequence is introduced at the 3’ end of the V3-V4 region. At the end of the RT1 step, isothermal Exonuclease I is used to degrade the primers. 5. After a purification step, the cDNA is used in the NASBA step for isothermal amplification via single stranded RNA. In our protocol the V3-V4 region of the 16S rRNA gene is targeted for amplification (described in Figure 8) using primers targeting the T7 promoter plus the Rd2 sequencing adaptor and the Rd1 sequencing adaptor are used to amplify the V3-V4 region of the 16S rRNA. 6. After heat inactivation of the enzymes performing the NASBA step, the amplified (-)RNA representing the V3-V416S ribosomal RNA region is translated into cDNA via an ReverseAttorney Docket No.: 673 / 16 PCT Transcriptase (RT2 step) using the 5’-Rd1 primer, after which Exonuclease I is used to degrade the non-incorporated primers. This enzyme won’t degrade the DNA:RNA hybrids synthesized during the RT2 step. Subsequently, libraries are amplified using indexing primers, after which the libraries are sequenced followed by data analysis.

[0056] In some embodiments of the invention, the processing and analysis of the amplified region of the microbial rRNA such as the 16S ribosomal RNA fragment sequences includes one or more of the following steps: 1. Processing of forward and reverse reads separately when 16S rRNA gene amplicons are too long for paired-end merging. 2. Assessment of read quality using FastQC (Andrews, 2010). 3. Trimming of adapter read-through with cutadapt (Martin, 2011). 4. Removal of reads that map to the human genome using Bowtie2 (Langmead et al, 2012). 5. Quality-based read filtering, error-correction and Amplicon Sequence Variant (ASV) creation using Dada2 (Callahan et al, 2016). 6. ASV classification using rdp classifier (Wang et al, 2007). 7. Community composition is calculated based on the percentage of reads assigned to a specific phylogenetic level, such as family, genus or species.

[0057] In addition to monitoring specific diseases, Ribosome Informed Phylogeny (RIP)- based analysis that targets the ribosomes present in EVs secreted by microbes for the analysis of the microbial signature can be useful as part of the general health screening. Unlike the stool microbiome, the microbiome of colonizing and infecting bacteria will be relatively stable, with changes occurring when the relation between host and microbes is changing. This includes situations of new invasions by infectious and colonizing microorganisms, such as the formation of stomach ulcers, the formation of intestinal polyps / adenomas and their progression into malignancies, gastrointestinal diseases including Irritable Bowel Disease (IBD), various tumors and their specific microbiomes including colorectal cancer, pancreatic cancer, lung cancer, prostate cancer, breast cancer and cervical cancer, Central Nervous System (CNS) diseases including multiple sclerosis (MS), Parkinson’s disease and Alzheimer’s disease (gut-brain axes), minimal residual disease (MRD) monitoring, monitoring of eye diseases such as uveitis where disease severity might be affected by a microbial compound (gut-eye axes) and other diseases characterizedAttorney Docket No.: 673 / 16 PCT by dysbiotic and inflammatory microbiomes such as cystic fibrosis or tuberculosis, and general risk monitoring of infections in patient populations with a compromised immune system, positioning RIP analysis of microbial signatures as an ideal tool for risk monitoring, early detection, prognostics and evaluation of disease progression.

[0058] Contrary to PCR based detection methods that monitor for the presence of specific bacteria, RIP analysis of microbial signatures provides an “open” diagnostics approach to detect any bacterium, archaea, or fungus based on the ribosomal RNA signatures in EVs present in peripheral blood and other liquid biopsy samples.

[0059] There are several diseases where a clear link has been reported between gut microbiome dysbiosis and disease onset and severity, without proof of translocation of disease- causing microorganisms to the affected organs. Instead of establishment of pro-inflammatory bacteria, their circulating BMVs might be important for triggering disease responses in distantly located parts of the body. BMVs have been shown to contain functions that contribute to disease progression and severity, with their bacterial origin clearly impacting the inflammatory response they elicited. Therefore, linking BMVs and their content via the RIP method to their bacterial origin should provide a more accurate diagnostics tool than mcfDNA analysis, especially when combining RIP-based microbial identification to functional microbial models that include virulence factors and other functions that contribute to disease severity, e.g. by eliciting an immune response. This would be especially true for diseases where BMVs crossing the gut-blood, blood-brain, blood- lung or blood-eye barriers can contribute to disease severity by delivering their content to the affected area in the body.

[0060] In one aspect of the invention, RIP analysis of microbial signatures from EVs present in peripheral blood samples can provide an important non-invasive method for (early) detection and identification of infectious and colonizing bacteria, even when present in parts of the body distant from the disease location, which can subsequently be linked to a broad range of diseases, including: screening for tuberculosis and other diseases caused by Mycobacterium species; determining pulmonary infection risks and causes in cystic fibrosis patients; determining the risk and onset of sepsis in patients with compromised immune systems; detection of opportunistic bacterial pathogens originating from the oral cavity that have been linked to Alzheimer’s disease, pancreatic cancer and other serious conditions such as endocarditis; women’s health issues including Chlamydia infection linked to mucopurulent cervicitis, pelvic inflammatory disease, tubal factor infertility, ectopic pregnancy and cervical cancer; detection and monitoring of progression of cancer; monitoring of minimal residual disease after oncology treatments; detection and monitoring of progression and minimal residual disease of breast cancer including tripleAttorney Docket No.: 673 / 16 PCT negative breast cancer; detection of esophageal cancer, precancerous colonic polyps and early stage colorectal cancer, and detection and monitoring of progression and minimal residual disease of gastrointestinal cancers in general; detection and monitoring of progression and minimal residual disease in lung cancer; non-invasive analysis of the microbiome in pancreatic cancer patients to propose treatment protocols and prognostics for long-term survival; detection of Clostridium difficile infections; post-transplant bloodstream infections and Graft versus Host Disease (GvHD); detection of hospital acquired infections by emerging pathogens of clinical concern; detection of an infection in an immune compromised person; detection of infection or inflammation of the gastrointestinal track in Irritable Bowel Disease (Crohn’s disease, Ulcerative colitis); and combinations thereof; or other auto immune disease such as sarcoidosis, ankylosing spondylitis (AS), Behçet Disease (BD), multiple sclerosis (MS), Vogt-Koyanagi-Harada syndrome (VKH), uveitis, etc. Therefore, RIP analysis of microbial signatures from EVs represents a quantum leap forward as a high-resolution, high-throughput and low-cost routine test in disease detection, patient monitoring, risk assessment and large-scale population screenings using BMV-informed biomarkers.

[0061] A method is provided of amplifying nucleic acid in a sample, that includes: isolating nucleic acids from a plurality of extracellular vesicles (EVs) from a sample; treating the isolated EV nucleic acids with DNase to essentially remove DNA; and amplifying a region of a microbial ribosomal RNA (rRNA) in the DNase-treated EV nucleic acids, thereby obtaining amplified rRNA. The method can further include converting the amplified microbial rRNA to cDNA. In some cases, the method further includes sequencing the cDNA. The method can include using a computer to search the sequenced cDNA against a database of genes corresponding to the microbial rRNA and assigning a microbial species based on the closest sequence match. In some cases, one or a combination of: (i) a presence, (ii) an absence, or (iii) a relative abundance of one or a combination of the microbial species is associated with a disease or disorder. In these cases, the one or the combination of the presence, the absence, or the relative abundance of the one or a combination of the microbial species in the sample can indicate: (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

[0062] In another aspect, a method is provided of diagnosing or monitoring a disease or disorder, that includes: isolating nucleic acids from a plurality of extracellular vesicles (EVs) from a sample; treating the isolated EV nucleic acids with DNase to essentially remove DNA; amplifying a region of a microbial ribosomal RNA (rRNA) in the DNase-treated EV nucleic acids, thereby obtaining amplified microbial rRNA; converting the amplified microbial rRNA to cDNA;Attorney Docket No.: 673 / 16 PCT sequencing or obtaining sequences of the cDNA; and assigning one of more of the cDNA sequences as belonging to a particular microbial species, or as comprising a microbial DNA signature, based on the closest sequence match to a database of genes corresponding to the microbial rRNA. In the case where one or a combination of: the microbial DNA signature, or (i) a presence, (ii) an absence, or (iii) a relative abundance of the one or a combination of the microbial species is associated with a disease or disorder, the one or a combination of the microbial DNA signature, the presence, the absence, or the relative abundance of the one or a combination of microbial species in the sample can indicate (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

[0063] In one case, a method is provided of diagnosing or monitoring a disease or disorder, that includes: providing cDNA sequences derived from a sample from a subject, wherein the cDNA sequences comprise sequence corresponding to a region of a microbial ribosomal RNA (rRNA), wherein the region of a microbial rRNA was amplified from nucleic acid present in a plurality of extracellular vesicles (EVs) from the sample after the nucleic acid was treated with DNase to essentially remove DNA; and assigning one or more of the cDNA sequences as comprising a microbial DNA signature or as belonging to a particular microbial species based on the closest sequence match to a database of genes corresponding to the microbial rRNA. In the case that one or a combination of: the microbial DNA signature, or (i) a presence, (ii) an absence, or (iii) a relative abundance of the one or a combination of the microbial species is associated with a disease or disorder, the one or a combination of the microbial DNA signature, the presence, the absence, or the relative abundance of the one or a combination of microbial species in the sample can indicate (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

[0064] Another method is provided for diagnosing or monitoring a disease or disorder, that includes: providing cDNA sequences derived from a sample from a subject, wherein the cDNA sequences comprise sequence corresponding to a region of a microbial ribosomal RNA (rRNA), wherein the region of a microbial rRNA was amplified from nucleic acid present in a plurality of extracellular vesicles (EVs) from the sample after the nucleic acid was treated with DNase to essentially remove DNA; and assigning one or more of the cDNA sequences as comprising a microbial DNA signature or as belonging to a particular microbial species based on the closest sequence match to a database of genes corresponding to the microbial rRNA. In the case that one or a combination of: the microbial DNA signature, or (i) a presence, (ii) an absence, or (iii) a relativeAttorney Docket No.: 673 / 16 PCT abundance of the one or a combination of the microbial species is associated with a disease or disorder, the one or a combination of the microbial DNA signature, the presence, the absence, or the relative abundance of the one or a combination of microbial species in the sample can indicate (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

[0065] In another aspect, a method is provided of diagnosing or monitoring a disease or disorder, that includes: isolating nucleic acids from a plurality of extracellular vesicles (EVs) from a sample; treating the isolated EV nucleic acids with DNase to essentially remove DNA; amplifying a region of a microbial ribosomal RNA (rRNA) in the DNase-treated EV nucleic acids, thereby obtaining amplified microbial rRNA; converting the amplified microbial rRNA to cDNA; and sequencing or obtaining sequences of the cDNA, wherein, the closer the sequence match between the cDNA sequences and nucleic acid sequences in a reference sample for a disease or disorder, the more likely the sample is to have the disease or disorder.

[0066] In the methods provided herein, the microbial rRNA can be a bacterial rRNA or a eukaryotic rRNA. The bacterial rRNA can comprise 16S rRNA, 23S rRNA, or 5S rRNA. The region of the 16S rRNA gene can comprise the V3-V4 region. The eukaryotic rRNA can comprises 5S rRNA, 18S rRNA, 28S rRNA, or 5.8S rRNA.

[0067] In the methods provided herein, the sample can comprise liquid biopsy samples, blood, spinal fluid, cerebral fluid, urine, saliva, sputum, lymph fluid, or stool, and combinations thereof.

[0068] In some embodiments of the methods provided herein, the plurality of EV’s include human EV’s and bacterial extracellular vesicles (BMVs) and the human EV’s are selectively lysed after which their DNA and RNA content can be degraded by nuclease activity prior to isolating nucleic acids from the plurality of EV’s.

[0069] In the methods provided herein, amplifying the region of a microbial rRNA can include: a first reverse transcription (RT) step to translate the region of a microbial rRNA into cDNA, wherein a 5’ end of the cDNA is extended by incorporation of a T7 promoter and a first sequencing adaptor and a 3’ end of the cDNA is extended by incorporation of a second sequencing adaptor in the first RT step, and wherein the first RT step is followed by isothermal amplification of the region of a microbial rRNA using primers that target the extended ends of the cDNA, thereby obtaining the amplified EV rRNA.

[0070] In the methods provided herein, a unique molecular identifier sequence (UMI) can be incorporated into the cDNA in the first RT step.Attorney Docket No.: 673 / 16 PCT

[0071] In the methods provided herein, the amplifying the region of a microbial rRNA comprises nucleic acid sequence-based amplification (NASBA).

[0072] The method provided herein can further include amplifying a region of one or more RNA transcripts in the DNase-treated EV nucleic acids, thereby obtaining one or more amplified transcript RNAs. In this case, when one or a combination of: (i) a presence, (ii) an absence, or (iii) a relative abundance of the one or more amplified transcript RNAs is associated with a disease or disorder, the one or a combination of the presence, the absence, or the relative abundance of the one or more amplified transcript RNAs in the sample can indicate (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

[0073] In the methods provided herein, the disease or disorder can include, but is not limited to, formation of stomach ulcers, formation of intestinal polyps / adenomas and their progression into malignancies, gastrointestinal diseases, Irritable Bowel Disease (IBD), tumors, colorectal cancer, pancreatic cancer, lung cancer, prostate cancer, breast cancer and cervical cancer, Central Nervous System (CNS) diseases, multiple sclerosis (MS), Parkinson’s disease, Alzheimer’s disease, minimal residual disease (MRD) monitoring, monitoring of eye diseases, uveitis, cystic fibrosis, tuberculosis, and general risk monitoring of infections in patient populations with a compromised immune system, and combinations thereof.

[0074] In another embodiment, a system is provided for amplifying nucleic acid in a sample that includes: a reaction vessel; a reagent dispensing module; and software to execute the method of any of the foregoing claims, wherein the method is executed at least partially robotically. EXAMPLES

[0075] The following Examples have been included to provide guidance to one of ordinary skill in the art for practicing representative embodiments of the presently disclosed subject matter. In light of the present disclosure and the general level of skill in the art, those of skill can appreciate that the following Examples are intended to be exemplary only and that numerous changes, modifications, and alterations can be employed without departing from the scope of the presently disclosed subject matter. EXAMPLE 1 Bacterial species found as contaminants in negative controls.Attorney Docket No.: 673 / 16 PCT

[0076] An important issue that hinders the analysis of microbial DNA signatures in liquid biopsy samples is the presence of contaminating DNA. For instance, the inventors have observed (results not shown) that skin associated bacteria like Cutibacterium species and Staphylococcus species are routinely identified among the bacteria that make up the microbial DNA signature in blood, this in addition to other bacterial species that are not known to be associated with a human host. To better understand the origin of these bacteria, a negative control experiment was performed on water blank samples instead of blood plasma that included the following steps: 1. Using the QIAGEN QIAamp ccfDNA / RNA Kit designed for the purification of cfDNA including vesicular and non-vesicular nucleic acids, DNA was extracted from 0.5 ml DNA-free water (i.e., water blank samples) and collected in 20 µl nucleotide free water. Theoretically, no DNA should be isolated. 2. The V3 – V4 region of the 16S rRNA gene was amplified using the standard Rd1-16SV3-F350 and Rd2-16SV4-R781 primers. Premix the following on ice: • 12.5 µl 2x KAPA HiFi HotStart ReadyMix • 8.5 µl Nuclease-Free Water • 1 µl Rd1-16SV3-F350 primer (10 µM) • 1 µl Rd2-16SV4-R781 primer (10 µM) Add 2 µl of the isolated DNA. Place the reactions in the thermocycler and run the following program: heated lid on, 95ºC for 10 min, [98ºC for 20 sec, 60ºC for 30 sec, 72ºC for 15 sec] for 35 cycles, followed by 72ºC for 1 min and 4ºC on hold. The PCR product is cleaned up with 2x AMPure beads and resuspend in 18 µl. 3. Indexing PCR for the construction of the sequencing libraries was performed by preparing for each sample the following reaction on ice: • 25 µl 2x KAPA HiFi HotStart ReadyMix • 10 µl UDI primers, using a unique set of Nextera UDI primers for each sample • 10 µl PCR product from step 2. Place the reactions in the thermocycler and run the following program: heated lid on, 95ºC for 10 min, [95ºC for 30 sec, 55ºC for 30 sec, 72ºC for 30 sec] for 7 cycles, 72ºC for 5 min, 4ºC on hold. The PCR product is cleaned up with 2x AMPure beads and resuspend in 20 µl. For the results described in EXAMPLES 1 and 2 herein below, the library pool was subsequentlyAttorney Docket No.: 673 / 16 PCT sequenced following the 150bp paired-end read sequencing protocol from Illumina using the Illumina iSEQ100 i1 Reagent v2 (300-cycle) kit on an Illumina iSEQ100 instrument according to the manufacturer’s instructions.

[0077] Processing and analysis of the V3 – V416S rRNA gene fragment sequences were performed using the following steps: 1. Processing of forward and reverse reads is done separately when 16S amplicons are too long for paired-end merging. 2. Assessment of read quality is done using FastQC (Andrews, 2010). 3. Cutadapt is used for trimming adapter read-through sequences (Martin, 2011). 4. Removal of reads that map to the human genome is performed using Bowtie2 (Langmead et al, 2012). 5. Quality-based read filtering, error-correction, and Amplicon Sequence Variant (ASV) creation is done using Dada2 (Callahan et al, 2016). 6. The rdp classifier is used for ASV classification (Wang et al, 2007) using a 50% confidence level. 7. Community composition is calculated based on the percentage of reads assigned to a specific phylogenetic level, such as family, genus, or species.

[0078] The results of two independent negative control experiments using the same materials are presented in Figures 1A and B and Figures 2A and B detail which contaminating species were found at the genus and family levels, respectively, when DNA was extracted from DNA-free water.

[0079] Figures 1A and B are graphs showing the relative abundance of bacterial contaminants detected in negative control water samples. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNA extracted from negative control water samples. After sequencing, bacteria were identified on the genus level.

[0080] Figures 2A and Bare graphs showing the relative abundance of bacterial contaminants detected in negative control water samples. The presence of bacterial contamination was evaluated using V3-V416S rRNA gene amplicon sequencing on DNA extracted from negative control water samples. After sequencing, bacteria were identified on the family level.Attorney Docket No.: 673 / 16 PCT

[0081] As shown in Figures 1A and B, Escherichia / Shigella, Pseudomonas, Cutibacterium, Streptococcus, Parcubacteria_genera_incertae_sedis, Streptophyta and Prevotella represent most of the classified bacteria at the genus level. At the family level (Figures 2A and B), Propionibacteriaceae, Enterobacteriaceae, Pseudomonadaceae, Streptococcaceae, Chloroplast, Parcubacteria_genera_incertae_sedis, Prevotellaceae and Corynebacteriaceae were the most dominant bacteria found in the negative control samples. This negative control experiment illustrates that bacterial DNA contamination is a major problem during DNA extraction, especially for samples with low amounts of bacterial DNA, such as blood plasma samples. The presence of the contaminating DNA will have major impacts on subsequent PCR amplification, construction of amplicon sequencing libraries, and the interpretation of their results. It should be noted that Pseudomonaceae and Corynebacteriaceae were previously reported as important bacteria that allowed, based on V3-V416S rRNA gene amplicon sequencing of DNA extracted from BMVs, to differentiate between healthy controls and biliary track cancer patients Lee et al (2020). The same bacterial families detected in this experiment were found when analyzing the amplicon profiles obtained from DNA-free water controls, confirming that they likely represent contamination. It should be noted that the Lee et al (2020) paper did not report using experimental controls, including negative DNA-free water controls that went through the same experimental procedure as blood plasma for BMV isolation, DNA extraction, and V3-V416S rRNA gene amplification and sequencing, or subsequent bioinformatics steps for in silico removal of sequences from contaminating bacterial DNA. As such the accuracy of the reported results is questionable.

[0082] The presence of the strong signal from contaminating DNA significantly complicates the analysis and interpretation of relatively weak microbial signatures from biopsy samples. In fact, when using amplicon sequencing on DNA of the V3-V4 region of the 16S rRNA gene, the signal from the contaminating bacterial DNA found in negative DNA-free water control samples often contributes 50% or more to the microbial signature in plasma (results not shown).

[0083] Considering that DNA-free water samples 1 and 2 were independently processed using the same materials (DNA extraction kit, PCR enzymes and primer stocks), it is surprising to see the differences in community composition obtained after amplicon sequencing of the V3-V4 region of the 16S rRNA gene. This indicates that the types and amounts of contaminating DNA are highly variable, further complicating the use of mcfDNA to accurately determine the microbial signature in blood plasma and other liquid biopsy samples. As such, different approaches are required to overcome the above identified limitations and shortcomings of methods that focus on mcfDNA for purposes of disease diagnosis and monitoring.Attorney Docket No.: 673 / 16 PCT EXAMPLE 2 Ribosomes as phylogenetic markers.

[0084] To overcome the problems of contaminating DNA, a new method referred to as Ribosome Informed Phylogeny (RIP) was developed that targets the RNA present in extracellular vesicles (EVs) secreted by microbes for the analysis of the microbial signature in liquid biopsy samples, such as blood plasma. The initial focus is on ribosomal RNAs (rRNAs), which represent the most abundant RNA species in the cell. Compared to an average of 5 copies of ribosomal RNA genes, around 15,000 ribosomes can be present in a single bacterial cell. The same approach can be used for any RNA, such as messenger RNA (mRNA).

[0085] The rRNA present in the EVs found in blood was isolated, processed and sequenced following the steps described below. The protocol steps and their objectives are listed in Table 2. Protocol step Objective , s e nAttorney Docket No.: 673 / 16 PCT Table 2: Overview of the protocol steps used in the Ribosome Informed Phylogeny (RIP) protocol and their specific objectives.

[0086] Purification of EVs: 1. Thaw 1 mL of blood frozen plasma samples, collected in EDTA or Streck tubes. 2. Add 100 µl Proteinase K (QIAGEN # 19131, 20 mg / mL) for protein degradation including RNase (Bender et al, 2020). Incubate at 50ºC for 1 hour. 3. Centrifuge twice at 3,000 rpm for 15 min at 4 °C to remove cells and debris. 4. Mix 0.95 mL supernatant with ice-cold 0.95 ml 1x PBS, pH 7.4, nuclease-free, in a 2-mL tube. 5. Centrifuge once at 13,000 rpm for 1 min at 4 °C. This centrifugation step leaves the EVs in the supernatant. 6. Filter the supernatant through a 0.45 µm filter to remove any remaining cells and debris from bacteria and large organelles while the EVs are smaller in size and stay in the filtrate. Work with samples on ice and collect flowthrough into a 2-mL tube. Use a 1-2 mL syringe. Once all the sample was pushed through the filter, load the syringe with 0.7 mL 1x PBS, pH 7.4 and push it through the filter to eject the remaining sample from the filter (dead volume).

[0087] Nucleic acid isolation: Total nucleic acid extraction, including DNA, RNA and small RNA, is performed using the QIAamp ccfDNA / RNA Kit (QIAGEN # 55184). Extraction is performed according to the manufacturer’s instruction, except for a boiling step to lyse the EVs and no RNAse A treatment. 7. Boil EV-filtrate (~1.8 mL) and incubate for 10 min at 50ºC 8. Boil EV-filtrate for 30 min at 100°C 9. Centrifugation for 30 min at 12,000 g and 4 °C to eliminate the remaining IVs and debris. 10. Transfer the supernatant to a new 5-mL tube. 11. Add 600 µl of buffer RPL, close the tube cap, and vortex for 5 sec. Incubate at 15-25ºC for 3 min. 12. Add 200 µl of buffer RPP and immediately mix by vortexing for 30 sec. Incubate on ice for 6 min.Attorney Docket No.: 673 / 16 PCT 13. Centrifuge at 4500 g for 6 min in the Allegra X-12R centrifuge; the supernatant should be clear and colorless. 14. Transfer the supernatant to a new 5-mL tube. Keep on ice. 15. Add 1 volume of ice-cold isopropanol and keep it on ice until all samples have been processed. No specific incubation time is required. 16. Transfer all the solution (may require several rounds), including any precipitate, onto a Rneasy Midi spin column in a 15 ml collection tube and centrifuge at 4500 g for 1 min at room temperature in the Allegra X-12R centrifuge. Discard the flow-through. 17. Add 4 ml of buffer RWT to the Rneasy Midi spin column and centrifuge at 4500 g for 1 min in the Allegra X-12R centrifuge. Discard the flow-through and keep the tube. 18. Add 2.5 ml of buffer RPE to the Rneasy Midi spin column and centrifuge at 4500 g for 5 min in the Allegra X-12R centrifuge. Discard the flow-through and keep the tube. 19. Place the Rneasy Midi spin column into a new 15 ml collection tube (supplied). Add 200 μl nuclease-free water directly to the center of the spin column membrane. Close the lid gently, incubate for 1 min, then centrifuge for 1 min at full speed in the Allegra X-12R centrifuge to elute the cfDNA. 20. Add 200 μl Buffer RPL and 800 μl ethanol (96–100%) and mix by vortexing. 21. Transfer 700 μl, including any precipitate, onto a Rneasy MinElute spin column in a 2 ml collection tube. Close the lid gently and centrifuge for 15 sec at ≥8000 g and room temperature in the Microfuge 22R centrifuge. Discard the flow-through and keep the tube. 22. Repeat step 15 using the remainder of the sample and make sure all liquid has passed through the column membrane completely, before proceeding to the next step. 23. Perform DNAse treatment: a. Combined 1 µl DNase I (2 U / µl, NEB #M0570) to 100 µl of 1X Buffer. Mix by gently inverting the tube. Centrifuge briefly to collect residual liquid from the sides of the tube. Note: DNase I is especially sensitive to physical denaturation. Mixing should only be carried out by gently inverting the tube. Do not vortex. b. Add the DNase I incubation mix (100 µl) directly to the RNeasy spin column membrane, and incubate at 37°C for 15 min.Attorney Docket No.: 673 / 16 PCT 24. Add 500 μl of buffer RPE onto the Rneasy MinElute spin column and centrifuge for 15 sec at ≥8000 x g (≥10,000 rpm) in the Microfuge 22R centrifuge. Discard the flow-through and keep the tube. 25. Add 500 μl of 80% ethanol to the Rneasy MinElute spin column. Close the lid and centrifuge for 15 sec at ≥8000 x g in the Microfuge 22R centrifuge. Discard the flow-through and the collection tube. 26. Place the Rneasy MinElute spin column in a new 2 ml collection tube. Open the lid of the spin column and place the spin columns into the centrifuge with at least one empty position between columns. Centrifuge at full speed for 5 min in the Microfuge 22R centrifuge to dry the membrane. Discard the flow-through and the collection tube. 27. Place the Rneasy MinElute spin column in a new 1.5 ml collection tube. Add 15 μl Rnase- free water directly to the center of the spin-column membrane. Close the lid gently, incubate for 1 min and centrifuge for 1 min at full speed in the Microfuge 22R centrifuge to elute the DNA / RNA.14 µl of EV-extract should be recovered.

[0088] NASBA amplification: Targeted reverse transcription (first and second strand) of the 16S rRNA V3V4 region followed by in vitro transcription (IVT) is performed using the NWK- 1 NASBA kit (Life Sciences Advanced Technologies Inc., https: / / lifesci.com / product / nwk-1 / ) containing Reverse transcriptase, RNaseH and T7 RNA polymerase enzyme mix. An alternative to the T7-16SV4-R785 primer, referred to as T7plus1-16SV4-R785 with an improved extended T7 promoter to maximize transcription (based on Conrad et al, 2020), was also validated (table 1). In addition, as an alternative to the T7plus1-16SV4-R785 primer, the T7plus2-16SV4-R785 (table 1) can be used. 28. Prepare the following reaction per sample: 4 µl of EV-extract + 1 µl primer mix 1 [T7-16SV4- R785 + 16SV3-F350] at 5 µM. 29. Incubate the reactions for 2 minutes at 65 °C and subsequently cooled to 41°C for 10 min in a thermocycler. This step is important to dissociate complementary RNA regions in the ribosomal RNA. 30. To each reaction, add 10 µl of [6.7 µl NECB-24 + 3.3 µl NECN-24] pre-heated at 41ºC for 5 min. 31. To each reaction, add 5 µl NEC 1-24 enzyme mixtureAttorney Docket No.: 673 / 16 PCT 32. Incubate at 41 °C for 90 minutes 33. To each reaction, add 1 µl DNase I (2 U / µl, NEB #M0570) in 9 µl of 1X Buffer. 34. Incubate at 95 °C for 10 minutes to stop the reactions 35. Clean up using the RNA Clean XP kit (Beckman Coulter) and elute in 20 µl RNAse-free water

[0089] Reverse transcription (RT) of RNA into cDNA: RNA transcripts of the V3-V4 16S rRNA region were reverse transcribed into cDNA using the SuperScript™ IV First-Strand Synthesis with the 16SV3-F350 primer (Table 1). 36. Combine 11 µl RNA sample (max 200 ng) with 1 µl dNTP (10 mM), 1 µl Rd1-16SV3-F350 primer (2 µM). 37. Mix at 65ºC for 5 minutes, and then on ice for at least 1 minute. 38. Prepare the following RT mix and add it to the reaction. • 4 µl 5X Buffer • 1 µl DTT (100mM) • 1 µl Ribonuclease Inhibitor • 1 µl SuperScriptTM IV Reverse Transcriptase (200 U / μL) 39. Incubate at 50°C for 20 minutes 40. Incubate at 80°C for 10 minutes to stop the reaction 41. Add 1 μL E. coli RNase H, and incubate at 37°C for 20 minutes to remove the (-)RNA

[0090] PCR amplification to introduce sequencing adapters (A-PCR): cDNA is amplified by PCR using the Rd1-16SV3-F350 and Rd2-16SV4-R781 primers that will introduce the adapters for sequencing library construction. It should be noted that compared to the T7-16SV4-R785 or T7plus1-16SV4-R785 primers used during NASBA, the Rd2-16SV4-R781 primer used for this PCR step contains a four-nucleotide extension at its 3’ end, allowing for enrichment PCR to further reduce the risk of converting and amplifying non-specific RNA fragments into cDNA. 42. Premix the following on ice: • 25 µl 2x KAPA HiFi HotStart ReadyMixAttorney Docket No.: 673 / 16 PCT • 3 µl Nuclease-Free Water • 1 µl Rd1-16SV3-F350 primer (10 µM) • 1 µl Rd2-16SV4-R781 primer (10 µM) 43. Add 20 µl of RT product + water (100 ng estimate) 44. Place the reactions in the thermocycler and run the following program: heated lid on, 95ºC for 10 min, [98ºC for 20 sec, 60ºC, 72ºC for 15 sec] for 14 cycles, 72ºC for 1 min, 4ºC on hold. 45. Quantify 1 µl to assess concentration and add more amplification cycles if necessary 46. Clean up with 2x AMPure beads and resuspend in 18 µl

[0091] Indexing PCR: This step allows for subsequent sequencing and multiplexing of the RIP sequencing libraries. 47. Prepare the following reaction on ice: • 25 µl 2x KAPA HiFi HotStart ReadyMix • 10 µl UDI primers, using a unique set of Nextera UDI primers for each sample • 15 µl of A-PCR product 48. Place the reactions in the thermocycler and run the following program: heated lid on, 95ºC for 10 min, [95ºC for 30 sec, 55ºC for 30 sec, 72ºC for 30 sec] for 7 cycles, 72ºC for 5 min, 4ºC on hold. 49. Clean up with 2x AMPure beads. 50. After quantification, the libraries are ready to be sequenced, e.g. on an iSeq100 sequencer according to the manufacturer’s instructions.

[0092] The processing and analysis of the V3-V416S rRNA gene fragment sequences: this step was performed as described in EXAMPLE 1.

[0093] To validate the critical protocol steps described in Table 2, 17 different iterations of the RIP protocol were performed represented by sequencing libraries 1 to 17. These iterations are summarized in Table 3. For libraries 1 to 11 blood plasma blood plasma was used, while libraries 12 to 17 represented negative controls using nucleotide-free water. Table 4 and Figure 5Attorney Docket No.: 673 / 16 PCT summarize the results obtained after sequencing and analysis of the microbial signatures obtained for libraries 1 to 17. To simplify the interpretation of the results, only the 30 most abundant genera found in libraries 1 to 17 were reported. RIP library number 17 X X X Xprotocol to confirm the importance of critical protocol steps. X indicates that a step was performed on a sample. In total 17 different iterations of the RIP protocol including negative water controls were evaluated.Attorney Docket No.: 673 / 16 PCT

[0094] Table 4 shows microbial signatures obtained after different iterations of the Ribosome Informed Phylogeny (RIP) protocol, and Figure 5 is a graph depicting the data from Table 4. Abundance of sequencing reads for the 30 most abundant genera observed after analysis of the sequences obtained for RIP libraries 1 to 17 are reported. The differences in protocol iterations for the samples is described in Table 3.3126082 751 52 1 929 603 43 1 2 763669 953 22 41 907 03 49038 0 15 3 5 6 1 339 6 4 4 0 00 1 1 9 1 008 0 4033281 021 827 63 198 1 31 22 70 5120 363 41001 3191 04 49 291 12861 8 1188 1242 0729 01 5 1391 135 76 64 91 01 8951 63 7671 50 6523 53294 0 421 55 9321 81 68041 50225 631 000 0 0 01 00 00 0 4 1 1 0067 272478947 4 5 06893 5 6431 1 52 070884 1644 0 01 1 631 345 21 54 4 1 032 0 2 9 7 63 1 0 7 1 070 1 75925T 22 8 8 07884 2 9 1 7 8 6 8C 797 0531 65 124042 3 0 0 5 939 3P17 7 1 55 2 1 30 1 1 4 3 1 30 32436 46 2 03 0 3 1 1 290 2 2 1 32 2 31 / 3s7 mia6u:.ir s suoetuc c alseamsaatsanare m fssim uasN cc clenoustp yogis ah ia rrnirrec osap uar deuirireehpsuecpoettae sbeococn olnaaim r oetoetncmIm oca aalorsenomcc eS etai tcdioelyalylemieaocnee capaboiirocmlneecgabkcunoon yty rpeh meg p ussni abitdb uucresilehtpak resdireorGa1trb o eosrtordraalclioro ntiovoraiisaidtfolhtDeo G CrtSathScieh ueS Np S CsrPanedP GemeopS Go P ReteciRcnIuuaciocacitceryleu oyT P M S F L M N B M A P H D G RenrottA8 4 974 7 1 66 52e1 1 1 1htf992 5 8 1 21 11 0 0 4 4 1 0010 065 2e1vres37 9 b 0 0 1 o a 5re86ne1 g 0 0 7 1tna0 4 d 4 4 n 8 52 5 uba3 9tso 6961 55 3 2 88m03 8 0e87 0ht2 06 2 1 77rofs0 046 02 dae2r9 g - 7 n 63 01577 5i2cne -3 u 3 q 04es38 8 0 1 42fo 2ec0 0 6 5nadn 0 088 2 u 4 ba. 7 9 d 7 n10 081 24ao det1s204 6ni e6atir03311 02 boarsbT 0eirlC 1 45 uPIP6 0 771 8t4an R gr1 / isof3 ml7d6 a:iu.leird ot ai en ov eteidebiatNafit l cas iforb bse csa isciokaolcnicssruaMslc s :4eco moht ene tuneDh nhtgonTaJo *otneelbauqy* geTesnrottAAttorney Docket No.: 673 / 16 PCT

[0095] Based on the results presented in Figure 5 the following conclusions can be drawn. Nearly all negative control samples (represented by libraries 12 to 17), especially the negative controls samples 15 to 17 where nucleotide-free water instead of plasma was used for the nucleotide extraction step, provided a microbial signature. This signature represents the bacteria whose nucleotides ended up as contaminants. Since RNases are omni present, it is assumed that these signatures are obtained from contaminating bacterial DNA that enters the samples during the nucleotide extraction step. The dominant genera observed include Schumanella, Cutibacterium, Corynebacterium, Streptococcus and Staphylococcus. This result is similar to the result reported in EXAMPLE 1, Figures 1A and B, confirming that these bacterial genera represent contaminating DNA. It was also noticed that during the NASBA step, which should specifically result in the amplification of RNA templates, contaminating DNA that contains the 16S rRNA gene region was also amplified, as illustrated by library 12

[0096] Plasma samples that were processed without the NASBA step (represented by libraries 3, 8, 11) showed no significant microbial signature, this in contrast to library 17, the water extracted sample that was processed without the NASBA step. The major difference between the extracted plasma and water samples is the presence of significant amounts of human cfDNA in the extracted plasma samples, which might inhibit the PCR amplification (A-PCR step) of the V3-V4 16S rRNA gene fragment. It should also be noted that samples that were processed without the NASBA step were amplified using an A-PCR step with 35 cycles instead of 14 cycles. As such these libraries were obtained using the same protocol for amplicon sequencing as detailed in EXAMPLE 1.

[0097] The microbial signatures for plasma samples that were extracted without the DNase treatment step (represented by libraries 1, 2, 9, 10) were dominated by the same contaminating bacteria as observed in the negative control samples (libraries 15 to 17), indicating that this step is critical to remove contaminating DNA. It also confirms that DNA contamination represents a major problem for the analysis and interpretation of microbial signatures in blood plasma and other liquid biopsy samples. On the other hand, including the DNase treatment step resulted in nearly complete elimination of contaminating bacterial DNA as well as human cfDNA. This is illustrated by the results for libraries 6 and 7, whose only difference is that the sample used for library 7 underwent a second DNase treatment after the NASBA step. For libraries 6 and 7, when compared to libraries for samples that didn’t undergo a DNase treatment step during the extraction of their nucleic acids, very low levels of bacteria representing contaminating DNA were observed as part of their microbial signature. Instead, Sphingomonas bacteria were detected, representing a genus that is well documented to reside inside the human host. Sphingomonas wereAttorney Docket No.: 673 / 16 PCT also detected in plasma samples that were processed without the DNase treatment step, but at a very low level compared to the levels of the bacteria representing contaminating bacterial DNA. Compared to libraries representing samples that didn’t receive the DNase treatment during the extraction of the nucleotides from blood plasma, relatively low levels of Streptococcus and Staphylococcus were reported. Although it can’t be excluded that these bacteria represent contaminating DNA, they represent species that are well documented to have associations with the human host. Overall, it can be concluded that without removing the contaminating bacterial DNA, it is extremely difficult to decipher a plasma derived microbial signature. On the other hand, the method that includes the DNase treatment step results in libraries representative for the microbial signature present in the liquid biopsy sample, with less interference of contaminating bacterial DNA. This represents a major improvement over the current state of the art where bacterial cfDNA or DNA isolated from BMVs is used for determining microbial signatures in liquid biopsy samples.

[0098] After the DNase treatment step, the amount of RNA available for the NASBA step is very low. Using the T7plus1-16SV4-R785 primer (samples 6 and 7), which has an extended T7 promoter to maximize transcription (based on Conrad et al, 2020) instead of the T7-16SV4-R785 (samples 4 and 5), significantly improved the efficiency of the NASBA reaction. Therefore, it is essential to use the T7plus1-16SV4-R785 primer for the RIP protocol.

[0099] With the T7plus1-16SV4-R785 primer the formation of primer concatemers was observed, something that likely happens as the result of template jumping by the reverse transcriptase enzyme during the NASBA step. This was observed to a much lesser degree when using the T7-16SV4-R785 primer, indicating that this is primer specific. To address this issue, the T7plus2-16SV4-R785 primer (Table 1) was tested as an alternative to the T7plus1-16SV4-R785 primer. Using the T7plus2-16SV4-R785 primer resulted in a 30% reduction of fragments comprised of primer concatemers (results not shown). Therefore, it is essential to use the T7plus2- 16SV4-R785 primer for the RIP protocol.

[0100] It should be noted that various alternative steps can be used for the RIP protocol that are obvious to someone trained in the field. For instance, alternative short read sequencing methods can be used besides Illumina short read sequencing, such as the short fragment mode (SFM) on the Oxford Nanopore Technologies sequencing platform, or PacBio sequencing by binding using the ONSO system’s short-read sequencing chemistry. Such modifications would require changes to the protocol steps “PCR amplification to introduce sequencing adapters” and “Indexing PCR”.

[0101] Several alternatives are also available to NASBA for the isothermal amplification of RNA, such as loop mediated isothermal amplification (LAMP), self-sustained sequenceAttorney Docket No.: 673 / 16 PCT replication (3SR), rolling circle amplification (RCA), strand displacement amplification (SDA), ligase chain reaction (LCR), transcription mediated amplification (TMA) and modifications thereof. These alternatives to NASBA would require changes to the protocol step “NASBA amplification” and possibly “Reverse transcription (RT) of RNA into cDNA” if isothermal amplification of the RNA already includes a step that converts the RNA template into cDNA. EXAMPLE 3 Variation of the RIP method to further reduce amplification of contaminating background DNA.

[0079] Although the Ribosome Informed Phylogeny (RIP) protocol from EXAMPLE 2 significantly reduces the signal caused by contaminating DNA fragments compared to other methods, there are still protocol steps that will result in the amplification of V3-V416S rRNA gene fragments; this is the result of using primers during NASBA and A-PCR amplification steps that specifically recognize the conserved DNA sequences flanking the V3-V416S rRNA gene region. Therefore, a variation of the RIP method was developed (Figure 6), referred to as RIP2, that compared to the method described in EXAMPLE 2 has the following modifications: 1) Before the NASBA amplification step, a Reverse Transcriptase step (described in Figure 7) is used to translate the V3-V4 region of the 16S ribosomal RNA into cDNA. During this process, a T7 promoter sequence, the Rd2 sequencing adaptor sequence, and a unique molecular identifier (UMI) are introduced at the 5’ end of the cDNA fragment covering the V3- V4 region. In addition, during the synthesis of the complementary cDNA strand the Rd1 sequencing adaptor sequence is introduced at the 3’ end of the V3-V4 region. 2) In the subsequent NASBA amplification step (described in Figure 8), a primer covering the T7 promoter plus the Rd2 sequencing adaptor and a primer covering the Rd1 sequencing adaptor are used to amplify the V3-V4 region of the 16S rRNA via amplified RNA targets that are converted into dsDNA, which is subsequently used as the template for RNA synthesis to create additional copies of the amplified RNA targets. 3) After a second (optional) RT step to convert the amplified RNA into cDNA, the library is immediately amplified (I-PCR) using indexing primers without performing the amplification PCR (A-PCR) step.

[0080] Figure 6 is a schematic representation of a modified protocol for Ribosome Informed Phylogeny (RIP)-based analysis of microbial signatures in liquid biopsy samples. Extracellular vesicles (EVs) are isolated from liquid biopsy samples, including but not limited to blood plasma. Once isolated, the EVs are heat treated and subsequently total RNA is isolated. The nucleotide sample is treated with DNase to remove any DNA from the sample. Subsequently, aAttorney Docket No.: 673 / 16 PCT Reverse Transcriptase step (RT1) is performed to convert the V3-V4 region of the 16S rRNA ribosomal subunit into cDNA. During this step the 5’ and 3’ ends of the cDNA are extended with sequences for the T7 promoter, the Rd2 sequencing adaptor and a Unique Molecular Identifier (UMI) sequence, and the Rd1 sequencing adaptor, respectively. This region is subsequently amplified by nucleic acid sequence-based amplification (NASBA) using primers that target the extended 5’ and 3’ ends of the cDNA. Once completed, the enzymes for the NASBA step are inactivated, and in a second Reverse Transcriptase step (RT2) the amplified RNA representing the V3-V4 region of the 16S rRNA ribosomal subunit is converted into single stranded cDNA. During the following indexing PCR (I-PCR) step, single stranded cDNA is converted into double stranded DNA and the library is amplified using indexing primers and sequenced.

[0081] Figure 7 is a schematic representation of the Reverse Transcriptase protocol for the synthesis of cDNA of the V3-V4 region of the 16S rRNA ribosomal subunit. During this step, the T7 promoter sequence, the Rd2 sequencing adaptor sequence and a unique molecular identifier sequence (UMI) are introduced at the 5’ end of the V3-V4 region. In addition, the Rd1 sequencing adaptor sequence is introduced at the 3’ end of the V3-V4 region.

[0082] Figure 8 is a schematic representation of the nucleic acid sequence-based amplification (NASBA) protocol for amplification of the V3-V4 region of the 16S rRNA ribosomal subunit. For this step, a primer covering the T7 promoter plus the Rd2 sequencing adaptor and a primer covering the Rd1 sequencing adaptor are used. It should be noted that during T7 transcription the T7 promotor sequence is not transcribed into RNA. Therefore, it is important that the primer covering the T7 promoter plus the Rd2 sequencing adaptor efficiently anneals to the Rd2 sequencing adaptor region.

[0083] Thus, the key differences between this modified protocol and the RIP protocol described in EXAMPLE 2 are that during the NASBA amplification step no primers are being used that target the conserved sequences naturally present upstream and downstream of V3-V4 region of the 16S ribosomal RNA gene, and that the A-PCR step that also uses primers targeting the V3-V4 region is omitted from the protocol. The detailed protocol steps are as following:

[0084] Purification of EVs: 1. Thaw 1 mL of blood frozen plasma samples, collected in EDTA or Streck tubes. 2. Add 100 µl Proteinase K (QIAGEN # 19131, 20 mg / mL) for protein degradation including RNase (Bender et al, 2020). Incubate at 50ºC for 1 hour. 3. Centrifuge twice at 3,000 rpm for 15 min at 4 °C to remove cells and debris.Attorney Docket No.: 673 / 16 PCT4. Mix 0.95 mL supernatant with ice-cold 0.95 ml 1x PBS, pH 7.4, nuclease-free, in a 2-mLtube.5. Centrifuge once at 13,000 rpm for 1 min at 4 °C. This centrifugation step leaves the EVs inthe supernatant.6. Filter the supernatant through a 0.45 µm filter to remove any remaining cells and debris frombacteria and large organelles while the EVs are smaller in size and stay in the filtrate. Work with samples on ice and collect flowthrough into a 2-mL tube. Use a 1-2 mL syringe. Once all the sample was pushed through the filter, load the syringe with 0.7 mL 1x PBS, pH 7.4 and push it through the filter to eject the remaining sample from the filter (dead volume).

[0085] Nucleic acid isolation: Total nucleic acid extraction, including DNA, RNA and small RNA, is performed using the QIAamp ccfDNA / RNA Kit (QIAGEN # 55184). Extraction is performed according to the manufacturer’s instruction, except for a boiling step to lyse the EVs and no RNAse A treatment. 7. Boil EV-filtrate (~1.8 mL) and incubate for 10 min at 50ºC8. Boil EV-filtrate for 30 min at 100°C9. Centrifugation for 30 min at 12,000 g and 4 °C to eliminate the remaining IVs and debris.10. Transfer the supernatant to a new 5-mL tube.11. Add 600 µl of buffer RPL, close the tube cap, and vortex for 5 sec. Incubate at 15-25ºC for 3min. 12. Add 200 µl of buffer RPP and immediately mix by vortexing for 30 sec. Incubate on ice for6 min. 13. Centrifuge at 4500 g for 6 min in the Allegra X-12R centrifuge; the supernatant should beclear and colorless. 14. Transfer the supernatant to a new 5-mL tube. Keep on ice.15. Add 1 volume of ice-cold isopropanol and keep it on ice until all samples have beenprocessed. No specific incubation time is required. 16. Transfer all the solution (may require several rounds), including any precipitate, onto aRneasy Midi spin column in a 15 ml collection tube and centrifuge at 4500 g for 1 min at room temperature in the Allegra X-12R centrifuge. Discard the flow-through.Attorney Docket No.: 673 / 16 PCT 17. Add 4 ml of buffer RWT to the Rneasy Midi spin column and centrifuge at 4500 g for 1 min in the Allegra X-12R centrifuge. Discard the flow-through and keep the tube. 18. Add 2.5 ml of buffer RPE to the Rneasy Midi spin column and centrifuge at 4500 g for 5 min in the Allegra X-12R centrifuge. Discard the flow-through and keep the tube. 19. Place the Rneasy Midi spin column into a new 15 ml collection tube (supplied). Add 200 μl nuclease-free water directly to the center of the spin column membrane. Close the lid gently, incubate for 1 min, then centrifuge for 1 min at full speed in the Allegra X-12R centrifuge to elute the cfDNA. 20. Add 200 μl Buffer RPL and 800 μl ethanol (96–100%), and mix by vortexing. 21. Transfer 700 μl, including any precipitate, onto a Rneasy MinElute spin column in a 2 ml collection tube. Close the lid gently and centrifuge for 15 sec at ≥8000 g and room temperature in the Microfuge 22R centrifuge. Discard the flow-through and keep the tube. 22. Repeat step 15 using the remainder of the sample and make sure all liquid has passed through the column membrane completely, before proceeding to the next step. 23. Perform DNAse treatment: a. Combined 1 µl DNase I (2 U / µl, NEB #M0570) to 100 µl of 1X Buffer. Mix by gently inverting the tube. Centrifuge briefly to collect residual liquid from the sides of the tube. Note: DNase I is especially sensitive to physical denaturation. Mixing should only be carried out by gently inverting the tube. Do not vortex. b. Add the DNase I incubation mix (100 µl) directly to the RNeasy spin column membrane, and incubate at 37°C for 15 min. 24. Add 500 μl of buffer RPE onto the Rneasy MinElute spin column and centrifuge for 15 sec at ≥8000 x g (≥10,000 rpm) in the Microfuge 22R centrifuge. Discard the flow-through and keep the tube. 25. Add 500 μl of 80% ethanol to the Rneasy MinElute spin column. Close the lid and centrifuge for 15 sec at ≥8000 x g in the Microfuge 22R centrifuge. Discard the flow-through and the collection tube. 26. Place the Rneasy MinElute spin column in a new 2 ml collection tube. Open the lid of the spin column and place the spin columns into the centrifuge with at least one empty position between columns. Centrifuge at full speed for 5 min in the Microfuge 22R centrifuge to dry the membrane. Discard the flow-through and the collection tube.Attorney Docket No.: 673 / 16 PCT 27. Place the Rneasy MinElute spin column in a new 1.5 ml collection tube. Add 16 μl Rnase- free water directly to the center of the spin-column membrane. Close the lid gently, incubate for 1 min and centrifuge for 1 min at full speed in the Microfuge 22R centrifuge to elute the DNA / RNA.15 µl of EV-extract should be recovered.

[0086] Reverse Transcription (RT1) into ss-cDNA: RNA transcripts of the 16S V3V4 region were reverse transcribed into cDNA using the SuperScript™ IV First-Strand Synthesis with the T7plus2-Rd2-UMI-16SV4-R781. During this process, which creates DNA / RNA hybrids, UMIs and T7plus2 promoter are introduced in each newly synthesized ss-cDNA molecule. The reaction is followed by RNAseH and second-strand synthesis of the cDNA using Rd1-16SV3-F350 primer to create ds-cDNA. Details of this RT step are shown in Figure 7. 28. Combine 15 µl RNA sample with a. 1 µl dNTP (10 mM) b. 1 µl T7plus2-Rd2-UMI-16SV4-R781 primer (2 µM) 29. Mix at 65ºC for 5 minutes, and then on ice for at least 1 minute. 30. Prepare the following RT mix and add it to the reaction. a. 5 µl 5X Buffer b. 1 µl DTT (100mM) c. 1 µl Ribonuclease Inhibitor d. 1 µl SuperScriptTM IV Reverse Transcriptase (200 U / μL) 31. Incubate at 50°C for 10 minutes. This will create ss-cDNA. 32. Add 1µl RNAseH and incubate at 37ºC for 10 min. This will remove RNA from ss-cDNA. 33. Add 1 µl Rd1-16SV3-F350 primer (2 µM) and incubate at 50ºC for 20 min. This step is important to produce ds-cDNA molecules and allow the usage of Exo-I for primer removal. 34. Incubate at 80°C for 10 minutes to stop the reaction 35. Add 1 µl of Thermolabile Exonuclease I (NEB # M0568S). Incubate at 37°C for 5 minutes followed by 80ºC for 10 minutes. Exo-I degrades ssDNA in the 3’ to 5’ direction, digesting non incorporated primers and trimming the single stranded 3’ end of the first cDNA strand. 36. AMPure bead cleanup – Elute in 20µl. This will further remove oligos and allow buffer change.Attorney Docket No.: 673 / 16 PCT 37. Incubate for 30 minutes at 75°C to dry sample and resuspend in 4 µl ultra-pure water.

[0087] NASBA amplification: Targeted reverse transcription (first and second strand) of the 16S rRNA V3V4 region followed by IVT is performed using the NWK-1 NASBA kit (Life Sciences Advanced Technologies Inc., https: / / lifesci.com / product / nwk-1 / ) containing Reverse transcriptase, RNaseH and T7 RNA polymerase enzyme mix. The sequences of the primers used in the protocol are shown in Table 1. 38. Incubate 4 µl of EV-extract for 10 minutes at 65 °C and subsequently cooled to 41°C for 10 min in a thermocycler. This step is important to dissociate complementary RNA regions in the ribosomal RNA. Note, the primer is only added after denaturation to prevent cross hybridization to eventual DNA contaminants. 39. Add 1 µl of 5 µM primers mix [T7plus2-5’-Rd2 primer and 5’-Rd1 primer] 40. Add 10 µl of [6.7 µl NECB-24 + 3.3 µl NECN-24] pre-heated at 41ºC for 5 min 41. Add 5 µl NEC 1-24 enzyme mixture. 42. Incubate at 41 °C for 90 minutes. 43. Incubate at 95 °C for 10 minutes to stop the reactions. 44. Clean up using the RNA Clean XP kit (Beckman Coulter) and elute in 10 µl RNAse-free water.

[0088] QC: 45. 1 µl RNA sample is quantified by using the Qubit RNA HS Assay 46.1 µl RNA sample is analyzed on TapeStation or Bioanalyzer

[0089] Reverse Transcription (RT2) into ss-cDNA: RNA transcripts of the 16S V3V4 region were reverse transcribed into cDNA using the SuperScript™ IV First-Strand Synthesis with the 5’-Rd1 primer (table 1). Since reverse transcription is also part of the NASBA protocol, this step can be optionally performed to increase the levels of cDNA for the indexing PCR amplification step. No RNaseH treatment is performed as part of the RT2 step. 47. Combine 8 µl RNA sample (max 200 ng) with 3 µl nuclease-free water, 1 µl dNTP (10 mM), 1 µl 5’-Rd1 primer (2 µM). 48. Mix at 65ºC for 5 minutes, and then on ice for at least 1 minute.Attorney Docket No.: 673 / 16 PCT 49. Prepare the following RT mix and add it to the reaction. a. 4 µl 5X Buffer b. 1 µl DTT (100mM) c. 1 µl Ribonuclease Inhibitor d. 1 µl SuperScriptTM IV Reverse Transcriptase (200 U / μL) 50. Incubate at 50°C for 20 minutes. 51. Add 1 µl of Thermolabile Exonuclease I (NEB # M0568S). Incubate at 37°C for 5 minutes. Exo-I degrades ssDNA oligo. 52. Incubate at 80°C for 10 minutes to stop the reaction.

[0090] Indexing PCR (I-PCR): This step allows for subsequent sequencing and multiplexing of the RIP sequencing libraries. 51. Prepare the following reaction on ice: • 25 µl 2x KAPA HiFi HotStart ReadyMix • 10 µl UDI primers, using a unique set of Nextera UDI primers for each sample • 15 µl of RT2 product 52. Place the reactions in the thermocycler and run the following program: heated lid on, 95ºC for 10 min, [95ºC for 30 sec, 55ºC for 30 sec, 72ºC for 30 sec] for 7 cycles, 72ºC for 5 min, 4ºC on hold. 53. Clean up and size selection with 0.7x AMPure beads to remove amplicons below 200 bp, elute in 25 µl. 54. Prepare the following reaction on ice: • 30 µl 2x KAPA HiFi HotStart ReadyMix • 5 µl of P5 and P7 primers (with additional water) • 25 µl of I-PCR products 55. Place the reactions in the thermocycler and run the following program: heated lid on, 95ºC for 10 min, [95ºC for 30 sec, 55ºC for 30 sec, 72ºC for 30 sec] for 10 cycles, 72ºC for 5 min, 4ºC on hold.Attorney Docket No.: 673 / 16 PCT 56. Clean up and Size selection with 0.7x AMPure beads, eluted in 25 µl. 57. After quantification, the libraries are ready to be sequenced, e.g. on an iSeq100 sequencer according to the manufacturer’s instructions.

[0091] Processing and analysis of the V3 – V416S rRNA gene fragment sequences obtained with the RIP2 protocol was performed using the following steps: 1. Processing of forward and reverse reads is done separately when 16S amplicons are too long for paired-end merging. 2. Assessment of read quality is done using FastQC (Andrews, 2010). 3. Extraction of the UMI with UMI-tools (Smith et al, 2017). 4. Cutadapt is used for trimming adapter read-through sequences (Martin, 2011). 5. Removal of reads that map to the human genome is performed using Bowtie2 (Langmead et al, 2012). 6. UMI-informed deduplication of reads. 7. Quality-based read filtering, error-correction, and Amplicon Sequence Variant (ASV) creation is done using Dada2 (Callahan et al, 2016). 8. The rdp classifier is used for ASV classification (Wang et al, 2007) using a 50% confidence level. 9. Community composition is calculated based on the percentage of reads assigned to a specific phylogenetic level, such as family, genus, or species.

[0092] Three plasma samples were processed using the RIP2 protocol. As a negative control sample, nucleotide free water was used; this negative control sample was treated the same way as the plasma samples. Subsequent comparison of the microbial signals showed the following results. The microbial signature from the negative control sample was comprised of bacteria belonging to the families Actinomycetaceae, Bacillales_Incertae Sedis XI, Corynebacteriaceae, Enterobacteriaceae, Lactobacillaceae, Lawsonellaceae, Moraxellaceae, Peptoniphilaceae, Propionibacteriaceae, Pseudomonadaceae, Pseudonocardiaceae, Staphylococcaceae and Streptococcaceae. These bacteria represent many of the contaminating bacterial families identified in EXAMPLE 1 using 16S rRNA gene sequencing on the negative control samples. These bacterial families were also present in the microbial signatures from the plasma samples, indicating thatAttorney Docket No.: 673 / 16 PCT despite the DNase treatment there is still contaminating DNA. It was also noticed that if nucleotide free water was included as a negative control sample starting with the RT1 step of the RIP2 protocol, no microbial signature was observed. This indicates that the major source of contaminating DNA is the step involving the isolation of the EVs and the subsequent isolation of nucleotides.

[0093] Importantly, several bacterial families including Brucellaceae, Clostridiaceae 1, Comamonadaceae, Hymenobacteraceae, Intrasporangiaceae, Lachnospiraceae, Methylobacteriaceae, Oxalobacteraceae, Peptoniphilaceae, and Prevotellaceae were uniquely found in the microbial signatures from the plasma samples, indicating that the RIP2 protocol can be successfully used to identify the presence of unique microbial signatures in blood plasma. The unique presence of bacteria belonging to the families Lachnospiraceae (genera Blautia, Agathobacter, Catonella) and Prevotellaceae (genus Prevotella) is interesting as they represent anaerobic bacteria commonly found as part of the gut microbiome.

[0094] Although the RIP2 protocol was initially designed for the amplification of ribosomal rRNA to determine microbial signatures in liquid biopsy samples, the method can also be used as an alternative to DNA amplicon sequencing, such as amplification of variable regions of the 16S rRNA gene like as the V1 -V2 or the V3 – V4 regions, to determine the composition of microbial communities. The advantage of the RIP2 protocol is that it allows for introduction of UMI sequences, something that is not possible with standard 16S rRNA gene amplicon sequencing. UMI sequences allow for correction of PCR biases caused by preferential amplification, resulting in a more accurate determination of the microbial community composition.

[0095] In case the RIP2 protocol is being used for the amplification of DNA targets, the nucleotide isolation protocol can include an RNase treatment step. The use of the RIP2 protocol would be beneficial to analyze the microbial community composition of samples with relatively low amounts of biomass, with applications including the detection of pathogens during infection. EXAMPLE 4 Combining the Single Point Amplicon (SPA) fragment sequencing protocol and the Ribosome Informed Phylogeny (RIP) protocol for the analysis of microbial signatures in liquid biopsy samples.

[0096] Although the Single Point Amplicon (SPA) fragment sequencing protocol and the Ribosome Informed Phylogeny (RIP) protocol both target microbial signatures in biopsy samples, there are some important differences. The SPA fragment sequencing protocol targets mcfDNA, which is derived from bacteria that perished and whose DNA was released in the environment. ThisAttorney Docket No.: 673 / 16 PCT represents all bacteria that reside in the body or that accidentally enter the body, after which they are killed by the immune system, resulting in the release of their DNA. On the other hand, the RIP protocols specifically targets bacteria that reside in the body and that release BMVs that can trigger disease responses in distantly located parts of the body. As such both methods provide different but likely synergistic information. Both methods can be performed in parallel, as shown by the workflow schematically presented in Figure 9.

[0097] The term “RIP protocol” as used herein generally refers to either version of the Ribosome Informed Phylogeny protocol as described in Example 2 and Example 3. Initially, a sample is processed following the steps described in Figure 3 for the RIP protocol. However, the nucleotide isolation step is performed without the DNase treatment. Once this step is completed, the nucleotides are recovered in 24 μl water. The sample is subsequently divided to provide material for the RIP protocol and the SPA protocol. For the RIP protocol, the sample is treated with DNase followed by a heat inactivation step, after which the sample continues to be processed as described in EXAMPLE 2 or EXAMPLE 3 for the RIP2 protocol. For the SPA protocol, the sample is processed as described in PCT / US2023 / 011406. After the biotin purification step, the sample enters the same workflow as for the RIP protocol.

[0098] In addition to being combined with the SPA protocol, the RIP protocols can be combined with methods to obtain additional information for disease diagnostics and monitoring. For instance, after the RNA isolation step, capture probes that target ribosomes can be used to recover the bacterial and human ribosomal rRNA, resulting in a solution enrichment with non- ribosomal RNA such as mRNA fragments. The different RNA fractions can subsequently be separately processed to provide complementary information informing disease prognosis and monitoring.

[0099] (a) The captured bacterial rRNA fragments can be sequenced to determine the microbial signature in the biopsy sample. This information can be used to detect differences in microbial signatures between healthy individuals and patients and the discovery of disease specific bacterial species, information that can be used for disease diagnostic and monitoring purposes, including minimal residual disease.

[0100] (b) The human mRNA fragments can be enriched using poly-T capture probes that target the poly-A tail of human mRNA fragments. This fraction can be processed (e.g. using a NASBA step or via a Reverse Transcriptase reaction to produce cDNA) and sequenced to determine the presence of disease specific mRNAs in the human EVs, including mRNAs that have mutations or underwent alternative splicing, both of which could point to specific diseases including inflammation or cancer. The information can be used to detect differences in geneAttorney Docket No.: 673 / 16 PCT expression levels for disease specific target genes between healthy individuals and patients, information that can be used for disease diagnostic and monitoring purposes, including minimal residual disease, especially when combined with disease-specific microbial signatures provided by the RIP protocol.

[0101] (c) The fraction with the remaining RNA (after depletion of rRNA and human mRNA) can be converted into cDNA using random primers plus reverse transcriptase and analyzed, e.g. via sequencing, and can provide valuable information on the genes that were actively transcribed in the microbes from which the EVs originated. Furthermore, mapping both the V3-V4 16S rRNA and the mRNA informed gene sequences to reference genomes will facilitate a more precise identification of the bacteria at the origin of the BMVs in a biopsy sample.

[0102] The phylogenic identity of the microbes can be linked to probabilistic models describing their functioning, including their therapeutic functions such as the synthesis of beneficial secondary metabolites and functions involved in virulence, including functions that elucidate a response by the host’s immune system. This response is of particular interest when considering the effects of bacteria on disease responses in distantly located parts of the body, as secondary metabolites, virulence factors and other functions that elucidate a response by the host’s immune system will also be present in their BMVs. This is in addition to the immunogenic properties of the BMVs’ elicited by their associated lipopolysaccharides (LPS). As such the present invention allows using microbial signatures informed by the ribosomes present in BMVs for making microbe- specific predictions on disease responses elicited by their BMVs and cargo.

[0103] Furthermore, by combining microbial rRNA profiles to identify the species that contribute to the BMVs in a biopsy sample (“who is there?”), by linking this to gene transcription patterns in these organisms based on microbial mRNA profiles in the BMVs (“which genes are active?”), and combining this with probabilistic models of these organisms that predict their metabolic activity (“which secondary metabolites are being produced, or which metabolites are being converted?”), a unique understanding of these microbes in disease development and progression is obtained. This understanding supported by predictive modeling is critical for accurate disease prognostics and disease management. To illustrate this, in the case of colorectal cancer (CRC), published results indicate that increased numbers of Escherichia coli have been observed as part of the microbial signature in blood plasma of some CRC patients. Using predictive modeling on E. coli, a cytidine deaminase gene was identified capable of converting the chemotherapeutic drug gemcitabine into its inactive form 2’,2’-difluorodeoxyuridine, thereby contributing to tumor resistance to gemcitabine, the standard of care for treating non-small cell lungAttorney Docket No.: 673 / 16 PCT cancer, colorectal cancer, breast cancer and pancreatic ductal adenocarcinomas. These findings are in line with previously published work.

[0104] Figure 9 is a schematic representation of the workflow that integrates the two alternative versions of the RIP protocol described in EXAMPLE 2 and EXAMPLE 3 and the SPA fragment sequencing protocol for obtaining microbial signatures from liquid biopsy samples.Attorney Docket No.: 673 / 16 PCT REFERENCES Andrews, S. (2010). FastQC: A Quality Control Tool for High Throughput Sequence Data [Online]. Available online at: http: / / www.bioinformatics.babraham.ac.uk / projects / fastqc / Bender A. T., et al (2020). Enzymatic and Chemical-Based Methods to Inactivate Endogenous Blood Ribonucleases for Nucleic Acid Diagnostics. J Mol. Diagnostics 22: 1030-1040. https: / / doi.org / 10.1016 / j.jmoldx.2020.04.211. Callahan, B. J., et al. (2016). DADA2: High-resolution sample inference from Illumina amplicon data. Nature methods 13: 581-583. Conrad, T., et al.2020. Maximizing transcription of nucleic acids with efficient T7 promoters. COMMUNICATIONS BIOLOGY 3: 439. https: / / doi.org / 10.1038 / s42003-020-01167-x Langmead B, Salzberg S. (2012). Fast gapped-read alignment with Bowtie 2. Nature Methods.9: 357-359. Lee H., et al. (2020).16S rDNA microbiome composition pattern analysis as a diagnostic biomarker for biliary tract cancer. World Journal of Surgical Oncology 18: 19. https: / / doi.org / 10.1186 / s12957-020-1793-3. Martin M. (2011) Cutadapt removes adapter sequences from high-throughput sequencing reads. https: / / journal.embnet.org / index.php / embnetjournal / article / viewFile / 200 / 458 Smith T., Heger A., Sudbery I. (2017). UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy. Genome Research 27: 491-499. Wang Q., Garrity G. M., Tiedje J. M., Cole J. R. (2007). Naive Bayesian classifier for rapid assignment of rRNA sequences into the new bacterial taxonomy. Appl Environ Microbiol.73: 5261-5267.

Claims

Attorney Docket No.: 673 / 16 PCT CLAIMED:

1. A method of amplifying nucleic acid in a sample, comprising: isolating nucleic acids from a plurality of extracellular vesicles (EVs) from a sample; treating the isolated EV nucleic acids with DNase to essentially remove DNA; and amplifying a region of a microbial ribosomal RNA (rRNA) in the DNase-treated EV nucleic acids, thereby obtaining amplified microbial rRNA.

2. The method of claim 1, further comprising converting the amplified microbial rRNA to cDNA.

3. The method of claim 2, further comprising sequencing the cDNA.

4. The method of claim 3, further comprising using a computer to search the sequenced cDNA against a database of genes corresponding to the microbial rRNA and assigning a microbial species based on the closest sequence match.

5. The method of claim 4, wherein one or a combination of: (i) a presence, (ii) an absence, or (iii) a relative abundance of one or a combination of the microbial species is associated with a disease or disorder, and wherein the one or the combination of the presence, the absence, or the relative abundance of the one or a combination of the microbial species in the sample indicates (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

6. A method of diagnosing or monitoring a disease or disorder, comprising: isolating nucleic acids from a plurality of extracellular vesicles (EVs) from a sample; treating the isolated EV nucleic acids with DNase to essentially remove DNA; amplifying a region of a microbial ribosomal RNA (rRNA) in the DNase-treated EV nucleic acids, thereby obtaining amplified microbial rRNA; converting the amplified microbial rRNA to cDNA; sequencing or obtaining sequences of the cDNA; andAttorney Docket No.: 673 / 16 PCT assigning one of more of the cDNA sequences as belonging to a particular microbial species, or as comprising a microbial DNA signature, based on the closest sequence match to a database of genes corresponding to the microbial rRNA, wherein one or a combination of: the microbial DNA signature, or (i) a presence, (ii) an absence, or (iii) a relative abundance of the one or a combination of the microbial species is associated with a disease or disorder, and wherein the one or a combination of the microbial DNA signature, the presence, the absence, or the relative abundance of the one or a combination of microbial species in the sample indicates (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

7. A method of diagnosing or monitoring a disease or disorder, comprising: providing cDNA sequences derived from a sample from a subject, wherein the cDNA sequences comprise sequence corresponding to a region of a microbial ribosomal RNA (rRNA), wherein the region of a microbial rRNA was amplified from nucleic acid present in a plurality of extracellular vesicles (EVs) from the sample after the nucleic acid was treated with DNase to essentially remove DNA; and assigning one or more of the cDNA sequences as comprising a microbial DNA signature or as belonging to a particular microbial species based on the closest sequence match to a database of genes corresponding to the microbial rRNA, wherein one or a combination of: the microbial DNA signature, or (i) a presence, (ii) an absence, or (iii) a relative abundance of the one or a combination of the microbial species is associated with a disease or disorder, and wherein the one or a combination of the microbial DNA signature, the presence, the absence, or the relative abundance of the one or a combination of microbial species in the sample indicates (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

8. A method of diagnosing or monitoring a disease or disorder, comprising: sequencing cDNA derived from a sample from a subject, wherein the cDNA comprises sequences corresponding to a region of a microbial ribosomal RNA (rRNA) that wasAttorney Docket No.: 673 / 16 PCT amplified from nucleic acid present in a plurality of extracellular vesicles (EVs) from the sample after the nucleic acid was treated with DNase to essentially remove DNA; and assigning one or more of the cDNA sequences as comprising a microbial DNA signature or as belonging to a particular microbial species based on the closest sequence match to a database of genes corresponding to the microbial rRNA, wherein one or a combination of: the microbial DNA signature, or (i) a presence, (ii) an absence, or (iii) a relative abundance of the one or a combination of the microbial species is associated with a disease or disorder, and wherein the one or a combination of the microbial DNA signature, the presence, the absence, or the relative abundance of the one or a combination of microbial species in the sample indicates (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

9. A method of diagnosing or monitoring a disease or disorder, comprising: isolating nucleic acids from a plurality of extracellular vesicles (EVs) from a sample; treating the isolated EV nucleic acids with DNase to essentially remove DNA; amplifying a region of a microbial ribosomal RNA (rRNA) in the DNase-treated EV nucleic acids, thereby obtaining amplified microbial rRNA; converting the amplified microbial rRNA to cDNA; and sequencing or obtaining sequences of the cDNA, wherein, the closer the sequence match between the cDNA sequences and nucleic acid sequences in a reference sample for a disease or disorder, the more likely the sample is to have the disease or disorder.

10. The method of any of the foregoing claims, wherein the microbial rRNA is a bacterial rRNA or a eukaryotic rRNA.

11. The method of claim 10, wherein the bacterial rRNA comprises 16S rRNA, 23S rRNA, or 5S rRNA.

12. The method of claim 11, wherein the region of the 16S rRNA gene comprises the V3-V4 region.Attorney Docket No.: 673 / 16 PCT 13. The method of claim 10, wherein the eukaryotic rRNA comprises 5S rRNA, 18S rRNA, 28S rRNA, or 5.8S rRNA.

14. The method of any of the foregoing claims, wherein the sample comprises liquid biopsy, blood, spinal fluid, cerebral fluid, urine, saliva, lymph fluid, sputum, or stool, and combinations thereof.

15. The method of any of the foregoing claims, wherein the plurality of EV’s include human EV’s and bacterial extracellular vesicles (BMVs) and the human EV’s are selectively degraded prior to isolating nucleic acids from the plurality of EV’s.

16. The method of any one of the foregoing claims, wherein the amplifying the region of a microbial rRNA comprises: a. a first reverse transcription (RT) step to translate the region of a microbial rRNA into cDNA, wherein a 5’ end of the cDNA is extended by incorporation of a T7 promoter and a first sequencing adaptor and a 3’ end of the cDNA is extended by incorporation of a second sequencing adaptor in the first RT step; and b. wherein the first RT step is followed by isothermal amplification of the region of a microbial rRNA using primers that target the extended ends of the cDNA, thereby obtaining the amplified EV rRNA.

17. The method of claim 16, wherein a unique molecular identifier sequence (UMI) is incorporated into the cDNA in the first RT step.

18. The method of any one of the foregoing claims, wherein the amplifying the region of a microbial rRNA comprises nucleic acid sequence-based amplification (NASBA).

19. The method of any of the foregoing claims, further comprising amplifying a region of one or more RNA transcripts in the DNase-treated EV nucleic acids, thereby obtaining one or more amplified transcript RNAs.

20. The method of claim 19, wherein one or a combination of: (i) a presence, (ii) an absence, or (iii) a relative abundance of the one or more amplified transcript RNAs is associated with a disease or disorder, and wherein the one or a combination of the presence, the absence, orAttorney Docket No.: 673 / 16 PCT the relative abundance of the one or more amplified transcript RNAs in the sample indicates (i) that the subject has the disease or disorder or the propensity for the disease or disorder, (ii) the stage or severity of the disease or disorder, or (iii) a type of treatment for the disease or disorder, and combinations thereof.

21. The method of any one of the foregoing claims, where the disease or disorder comprises formation of stomach ulcers, formation of intestinal polyps / adenomas and their progression into malignancies, gastrointestinal diseases, Irritable Bowel Disease (IBD), tumors, colorectal cancer, pancreatic cancer, lung cancer, cervical cancer, prostate cancer, breast cancer, Central Nervous System (CNS) diseases, multiple sclerosis (MS), Parkinson’s disease, Alzheimer’s disease, minimal residual disease (MRD) monitoring, monitoring of eye diseases, uveitis, cystic fibrosis, tuberculosis, and general risk monitoring of infections in patient populations with a compromised immune system, and combinations thereof.

22. A system for amplifying nucleic acid in a sample, comprising: a. a reaction vessel; b. a reagent dispensing module; and c. software to execute the method of any of the foregoing claims, wherein the method is executed at least partially robotically.

Citation Information

Patent Citations

  • Sequencing and analysis of exosome associated nucleic acids

    US20200208213A1

  • Single-LOCI and multi-LOCI targeted single point amplicon fragment sequencing

    WO2023141347A2