Methods for sequencing samples from a same collection site
By sequencing samples from a series of blood draws at the same collection site and analyzing changes in abundance, the method addresses contamination issues in metagenomic sequencing, enhancing accuracy and precision in identifying target nucleic acids.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KARIUS INC
- Filing Date
- 2025-11-25
- Publication Date
- 2026-06-04
AI Technical Summary
Metagenomic sequencing is highly sensitive to contamination from environmental and laboratory sources, making it difficult to distinguish between target nucleic acids from microbes of interest and contaminant nucleic acids, particularly in non-sterile sample collection environments.
Perform a sequencing assay on a sample from a later blood draw in a series of blood draws taken from the same collection site, comparing sequencing reads across draws to identify and discard contaminant nucleic acids by analyzing changes in abundance and mapping to microbial genomes, thereby enhancing the signal-to-background ratio.
This method improves the accuracy of metagenomic sequencing by distinguishing target nucleic acids from contaminants, increasing the signal-to-background ratio and enabling more precise identification of microbial species, particularly in non-sterile sample collection settings.
Smart Images

Figure US2025057148_04062026_PF_FP_ABST
Abstract
Description
WSGR Docket No.: 47697-751601METHODS FOR SEQUENCING SAMPLES FROM A SAME COLLECTION SITECROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 726,064, filed November 27, 2024, which application is incorporated herein by reference in its entirety.BACKGROUND
[0002] Metagenomic sequencing can detect microbial cell-free nucleic acids (mcfNA) shed into the extracellular milieu of a host (e.g., bloodstream or cerebrospinal fluids) and represents a new frontier in infectious disease diagnostics. While cell-free metagenomic sequencing is a powerful tool for characterizing microbes across all kingdoms, it is highly sensitive to contamination, particularly from environmental and laboratory sources.SUMMARY
[0003] There remains a need for developing improved methods of sequencing and deconvoluting target nucleic acid sequences from environmental contaminants from the sample collection site.
[0004] This Summary introduces a selection of concepts that are described further below in the Detailed Description. This Summary is not intended to limit the scope of the claimed subject matter.
[0005] Described herein in some embodiments is a method comprising performing a sequencing assay on a sample from a subject to generate sequencing reads, wherein: a) the sample from the subject: comprises a nucleic acid, and is from a later blood draw in a series of blood draws; and b) the series of blood draws is taken from a same sample collection site of the subject. In some embodiments, the subject is an animal. In some embodiments, the animal comprises a mammal. In some embodiments, the mammal comprises a canine, a swine, a feline, or a nonhuman primate. In some embodiments, the subject is a human. In some embodiments, the same sample collection site comprises a same area of skin on the subject, a same blood vessel of the subject, or an access to a same blood vessel of the subject. In some embodiments, the series of blood draws comprises at least two, at least three, or at least four blood draws. In some embodiments, the series of blood draws comprises two blood draws. In some embodiments, the series of blood draws comprises blood draws taken sequentially from the same collection site of the subject. In some embodiments, the series of blood draws comprises blood draws taken from the same sample collection site of the subject. In some embodiments, the later blood draw was taken at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at leastWSGR Docket No.: 47697-751601 about 45 seconds, or at least about 1 minute later than the earlier blood draw in the series of blood draws. In some embodiments, the later blood draw was taken at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at least about 45 seconds, at least about 1 minute, at least about 2 minutes, at least about 3 minutes, at least about 4 minutes, or at least about 5 minutes later than the first blood draw in the series of blood draws. In some embodiments, each blood draw in the series of blood draws was taken at intervals of at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at least about 45 seconds, or at least about 1 minute. In some embodiments, the series of blood draws was taken over the period of about 5 seconds, about 10 seconds, about 30 seconds, about 45 seconds, about 1 minute, about 1.5 minutes, about 2 minutes, about 2.5 minutes, about 2.5 minutes, about 3 minutes, about 3.5 minutes, about 4 minutes, about 4.5 minutes, or about 5 minutes. In some embodiments, the series of blood draws was taken using a vacuum blood collection system. In some embodiments, the series of blood draws was taken using one needle. In some embodiments, the series of blood draws was taken from a same vein of the subject. In some embodiments, the series of blood draws was taken from a single puncture to the vein of the subject. In some embodiments, the collection site of the subject comprises a vein or a venous access of the subject. In some embodiments, the venous access of the subject comprises skin above the vein or a catheter. In some embodiments, the vein comprises a jugular vein or an ear vein. In some embodiments, the sequencing reads comprise sequencing reads generated from microbial cell free nucleic acids (mcfNA) in the sample. In some embodiments, the sequencing reads comprise at most 50, at most 100, at most 150, at most 200, at most 250, at most 300, at most 350, at most 400, at most 450, or at most 500 sequencing reads generated from the mcfNA in the sample. In some embodiments, the sequencing reads comprise at least 500, at least 1500, at least 2500, or at least 3500 sequencing reads generated from the mcfNA in the sample. In some embodiments, the method further comprises applying a control nucleic acid to the same sample collection site of the subject prior to collecting the series of blood draw. In some embodiments, the sequencing reads generated from the nucleic acids in the sample comprise a sequencing read generated from the control nucleic acid applied to the collection site. In some embodiments, the method further comprises mapping a sequencing read to a reference sequence. In some embodiments, the reference sequence comprises a region of a microbial genome. In some embodiments, the method further comprises assigning a sequencing read to a microbial sequence. In some embodiments, the method further comprises adding a known amount of synthetic spikein molecules to the sample and generating sequencing reads from the synthetic spike-in molecules. In some embodiments, the method further comprises calculating an abundance of a sequencing read generated from the nucleic acids in the sample. In some embodiments, theWSGR Docket No.: 47697-751601 abundance of the sequencing reads comprises an absolute abundance. In some embodiments, the abundance of the sequencing reads comprises a relative abundance or a normalized abundance. In some embodiments, the method further comprises calculating the relative abundance of a sequencing read at least in part by comparing the abundance of the sequencing read to the abundance of another sequencing read. In some embodiments, the method further comprises calculating the normalized abundance of a sequencing read at least in part by comparing the abundance of the sequencing read to the abundance of sequencing reads generated from the synthetic spike-in molecules in the sample. In some embodiments, the method further comprises discarding a sequencing read generated from the later blood draw. In some embodiments, the discarding the sequencing read is performed based in part on the following criteria being met: a) the sequencing read assigned to a single species of microbe, and b) the abundance of sequencing reads generated from the later blood draw assigned to the single species of microbe is lower than the abundance of sequencing reads generated from an earlier blood draw assigned to the same species of microbe. In some embodiments, the abundance of sequencing reads generated from the later blood draw assigned to the single species of microbe is at least 1 molecule per milliliter (MPM), at least 2 MPM, at least 3 MPM, at least 4 MPM, at least 5 MPM, at least 10 MPM, at least 15 MPM, or at least 20 MPM lower than the abundance of sequencing reads generated from the earlier blood draw assigned to the same species of microbe . In some embodiments, the abundance of sequencing reads generated from the later blood draw assigned to the single species of microbe is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, or at least 60% lower than the abundance of sequencing reads generated from the earlier blood draw assigned to the same species of microbe. In some embodiments, the method further comprises comparing an abundance of sequencing reads generated from the later blood draw to an abundance of sequencing reads generated from the earlier blood draw. In some embodiments, the method further comprises calculating a variance of the abundance of the sequencing reads. In some embodiments, the method comprises comparing a variance of abundance of the sequencing reads from the later blood draw to a variance of the abundance of sequencing reads from the earlier blood draw. In some embodiments, the variance of sequencing reads from the later blood draw is lower than the variance of sequencing reads from the earlier blood draw. In some embodiments, the method further comprises identifying a contaminant nucleic acid sequence, at least in part based on a change in an abundance of the contaminant nucleic acid sequence from the earlier blood draw to the later blood draw. In some embodiments, the method further comprises identifying a contaminant nucleic acid sequence, at least in part based on a reduced abundance of the contaminant nucleic acid sequence from the earlier blood draw to the later blood draw. In some embodiments, the method further comprises identifying aWSGR Docket No.: 47697-751601 contaminant nucleic acid sequence by comparing an abundance of the contaminant nucleic acid to an abundance of the control nucleic acid applied to the collection site. In some embodiments, the method further comprises identifying a contaminant microbe based on the contaminant nucleic acid sequence, wherein the contaminant microbe comprises a commensal microbe from the subject or a microbe of the subject’s general environment. In some embodiments, the method further comprises identifying a contaminant nucleic acid sequence from a soil microbe, a gastrointestinal tract microbe, or a skin microbe. In some embodiments, the method further comprises identifying a contaminant nucleic acid sequence from one or more Staphylococcus spp., Escherichia coli, Salmonella spp., Pseudomonas spp., Clostridium spp., Candida spp., Aspergillus spp., Penicillium spp., Rodent Parvovirus, Mouse Hepatitis Virus (MHV), Sendai Virus, helminths (e.g., Syphacia obvelata), protozoa (e.g., Giardia spp.), various soil microorganisms, or any combination thereof. In some embodiments, the method increases the signal -to-background ratio of the sequencing assay. In some embodiments, the method further comprises generating a report of the species of microbes identified in the sample. In some embodiments, the microbe comprises a virus, a bacterium, a protozoa, or a fungus. In some embodiments, the nucleic acid comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the nucleic acid comprises a cell-free nucleic acid (cfNA). In some embodiments, the nucleic acid comprises a microbial cell-free nucleic acid (mcfNA). In some embodiments, the sample comprises mcfNA from at least one, at least two, or at least three blood draws in the series of blood draws. In some embodiments, the sample comprises mcfNA from the later blood draws in the series of blood draws. In some embodiments, the sample comprises mcfNA from at least one, at least two, or at least three microbial species. In some embodiments, the sample comprises whole blood or plasma. In some embodiments, the sequencing assay comprises a high throughput sequencing assay. In some embodiments, the sequencing assay comprises a next generation sequencing assay. In some embodiments, the sequencing assay comprises a sequencing-by-synthesis assay. In some embodiments, the method described herein further comprises preparing the collection site prior to taking the series of blood draws from the subject. In some embodiments, preparing the collection site comprises shaving, scrubbing or cleaning. In some embodiments, the method further comprises identifying a commensal microbe of the subject at least in part based on a result of the sequencing assay. In some embodiments, the method further comprises identifying a pathogen infecting the subject at least in part based on a result of the sequencing assay. In some embodiments, the pathogen infecting the subject comprises a virus, a bacterium, a protozoa, or a fungus. In some embodiments, the method further comprises administering a treatment to the subject, wherein theWSGR Docket No.: 47697-751601 subject is infected by the pathogen. In some embodiments, the treatment comprises an antimicrobial agent.
[0006] Described herein in some embodiments is a method of distinguishing between a contaminant nucleic acid from a target nucleic acid in a sample using the method provided herein. In some embodiments, the target nucleic acid comprises mcfDNA from a pathogen.
[0007] Described herein in some embodiments is a method of determining an eligibility of a subject for transplantation, comprising analyzing a sample from the subject using the method provided herein. In some embodiments, the subject is a xenotransplant animal donor. In some embodiments, the subject comprises a pig or a non-human primate.INCORPORATION BY REFERENCE
[0008] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0010] FIGs. 1A-1B show the bacteria families identified in human and porcine samples.
[0011] FIG. 2 shows the bacteria species identified in the pilot porcine samples.
[0012] FIG. 3 is a table summarizing the taxa observed in the porcine donors. The table also lists the donors’ diagnosis status of porcine circovirus 3 (PCV3) and S. suis infections.
[0013] FIGs. 4A-4E depict the normalized molecules per microliter (MPM) of microbial cell- free DNA (mcfDNA) for microbes of different kingdoms, e.g., bacteria, virus, and eukaryotes.
[0014] FIG. 5A depicts normalized MPM of mcfDNA for different kingdoms. FIG. 5B depicts MPM and normalized MPM of Staphlococcus epidermidis mcfDNA.
[0015] FIGs. 6A-6B depict the numbers of significant observations in all blood draws from the donors.
[0016] FIGs. 7A-7B show the MPM and normalized MPM of mcfDNA of porcine type-c oncovirus in all blood draws from the donors.
[0017] FIGs. 8A-8D depict the correlation heatmaps with different p-values. The in a cell indicates p-value <= 0.0003.WSGR Docket No.: 47697-751601
[0018] FIGs. 9A-9C depict the bioinformatic analyses performed to investigate the species of microbes that contributed to the decrease of the total mcfDNA.
[0019] FIG. 10 depicts the change in mcfDNA concentrations for Lactobacillus amylovorous and Corynebacterium xerosis in all donors.
[0020] FIG. 11 depicts the bioinformatic analyses performed to track the changes in mcfDNA concentration between blood draws.
[0021] FIG. 12A shows mcfDNA concentrations of various microbes identified in the blood samples obtained from Subject ID 22206. FIGs. 12B-12E show mcfDNA concentrations of various microbes identified in the 15852 EC samples.
[0022] FIG. 13 is a bar graph of the distribution of microbes called from different environments.
[0023] FIG. 14 is an embodiment of the system described herein.DETAILED DESCRIPTION
[0024] The following passages describe different aspects of the disclosure in greater detail. Each aspect, embodiment, or feature of the disclosure can be combined with any other aspect, embodiment, or feature of the disclosure unless clearly indicated to the contrary.
[0025] Metagenomic sequencing, while a powerful tool for characterizing microbial communities, is highly sensitive to contamination, particularly from environmental and laboratory sources. This sensitivity stems from the technology’s ability to detect single microbial nucleic acid molecules, including degraded molecules. Contaminant nucleic acids can be introduced during sample collection, processing, or sequencing, often masking true biological signals. To avoid such contamination, researchers can use nucleic acid-free reagents and equipment in designated nucleic acid-free areas. Furthermore, researchers can exercise caution when collecting and handling samples containing nucleic acids for metagenomic sequencing. This includes wearing protective gear and preparing the collection site on the subject by scrubbing or shaving. Unfortunately, this level of preparation may not be feasible under some circumstances. For example, blood draws from animals are mostly collected in non-sterile environments. Animal skin, hair, and housing surroundings may harbor nucleic acids from environmental microbes or microbes originating from the subject’s microbiome that could compromise the sequencing analysis. It may be difficult to fully prepare the sample collection site (e.g., the body of an animal or geographic sample collection site such a housing facility) before blood draws. For example, some animals may not tolerate rigorous cleaning protocols without anesthesia. In some cases, sample collection from human subjects may present similar challenges. Environmental contaminants can confound the analysis of cell-free nucleic acids in the blood, making it difficult to distinguish between the true signals of interest and backgroundWSGR Docket No.: 47697-751601 noise. Therefore, more robust methods for metagenomic sequencing are needed to minimize environmental contamination and accurately detect target nucleic acids from microbes of interest.
[0026] The methods described herein can be used for various applications related to sequencing nucleic acids in a sample. The nucleic acids in the sample may include target nucleic acids (e.g., cell-free nucleic acids from a subject) and / or contaminant nucleic acids from sample collection and / or sample processing. The methods described herein can be used to identify these contaminant nucleic acids. The methods can also be used to increase signal-to-background ratios in a sequencing assay, thereby achieving accurate analysis of the target nucleic acids.
[0027] In some cases, the method can distinguish the target nucleic acids from the contaminant nucleic acids by sequencing a sample from a later blood draw in a series of blood draws taken from a same collection site of the subject. In some cases, the methods comprise sequencing samples from both an earlier blood draw and a later blood draw in the series of blood draws. In some embodiments, the methods can further include generating, detecting, mapping, or quantifying sequencing reads from the nucleic acids in the sample, and performing various bioinformatic analyses on the sequencing reads. Some sequencing reads from microbial cell-free nucleic acids (mcfNA) may be detected in all the samples collected from the single collection site. These mcfNA fragments likely originated from the subject’s blood. Some sequencing reads may be detected only from nucleic acids in the earlier blood draw, without detection from the later blood draw. Some sequencing reads may be detected at a reduced concentration in the later blood draw. In this case, the nucleic acids can potentially be contaminants originating from the sample collection. The methods herein may distinguish sequencing reads from nucleic acids in the later blood draw from those in an earlier blood draw in the series. These distinctions may facilitate the identification of contaminant nucleic acids from the sample collection site. The contaminant nucleic acids can be discarded if desired. Moreover, the methods improve the accuracy of the sequencing assays by increasing the signal-to-background ratio and enhancing the identification of the target nucleic acids.
[0028] The present disclosure provides methods comprising performing a sequencing assay on a sample from a subject to generate sequencing reads. In some embodiments, the sample comprises nucleic acids and can be obtained from a later blood draw in a series of blood draws from the subject. In some embodiments, the nucleic acids comprise cell-free nucleic acids (cfNA). In some embodiments, all blood draws (e.g., later and earlier blood draws) in the series of blood draws are taken from the same collection site of the subject. In some embodiments, the method further comprises mapping the sequencing reads generated from a later blood draw to a genome of a microbe to identify the microbe. In some embodiments, the method further comprises quantifying an abundance of the sequencing reads of cfNA from a species of microbe. In some embodiments,WSGR Docket No.: 47697-751601 the method further comprises detecting decreasing concentrations of nucleic acids of certain species of microbes from an earlier blood draw to a later blood draw. In some embodiments, such species of microbes comprise contaminant species from the single or same sample collection site. In some embodiments, such species of microbes comprise contaminant species from the same body sample collection site. In some embodiments, such species of microbes comprise contaminant species from the same geographic sample collection location. In some embodiments, when the abundance of sequencing reads generated from a later blood draw mapping to a species of microbe is lower than that from an earlier blood draw in the series obtained from the same collection site, the species of microbe and its nucleic acids are deemed contaminants of the same collection site. In some embodiments, the methods comprise identifying nucleic acids as environmental contaminants from the same collection site of the subject, at least in part based on the abundances of the sequencing reads mapping to a contaminant microbe. In some embodiments, the method further comprises identifying sequencing reads mapping to a species of microbe, wherein the abundance of said sequencing reads decreases from the earlier blood draw to the later blood draw. In some embodiments, identifying nucleic acids as environmental contaminants does not require quantifying the abundances of the sequencing reads. In some embodiments, the method further comprises discarding a sequencing read mapping to a species of microbe generated from an earlier blood draw. In some embodiments, the method further comprises discarding a sequencing read mapping to a species of microbe generated from a later blood draw. In some embodiments, the method further comprises discarding sequencing reads from nucleic acids originated from contaminants at the sample collection site. In some embodiments, the method further comprises enhancing detection and identification of the target nucleic acids in the sample.Applications
[0029] In some embodiments, the methods described herein are for various purposes where identifying a microbe in a subject is necessary. In some embodiments, the methods are for detecting or identifying a pathogen to the subject. In some embodiments, the methods are for detecting, diagnosing, treating, monitoring, staging, or prognosing a disease or a disorder provided herein. In some embodiments, the disease or disorder comprises an infection or another medical indication related to an infection. In some embodiments, the methods are for determining an infection site or identifying a source of an infection. In some embodiments, the methods are for determining the biological relationship between a microbe and a host (e.g., the subject herein). In some embodiments, the methods are for detecting or identifying a commensal microbe of the subject. In some cases, a commensal microbe of a first subject may be a potential pathogen to a second subject. For example, a commensal microbe of an animal may be a potential pathogen toWSGR Docket No.: 47697-751601 a human. In some embodiments, methods disclosed herein can be used in conjunction with one or more medical tests.
[0030] In some embodiments, the methods can be used for determining an eligibility of a subject in transplantation. The eligibility of a subject in transplant can be determined in part based on whether the subject has an infection. In some embodiments, the subject can be eligible for transplantation when the subject does not have an infection. The subject described herein can be a donor or a recipient of the transplantation. In some embodiments, the transplantation comprises organ transplantation, tissue transplantation, composite tissue transplantation, living donor transplantation, or xenotransplantation. In some embodiments, the methods can be used for determining the eligibility of an animal donor for xenotransplant. In some embodiments, the animal subject is eligible for xenotransplantation when the animal does not have a zoonotic infection that can infect a human recipient. In some embodiments, the methods can be used for determining the eligibility of a human donor for organ transplant. In some embodiments, the methods can be used for determining the eligibility of a human transplant recipient.
[0031] In some embodiments, the methods can be used for detecting, diagnosing, treating, monitoring, staging, prognosing an infection, or any combination thereof. The infection may be caused by any microbe described herein, including but not limited to pathogenic or commensal microbes. In some embodiments, the methods can be used for detecting, diagnosing, treating, monitoring, staging, prognosing a medical indication related to an infection, or any combination thereof. In some embodiments, a medical indication related to an infection comprises any disease, disorder or procedure (e.g., medical or surgical) that renders a subject immunocompromised or susceptible to infections. In some embodiments, the medical indication related to an infection may cause a new or reoccurring infection in a subject. For example, the medical indication related to an infection can comprise any disease for which an immunosuppressant (e.g., chemotherapy, radiation, corticosteroids, transplant medications, or certain biologies) or an anti-infective agent is required as a treatment. In some embodiments, a medical indication related to an infection may comprise cancers, transplantation, surgeries, bums, infections, malnourishment, chronic kidney diseases, diabetes mellitus, autoimmune diseases, or immune disorders (e.g., acquired immunodeficiency syndrome (AIDS)).
[0032] In some embodiments, the methods described herein can be used for individualized treatment for an infected subject or a subject who is susceptible or at risk for infections (e.g., immunosuppressed, immunocompromised, living conditions, or genetic variations resulting in increased susceptibility for infection). In some embodiments, the subject described herein comprises a human or an animal. In some embodiments, individualized treatment can include predicting if an infection will progress to an invasive disease stage, monitoring the efficacy of aWSGR Docket No.: 47697-751601 therapy in a subject, modifying a therapeutic regimen depending on the subject's response to the therapy, and determining a pathogen's resistance to a particular therapeutic. In some embodiments, the methods disclosed herein can be used to detect, diagnose, predict, or prognose a pathogen's resistance to a particular therapeutic. In some cases, the methods can further comprise sequencing of the subject's DNA for genetic variations that are associated with therapeutic resistance to therapeutics or to a particular therapeutic. In some cases, samples can be collected serially at various times before or during the course of the infection to determine the pathogen's and subject's response to a treatment, thereby providing a regimen that is individually tailored. In some cases, the serially collected samples can be compared to each other to determine whether the infection is improving or worsening in the subj ect. In some embodiments, a treatment can involve administering a drug or other therapy to reduce or eliminate the colonization or invasive disease associated with an infection. In some cases, the subject can be treated prophylactically to prevent the development of an infection. Any medical procedure or treatment including administration of a drug can be used to improve or reduce the symptoms of an infection. Some nonlimiting exemplary drugs that can be used are antibiotics (such as ampicillin, sulbactam, penicillin, vancomycin, gentamycin, aminoglycosides, clindamycin, cephalosporin, metronidazole, timentin, ticarcillin, clavulanic acid, cefoxitin), antiretroviral drugs (e.g., highly active antiretroviral therapy (HAART), reverse transcriptase inhibitors, nucleoside / nucleotide reverse transcriptase inhibitors (NRTIs), Non-nucleoside RT inhibitors, and / or protease inhibitors), immunoglobulins, or any variant or combination thereof.
[0033] In some embodiments, the methods can be used for adjusting a therapeutic regimen. For example, the subject can be administered a drug to treat an infection. In some embodiments, methods can be used to track or monitor the efficacy of the drug treatment. In some cases, the therapeutic regimen can be adjusted, depending on upward or downward course of the infection. For example, if the methods provided herein indicate that an infection is not improving with drug treatment, the therapeutic regimen can be adjusted by changing the type of drug or treatment, discontinuing the use of the drug, continuing the use of the drug, increasing the dose of the drug, or adding a new drug or treatment to the subject's therapeutic regimen.
[0034] In some embodiments, the methods described herein can be used for detection, monitoring, diagnosis, prognosis, treatment, prediction, or prevention of colonization by the microbes described herein. The disclosure also provides methods to detect, monitor, diagnose, prognose, treat, predict, or prevent invasive disease caused by the microbes described herein. The methods of the disclosure can be applied to any pathogen that has various stages of infection. The methods can be especially useful for pathogens that have a colonization stage and an invasiveWSGR Docket No.: 47697-751601 disease stage. In some cases, the invasive disease stage can be caused by the pathogen infection. In some cases, the invasive disease stage can be associated with the pathogen infection.
[0035] In some embodiments, the methods can be used to distinguish populations of nucleic acids or for detecting a microbe in a subject. In some embodiments, the methods provide a more comprehensive view of the state and diversity of the infection or symbiotic microbes in a subject. For example, the identification of both RNA and DNA in a sample can be useful to detect RNA and DNA type viruses, or to detect bacterial, protist, parasitic or fungal genomic DNA and / or gene expression products, e.g., mRNA. Such process can also be able to differentiate between latent infection (e.g., which might be indicated by the presence of integrated retroviral DNA) versus active infection (e.g., which might be indicated by the presence of viral RNA from intact viral particles). Such processes can also be able to detect drug resistance and / or the origin of infection. Such processes can also be used to analyze host response. Such analyses can include analysis of cell-free, circulating nucleic acids, e.g., for microbial or viral infection identification. Blood Draws from a Single Collection Site
[0036] Provided herein are methods and systems for identifying target nucleic acids originated from a microbe (e.g., cell-free nucleic acids from a pathogen) in a subject (e.g., an animal), while reducing interference from contaminant nucleic acids, such as those from contaminating microbes present at the body sample collection site (e.g., skin of the subject). When a series of blood draws is collected from a subject, a single vacutainer system can be used. After puncturing the subject’s skin with a single needle, multiple vials of blood can be drawn from the same puncture. Sometimes the initial aliquots of blood can be discarded because they often contain skin flora flushed off the skin by the subject's blood. The methods and systems described herein retain all aliquots of blood collected from the same puncture and waste little to no blood from the subject, allowing for a more comprehensive analysis. The methods and systems comprise collecting a series of blood draws from a same sample collection site of the subject (e.g., body sample collection site) and analyze both the later aliquots of blood (e.g., later blood draws) and the earlier / initial aliquots of blood (e.g., earlier blood draws). By including the earlier / initial blood draws in the analysis, the methods and systems enable tracking nucleic acid contaminants across samples and more accurately distinguishing target nucleic acids from contaminant nucleic acids originated from microbes present at the same collection site. By comparing sequencing reads from nucleic acids across blood draws, contaminant nucleic acids can be identified and discarded if needed. Therefore, the methods and systems described increase the signal-to- background ratio of sequencing assays, thereby enhancing the accuracy of microbial cell-free nucleic acids detection.WSGR Docket No.: 47697-751601
[0037] The sample provided herein comprises nucleic acid molecules (e.g., cell-free nucleic acids) to be sequenced by the methods described herein. In some embodiments, the sample can be from a blood draw obtained from a subject. In some embodiments, the sample comprises a blood product. In some embodiments, the sample comprises whole blood or plasma of the subject.
[0038] In some embodiments, the series of blood draws can comprise at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, at least eighteen, at least nineteen, or at least twenty blood draws. In some embodiments, the series of blood draws comprises at most two, at most three, at most four, at most five, at most six, at most seven, at most eight, at most nine, at most ten, at most eleven, at most twelve, at most thirteen, at most fourteen, at most fifteen, at most sixteen, at most seventeen, at most eighteen, at most nineteen, or at most twenty blood draws. In some embodiments, the series of blood draws comprises two blood draws. In some embodiments, the series of blood draws comprises three blood draws. In some embodiments, the later blood draw can be the third blood draw in a series of three blood draws collected from a same sample collection site.
[0039] In some embodiments, the series of blood draws can be collected from a same sample collection site of a subject. In some embodiments, a later blood draw in a series of blood draws can be collected from a same sample collection site of a subject as a previous blood draw in a series of blood draws. In some embodiments, a choice of a sample collection site may be determined based on a few factors, such as an amount of blood required, a condition of a subject, and a type of downstream testing to be performed, or any combination thereof. In some embodiments, a sample collection site can comprise a body sample collection site, a geographic sample collection site, or a combination thereof. In some embodiments, a geographic sample collection site can comprise a location, a room, or a facility where the sample collection occurs. In some embodiments, a geographic sample collection site can comprise a clinic, a research lab, or a farm. In some embodiments, a body sample collection site can comprise a location of a subject’s body from which a sample is collected. In some embodiments, a body sample collection site can comprise an area of skin, a blood vessel (e.g., a vein or an artery), or an access to a blood vessel (e.g., vascular access). In some embodiments, a skin can comprise a skin covering a figure tip, a heel, or an ear lobe of a subject. In some embodiments, the sample collection site can be the same area of the subject’s skin. In some embodiments, the same area of the subject’s skin can be less than 2 cm2, less than 1.5 cm2, less than 1 cm2, or less than 0.5 cm2. In some embodiments, the series of blood samples can be collected from one or more puncture of the same area of the subject’s skin. In some embodiments, the series of blood samples can be collected from a singleWSGR Docket No.: 47697-751601 puncture of the same area of the subject’s skin. In some embodiments, the series of blood samples can be collected from one or more puncture of the same vein. In some embodiments, the vein can comprise a median cubital vein, a cephalic vein, a basilic vein, a jugular vein, an anterior vena cava, an ear vein, a tail vein, a femoral vein, or a combination thereof. In some embodiments, an artery can comprise a radial artery, a brachial artery, a femoral artery, a dorsalis pedis artery, or any combination thereof. In some embodiments, a vascular access can comprise an area of skin covering a blood vessel, a catheter, a port, a fistula, a graft, or any combination thereof.
[0040] In some embodiments, the series of blood draws can comprise capillary blood draws collected from a single prick of a same area of the subject’s skin. In some embodiments, the series of blood draws can comprise capillary blood draws collected from more than one pricks of the same area of the subject’s skin. In some embodiments, the series of blood draws can be collected from a cut or prick of an ear vein of the subject. In some embodiments, the series of blood draws can be collected by one or more venipuncture from a same vein. In some embodiments, the series of blood draws can be collected by a single needle puncture of the same vein. In some embodiments, the series of blood draws can be collected by a single puncture of a jugular vein of an animal subject. In some embodiments, the series of blood draws can be collected by a single puncture of an ear vein of an animal subject. In some embodiments, the series of blood draws can be collected by two or more needle punctures of the same vein through a same area of the skin (e.g., jugular or ear) covering the vein. In some embodiments, the series of blood draws can be collected by one or more needle punctures from a same artery. In some embodiments, the series of blood draws can be collected by a single puncture of the same artery. In some embodiments, the series of blood draws can be collected by two or more punctures of the same artery through a same area of skin covering the artery.
[0041] In some embodiments, the blood draws described herein may be collected using a needle. In some embodiments, the needle is a butterfly needle. In some embodiments, the needle may be comprised in a blood collection system. In some embodiments, the blood collection system can comprise a vacuum blood collection system (e.g., Vacutainer®), wherein the system can comprise a collection tube with a vacuum inside. In some embodiments, the blood collection system can further comprise a plastic sheath with a needle. In some embodiments, the series of blood draws can be obtained by puncturing a vein of the subject with one needle (e.g., a butterfly needle) and contacting one or more vacuum collection tubes sequentially with a single needle in a plastic sheath of the collection system, thereby releasing the vacuum in the one or more vacuum collection tubes. In some embodiments, the blood draws may be collected through a catheter. In some embodiments, the catheter comprises a peripheral venous catheter (PVC), a central venousWSGR Docket No.: 47697-751601 catheter (CVC), an arterial catheter, a pulmonary artery catheter (Swan-Ganz Catheter), a hemodialysis catheter, or an umbilical catheter. The location of the catheter placed on a subject (e.g., body sample collection site) may be decided based on the conditions of the subject. In some embodiments, the catheter can comprise a PCV catheter inserted to an ear vein of an animal (e.g., pig)-
[0042] In some embodiments, the sample collection site may be prepared prior to collecting the series of blood draws. In some embodiments, the geographic sample collection site can be cleaned prior to collecting the series of blood draws. In some embodiments, preparing can comprise shaving, scrubbing or cleaning the body sample collection site or the skin covering the body sample collection site. In some embodiments, the sample collection site does not need to be prepared prior to collecting the series of blood draws. In some embodiments, the body sample collection site is not shaved or cleaned prior to collecting the series of blood draws.
[0043] In some embodiments, the series of blood draws can comprise blood draws collected sequentially from a same sample collection site. In some embodiments, the later blood draw has been collected later in time than an earlier blood draw in the series of blood draws. In some embodiments, the sample can be from a later blood draw in a series of blood draws obtained from a subject. In some embodiments, the later blood draw can comprise the second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, eleventh, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, seventeenth, eighteenth, nineteenth, or twentieth blood draw in the series of blood draws obtained from the subject. In some embodiments, an earlier blood draw in the series of blood draws can comprise the first, the second, the third, the fourth, the fifth, the sixth, the seventh, the eighth, or the ninth blood draw in the series of blood draws obtained from the subj ect. In some embodiments, the earlier blood draw can be the first blood draw in the series of blood draws. In some embodiments, the later blood draw can be the second blood draw in the series of blood draws. In some embodiments, the later blood draw can be the third blood draw in the series of blood draws.
[0044] In some embodiments, the later blood draw can be collected at most about 1 second, at most about 5 seconds, at most about 10 seconds, at most about 20 seconds, at most about 30 seconds, at most about 40 seconds, at most about 50 seconds, at most about 60 seconds, at most about 70 seconds, at most about 80 seconds, at most about 90 seconds, at most about 100 seconds, at most about 110 seconds, or at most about 2 minutes, at most about 2.5 minutes, at most about 3 minutes, at most about 3.5 minutes, at most about 4 minutes, at most about 4.5 minutes, at most about 5 minutes later in time than the earlier blood draw in the series of blood draws. In some embodiments, the later blood draw can be collected at most about 10 minutes, at most about 30 minutes, at most about 45 minutes, at most about 60 minutes, at most about 90 minutes, at mostWSGR Docket No.: 47697-751601 about 2 hours, at most about 3 hours, at most about 4 hours, at most about 6 hours, at most about 8 hours, at most about 10 hours, at most about 12 hours, at most about 18 hours, at most about 24 hours, at most about 1.5 days, at most about 2 days, at most about 2.5 days, at most about 3 days, at most about 3.5 days, at most about 4 days, at most about 4.5 days, at most about 5 days, at most about 5.5 days, at most about 6 days, at most about 6.5 days, at most about 7 days, at most about 7.5 days, at most about 8 days, at most about 8.5 days, at most about 9 days, at most about 9.5 days, at most about 10 days, at most about 10.5 days, at most about 11 days, at most about 11.5 days, at most about 12 days, at most about 2 weeks, at most about 2.5 weeks, or at most about 3 weeks later in time than the earlier blood draw in the series of blood draws.
[0045] In some embodiments, the blood draw in the series of blood draws can be collected from the same sample collection site at an interval of at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at least about 45 seconds, at least about 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 10 minutes, at least 15 minutes, at least 20 minutes, at least 25 minutes, at least 30 minutes, at least 40 minutes, at least 50 minutes, at least 60 minutes, at least 70 minutes, at least 80 minutes, at least 90 minutes, at least 100 minutes, at least 110 minutes, at least 2 hours, at least 3 hours, at least 4 hours, at least 6 hours, at least 8 hours, at least 10 hours, at least 12 hours, at least 18 hours, at least 24 hours, at least 1.5 days, at least 2 days, at least 2.5 days, at least 3 days, at least3.5 days, at least 4 days, at least 4.5 days, at least 5 days, at least 5.5 days, at least 6 days, at least6.5 days, at least 7 days, at least 7.5 days, at least 8 days, at least 8.5 days, at least 9 days, at least9.5 days, at least 10 days, at least 10.5 days, at least 11 days, at least 11.5 days, at least 12 days, at least 2 weeks, at least 2.5 weeks, or at least 3 weeks.
[0046] In some embodiments, the series of blood draws can be collected from the same sample collection site over the period of at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at least about 45 seconds, at least about 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 10 minutes, at least 15 minutes, at least 20 minutes, at least 25 minutes, at least 30 minutes, at least 40 minutes, at least 50 minutes, at least 60 minutes, at least 70 minutes, at least 80 minutes, at least 90 minutes, at least 100 minutes, at least 110 minutes, at least 2 hours, at least 3 hours, at least 4 hours, at least 6 hours, at least 8 hours, at least 10 hours, at least 12 hours, at least 18 hours, at least 24 hours, at least 1.5 days, at least 2 days, at least 2.5 days, at least 3 days, at least 3.5 days, at least 4 days, at least 4.5 days, at least 5 days, at least 5.5 days, at least 6 days, at least 6.5 days, at least 7 days, at least 7.5 days, at least 8 days, at least 8.5 days, at least 9 days, at least 9.5 days, at least 10 days, at least 10.5 days, at least 11 days, at least 11.5 days, at least 12 days, at least 2 weeks, at least 2.5 weeks, or at least 3 weeks.WSGR Docket No.: 47697-751601Subjects
[0047] The methods and systems described herein can analyze one or more samples obtained from a subject. A sample or the blood draws provided herein can be from any subject with blood. In some embodiments, the subject can comprise a human or an animal. In some embodiments, the animal can comprise a mammal, a bird, a reptile, an amphibian, a fish, or an arachnid. In some embodiments, the animal can comprise a research animal, an animal for medical use (e.g., xenotransplant donor), a companion animal, a farm animal, a working animal, a performance animal, or a wild animal. In some embodiments, the mammal can comprise a non-human primate (e.g., a macaques or rhesus monkey), a rodent, a carnivore (e.g., a canine or a feline), a bat, a cetacean (e.g., a dolphin), an ungulate, or an insectivore (e.g., hedgehog). In some embodiments, the ungulate comprises a swine, a sheep, a cow, a deer, or a horse.
[0048] In some embodiments, the subject can comprise a male or a female. In some embodiments, the subject can be of any age. In some embodiments, the subject can comprise an embryo or a fetus.
[0049] In some embodiments, the subject can be a healthy subject. In some embodiments, the subject can comprise a subject having, suspected of having, or is at risk of having a disease or a disorder described herein. In some embodiments, the disease or disorder can comprise an infection. In some embodiments, the infection may be caused by any microbe described herein, including but not limited to pathogenic or commensal microbes. In some embodiments, the disease or disorder can comprise a medical indication related to an infection. In some embodiments, a medical indication related to an infection can comprise any disease, disorder or procedure (e.g., medical or surgical) that renders a subject immunocompromised or susceptible to infections. In some embodiments, the medical indication related to an infection may cause a new or recurring infection in a subject. For example, the medical indication related to an infection can comprise any disease for which an immunosuppressant (e.g., chemotherapy, radiation, corticosteroids, transplant medications, or certain biologies) or an anti-infective agent is required as a treatment. In some embodiments, a medical indication related to an infection may comprise cancers, transplantation, surgeries, bums, infections, malnourishment, chronic kidney diseases, diabetes mellitus, autoimmune diseases, or immune disorders (e.g., acquired immunodeficiency syndrome (AIDS)).
[0050] In some embodiments, the subject has or is at risk for developing an infection. In some embodiments, the subject has or is at risk for developing a cancer. In some embodiments, the subject receives an immunosuppressant (e.g., chemotherapy, radiation, corticosteroids, transplant medications, or certain biologies) or an anti-infective agent. In some embodiments, the subject may be eligible as a recipient of transplantation. In some embodiments, the subject may beWSGR Docket No.: 47697-751601 eligible as a donor of transplantation. In some embodiments, the subject can be an animal donor for xenotransplant.Microbes
[0051] The nucleic acids provided herein can be derived from a plurality of microbes. In some embodiments, the microbe can comprise a virus, a bacterium, a protozoa, a fungus or any other microorganisms.
[0052] In some embodiments, the microbe can comprise a pathogenic microbe (e.g., a pathogen) to the subject, a commensal microbe of the subject, or a microbe of the general environment. In some embodiments, a pathogen can comprise any pathogenic or virulent microbe. In some embodiments, a commensal microbe of the subject can include a plurality of microbes that inhabit any location in or on the subject without causing any symptom of a disease or disorder. In some embodiments, a microbe of the general environment can comprise a microbe at or near the sample collection site or a microbe at or near the housing of the subj ect. In some embodiments, a microbe of the general environment can comprise a commensal microbe. In some embodiments, a commensal microbe of a subject may become a pathogen to the subject. In some embodiments, a commensal microbe of a first subject may be a pathogen to a second subject. In some embodiments, a microbe of the general environment of a subject may become a pathogen to the subject. In some embodiments, a microbe of the general environment of a first subject may be a pathogen to a second subject.
[0053] In some embodiments, commensal microbes may inhabit the gastrointestinal tract, skin, respiratory tract, urogenital tract, or the oral cavity of the subject. In some embodiments, commensal microbes may comprise endogenous viruses (e.g., endogenous retroviruses (ERVs)) of the subject. In some embodiments, commensal microbes can comprise Bacteroides fragilis, Lactobacillus acidophilus, Bifidobacterium bifidum, Escherichia coli (non-pathogenic strains), Staphylococcus epidermidis, Streptococcus salivarius, Propionibacterium acnes, Candida albicans (under normal conditions), Enterococcus faecalis, Clostridium difficile (non-toxigenic strains), Rothia mucilaginosa, Fusobacterium nucleatum, Peptostreptococcus anaerobius, or Prevotellamelaninogenica. In some embodiments, commensal microbes comprise Lactobacillus spp., Bacteroides spp., Faecalibacterium prausnitzii, Escherichia coli, Clostridium spp., Enterococcus spp., Staphylococcus spp., Candida spp., Aspergillus spp., Porcine Endogenous Retroviruses (PERVs), Eimeria spp., Bifidobacterium spp., Fusobacterium spp., Simian Immunodeficiency Virus (SIV), Simian Retrovirus (SRV) or Entamoeba spp.
[0054] Microbes of the general environment may vary depending on the subject and the living condition of the subject. In some embodiments, microbes of general environment can comprise microbes in the dirt housing the subject, wherein the subject is an animal. In some embodiments,WSGR Docket No.: 47697-751601 microbes of general environment comprise a plurality of bacteria, fungi, viruses, or parasites. In some embodiments, microbes of general environment comprise Staphylococcus spp., Escherichia coli, Salmonella spp., Pseudomonas spp., Clostridium spp., Candida spp., Aspergillus spp., Penicillium spp., Rodent Parvovirus, Mouse Hepatitis Virus (MHV), Sendai Virus, helminths (e.g., Syphacia obvelatd), protozoa (e.g., Giardia sppfi various soil microorganisms, or any combination thereof.
[0055] In some embodiments, the microbe described herein can comprise one or more microbes selected from the group consisting of: Streptococcus suis, Coniosporium, Hantavirus, Talaromyces, Machlomovirus, Betatetravirus, Raoultella, Aeromonas, Ephemerovirus, Empedobacter, Loa, Macluravirus, Stenotrophomonas, Alfamovirus, Rosavirus, Emmonsia, Aggregatibacter, Orthopneumovirus, Weeksella, Nairovirus, Salivirus, Weissella, Mosavirus, Gammapartitivirus, Strongyloides, Passerivirus, Erysipelatoclostridium, Bacillarnavirus, lotatorquevirus, Taenia, Trypanosoma, Olsenella, Cladosporium, Rhizobium, Prevotella, Leclercia, Paracoccus, liarvirus, Lagovirus, Rasamsonia, Plasmodium, Acremonium, Chlamydia, Clonorchis, Vibrio, Bartonella, Nakazawaea, Franconibacter, Anisakis, Norovirus, Nocardia, Solobacterium, Parechovirus, Avenavirus, Orthohepevirus, Aphthovirus, Hepandensovirus, Microbacterium, Lichtheimia, Lomentospora, Achromobacter, Ipomovirus, Tsukamurella, Elizabethkingia, Hepevirus, Seadornavirus, Alternaria, Trueperella, Gammatorquevirus, Bifidobacterium, Chrysosporium, Thogotovirus, Curtovirus, Deltatorquevirus, Balamuthia, Mastrevirus, Bdellomicrovirus, Mupapillomavirus, Pseudozyma, Wickerhamiella, Aquamavirus, Alloscardovia, Thielavia, Idaeovirus, Henipavirus, Coxiella, Haemophilus, Gammacoronavirus, Negevirus Brevibacterium, Peptoniphilus, Alphacarmotetravirus, Nosema, Trichovirus, Arenavirus, Thermomyces, Necator, Waikavirus, Blosnavirus, Jonesia, Tetraparvovirus, Emaravirus, Plectrovirus, Sclerodamavirus, Toxocara, Umbravirus, Burkholderia, Chromobacterium, Paracoccidioides, Brugia, Eragrovirus, Macrococcus, Absidia, Colletotrichum, Inovirus, Phycomyces, Wickerhamomyces, Acidaminococcus, Moraxella, Rothia, Phlebovirus, Slackia, Purpureocillium, Betapapillomavirus, Tupavirus, Cryspovirus, Saksenaea, Erysipelothrix, Kobuvirus, Mimoreovirus, Echinococcus, Mannheimia, Bergeyella, Cyclospora, Xylanimonas, Leptospira, Finegoldia, Curvularia, Cryptosporidium, Babuvirus, Pecluvirus, Lambdatorquevirus, Pythium, Carlavirus, Entomobimavirus, Kocuria, Anaplasma, Ampelovirus, Avihepatovirus, Nepovirus, Rhodococcus, Bordetella, Mischivirus, Scedosporium, Gardnerella, Maculavirus, Trichoderma, Aveparvovirus, Salmonella, Avastrovirus, Copiparvovirus, Trachipleistophora, Clostridioides, Nanovirus, Siccibacter, Leptotrichia, Citrivirus, Odoribacter, Sanguibacter, Novirhabdovirus, Acremonium, Hafnia, Chaetomium, Tenuivirus, Yokenella, Rubulavirus, Varicellovirus,WSGR Docket No.: 47697-751601Alphamesonivirus, Sicinivirus, Leuconostoc, Microvirus, Gallantivirus, Morbillivirus, Lolavirus, Pantoea, Hepatovirus, Nupapillomavirus, Metschnikowia, Bamavirus, Kytococcus, Tritimovirus, Tannerella, Respirovirus, Pneumocystis, Dirofilaria, Pediococcus, Lactococcus, Blastomyces, Dianthovirus, Actinobacillus, Teschovirus, Oscivirus, Begomovirus, Potyvirus, Byssochlamys, Alphacoronavirus, Molluscipoxvirus, Lymphocryptovirus, Sapelovirus, Parabacteroides, Pyrenochaeta, Listeria, Senecavirus, Brevidensovirus, Potexvirus, Parvimonas, Flavivirus, Recovirus, Toxoplasma, Yatapoxvirus, Opisthorchis, Trichuris, Cyphellophora, Morganella, Perhabdovirus, Micrococcus, Pequenovirus, Mastadenovirus, Anaeroglobus, Tropheryma, Dolosigranulum, Wolbachia, Lelliottia, Mycoplasma Tobravirus, Shewanella, Paeniclostridium, Erythroparvovirus, Sutterella, Sporopachydermia, Namavirus, Nyavirus, Francisella, Arthroderma, Epsilontorquevirus, Sigmavirus, Amdoparvovirus, Actinomyces, Alphapermutotetravirus, Cardiobacterium, Influenzavirus C, Orthopoxvirus, Poacevirus, Phialophora, Lactobacillus, Polyomavirus, Debaryomyces, Foveavirus, Bymovirus, Mycoflexivirus, Grimontia, Mucor, Rhytidhysteron, Quadrivirus, Thermoascus, Aureusvirus, Trichosporon, Myceliophthora, Dermacoccus, Dysgonomonas, Pseudoramibacter, Becurtovirus, Gordonia, Sapovirus, Orthobunyavirus, Spiromicrovirus, Pomovirus, Exophiala, Sneathia, Helicobacter, Photorhabdus, Mogibacterium, Betapartitivirus, Avibirnavirus, Ambidensovirus, Oleavirus, Orientia, Deltacoronavirus, Anulavirus, Trichomonasvirus, Budvicia, Geotrichum, Enamovirus, Lachnoclostridium, Schistosoma, Paecilomyces, Panicovirus, Rhizoctonia, Brevibacillus, Beauveria, Pestivirus, Tombusvirus, Cilevirus, Cokeromyces, Peptostreptococcus, Phanerochaete, Proteus, Idnoreovirus, Aspergillus, Pasteurella, Malassezia, Hanseniaspora, Endornavirus, Azospirillum, Velarivirus, Cystovirus, Avisivirus, Bacteroides, Picobirnavirus, Myroides, Cir covirus, Arterivirus, Aquaparamyxovirus, Onchocerca, Cosavirus, Kluyveromyces, Fijivirus, Candida, Hepacivirus, Dermabacter, Ourmiavirus, Allexivirus, Enterobacter, Acidovorax, Bracorhabdovirus, Carmovirus, Pluralibacter, Coltivirus, Fonsecaea, Streptobacillus, Corynebacterium, Macrophomina, Marburgvirus, Comovirus, Fabavirus, Alphanodavirus, Cellulomonas, Enter obius, Catabacter, Moellerella, Nakaseomyces, Cucumovirus, Valsa, Deltapartitivirus, Plesiomonas, Pseudomonas, Torovirus, Cuevavirus, Hypovirus, Trichomonas, Influenzavirus D, Giardiavirus, Crinivirus, Tepovirus, Sakobuvirus, Cyberlindnera, Paenalcaligenes, Bafmivirus, Rymovirus, Pegivirus, Yarrowia, Treponema, Borreliella, Rubivirus, Aureobasidium, Angiostrongylus, Filobasidium, Photobacterium, Rhizopus, Orthoreovirus, Ustilago, Simplexvirus, Aquareovirus, Protoparvovirus, Propionibacterium, Sprivivirus, Hunnivirus, Apophysomyces, Meyerozyma, Alphapapillomavirus, Candida, Brucella, Gallivirus, Dinovernavirus, Anaerobiospirillum, Eubacterium, Tatlockia, Terrisporobacter, Quaranjavirus, Sobemovirus, Dicipivirus,WSGR Docket No.: 47697-751601Arcanobacterium, Macanavirus, Atopobium, Vesivirus, Lodderomyces, Dinornavirus, Betatorquevirus, Kerstersia, Aparavirus, Neisseria, Agrobacterium, Edwardsiella, Labyrnavirus, Totivirus, Actinomadura, Tobamovirus, Influenzavirus B, Mandarivirus, Anaerococcus, Kunsagivirus, Naegleria, Campylobacter, Veillonella, Yamadazyma, Filobasidiella, Oerskovia, Penicillium, Anncaliia, Leptosphaeria, Pneumovirus, Psychrobacter, Isavirus, Granulicatella, Torradovirus, Cladophialophora, Influenzavirus A, Ophiostoma, Aerococcus, Ureaplasma, Etatorquevirus, Bocaparvovirus, Megasphaera, Reptarenavirus, Comamonas, Capnocytophaga, Alphatorquevirus, Syncephalastrum, Wallemia, Betacoronavirus, Hyphopichia, Nocardiopsis, Legionella, Trichinella, Paraburkholderia, Mammarenavirus, Echinostoma, Sphingobacterium, Enterovirus, Methanobrevibacter, Ochroconis, Cheravirus, Pasivirus, Enterococcus, Mycoreovirus, Tospovirus, Betanodavirus, Phytoreovirus, Enterocytozoon, Ferlavirus, StemphyliumFilif actor, Leishmaniavirus, Gemella, Bromovirus, Alloiococcus, Cunninghamella, Cronobacter, Oribacterium, Orbivirus, Chrysovirus, Cripavirus, Tatumella, Pandoraea, Ogataea, Dracunculus, Volvariella, flavirus, Benyvirus, Rhadinovirus, Histoplasma, Rahnella, Morococcus, Verticillium, Janibacter, Gyrovirus, Alphapartitivirus, Mycobacterium, Roseomonas, Varicosavirus, Chryseobacterium, Parapoxvirus, Rhizomucor, Aureimonas, Levivirus, Leishmania, Luteovirus, Cypovirus, Ochrobactrum, Microsporum, Piscihepevirus, Ceratocystis, Sporothrix, Vesiculovirus, Cupriavidus, Cryptococcus, Metapneumovirus, Alphanecrovirus, Eikenella, Brevundimonas, Escherichia, Leifsonia, Schizophyllum, Granulibacter, Gordonibacter, Lachancea, Madurella, Ophiovirus, Phellinus, Nebovirus, Acanthamoeba, Fusobacterium, Pichia, Verruconis, Ehrlichia, Tibrovirus, Higrevirus, Wohlfahrtiimonas, Rhinocladiella, Neorickettsia, Sadwavirus, Roseobacter, Sequivirus, Pannonibacter, Rotavirus, Turicella, Cardiovirus, Propionimicrobium, Furovirus, Naumovozyma, Closterovirus, Fluoribacter, Zeavirus, Clavispora, Megrivirus, Gammapapillomavirus, Rickettsia, Polemovirus, Corynespora, Encephalitozoon, Shimwellia, Fusarium, Yersinia, Capronia, Delftia, Victorivirus, Marafivirus, Kluyvera, Iteradensovirus, Isoptericola, Vitivirus, Roseolovirus, Conidiobolus, Abiotrophia, Babesia, Phoma, Sanguibacteroides, Staphylococcus, Rhodotorula, Zetatorquevirus, Hymenolepis, Fasciola, Cytorhabdovirus, Cardoreovirus, Memnoniella, Trichophyton, Mitovirus, Phaeoacremonium, Providencia, Lysinibacillus, Giardia, Oligella, Streptomyces, Paraclostridium, Ralstonia, Coccidioides, Brambyvirus, Biatriospora, Allolevivirus, Acinetobacter, Starmerella, Omegatetravirus, Porphyromonas, Avulavirus, Streptococcus, Arcobacter, Topocuvirus, Mamastrovirus, Ancylostoma, Bomavirus, Capillovirus, Alphavirus, Tymovirus, Nucleorhabdovirus, Diaporthe, Chlamydiamicrovirus, Tumcurtovirus, Saccharomyces, Riemerella, Betanecrovirus, Clostridium, Mobiluncus, Cercospora, Mamavirus, Mortierella,WSGR Docket No.: 47697-751601Aquabimavirus, Xanthomonas, Dependoparvovirus, Ebolavirus, Neofusicoccum, Borrelia, Leminorella, Klebsiella, Blastocystis, Alcaligenes, Citrobacter, Eggerthella, Cedecea, Serratia, Penstyldensovirus, Bacillus, Laribacter, Wuchereria, Hordeivirus, Cytomegalovirus, Actinomucor, Ascaris, Shigella, Vittaforma, Torulaspora, Kingella, Oryzavirus, Polerovirus, Tremovirus, Erbovirus, Entamoeba, Lyssavirus, Paenibacillus, Facklamia, Kappatorquevirus, Metarhizium, Stachybotrys, Okavirus, Botrexvirus, Thetatorquevirus, andBasidiobolus. In some embodiments, the microbe described herein comprises one or more microbes disclosed in the Drawings or the Examples.Sample
[0056] The sample provided herein comprises nucleic acid molecules (e.g., cell-free nucleic acids) to be sequenced by the methods described herein. In some embodiments, the sample can be a biological sample. In some embodiments, the biological sample contains cells. In some embodiments, the biological sample can be cell-free. In some embodiments, the biological sample comprises a biological fluid. The biological fluid may comprise a bodily fluid of a subject (e.g., blood), a fluid obtained from the subject via a medical procedure (e.g., lavage), or any fluid obtained from processing a biopsy of the subject (e.g., serous fluid). In some embodiments, the biological fluid can comprise a bodily fluid. In some embodiments, the bodily fluid can comprise blood, plasma, serum, lymph, synovial fluid, cerebrospinal fluid (CSF), saliva, gastric juice, bile, pancreatic juice, intestinal fluid, respiratory tract mucosal secretions, semen, cervical mucus, vaginal secretions, urine, sebum, breast milk, amniotic fluid, pericardial fluid, pleural fluid, peritoneal fluid, or any combination thereof. In some embodiments, biological fluid has been processed from a bodily fluid. For example, blood from a subject can be processed to generate a plasma sample, a serum sample, or a platelet sample. In some embodiments, the biological fluid comprises a plasma sample. In some embodiments, the biological fluid comprises synovial fluid.
[0057] In some embodiments, the biological fluid can comprise a lavage from diagnosing, treating, or cleaning areas of the body of the subject. In some embodiments, the lavage comprises bronchoalveolar lavage (BAL), gastric lavage, peritoneal lavage, nasal lavage, bladder lavage, rectal lavage, wound lavage, joint lavage (arthrocentesis), eye lavage, sinus lavage, or any combination thereof. In some embodiments, the biological fluid comprises amniotic fluid. In some embodiments, the biological fluid comprises BAL. In some embodiments, the biological fluid comprises joint lavage. In some embodiments, the biological fluid comprises any fluid obtained from processing a biopsy of the subject. In some embodiments, the biological fluid comprises needle aspiration fluid, serous fluid, microdialysis fluid, exudate fluid, or any combination thereof.WSGR Docket No.: 47697-751601
[0058] The sample provided herein comprises cell-free nucleic acids (e.g., cfNAs). In some embodiments, the phrase “target nucleic acids” refers to the cfNAs. In some embodiments, the cfNAs can be derived from any source of nucleic acids provided herein. In some embodiments, the sample comprises cell-free nucleic acids originated from genomic nucleic acids (e.g., nuclear nucleic acids), mitochondria nucleic acids, exosomal nucleic acids, fetal nucleic acids, or any combination thereof. In some embodiments, the sample comprises a mixture of nucleic acids. In some embodiments, the sample comprising the target nucleic acids (e.g., cfNAs) can further comprise any nucleic acids provided herein. In some embodiments, the sample further comprises contaminant nucleic acids. In some embodiments, contaminant nucleic acids comprise nucleic acids from the general environment (e.g., sample collection site). In some embodiments, contaminant nucleic acids comprise nucleic acids from the body sample collection site or the geographic sample collection site.
[0059] In some embodiments, the sample comprises microbial nucleic acids. In some embodiments, the sample comprises microbial cell-free nucleic acids (mcfNAs). In some embodiments, the phrase “target nucleic acids” refers to the mcfNAs. In some embodiments, the mcfNAs can be from one or more species of microbes described herein. In some embodiments, the mcfNAs comprise prokaryotic or eukaryotic mcfNAs. In some embodiments, the mcfNAs comprise bacterial cfNAs, fungal cfNAs, viral cfNAs, protozoan cfNAs, archaeal cfNAs, algal cfNAs, or any combination thereof. In some embodiments, the sample comprises non-microbial nucleic acids (e.g., non-microbial cell-free nucleic acids). In some embodiments, the sample comprises host nucleic acids (e.g., host cell-free nucleic acids). In some embodiments, the host comprises any subject provided herein.
[0060] In some embodiments, the sample comprises mcfNAs from one or more species of microbes. In some embodiments, the sample comprises mcfNAs from at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 species of microbes.
[0061] Disclosed herein in some embodiments are samples derived from subjects. In some embodiments, a sample provided herein can comprise a nucleic acid molecule to be sequenced by a method described herein. As used herein, a “sample” generally refers to any material comprising nucleic acids that has been derived from a subject. A sample can comprise a raw biological sample, such as whole blood.
[0062] As used herein, the phrase “raw biological sample” refers to an unmanipulated or unprocessed sample obtained from a subject, e.g., a host, containing or presumed to contain targetWSGR Docket No.: 47697-751601 nucleic acids. In some embodiments, a raw biological sample has not been subjected to any extraction methods after being obtained from a subject In some embodiments, a raw biological sample can be processed or manipulated to produce an initial sample. For example, a raw biological sample can comprise whole blood which is centrifuged to produce an initial sample of plasma for a sequencing assay.
[0063] As used herein, the term “initial sample” refers to a sample comprising nucleic acids derived from a raw biological sample. In some embodiments, an initial sample can comprise a sample that has been processed or manipulated, such as plasma or serum. In some embodiments, an initial sample can comprise target or desired nucleic acids obtained or extracted from a raw biological sample. In some embodiments, an initial sample can be subjected to a sequencing assay. In some embodiments, a raw biological sample or an initial sample can be used directly in a sequencing assay without extraction of a nucleic acid. In some embodiments, a nucleic acid as described herein can be extracted from a raw biological sample or an initial sample for use in a sequencing assay. In some embodiments, an extraction method can comprise an alcohol-based extraction, a column purification, a filtration, a size separation, or any combination thereof. As used herein, “removal” or “extraction,” and their cognates, of nucleic acids refers to steps prior to the start of generating or preparing a nucleic acid library that separate nucleic acids from at least one component with which they are normally associated. In some embodiments, removal or extraction of nucleic acids can refer to the process of creating an initial sample from a raw biological sample. For example, without limitation, the fractionation of whole blood into its component parts, such as plasma, can be considered to involve removal or extraction. Similarly, purification or isolation of DNA from a sample (e.g., plasma sample) can be considered extraction. In some embodiments, a nucleic acid extracted from a sample can be subjected to a sequencing assay. In some embodiments, a raw biological sample or an initial sample can comprise a biological sample.
[0064] In some embodiments, a sample can comprise a biological sample obtained or collected from a subject. In some embodiments, a biological sample can comprise cells. In some embodiments, a biological sample can be substantially cell-free. In some embodiments, a biological sample can comprise a biological fluid. In some embodiments, a biological fluid can comprise a bodily fluid of a subject (e.g., blood), a fluid obtained from the subject via a medical procedure (e.g., lavage), or any fluid obtained from processing a biopsy of the subject (e.g., serous fluid). In some embodiments, a biological fluid can comprise a bodily fluid. In some embodiments, a bodily fluid can comprise a whole blood, a plasma, a serum, a lymph, a synovial fluid, a cerebrospinal fluid (CSF), a saliva, a gastric juice, a bile, a pancreatic juice, an intestinal fluid, a respiratory tract mucosal secretion, a semen, a cervical mucus, a vaginal secretion, aWSGR Docket No.: 47697-751601 urine, a sebum, a breast milk, an amniotic fluid, a pericardial fluid, a pleural fluid, a peritoneal fluid, or any combination thereof. In some embodiments, a biological fluid can be processed from a bodily fluid. For example, blood from a subject can be processed to generate a plasma sample, a serum sample, or a platelet sample. In some embodiments, a biological fluid can comprise a plasma sample. In some embodiments, a biological fluid can comprise a lavage from diagnosing, treating, or cleaning an area of a body of a subject. In some embodiments, a lavage can comprise a bronchoalveolar lavage (BAL), a gastric lavage, a peritoneal lavage, a nasal lavage, a bladder lavage, a rectal lavage, a wound lavage, a joint lavage (arthrocentesis), an eye lavage, a sinus lavage, or any combination thereof. In some embodiments, a biological fluid can comprise an amniotic fluid. In some embodiments, a biological fluid can comprise a BAL. In some embodiments, a biological fluid can comprise a joint lavage. In some embodiments, a biological fluid can comprise a fluid obtained from processing a biopsy of a subject. In some embodiments, a biological fluid can comprise a needle aspiration fluid, a serous fluid, a microdialysis fluid, an exudate fluid, or any combination thereof. As used herein, “plasma” or “blood plasma” refers to the liquid component or fraction of blood. Plasma is generally obtained by spinning a whole blood sample and removing the liquid component.Sample Preparation and Processing
[0065] In some embodiments, a sample comprising nucleic acids can be prepared prior to a sequencing assay. In some embodiments, a raw biological sample comprising whole blood can be processed by centrifugation to generate an initial sample of plasma.
[0066] In some embodiments, whole blood can be collected in a K2-EDTA tube. In some embodiments, whole blood draws are not pooled. In some embodiments, a tube can be gently inverted multiple times after draw. In some embodiments, a tube can be centrifuged for about 1200 RCF (g), about 1400 RCF (g), about 1600 RCF (g), or more after draw in order to separate plasma from the blood. The centrifugation can occur at ambient temperature. In some cases, the centrifugation occurs for greater than 5 minutes, 7 minutes, 10 minutes, 15 minutes or 20 minutes. In some embodiments, for tubes containing less than 4 mL a tube manufacturer’s instruction and centrifugation speed and time can be used. In some embodiments, the plasma fraction can be transferred into a new tube. In some embodiments, the plasma is subjected to centrifugation a second time to remove residual cells (e.g., mammalian cells and microbial cells). The additional centrifugation can be conducted at, e.g., about 1400 RCF (g), about 1600 RCF (g), about 1800 RCF (g), about 2000 RCF (g), or more.
[0067] In some embodiments, at least 0.4 ml, at least 0.5 ml, at least 0.7 ml, or at least 1.0 ml of plasma can be transferred into a sterile polypropylene tube, with care taken to not disturb a buffy coat when transferring. In some embodiments, a tube can be labelled with a patient’s first andWSGR Docket No.: 47697-751601 last name, a unique identifier (DOB or MRN), and / or a date and time of specimen collection. In some embodiments, if a specimen is unlikely to reach a testing facility within 96 hours of collection it can be frozen directly in K2-EDTA after centrifugation. In some embodiments, if a gel plug does not rise to separate cells from plasma, then a tube can be re-centrifuged at a higher speed. In some embodiments, a specimen tube can be shipped to a testing facility.
[0068] In some embodiments, the methods comprise analyzing a cell-free sample. In some embodiments, the methods comprise performing a sequencing assay on the cell-free sample. In some embodiments, the methods comprise preparing a cell-free sample from a biological sample for the sequencing assay. In some instances, the methods comprise centrifuging a biological sample to generate a cell-free sample. In some embodiments, the cell free sample is plasma. In some instances, the methods comprise centrifuging a biological sample to generate a cell-free sample devoid of or almost devoid of human cells. In some instances, the methods comprise centrifuging a biological sample to generate a cell-free sample devoid of or almost devoid of non-human cells, such as a microbe. In some embodiments, the methods described herein comprise processing the sample comprising nucleic acids to maximize collection of target nucleic acids. In some embodiments, the methods can enrich for or maximize the collection of degraded nucleic acids, ultra-short nucleic acids, single stranded nucleic acids, double stranded nucleic acids, nicked nucleic acids, or rare nucleic acids.
[0069] In some embodiments, a nucleic acid can be extracted from a sample. In some embodiments, an extraction can comprise separating nucleic acids from other cellular components and contaminants that can be present in a sample. In some embodiments, a nucleic acid can be extracted from a sample using a liquid extraction (e.g., a Trizol, a DNAzol) technique. In some embodiments, an extraction can be performed by phenol chloroform extraction or precipitation by organic solvents (e.g., ethanol, or isopropanol). In some embodiments, an extraction can be performed using a nucleic acid-binding column, a nucleic acid-binding spin column, or a combination thereof. In some cases, an extraction of a cell-free nucleic acid can involve filtration or ultra-filtration. In some embodiments, a nucleic acid can be extracted or purified by use of magnetic beads that bind nucleic acids. In some embodiments, compositions of a binding buffer can be adjusted to control a strength of bonds between functional groups and a nucleic acid, allowing for controlled and reversible binding. In some embodiments, a nucleic acid can be released from a magnetic particle with an elution buffer.
[0070] In some embodiments, the methods described herein comprise enriching a population of cfNA. In some embodiments, enriching a population of cfNA comprises bioinformatically enriching or physically enriching. In some embodiments, enriching a population of cfNA comprises isolating, extracting, or selectively amplifying a desired population of cfNA (e.g.,WSGR Docket No.: 47697-751601 target cfNA) from the initial sample. In some embodiments, enriching a population of cfNA does not comprise isolating or extracting the desired population of cfNA from the initial sample. In some embodiments, enriching a population of cfNA does not comprise amplifying the desired population of cfNA. In some embodiments, enriching a population of cfNA comprises removing the undesired cfNA population (e.g., non-target cfNA or contaminants) from the initial sample. In some embodiments, enriching cfNA can comprise differentiating and / or selecting the cfNA by one or more characteristics comprising size, sequence, GC content, secondary structure, biological source, or protein-binding.
[0071] In some embodiments, the methods comprise enriching microbial cfNA (mcfNA) in the biological sample. In some embodiments, the methods provided herein comprise enriching for at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of the mcfNA in the biological sample. In some embodiments, enriching mcfNA comprises enriching cfNA that are less than about 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, or 250 bp in length. In some embodiments, enriching mcfNA comprises amplifying the nucleic acids in the initial sample with primers containing non-human nucleic acid sequences. In some embodiments, enriching mcfNA comprises removing non-microbial cfNA. In some embodiments, enriching mcfNA comprises removing nucleosome-bound cfNA. In some embodiments, the methods comprise enriching non-microbial cfNA (e.g., host cfNA) in the biological sample. In some embodiments, the methods comprise enriching mammalian cfNA in the biological sample.
[0072] In some embodiments, the methods comprise enriching cfNA by size selection. In some embodiments, size selection can comprise removing nucleic acids not in the desired size range. In some embodiments, the desired size range comprises an artificial or engineered threshold. In some embodiments, size selection comprises separating nucleic acids by size via chromatography (e.g., size-exclusion chromatography), electrophoresis (e.g., gel or capillary electrophoresis), centrifugation (e.g., density -gradient centrifugation), filtration (e.g., membrane ultrafiltration), magnetic bead-based methods (e.g., SPRI beads), affinity-based methods (e.g., streptavidin beads), or any combination thereof. In some embodiments, the methods described herein can comprise discriminating or differentiating cfNA by size and / or selectively isolating or extracting the cfNA of the desired size range. In some embodiments, the methods comprise differentiating microbial cfNA from the non-microbial cfNA (e.g., host cfNA) in the sample and selectively removing the non-microbial cfNA in the sample by size. In some embodiments, the non- microbial cfNA comprises mammalian cfNA (e.g., human or animal cfNA).WSGR Docket No.: 47697-751601
[0073] In some embodiments, the methods can comprise selectively removing nucleic acid fragments greater than about 500 bp, about 450 bp, about 400 bp, about 350 bp, about 300 bp, about 250 bp, about 200 bp, about 150 bp, about 140 bp, about 130 bp, about 120 bp, about 110 bp, about 100 bp, about 90 bp, about 80 bp, about 70 bp, or about 60 bp in length. In some embodiments, the methods can comprise selectively enriching nucleic acid fragments at most about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp in length.
[0074] In some embodiments, the methods can comprise selectively enriching nucleic acid fragments of about 10 bp to about 20 bp, about 10 bp to about 30 bp, about 10 bp to about 40 bp, about 10 bp to about 50 bp, about 10 bp to about 60 bp, about 10 bp to about 70 bp, about 10 bp to about 80 bp, about 10 bp to about 90 bp, about 10 bp to about 100 bp, about 10 bp to about 110 bp, about 10 bp to about 120 bp, about 10 bp to about 130 bp, about 10 bp to about 140 bp, about 10 bp to about 150 bp, about 10 bp to about 160 bp, about 10 bp to about 170 bp, about 10 bp to about 180 bp, about 10 bp to about 190 bp, about 10 bp to about 200 bp, about 10 bp to about 210 bp, about 10 bp to about 220 bp, about 10 bp to about 230 bp, about 10 bp to about 240 bp, or about 10 bp to about 250 bp in length. In some embodiments, the methods can comprise selectively enriching nucleic acid fragments of about 20 bp to about 250 bp, about 20 bp to about 200 bp, about 20 bp to about 150 bp, about 20 bp to about 100 bp, about 20 bp to about 90 bp, about 20 bp to about 80 bp, about 20 bp to about 70 bp, about 20 bp to about 60 bp, about 20 bp to about 50 bp, about 30 bp to about 250 bp, about 30 bp to about 200 bp, about 30 bp to about 150 bp, about 30 bp to about 100 bp, about 30 bp to about 90 bp, about 30 bp to about 80 bp, about 30 bp to about 70 bp, about 30 bp to about 60 bp, about 30 bp to about 50 bp, about 40 bp to about 250 bp, about 40 bp to about 200 bp, about 40 bp to about 150 bp, about 40 bp to about 100 bp, about 40 bp to about 90 bp, about 40 bp to about 80 bp, about 40 bp to about 70 bp, about 40 bp to about 60 bp, or about 40 bp to about 50 bp in length.Exemplary Process 1 - double-stranded cfDNA
[0075] In some embodiments, the methods provided herein comprise performing process 1 to prepare a sample for high throughput sequencing assay. Process 1 provides an example of preparing a sequencing library from the double-stranded cfDNA in the original sample. In some embodiments, a control molecule can be added to an initial sample of plasma to generate a spiked plasma sample. In some embodiments, nucleic acid extraction can be performed on a spiked plasma sample to generate purified and concentrated cfDNA.WSGR Docket No.: 47697-751601
[0076] In some embodiments, a library preparation process can be performed on a purified and concentrated cfDNA sample. In some embodiments, the library preparation can comprise attaching (e.g., by ligation) double-stranded adapters to double-stranded cfDNA. In some embodiments, library preparation can comprise performing unbiased amplification on the sample.
[0077] In some embodiments, a sample preparation method does not comprise extracting nucleic acids from a raw or initial sample. For example, in some cases, nucleic acids can be extracted during or following library preparation, if at all.
[0078] In some embodiments, an adapter pair can be attached to cfDNA fragments in a sample such as by ligation or PCR amplification. In some embodiments, a pair of adapters can comprise a p5 adapter that is attached to a 5’ end of a molecule and a p7 adapter that is attached to a 3’ end of a molecule. In some embodiments, a p5 and p7 sequence can allow a nucleic acid library to bind and generate clusters on a flow cell.
[0079] In some cases, the cfDNA can be attached to adapters comprising identifier sequences that can differentiate between multiple samples. In some embodiments, the multiple samples comprise a plurality of patient samples and / or control samples (e.g., positive control, negative control). In some embodiments, samples can be pooled after barcoding, then sequenced, then demultiplexed to assign each cluster to its sample.Exemplary Process 2- double-stranded cfDNA and single-stranded cfDNA
[0080] In some embodiments, the methods provided herein comprise performing process 2 to prepare a sample for high throughput sequencing assay. Process 2 provides an example of preparing a sequencing library from the double-stranded and single-stranded cfDNA in the original sample. In some embodiments, a control molecule can be added to an initial sample of plasma to generate a spiked plasma sample. In some embodiments, nucleic acid extraction can be performed on a spiked plasma sample to generate purified and concentrated cfDNA. In some embodiments, a sample preparation method does not comprise extracting nucleic acids from a raw or initial sample. For example, in some cases, nucleic acids can be extracted during or following library preparation, if at all.
[0081] In some embodiments, Process 2 comprises physically enriching for degraded nucleic acids, ultra-short nucleic acids, single stranded nucleic acids, double stranded nucleic acids, nicked nucleic acids, or rare nucleic acids.
[0082] In some embodiments, the cfDNA can be denatured in order to separate strands of doublestranded cfDNA, converting the double-stranded cfDNA into single-stranded cfDNA. In some embodiments, the resulting sample can comprise single-stranded cfDNA derived from both double-stranded and single-stranded cfDNA in the original sample.WSGR Docket No.: 47697-751601
[0083] In some embodiments, the methods described herein can comprise attaching a splint adapter (e.g., a splint oligonucleotide) to the single-stranded cfDNA in the sample. In some embodiments, the splint adapter can comprise dsDNA with a ssDNA overhang. In some cases, the ssDNA overhang can comprise random or degenerate nucleotides (e.g., NNNNNN). In some embodiments, the ssDNA overhang can comprise a specific sequence capable of hybridizing to a target molecule.
[0084] In some embodiments, the splint adapter can randomly hybridize to cfDNA via the singlestranded DNA overhang within the splint adapter. In some embodiments, attaching the splint adapter can further comprise ligating the hybridized adapter to cfDNA, e.g., using an enzyme (e.g., ligase). In some embodiments, the ligase comprises T4 ligase, CircLigase II, CircLigase ssDNA Ligase, Splint ligase, any engineered variants thereof, any natural variants thereof, or any combination thereof.
[0085] In some embodiments, an adapter pair can be attached to cfDNA fragments in a sample such as by ligation or PCR amplification. In some embodiments, a pair of adapters can comprise a p5 adapter that is attached to a 5’ end of a molecule and a p7 adapter that is attached to a 3’ end of a molecule. In some embodiments, a p5 and p7 sequence can allow a nucleic acid library to bind and generate clusters on a flow cell.
[0086] In some cases, the cfDNA can be attached to adapters comprising identifier sequences that can differentiate between multiple samples. In some cases, the sample identifiers attached to the cfDNA via ligation or PCR amplification. In some embodiments, the sample identifiers are incorporated into the splint oligonucleotides. In some embodiments, the multiple samples comprise a plurality of patient samples and / or control samples (e.g., positive control, negative control). In some embodiments, multiple samples can be pooled after barcoding, then sequenced, then de-multiplexed to assign each cluster to its sample.Sequencing Methods
[0087] In some embodiments, the methods described herein comprises performing a sequencing assay of nucleic acids in a sample to generate sequencing reads. In some embodiments, the methods further comprise performing any necessary steps, methods, or techniques described herein to prepare the sample for the sequencing assay. For example, the methods may comprise adding process control molecules (or synthetic spike-in molecules) to the sample, ligating adapters to the nucleic acids, or generating a nucleic acid library. In some embodiments, the methods further comprise analyzing the sequencing reads generated from the sequencing assay. In some embodiments, analyzing the sequencing reads comprises performing the bioinformatic analyses described herein. For example, the methods may comprise calculating an abundance of the sequencing reads or generating fragment lengths profiles of the sequencing reads. In someWSGR Docket No.: 47697-751601 embodiments, the methods further comprise generating a report of the results of the sequencing assay. For more steps related to methods for performing sequencing assays of nucleic acids (e.g., cell-free nucleic acids) in a sample, see US Patents 9976181, 10450620, 11111520, 10697008, and 11674167 and US Patent Applications 17 / 323834 and 17 / 323843, each of which is incorporated by reference in their entirety herein, including any drawings.
[0088] In some embodiments, the methods can comprise calculating an abundance or concentration of sequencing reads generated from the nucleic acids in the sample. In some embodiments, the methods can comprise calculating an abundance of nucleic acids (e.g., cell- free nucleic acids (cfNA)) in the sample. The abundance of nucleic acids in the sample may be calculated based on the abundance of sequencing reads generated from the nucleic acids in the sample. In some embodiments, the abundance or concentration described herein can comprise an absolute abundance, a relative abundance, or a normalized abundance. A relative abundance or a normalized abundance can be calculated based on a reference value. The reference value may comprise an abundance of sequencing reads generated from any nucleic acids in the sample. For example, the reference value may be an abundance of sequencing reads from microbial cell-free DNA (mcfDNA) or of sequencing reads from a synthetic spike-in molecule (or process control molecule). In some embodiments, the methods further comprise calculating the relative abundance of a sequencing read at least in part by comparing the abundance of the sequencing read to the abundance of another sequencing read. In some embodiments, the methods further comprise calculating the normalized abundance of a sequencing read at least in part by comparing the abundance of the sequencing read from mcfDNA in the sample to the abundance of sequencing reads from synthetic spike-in molecules.
[0089] In some embodiments, the methods detect a presence or an absence of sequencing reads generated from the nucleic acids in the sample. In some embodiments, the methods detect a presence or an absence of sequencing reads generated from the nucleic acids in an earlier blood draw. In some embodiments, the methods detect a presence or an absence of sequencing reads generated from the nucleic acids in a later blood draw. In some embodiments, the methods do not comprise calculating an abundance of the sequencing reads generated from the nucleic acids in the sample.
[0090] In some embodiments, the methods can further comprise mapping a sequencing read to a reference sequence. In some embodiments, the reference sequence comprises a microbial sequence. In some embodiments, the reference sequence comprises a non-microbial sequence. In some embodiments, the reference sequence comprises a genomic sequence. In some embodiments, the method further comprise mapping a sequencing read to a species of microbe described herein. In some embodiments, the methods further comprise identifying the source ofWSGR Docket No.: 47697-751601 nucleic acids in the sample from which the sequencing read originated. For example, the nucleic acids in the sample may be from the mcfDNA in the sample or may be from contaminant nucleic acids from the same sample collection site, from which the series of blood draws were collected.
[0091] In some embodiments, the methods can further comprise discarding a sequencing read generated from the later blood draw when the following criteria are met: a) the sequencing read can be attributed or assigned to a single species of microbe, and b) the abundance or concentration of sequencing reads generated from the later blood draw that attributed or assigned to the single species of microbe is lower than the abundance of sequencing reads generated from the earlier blood draw attributed or assigned to the same species of microbe. In some embodiments, the abundance or concentration of sequencing reads is a normalized abundance, a relative abundance, or an absolute abundance of sequencing reads that are attributed or assigned to a single species of microbe.
[0092] In some embodiments, the abundance of sequencing reads generated from the later blood draw that attributed or assigned to the single species of microbe is at least 1 molecule per milliliter (MPM), at least 2 MPM, at least 3 MPM, at least 4 MPM, at least 5 MPM, at least 6 MPM, at least 7 MPM, at least 8 MPM, at least 9 MPM, at least 10 MPM, at least 15 MPM, at least 20 MPM, at least 25 MPM, at least 30 MPM, at least 35 MPM, at least 40 MPM, at least 45 MPM, at least 50 MPM, at least 65 MPM, at least 60 MPM, at least 65 MPM, at least 70 MPM, at least 75 MPM, at least 80 MPM, at least 85 MPM, at least 90 MPM, at least 95 MPM, or at least 100 MPM lower than the abundance of sequencing reads generated from the earlier blood draw that attributed or assigned to the same species of microbe.
[0093] In some embodiments, the normalized abundance of sequencing reads generated from the later blood draw that attributed or assigned to the single species of microbe is at least 1 molecule per milliliter (MPM), at least 2 MPM, at least 3 MPM, at least 4 MPM, at least 5 MPM, at least 6 MPM, at least 7 MPM, at least 8 MPM, at least 9 MPM, at least 10 MPM, at least 15 MPM, at least 20 MPM, at least 25 MPM, at least 30 MPM, at least 35 MPM, at least 40 MPM, at least 45 MPM, at least 50 MPM, at least 65 MPM, at least 60 MPM, at least 65 MPM, at least 70 MPM, at least 75 MPM, at least 80 MPM, at least 85 MPM, at least 90 MPM, at least 95 MPM, or at least 100 MPM lower than the normalized abundance of sequencing reads generated from the earlier blood draw that attributed or assigned to the same species of microbe.
[0094] In some embodiments, the abundance of sequencing reads generated from the later blood draw attributed to the single species of microbe is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at leastWSGR Docket No.: 47697-75160190%, or at least 95% lower than the abundance of sequencing reads generated from the earlier blood draw attributed or assigned to the same species of microbe.
[0095] In some embodiments, the methods provided herein further comprises performing a bioinformatics analysis. In some embodiments, the bioinformatics analysis can comprise assembling sequence data, detecting and quantifying sequencing reads, distinguish microbial nucleic acids from nucleic acids of a non-microbial host, removing sequencing reads from the nucleic acids of a non-microbial host from the analysis, detecting the presence and measuring the abundance of microbial nucleic acids, comparing sequencing reads, calculating abundances of sequencing reads, determining the variances of multiple statistics (e.g., abundances, number of species of microbes identified), comparing abundances or variances of sequencing reads, identifying contaminant nucleic acids from the sample collection site, identifying target nucleic acids (e.g., cell-free nucleic acids), identifying host nucleic acids, generating fragment lengths profiles of microbial nucleic acids, generating fragment lengths profiles of control process molecules, comparing fragment lengths profiles of the microbial nucleic acids, detecting site of infection, detecting the state of infection, detecting the risk of organ rejection in a transplant patient, determining the eligibility of a subject for a transplant, and / or detecting potential for drug resistance.
[0096] In some embodiments, the methods can further comprise determining the variances of multiple statistics. In some embodiments, the methods can comprise determining the variance in the abundance of the sequencing reads generated from one or more microbes. In some embodiments, the methods can comprise determining the variance in the abundance of the sequencing reads generated from one or more blood draws in the series of blood draw. In some embodiments, the methods can comprise determining the variance in the number of microbial species identified across different blood draws in the series of blood draws. In some embodiments, the methods can comprise detecting lower variances of multiple statistics in a later blood draw in the series of blood draws. In some embodiments, the methods can comprise detecting a lower variance of the abundance of the sequencing reads in a later blood draw. In some embodiments, the methods can comprise detecting a lower variance of the number of microbial species identified in a later blood draw.
[0097] In some embodiments, the methods can further comprise applying a control nucleic acid to the collection site of the subject prior to collecting a blood draw. In some embodiments, the control nucleic acid can comprise the synthetic spike-in molecules or the process control molecules described herein. In some embodiments, a known amount of the control nucleic acids can be applied to the sample collection site. In some embodiments, a known amount of the control nucleic acids can be applied to the body sample collection site. In some embodiments, a knownWSGR Docket No.: 47697-751601 amount of the control nucleic acids can be applied to the geographic sample collection site. In some embodiments, the control nucleic acids are applied prior to preparing the sample collection site for the series of blood draws obtained from the sample collection site. The control nucleic acids applied to the sample collection site may be transferred to the blood draws described herein in the sample collection process. In some embodiments, the control nucleic acids applied to the sample collection site may be present in the sample to be sequenced. In some embodiments, the methods can further comprise generating sequencing reads from the control nucleic acids applied to the sample collection site.
[0098] In some embodiments, the methods can further comprise identifying a contaminant nucleic acid sequence, at least in part based on a change in an abundance of the contaminant nucleic acid from the earlier blood draw to the later blood draw. In some embodiments, the methods can comprise identifying a contaminant nucleic acid sequence, at least in part based on a change in an abundance of the sequencing reads from the contaminant nucleic acid from the earlier blood draw to the later blood draw. In some embodiments, the methods can comprise identifying a contaminant nucleic acid sequence by comparing an abundance of the contaminant nucleic acid to an abundance of the control nucleic acid applied to the collection site. In some embodiments, the methods can comprise identifying a species of microbe as a contaminant species, at least in part based on the abundance of sequencing reads generated from nucleic acids of the contaminant species in the sample. In some embodiments, the methods can comprise identifying a pathogen infecting the subject at least in part based on a result of the sequencing assay. In some embodiments, the methods can comprise distinguishing between a contaminant microbe from a target microbe in a sample.
[0099] In some embodiments, the methods can comprise performing a bioinformatic analysis. In some embodiments, performing the bioinformatics analysis comprises assembling sequence data, detecting and quantifying sequencing reads, distinguish populations of nucleic acids, detecting the presence and measuring the abundance of microbial nucleic acids, comparing sequencing reads, comparing abundances of sequencing reads, identifying contaminant nucleic acids from the sample collection site, identifying target nucleic acids (e.g., cell-free nucleic acids), identifying host nucleic acids, generating fragment lengths profiles of microbial nucleic acids, generating fragment lengths profiles of control process molecules, comparing fragment lengths profiles of the microbial nucleic acids, detecting site of infection, detecting the state of infection, detecting the risk of organ rejection in a transplant patient, determining the eligibility of a subject for a transplant, and / or detecting potential for drug resistance.
[0100] In some embodiments, the methods can comprise administering a treatment to the subject, wherein the subject is infected by a microbe described herein. In some embodiments, theWSGR Docket No.: 47697-751601 treatment can comprise an antimicrobial agent. In some embodiments, the methods can comprise determining an eligibility of a subject for transplantation described herein.
[0101] In some embodiments, the methods do not comprise extracting the nucleic acids from the sample or the blood draws described herein. In some embodiments, the methods can comprise extracting nucleic acids from the sample or blood draws. In some embodiments, methods comprise extracting nucleic acids or target nucleic acids from the sample or purification of nucleic acids or target nucleic acid from unwanted components in a reaction mixture (e.g., ligation, amplification, restriction enzyme, end repair, etc.). Any means of extracting nucleic acids can be used in the methods of the application. In some embodiments, the extraction can comprise separating the nucleic acids from other cellular components and contaminants that can be present in the sample. Nucleic acids can be extracted from a sample using liquid extraction (e.g., Trizol®, DNAzol™) techniques. In some cases, the extraction can be performed by phenol chloroform extraction or precipitation by organic solvents (e.g., ethanol, or isopropanol). In some cases, the extraction is performed using nucleic acid-binding columns. In some cases, the extraction is performed using commercially available kits such as the Qiagen Qiamp Circulating Nucleic Acid Kit Qiagen Qubit dsDNA HS Assay kit, Agilent™ DNA 1000 kit, TruSeq™ Sequencing Library Preparation, QIAamp Circulating Nucleic Acid Kit, Qiagen DNeasy kit, QIAamp kit, Qiagen Midi kit, QIAprep spin kit) or nucleic acid-binding spin columns (e.g., Qiagen DNA mini-prep kit). In some cases, extraction of cell-free nucleic acids can involve filtration or ultra-filtration. In some embodiments, nucleic acids can be extracted or purified by the use of magnetic beads. For example, magnetic beads with an iron-oxide core and a surface coated with molecules containing a free carboxylic acid or a synthetic polymer can be used. The salt concentration or polyalkylene glycol can be adjusted to control the strength of the bonds between functional groups and nucleic acid, allowing for controlled and reversible binding. Finally, nucleic acids can be released from the magnetic particles with an elution buffer. In some cases, the extraction or purification is performed using commercially available kits such as Omega Bio-tek Mag-Bind® magnetic bead kit, Agencourt®, RNAClean®, and / or XP magnetic beads.
[0102] In some embodiments, the methods can comprise providing a sample (e.g., a blood draw in a series of blood draws collected from a same site of the subject) from an animal subject. In some embodiments, a microbe can be present in an animal subject or suspected of being present in an animal subject. In some embodiments, a method can comprise generating sequence reads associated with microbial cell-free DNA (mcfDNA) from a sample by performing massively parallel sequencing on cell-free nucleic acids in a sample. In some embodiments, aWSGR Docket No.: 47697-751601 method can comprise aligning a sequence reads corresponding to sequences associated with microbial cell-free DNA (mcfDNA) with a reference sequence.
[0103] The sequencing assay provided herein can be performed by any sequencing methods suitable for sequencing the nucleic acids provided herein. In some embodiments, the sequencing assay herein can comprise massively parallel sequencing. In some embodiments, the massively parallel sequencing can comprise whole genome sequencing. In some embodiments, a massively parallel sequencing can comprise Next Generation Sequencing or a Next Next Generation sequencing. In some embodiments, the methods provided herein comprise determining a concentration or quantity of a mcfDNA. In some embodiments, the methods comprise monitoring a concentration or quantity of a mcfDNA over time. In some embodiments, the method comprises identifying fragments of mcfDNA that vary during a course of treatment.
[0104] In some embodiments, the sequencing assay or the sequencing method can comprise sequencing-by-synthesis. In some embodiments, sequencing methods provided herein comprise Maxam-Gilbert sequencing-based techniques, chain-termination-based techniques, shotgun sequencing, bridge PCR sequencing, single-molecule real-time sequencing, ion semiconductor sequencing (e.g., Ion Torrent sequencing), nanopore sequencing, pyrosequencing (454), sequencing by synthesis, sequencing by ligation (SOLiD sequencing), sequencing by electron microscopy, dideoxy sequencing reactions (Sanger method), massively parallel sequencing, polony sequencing, DNA nanoball sequencing and any variation thereof. The term “Next Generation Sequencing (NGS)” herein refers to sequencing methods that allow for massively parallel sequencing of nucleic acid molecules during which a plurality, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced simultaneously. Non-limiting examples of NGS include sequencing-by-synthesis, sequencing- by-ligation, real-time sequencing, and nanopore sequencing. In some embodiments, sequencing involves hybridizing a primer to the template to form a template / primer duplex, contacting the duplex with a polymerase enzyme in the presence of detectably labeled or unlabeled nucleotides under conditions that permit the polymerase to add labeled or unlabeled nucleotides to the primer in a template-dependent manner, detecting a signal from the incorporated labeled nucleotide or detecting a signal resulting from the process of incorporating labeled or unlabeled nucleotide (e.g., proton release), and sequentially repeating the contacting and / or detecting steps at least once, wherein sequential detection of incorporated labeled or unlabeled nucleotide determines the sequence of the nucleic acid. In some embodiments, exemplary detectable labels include radiolabels, fluorescent labels, protein labels, dye labels, enzymatic labels, etc. In some embodiments, the detectable label can be an optically detectable label, such as a fluorescent label.WSGR Docket No.: 47697-751601Exemplary fluorescent labels include cyanine, rhodamine, fluorescein, coumarin, BODIPY, Alexa Fluor™, or conjugated multi-dyes.
[0105] Disclosed herein in some embodiments are methods for identifying sequence reads obtained through sequencing as host or non -host. The host can be any subject described herein. In some embodiments, sequence reads identified as non-host can then be aligned to a nucleotide database. In some embodiments, the nucleotide database comprises microbial reference sequences. In some embodiments, the database can be selected for those microbial sequences known to be associated with the host, e.g., the set of commensal and pathogenic microorganisms of the subject (e.g., animal or human). In some embodiments, the microbial database can be optimized to mask or remove contaminating sequences. For example, many public database entries include artifactual sequences not derived from the microorganism, e.g., primer sequences, host sequences, and other contaminants. In some embodiments, sequence reads can be aligned to a reference sequence comprising artifactual sequences. In some embodiments, regions that show irregularities in read coverage when multiple samples are aligned can be masked or removed as an artifact. In some embodiments, the detection of such irregular coverage can be done by various metrics, such as the ratio between coverage of a specific nucleotide and the average coverage of the entire contig within which this nucleotide is found. In some embodiments, a sequence that is represented as greater than about 5*, about 10*, about 25*, about 50*, about 100* the average coverage of the reference sequence comprising artifactual sequences can be artifactual. In some embodiments, a binomial test can be applied to provide a per-base likelihood of coverage given the overall coverage of the contig. In some embodiments, each high confidence read can align to multiple organisms in the given microbial database. In some embodiments, to correctly assign organism abundance based upon this possible mapping redundancy, an algorithm can be used to compute the most likely organism (for example, see Lindner et al. Nucl. Acids Res. (2013) 41 (1): elO, which is referenced herein in its entirety). For example, GRAMMy or GASiC algorithms can be used to compute the most likely organism that a given read came from. In some embodiments, alignments and assignment to a host sequence or to a non-host (e.g., microbial) sequence can be performed in accordance with art-recognized methods. For example, a read of 50 nt. can be assigned as matching a given genome if there is not more than 1 mismatch, not more than 2 mismatches, not more than 3 mismatches, not more than 4 mismatches, not more than 5 mismatches, etc. over the length of the read. In some embodiments, publicly available algorithms can be used for alignments and identification. A non-limiting example of such an alignment algorithm is the bowtie2 program (Johns Hopkins University). In some embodiments, these assignments of reads to an organism (e.g., host organism, non-host organism, microbe, pathogen, etc.) can then be totaled and used toWSGR Docket No.: 47697-751601 compute the estimated number of reads assigned to each organism in a given sample, in a determination of the prevalence of the organism in the sample (for example, a cell-free nucleic acid sample). In some embodiments, this information can be used to determine an origin of a pathogen or contaminant. In some embodiments, the analysis described herein can be used to normalize the counts for the size of the microbial genome to provide a calculation of coverage for a microbe. In some embodiments, the normalized coverage for each microbe can be compared to the host sequence coverage in the same sample to account for differences in sequencing depth between samples. In some embodiments, a dataset of microbial organisms represented by sequences in the sample, and the prevalence of those microorganisms can be optionally aggregated and displayed for ready visualization, e.g., in the form of a report.
[0106] In some aspects, the methods provided herein can comprise generating sequencing reads (or sequence reads) from the nucleic acids in the sample. “Sequencing read” and “sequence read” can be used interchangeably herein. In some embodiments, the methods further comprise detecting, identifying, analyzing, processing, comparing, aligning, or mapping the sequencing reads generated from the nucleic acids in the sample. In some embodiments, the sequencing reads are generated from the microbial cell-free nucleic acids (mcfNA) in the sample as described herein. In some embodiments, the methods comprise mapping the sequencing reads generated from the nucleic acids in the sample to a reference sequence. In some embodiments, the reference sequence can comprise a microbial sequence. In some embodiments, the microbial sequence comprises a region of a microbial genome. In some embodiments, the reference sequence comprises an artifactual sequence.
[0107] In some embodiments, the sequencing reads can comprise at least 100, at least 250, at least 500, at least 750, at least 1000, at least 1500, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, at least 6000, at least 7000, at least 8000, at least 9000, at leastlOOOO, at least 12500, at least 15000, at least 17500, at least 20000, at least 30000, at least 40000, at least 50000, at least 60000, at least 70000, at least 80000, at least 90000, or at least 100,000 sequencing reads generated from the mcfNA in the sample. In some embodiments, the sequencing reads can comprise at most 100, at most 250, at most 500, at most 750, at most 1000, at most 1500, at most 2500, at most 3000, at most 3500, at most 4000, at most 4500, at most 5000, at most 5500, at most 6000, at most 7000, at most 8000, at most 9000, at mostlOOOO, at most 12500, at most 15000, at most 17500, at most 20000, at most 30000, at most 40000, at most 50000, at most 60000, at most 70000, at most 80000, at most 90000, or at most 100,000 sequencing reads generated from the mcfNA in the sample.DenaturationWSGR Docket No.: 47697-751601
[0108] The methods described herein can include steps of denaturing nucleic acids. Denaturation can cause all, most, part, or a sufficient part for detection, of the double-stranded nucleic acids to become single-stranded. Denaturation can occur at any step in the process. In some embodiments, denaturation can remove all, most, or part of the secondary, tertiary, or quaternary structure of double-stranded or single-stranded nucleic acids. As such, any type of initial sample can be subjected to the denaturation step, including samples that contain, or are suspected to contain, only double-stranded nucleic acids, only single-stranded nucleic acids, a mixture of double-stranded and single-stranded nucleic acids, or any higher order nucleic acid structure.
[0109] In some embodiments, the nucleic acids can be denatured using heat. In some embodiments, single-stranded nucleic acids in the sample arise as a result of being subjected to denaturation. In some embodiments, however, the nucleic acids in the sample are single-stranded because they were originally single-stranded when they were obtained from the subject, e.g., without limitation, as single-stranded viral genomic RNA, or single-stranded DNA or as a result of shipping and handling conditions.
[0110] In some embodiments, denaturation can be accomplished by applying heat to the sample for an amount of time sufficient to denature double-stranded nucleic acids of interest or to denature secondary, tertiary, or quaternary structures of double-stranded or single- stranded nucleic acids. In general, the sample can be denatured by heating at 95 °C, or within a range from about 65 to about 110 °C, such as from about 85 to about 100 °C. Similarly, the sample can be heated at any temperature between about 50 °C and about 110 °C for any length of time sufficient to effectuate the denaturation, e.g., from about 1 second to about 60 minutes. In some embodiments, long nucleic acids such as intact dsRNA viruses can require longer denaturation times. In general, denaturation is performed in order to ensure that all, most, or part of the nucleic acids or nucleic acids of interest within a sample are present in single-stranded form.
[0111] In some embodiments, denaturation comprises denaturation to enrich certain nucleic acids. In some embodiments, selective denaturation comprises one or more denaturation steps effective for the selection of fragments of a certain length and / or GC- content. In some embodiments, selective denaturation comprises incubation at selected or elevated temperatures.
[0112] In some embodiments, denaturation can remove all, most, part, or a sufficient part for detection of the secondary, tertiary, or quaternary structures in single-stranded DNA and / or RNA molecules. Non-limiting examples of domains of secondary structure that can be removed during the denaturation step include hairpin loops, hairpin stems, bulges, internal loops, and complexes of complementary nucleic acid sequences and any element contributing to folding of the molecule or complexes. In some embodiments, denaturation is not performed, for exampleWSGR Docket No.: 47697-751601 when the sample is known to contain only single-stranded nucleic acids or when there is a desire to restrict the ultimate analysis to only the single-stranded and not the double-stranded nucleic acids in the sample.
[0113] In some embodiments, denaturation comprises adding one or more denaturing agents for a selective or controlled denaturation. In some embodiments, denaturation comprises a selective or controlled denaturation. Depending on the application, chemical or mechanical denaturation can be used (e.g., sonication, mechanical force applied by magnetic field (e.g., magnetic tweezers) or optical traps (e.g., optical tweezers) or the like) with the methods. Chemical denaturation agents that can be used with the methods of the disclosure include but are not limited to, alkaline agents (e.g., NaOH), formamide, guanidinium chloride, guanidine, sodium salicylate, dimethyl sulfoxide (DMSO), propylene glycol, betaine, or urea. In some embodiments, the one or more denaturing agents comprises for example, without limitation, one or more of formamide, urea, guanidinium chloride, salts, betaine, detergents, surfactants, and / or DMSO. Salts can comprise for example, without limitation, NaCl and MgCh.
[0114] In some embodiments, the sample comprising the cell-free nucleic acids can be treated to denature other components in the sample that are not nucleic acids. In some embodiments, the sample can be treated with an enzyme. In some embodiments, the enzyme can comprise a protease. In some embodiments, the protease can comprise a serine protease, a cysteine protease, an aspartic protease, a metalloprotease, or any combination thereof. In some embodiments, the protease can be a broad-spectrum protease. In some embodiments, the protease can comprise trypsin, chymotrypsin, elastase, papain, carboxypeptidase A, proteinase K, or a combination thereof.Nucleic acid library preparation
[0115] The methods described herein can comprise preparing a nucleic acid library from a sample comprising microbial nucleic acids. The nucleic acids in the sample may comprise microbial nucleic acids or non-microbial nucleic acids (e.g., nucleic acids from a subject). In some embodiments, preparing a nucleic acid library may comprise 1) attaching a blocking moiety to the 5 ’-end phosphate moiety to the non-microbial nucleic acids or degrading the non-microbial nucleic acids using a 5’ phosphorylation-specific exonuclease; 2) phosphorylating the microbial nucleic acids to produce 5’ phosphorylated microbial nucleic acids; and 3) attaching 5’- end adapters to the 5’ phosphorylated microbial nucleic acids to produce a nucleic acid library enriched for microbial nucleic acids.
[0116] In one embodiment, the method further comprises attaching 3 ’-end adapters to nucleic acids. In another embodiment, 5’-end phosphorylation-blocking adapters can be attached via the 5 ’-end phosphate moi eties and impede a reaction, resisting ligation, resistingWSGR Docket No.: 47697-751601 phosphorylation, resisting amplification, or resisting sequencing. In another embodiment, the blocking adapters can be splint oligonucleotides that comprise a double-stranded region preferably comprising an uracil base situated not directly connected to the single- stranded region; and a single-stranded region preferably comprising random nucleotides, or a sequence that hybridizes at an end of the non-microbial nucleic acids.
[0117] In another embodiment the nucleic acids can be denatured; the 5 ’-end adapters can be ligated to the 5 ’-end decoy adapters; and a uracil base or DNA backbone in the 5 ’-end decoy adapters can be cleaved to disconnect the adapters from the non-microbial nucleic acids. In another embodiment the non-microbial nucleic acids can be treated by a 5’ phosphorylationspecific exonuclease, preferably a lambda exonuclease or terminator exonuclease.
[0118] In another embodiment, the method further comprises amplifying the microbial nucleic acids attached to 5 ’-end adapters to enrich for the mcfNA, preferably using primers that hybridize to the 3 ’-end adapters or the 5 ’-end adapters, preferably adapters that can be doublestranded oligonucleotides, or splint oligonucleotides that comprise a double-stranded region and a single-stranded region, preferably where the single-stranded region comprises random nucleotides or a sequence that hybridizes to an end of the microbial nucleic acids.
[0119] In another embodiment, the 5’- end adapters can be ligated to the 5’ phosphorylated microbial nucleic acids by a T4 DNA ligase, a SplintR ligase, a PBCV-1 DNA ligase or a Chlorella virus DNA ligase. In another embodiment the method further comprises denaturing the nucleic acids of the sample to produce denatured nucleic acids prior to addition of the adapters; and results in a 2-fold enrichment of the mcfNA.
[0120] Another aspect of the disclosure is a method of preparing a nucleic acid library from a sample comprising nucleic acids, wherein the method comprises: a) providing a sample comprising 5 ’-end phosphorylated nucleic acids and 5 ’-end non-phosphorylated nucleic acids; b) attaching a decoy 5' -end adapter to the 5’-end phosphorylated nucleic acids to produce nucleic acids attached to the 5’-end adapter, wherein the decoy 5' -end adapter is 5’-end nonphosphorylated and is modified to resist 5 ’-end phosphorylation; c) phosphorylating the 5 ’-end non-phosphorylated nucleic acids with a kinase to produce phosphorylated nucleic acids; d) attaching a 5'-end adapter to the 5 ’-end phosphorylated nucleic acids to produce nucleic acids attached to the 5 ’-end adapters; and e) amplifying the nucleic acids attached to the 5 ’-end adapters.
[0121] In one embodiment of the method, 3 '-end adapters can be attached to the nucleic acids that have been modified by 5 ’-end phosphorylation. In some embodiments, the 3’ end adapters can be also attached to nucleic acids that have not been modified by 5’-end nonphosphorylation.WSGR Docket No.: 47697-751601
[0122] Adapters, full length or partial, can be attached to the nucleic acids in a sample at one or more points during the sample preparation process. In some embodiments, adapters can be attached by ligation, by primer extension, by non-templated extension, by template switching, by the addition of nucleotides to the 3' terminus of a nucleic acid molecule, by hybridization, by amplification (e.g., PCR) or a combination of any of these reaction types. In some embodiments, adapters can be attached by a ligation reaction method using a ligase enzyme that recognizes a particular nucleic acid form. In some embodiments, adapters can be attached by a primer extension reaction method using, e.g., a PCR reaction, where the adapter also acts as a primer for a polymerase which acts on a particular nucleic acid form. In some embodiments, adapters can be attached with a combination of a non-templated nucleic acid polymerase and primer extension off of non-templated sequences (e.g., template switching or template switching PCR).
[0123] Depending on the type of nucleic acid molecule in the sample, the adapter attached can be either double-stranded or single-stranded such that the adapter is compatible with the nucleic acid molecules in the sample. For example, in some embodiments a double-stranded adapter is attached to a double-stranded nucleic acid. In some embodiments, it is desirable to protect adapter ends, for example by adding 5'-end and / or 3'-end protective groups, such as amino modifiers, C3 spacers, dideoxy nucleotides, and / or inverted nucleotides or by providing an adapter that is duplexed on one end (or double-stranded) and single-stranded on the other end. Any combination of protective methods and / or groups set forth herein can be used.
[0124] Primer extension reactions can be carried out with a DNA-dependent polymerase, an RNA-dependent polymerase, polymerase with non-templated activity, a reverse transcriptase, or a combination thereof. In some embodiments, the primer extension reaction can be carried out by a DNA or RNA polymerase having strand displacing activity. In some embodiments, the primer extension reaction is carried out by a DNA or RNA polymerase that has non-templated activity. In some other embodiments, the primer extension reaction can be carried out by a DNA or RNA polymerase having strand displacing activity and a DNA or RNA polymerase that has non-templated activity. In some embodiments, primer extension is carried out with a Klenow fragment.Adapters
[0125] Particular adapters can be used with the present disclosure. In general, the adapter compositions allow for the detection of different nucleic acid forms in a sample. Depending on the starting sample type, what nucleic acid(s) are being analyzed, the method, and what detection system is being used, an appropriate adapter can be employed (e.g., particular functional elements or modifications).WSGR Docket No.: 47697-751601
[0126] In some embodiments, an adapter can comprise a polymerase priming sequence, a sequence required to initiate reading of a nucleic acid sequence in sequencing, a sequence required to initiate reading of identifying sequences, and / or one or more identifying sequences (e.g., such as an index, a barcode, a non-templated overhang, a random sequence, unique molecular identifiers, or a combination thereof). For other applications, an adapter can comprise at least one functional element selected from polymerase priming sequence, a sequencing priming sequence, binding sites for amplification primers, a recognition sequence or structural elements required by the sequencing method utilized, one or more identifying sequences, and a label (e.g., radioactive phosphates, biotin, fluorophores, or enzymes). Labels can be added to an adapter if a purification step or particular detection system is desired (e.g., digital PCR, ddPCR, quantitative PCR, microfluidic device, microarray, etcetera).
[0127] The adapter can be single-stranded or double-stranded or can have both singlestranded and double-stranded regions. In some embodiments, the adapter comprises an RNA molecule, a DNA molecule, or a molecule that contains both DNA and RNA sections and / or strands, and / or a single strand that has both RNA and DNA components. In some embodiments, a double-stranded adapter can be blunt- ended. In some embodiments, a double-stranded adapter can contain nucleic acid residue overhang(s). In some embodiments, the adapter can comprise a splint nucleic acid molecule.
[0128] Such nucleic acid residue overhangs (or tails) can be used to mark a molecule as originating from DNA or RNA in the starting sample, particularly when the overhangs are complementary to an overhang sequence deposited by a DNA nucleotidylexotransferase (e.g. TdT), Poly (A) Polymerase, a RT (e.g., SMART er RT, HIV RT), RNA-dependent polymerase (e.g. RdRP from turnip crinkle virus), and / or a DNA-dependent polymerase (e.g., Bst 2.0 DNA polymerase). For example, the adapter overhang can contain one or more T residues in order to hybridize to one or more overhang residues deposited by a DNA polymerase (e.g., Bst 2.0 DNA polymerase, TdT or the like). Similarly, the adapter overhang can contain one or more C residues in order to hybridize to one or more overhang residues deposited by an RT (e.g., SMART er RT, reverse transcriptases derived from Moloney Murine Leukemia Virus, or the like). HIV reverse transcriptase and the long terminal repeat retrotransposon also have non-templated activity but can add a different nucleotide other than C.
[0129] An adapter can comprise an amplification primer that is a primer used to carry out a polymerase chain reaction (PCR). In some embodiments, the amplification primer comprises a random primer. In some embodiments, the amplification primer comprises a template-specific primer. In some embodiments, the amplification primer comprises a primer complementary to a known non-templated overhang known to be added by the polymerase. In some embodiments,WSGR Docket No.: 47697-751601 the amplification primer comprises a standardized flow cell adapter sequence or a part thereof; standardized flow cell adapter sequences are known in the art and include but are not limited to P5 and P7. In some embodiments, the amplification primer comprises a P5 primer. In some embodiments, the amplification primer comprises a P7 primer. In some embodiments, the amplification primer comprises only part of a P5 or P7 primer. In some embodiments, depending on the method of detection, the amplification primer comprises one or more additional functional elements.
[0130] Identifying sequences (e.g., barcode, index, or a combination thereof) can comprise a unique sequence. The identifying sequences can be added to a particular nucleic acid form by the methods provided herein (e.g., ligation, primer extension, amplification, non- templated extension, template switching, template switching PCR or a combination thereof) allowing the identification of each nucleic acid form in a sample or after sequencing. In some embodiments, the identifying sequences can also contain additional functional elements such as primer amplification sites, sequencing priming sites, or sample indexes.
[0131] The identifying sequences can be completely scrambled (e.g., randomers of A, C, G, and T for DNA or A, C, G, and U for RNA) or they can have some regions of shared sequence. For example, a shared region on each end can reduce sequence biases in ligation events. In some embodiments, the adapter comprises shared region and the shared region comprises about or at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 common base pairs.
[0132] Combinations of barcodes and / or indexes can be added to increase diversity. For example, barcodes and / or indexes can be used as identifiers for well position in a microtiter plate, array, or the like (e.g., 96 different barcodes for a 96-well plate), and another barcode can be used as an identifier for a plate number (e.g., 24 different barcodes for 24 different plates), giving 96x24 = 2,304 combinations using 96+24 = 120 sequences. Using three or more barcodes per sample can further increase achievable diversity.
[0133] In some embodiments, the adapter comprises barcodes and / or indexes. In some embodiments, the barcodes and / or indexes are linked to sequencing reads. In some embodiments, particular barcodes and / or indexes can be linked to particular sequencing reads. In some embodiments, particular barcodes and / or indexes can be linked to particular initial sample. In some embodiments, barcodes comprise about 2, 3, 4, 5, 6, 7, 8 ,9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 200, 250, 300, 350, or 400, 500, or 1000 nucleotides (or base pairs) in length.
[0134] In some embodiments, the adapter comprises one or more labels. Labels can be added to an adapter when purification is desired or for using particular detection. Examples of labels that can be used with the disclosure include, but are not limited to, any of those known inWSGR Docket No.: 47697-751601 the art, enzymes such as fluorophores, radioisotopes, stable free radicals, luminescers, such as chemliuminescers, bioluminescers, dyes, pigments, enzyme substrates, biotin, digoxigenin, antigens, antibodies, a His-tag, and other labels. One skilled in the art will choose a label that is compatible with the chosen detection method.
[0135] In some embodiments, attaching comprises using a ligase; in other embodiments attaching comprises using a polymerase. In some embodiments, attaching an adapter comprises attaching an adapter to both DNA and RNA target molecules. When multiple different ligases are used (e.g., a dual ligase system), the ligases can each be specific for a target (e.g., DNA- specific, or RNA-specific). In some embodiments, attaching comprises using a dual ligase system. In some embodiments, the dual ligase system comprises DNA-specific, RNA-specific, and / or ligases that ligate both DNA and RNA templates in any combination.
[0136] In some embodiments, the ligase comprises a ligase specific for double-stranded nucleic acids (e.g., dsDNA, dsRNA, RNA / DNA duplex). An example of a ligase specific for double-stranded DNA and DNA / RNA hybrids is T4 DNA ligase. In some embodiments, the ligase is specific for single-stranded nucleic acids (e.g., ssDNA, ssRNA). An example of such ligase is CircLigase IL In some embodiments, the ligase comprises a ligase specific for RNA / DNA duplexes. In some embodiments, the ligase comprises a ligase that is able to work on single-stranded, double-stranded, and / or RNA / DNA nucleic acids in any combination.
[0137] Both DNA or / and RNA ligases can be used with the disclosure. Ligases that can be used in the methods provided herein can include, but are not limited to, T4 DNA Ligase, T3 DNA Ligase, T7 DNA Ligase, E. coli DNA Ligase, HiFi Taq DNA Ligase, 9°N™ DNA Ligase, Taq DNA Ligase, SplintR® Ligase (also known as Splint-R ligase or PBCV-1 DNA Ligase or Chiarella virus DNA Ligase), Thermostable 5' AppDNA / RNA Ligase, T4 RNA Ligase, T4 RNA Ligase 2, T4 RNA Ligase 2 Truncated, T4 RNA Ligase 2 Truncated K227Q, T4 RNA Ligase 2, Truncated KQ, RtcB Ligase, CircLigase II, CircLigase ssDNA Ligase, CircLigase RNA Ligase, Ampligase® Thermostable DNA Ligase, T4 RNA ligase II and its modified or truncated derivatives, or a combination thereof.
[0138] In some embodiments, the adapters can be attached to nucleic acids, such as, for example, without limitation, a single-stranded RNA, comprising a 5'-end modification such as App (e.g., pre-adenylation). The presence of the 5' App modification can enable oligonucleotides to act as direct substrates for certain ligases and remove the need for ATP. Adapters to singlestranded RNA can contain a 5' adenylation (5' App) modification and / or an RNA-identifying code.
[0139] Alternatively, or additionally, DNA and RNA in a sample can be specifically marked during an adapter attachment step. In some embodiments the adapter attachment step canWSGR Docket No.: 47697-751601 involve template switching or ligation. In some embodiments, the ligase comprises a ligase specific for one type of nucleic acids. For example, a DNA-specific ligase can be used so that adapters are only ligated to the DNA molecules in the sample. In another example, an RNA- specific ligase can be used so that adapters are only ligated to the RNA molecules in the sample. In some embodiments, ligation comprises successive ligation with a first ligase specific to one type of nucleic acid and a second ligase not discriminating between nucleic acid types. For example, successive ligation first with a DNA-specific ligase (e.g., CircLigase ssDNA ligase) followed by a ligase that can act on a DNA or RNA template (e.g., CircLigase II) can be used. Sequential or concurrent first adapter attachment and / or sequential or concurrent second adapter attachment can provide the ability to distinguish between chemical forms of nucleic acids (e.g., DNA and RNA). The choice of ligation method can depend on the ligase specificities and reaction conditions for each ligase used.
[0140] In some embodiments, the ligase comprises ligase selected with an appropriate profile of contaminating nucleic acids so that the profile deters sufficiently from an expected signal of interest (e.g., endogenous microbe signal in cell-free nucleic acid pool) in order to recognize and filter contamination signal originating from ligase. In some embodiments, an appropriate profile of contaminating components (e.g., buffers, buffer components, oligonucleotides, enzymes, water, beads, etc.) is selected so that the profile deters sufficiently from an expected signal of interest in order to recognize and filter contamination signal originating from components.
[0141] Some methods can produce adapter dimers and adapter-derived by-products. Adapter dimers and adapter-derived by-products are two classes of unwanted products of a single-stranded library protocol that are generated by two distinct mechanisms. For example, the single-stranded nucleic acid library protocol developed by Gansauge et al. generates high concentration of adapter dimers and adapter-derived by- products, especially with input samples characterized by low nucleic acid concentration (See, Gansauge, MT and Meyer, M., Singlestranded DNA library preparation for the sequencing of ancient or damaged DNA, Nat Protoc. 2013 Apr;8(4):737-48 and Gansauge MT, Gerber T, Glocke I, Korlevic P, Lippik L, Nagel S, Riehl LM, Schmidt A, and Meyer M., Single- stranded DNA library preparation from highly degraded DNA using T4 DNA ligase, Nucleic Acids Res. 2017 Jun 2;45(10), each of which is incorporated by reference in their entirety herein, including any drawings).
[0142] One way to decrease adapter-derived by-products according to an embodiment of the disclosure comprises using an RNA splint oligonucleotide. In some embodiments, attaching a 3'-end adapter to the denatured nucleic acids and / or single-stranded nucleic acids comprises attaching said adapter with a splint oligonucleotide. In some embodiments, the splintWSGR Docket No.: 47697-751601 oligonucleotide comprises a DNA splint oligonucleotide. In some embodiments, the splint oligonucleotide comprises an RNA splint oligonucleotide or a partial RNA splint oligonucleotide. In some embodiments, attaching a 3 '-end adapter to the denatured nucleic acids and / or single-stranded nucleic acids can comprise ligating with a Splint-R ligase. In some embodiments, attaching a 3'-end adapter to the denatured nucleic acids further comprises adding an RNase inhibitor. In some embodiments, an adapter is attached through a primer extension reaction performed with a polymerase comprising DNA-dependent RNA-dependent polymerase, or a polymerase having non-templated activity.
[0143] Some embodiments comprise preventing ligation of the 5' -end adapter to a complement synthesized during the primer extension reaction (set forth above) during a second ligation step. Some embodiments comprise preventing ligation of the 5' -end adapter to an adapter-derived side product. Digoxigenin can be introduced to the 5' -end of the splint oligo and an anti-digoxigenin antibody can be added to the bead-binding buffer during immobilization of adapted products. An anti-digoxigenin antibody can be added at any point prior to the second ligation. This will produce a bulky moiety at the 5' -end of any splint oligo attached to the biotinylated 3'-end adapter. This moiety will reduce the ability of T4 DNA ligase in a second ligation step to ligate 5' -end adapter to splint oligo hybrid rendering it un-amplifiable in the final PCR step. It can also reduce the efficiency of primer-extension.
[0144] In some embodiments, the methods herein comprise adding an antibody, such as an anti-digoxigenin antibody. In some embodiments, the anti-digoxigenin antibody is added after the 3'-end adapter is attached to the denatured nucleic acids and before a 5' -end adapter is attached. Some embodiments further comprise using beads comprising anti-digoxigenin antibody. Beads can be removed by, for example, without limitation, pelleting on a magnet. For example, an anti-digoxigenin antibody-coated magnetic bead can be added to deplete digoxiginated splint oligos as well as any unhybridized digoxiginated splint oligos. This can be followed by streptavidin-coated magnetic bead. In some embodiments, the anti-digoxigenin antibody is added during a separation step, annealing step, primary extension step, or second ligation step.Process Control Molecules
[0145] One or more process control molecules can be added to the sample provided herein for various reasons, for example, without limitation, to facilitate accuracy in the process of distinguishing one population of nucleic acids e.g., cfDNA from another (or multiple populations of nucleic acids from each other). In some embodiments, the process control molecules comprise synthetic nucleic acids (e.g., synthetic spike-in molecules). In some embodiments, the process control molecules can have special features such as specific sequences,WSGR Docket No.: 47697-751601 lengths, GC content, degrees of degeneracy, degrees of diversity, secondary, tertiary, and quaternary structure, and / or known starting concentrations. In some embodiments, process control molecules can be used for normalizing signal in a sample described herein to account for variations in sample processing. In some embodiments, process control molecules can be added during the library process itself, e.g., without limitation, dephosphorylation controls can be added before and after dephosphorylation or attachment control before and / or after the 3'-end adapter attachment step. In some embodiments, process control molecules can comprise ID Spike(s), Spanks, Sparks, GC Spike-in Panel molecules, or any combination thereof.In some embodiments, the process control molecules can be applied to a sample collection site prior to collection of the sample described herein.
[0146] ID Spike(s) refers to identification spikes used for sample identification tracking, distinguishing different populations of nucleic acids, e.g., distinguishing quantitatively or qualitatively, cross-contamination detection, reagent tracking, and / or reagent lot tracking (See, for example, US Patent no. 9,976,181, which is incorporated by reference in their entirety herein, including any drawings). Spanks are degenerate pools of nucleic acids, or pools of nucleic acids with diverse sequences, used for diversity assessment and abundance calculation (id.). Sparks, “GC Spike-in Panel,” or “GC dSPARKS” are size or length markers which can be used for abundance, normalization, development and / or analysis purposes, process performance monitoring, and other purposes (id.). In some embodiments, the process control molecules can comprise identifying sequences, degenerate bases, various lengths, various GC contents, known starting concentrations, or any combination thereof.
[0147] Process control molecules can additionally include molecules designed to monitor individual steps of the process. Process control molecules can additionally include dephosphorylation control molecules, denaturation control molecules, ligation control molecules, and / or control molecules for non-templated extension or template switching. Partially or fully phosphorylated control molecules (i.e., phosphorylated 5'-end and / or 3'-ends of the control nucleic acids), control molecules with adapter sequences pre-attached (i.e., an example of a control molecule that is added after 3'-end adapter attachment step) can be added during the library process itself, e.g., dephosphorylation control post dephosphorylation step or adapter attachment control post 3'-end adapter attachment step. In some embodiments, process control molecules can comprise dephosphorylation control molecules, denaturation control molecules, and / or ligation control molecules.
[0148] As used herein, the phrase “process control molecules” refers to molecules that are added to a sample before or during nucleic acid library generation to aid in the identification or quantification of nucleic acids in a sample. In some embodiments, process control moleculesWSGR Docket No.: 47697-751601 can comprise nucleic acids. In some embodiments, process control molecules can comprise synthetic nucleic acids. In some embodiments, process control molecules can comprise synthetic nucleic acid sequences. In some embodiments, process control molecules can comprise naturally- occurring nucleic acids. In some embodiments, process control molecules can comprise naturally-occurring nucleic acid sequences. In some embodiments, process control molecules are separate from and not integrated in the target molecules. In some embodiments, process control molecules can have special features such as specific sequences, lengths, GC content, degrees of degeneracy, degrees of sequence diversity, different secondary, tertiary, or quaternary structures, and / or known starting concentrations. In some embodiments, process control molecules can be used for normalizing the signal in a sample to account for variations in sample processing or to control process performance. In some embodiments, process control molecules can include whole assay internal control (WINC) molecules. In some embodiments, at least 10,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 35,000, at least 40,000, at least 45,000, or at least 50,000 unique WINC molecules are spike in the sample. In some embodiments, process control molecules can include sample identifiers. In some embodiments, process control molecules can comprise dephosphorylation control molecules, denaturation control molecules, and / or ligation control molecules. In some embodiments, multiple different types or sets of control molecules can be added to a sample.
[0149] In some embodiments, the process control molecules comprise DNA (deoxyribonucleic acid), RNA (ribonucleic acid), or DNA / RNA hybrid. In some embodiments, the process control molecules comprise natural, synthetic, and / or artificial nucleotide bases or nucleotide analogues. In some embodiments, the nucleotide bases or nucleotide analogues comprise naturally occurring or artificial modifications at one or more of a deoxyribose moiety, ribose moiety, phosphate moiety, nucleoside moiety, or a combination thereof. The modification can comprise an H, OR, R, halo, SH, SR, NH2, NHR, NR2, or CN, wherein R is an alkyl moiety. For example, the nucleotide bases or nucleotide analogues can comprise 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), or a derivative thereof, or any combination thereof.
[0150] In some embodiments, the process control molecules comprise one or more nucleotide bases or nucleotide analogues selected from the group consisting of: 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), N1 -methylpseudouridine, 5-propynyluridine, 5-propynylcytidine, 6- methyladenine, 6- methylguanine, N, N, -dimethyladenine, 2-propyladenine, 2propylguanine, 2-aminoadenine, 1- methylinosine, 3 -methyluridine, 5-methyluridine, 5- (2- amino) propyl uridine, 5-halocytidine, 5-halouridine, 4-acetylcytidine, 1- methyladenosine, 2-methyladenosine, 3 -methylcytidine, 6-WSGR Docket No.: 47697-751601 methyluridine, 2- methylguanosine, 7-methylguanosine, 2, 2-dimethylguanosine, 5- methylaminoethyluridine, 5-methyloxyuridine, deazanucleotides (such as 7-deaza- adenosine, 6- azouridine, 6-azocytidine, or 6-azothymidine), 5-methyl-2-thiouridine, other thio bases (such as 2-thiouridine, 4-thiouridine, and 2-thiocytidine), dihydrouridine, pseudouridine, queuosine, archaeosine, naphthyl and substituted naphthyl groups, any O-and N-alkylated purines and pyrimidines (such as N6-methyladenosine, 5-methylcarbonylmethyluridine, uridine 5-oxyacetic acid, pyridine-4-one, or pyridine-2-one), phenyl and modified phenyl groups such as aminophenol or 2,4, 6-trimethoxy benzene, modified cytosines that act as G-clamp nucleotides, 8-substituted adenines and guanines, 5-substituted uracils and thymines, azapyrimidines, carboxyhydroxyalkyl nucleotides, carboxyalkylaminoalkyi nucleotides, and alkylcarbonylalkylated nucleotides.
[0151] As used herein, the phrase “adapter attachment control molecule” refers to a control molecule that allows monitoring of the efficiency of an adapter attachment reaction. An adapter attachment reaction can be ligation-based, TdT-based, template-switching-based, primer- extension-based, amplification-based, or a combination thereof.
[0152] As used herein, the phrase “degradation assessment molecules” refers to a control molecule used to evaluate sample and spiked sample integrity during processing.
[0153] As used herein, the phrase “spiked initial sample” refers to an initial sample to which process control molecules (or synthetic spike-ins) have been added prior to the start of generating a sequencing library.
[0154] As used herein, “sequence diversity controls” refers to degenerate pools, or pools of nucleic acids with diverse sequences, which degenerate pools can often be used for diversity assessment, abundance calculation, and / or determination of information transfer efficiency.
[0155] As used herein, “size controls,” “length controls,” “GC Spike-in Panel” or “GC size / length controls” refers to nucleic acids that are size or length or GC-content markers, which can be used for abundance normalization, development, and / or analysis purposes and other purposes.
[0156] As used herein, “ID Spike(s)” refers to identification spikes that can be used, for example without limitation, for sample identification tracking, cross-contamination detection, reagent tracking, and / or reagent lot tracking.Kits and Systems
[0157] In some aspects, this disclosure provides kits and systems for performing the methods described herein. In some cases, the kits and / or systems can be used to identify a particular population of nucleic acids (e.g., mcfDNA) present in a sample comprising of nucleicWSGR Docket No.: 47697-751601 acids. In some embodiments, the sample can comprise a mixture of nucleic acids (e.g., mcfDNA, host DNA and contaminant nucleic acids from the sample collection site).
[0158] In some embodiments, the kit may comprise an enzyme. In some instances, the kit may comprise a kinase. For example, the kit can comprise a PNK kinase. In some cases, the kit comprises a ligase (e.g., T4 ligase). In some cases, the kit further comprises an uracil DNA glycosylase (UDG). In some cases, the kit further comprises an endonuclease. In some cases, the endonuclease is DNA glycosylase-lyase Endonuclease VIII.
[0159] In some cases, the kit comprises a process control molecule described herein (e.g., SPANKs, SPARKs, ID SPIKEs, or other process control molecules). In some cases, the oligonucleotides and / or reagents in a kit provided herein are present in a buffer. In some cases, they can be lyophilized.
[0160] The kit or system can further comprise a software package for data analysis, which can include reference profiles for comparison with the test profile from a clinical sample, and in particular can include reference databases.
[0161] In some cases, the kit can include instructions on how to use the kit; the instructions can be recorded on any suitable recording medium, including but not limited to paper, electronic format, etc. Instructions can be present in the kit as a package insert, in the labeling of the container of the kit, or kit components thereof (i.e., associated with the packaging or sub packaging), etc. In some embodiments, the instructions can be obtained virtually or remotely and can be downloadable or printable including but not limited to via the internet, email, fax, etcetera, process.
[0162] The kit can comprise reagents and steps to isolate the different population of nucleic acids. The kit can comprise specific instructions regarding use of each of the oligos, reagents and handling of such. The kit can provide instructions to direct the purification or isolation, purification of nucleic acids herein, e.g., the purification of cfDNA or sample DNA and further instructions on how to proceed with any amplification reactions, sequencing reactions etcetera including any analytical procedures.
[0163] Such kits or systems can also include information, such as scientific literature references, package insert materials, clinical trial results, and / or summaries of these and the like. Such kits or systems can also include instructions to access a database. Kits or systems described herein can be provided, marketed and / or promoted to health providers, including physicians, nurses, pharmacists, formulary officials, and the like. Kits or system can also be marketed directly to the consumer.
[0164] The kit or system can further comprise an apparatus for detection and / or computer control systems with machine-executable instructions to implement the methods. In someWSGR Docket No.: 47697-751601 embodiments, computer control systems can be further programmed for conducting genetic analysis. Detection systems that can be used including, but are not limited to, sequencing, digital PCR, ddPCR, quantitative PCR (e.g., real-time PCR), or by a microfluidic device, microarray, or the like.Hardware and Software
[0165] A kit or system can comprise a nucleic acid sequencer for generating DNA or RNA sequence information. The kit or system can further comprise a computer comprising software that performs bioinformatics analysis on the DNA or RNA sequence information. Bioinformatics analysis can include, without limitation, assembling sequence data, detecting and quantifying sequencing reads, distinguish populations of nucleic acids, detecting the presence and measuring the abundance of microbial nucleic acids, comparing sequencing reads, calculating abundances of sequencing reads, determining the variances of multiple statistics (e.g., abundances, species of microbes identified), comparing abundances or variances of sequencing reads, identifying contaminant nucleic acids from the sample collection site, identifying target nucleic acids (e.g., cell-free nucleic acids), identifying host nucleic acids, generating fragment lengths profiles of microbial nucleic acids, generating fragment lengths profiles of control process molecules, comparing fragment lengths profiles of the microbial nucleic acids, detecting site of infection, detecting the state of infection, detecting the risk of organ rejection in a transplant patient, determining the eligibility of a subject for a transplant, and / or detecting potential for drug resistance.
[0166] The kit or system can also include computer control systems with machineexecutable instructions (e.g., software) to implement the methods. FIG. 14 shows a computer system 1401 that is programmed or otherwise configured to implement methods of the present disclosure. The computer system 1401 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 1405, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 1401 also includes memory or memory location 1410 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 1415 (e.g., hard disk), communication interface 1420 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1425, such as cache, other memory, data storage and / or electronic display adapters.
[0167] The memory 1410, storage unit 1415, interface 1420, and peripheral devices 1425 are in communication with the CPU 1405 through a communication bus (solid lines), such as a motherboard. The storage unit 1415 can be a data storage unit (or data repository) for storing data.WSGR Docket No.: 47697-751601
[0168] The computer system 1401 can be operatively coupled to a computer network (“network”) 1430 with the aid of the communication interface 1420. The network 1430 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 1430 in some embodiments is a telecommunication and / or data network. The network 1430 can include one or more computer servers, which can enable distributed computing, such as cloud computing.
[0169] The network 1430, in some embodiments with the aid of the computer system 1401, can implement a peer-to-peer network, which can enable devices coupled to the computer system 1401 to behave as a client or a server.
[0170] The CPU 1405 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory 1410. The instructions can be directed to the CPU 1405, which can subsequently program or otherwise configure the CPU 1405 to implement methods of the present disclosure. Examples of operations performed by the CPU 1405 can include fetch, decode, execute, and writeback. The CPU 1405 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1401 can be included in the circuit. In some embodiments, the circuit is an application specific integrated circuit (ASIC).
[0171] The storage unit 1415 can store files, such as drivers, libraries, and saved programs. The storage unit 1415 can store user data, e.g., user preferences and user programs. The computer system 1401 in some embodiments can include one or more additional data storage units that are external to the computer system 1401, such as located on a remote server that is in communication with the computer system 1401 through an intranet or the Internet.
[0172] The computer system 1401 can communicate with one or more remote computer systems through the network 1430. For instance, the computer system 1401 can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC's (e.g., APPLE® iPad, SAMSUNG® Galaxy Tab), telephones, Smart phones (e.g., APPLE® iPhone, Android- enabled device, BLACKBERRY®), or personal digital assistants. The user can access the computer system 1401 via the network 1430.
[0173] The kit or system can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 1401, such as, for example, on the memory 1410 or electronic storage unit 1415. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 1405. In some embodiments, the code can be retrieved from the storage unit 1415 and stored on the memory 1410 for ready access by the processorWSGR Docket No.: 47697-7516011405. In some situations, the electronic storage unit 1415 can be precluded, and machineexecutable instructions are stored on memory 1410. The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.
[0174] Parts of the kits and systems, such as the computer system 1401, can be embodied in programming. Various aspects of the technology can be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machineexecutable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which can provide non-transitory storage at any time for the software programming. All or portions of the software can at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, can enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that can bear the software elements includes optical, electrical, and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also can be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0175] Hence, a machine readable medium, such as computer-executable code, can take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as can be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readableWSGR Docket No.: 47697-751601 media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD- ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0176] The computer system 1401 can include or be in communication with an electronic display 1435 that comprises a user interface (UI) 1440 for providing, an output of a report, which can include a diagnosis of a subject or a therapeutic intervention for the subject. Examples of UI's include, without limitation, a graphical user interface (GUI) and web-based user interface. The analysis can be provided as a report. The report can be provided to a subject, to a health care professional, a lab-worker, or other individual.
[0177] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 1405. The algorithm can, for example, facilitate the enrichment, sequencing and / or detection of pathogen or microbe or other target nucleic acids.
[0178] Information about a patient or subject can be entered into a computer system, for example, patient background, patient medical history, or medical scans. The computer system can be used to analyze results from a method described herein, report results to a patient or doctor, or come up with a treatment plan.Nucleic acids
[0179] Disclosed herein comprises methods for sequencing nucleic acids in a sample. In some embodiments, the nucleic acids can comprise a plurality of chemical forms of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or DNA / RNA hybrid. In some embodiments, the nucleic acids comprise a plurality of structural forms of DNA, RNA, or DNA / RNA hybrid. In some embodiments, the nucleic acids comprise a mixture of nucleic acids from different sources. In some embodiments, the nucleic acids comprise microbial nucleic acids and / or non-microbial nucleic acids. In some embodiments, the nucleic acids comprise target nucleic acids (e.g., cell-free nucleic acids in the subject’s blood) and / or contaminant nucleic acids (e.g., from a collection site or general environment). In some embodiments, the nucleic acids in the sample comprise a mixture of nucleic acids. In some embodiments, the mixture comprises target nucleic acids (e.g., cell-free nucleic acids in the subject’s blood) and / or contaminant nucleic acids (e.g., from a collection site or general environment).WSGR Docket No.: 47697-751601
[0180] In some embodiments, the nucleic acids comprise cell-free nucleic acids (cfNA). In some cases, cfNA comprises circulating cfNA. In some embodiments, the nucleic acids comprise circulating cfDNA, circulating cfRNA, cfDNA, cfRNA, circulating DNA, circulating RNA, or any combination thereof. In some embodiments, cfNA comprises microbial cfNA (mcfNA) (e.g., mcfDNA or mcfRNA). In some embodiments, the nucleic acids comprise linear nucleic acids or circular nucleic acids. In some embodiments, the nucleic acids comprise single strand nucleic acids, double strand nucleic acids or hybrid nucleic acids. In some embodiments, the nucleic acids are members selected from the group consisting of genomic DNA, cDNA, mRNA, cRNA, tRNA, ribosomal RNA, miRNA, siRNA, nuclear DNA, nuclear RNA, plasmids, and vectors.
[0181] In some embodiments, the nucleic acids described herein can be from a plurality of sources. In some embodiments, the nucleic acids are from any one of the subjects provided herein. In some embodiments, the nucleic acids comprise human nucleic acids, animal nucleic acids, or microbial nucleic acids. In some embodiments, microbial nucleic acids comprise bacterial nucleic acids, fungal nucleic acids, parasite nucleic acids, viral nucleic acids, cell-free bacterial nucleic acids, cell-free fungal nucleic acids, cell-free parasite nucleic acids, or viral particle associated nucleic acids. In some embodiments, nucleic acids can be nucleic acids derived from microbes or pathogens including but not limited to viruses, bacteria, archaea, fungi, molds, protists, protozoa, parasites, and any other microbe, particularly an infectious microbe or potentially infectious microbe. In some embodiments, nucleic acids can be derived directly from the subject, as opposed to a microbe or pathogen. In some embodiments, the subject can have, or is suspected of having, a pathogenic infection. In some embodiments, the sample from the host subject comprises host DNA and RNA, as well as DNA and RNA from a pathogen or microbe which can be in the chemical or structural form of ssRNA, ssDNA, dsRNA, or dsDNA. In some embodiments, the nucleic acids comprise mcfDNA or mcfRNA.
[0182] In some embodiments, the nucleic acids can be from a genome of an organism or an organelle of a cell (e.g., an exosome or a mitochondria). In some embodiments, the nucleic acids comprise mitochondrial DNA, intercellular signal nucleic acids, exogenous nucleic acids, DNA enzymes, RNA enzymes, food-derived nucleic acids, any metabolic form of nucleic acidbased therapeutics, or any combination thereof.
[0183] In some embodiments, the nucleic acids comprise environmental nucleic acids. In some embodiments, environmental nucleic acids comprise any nucleic acid at or near the sample collection site comprise, or any nucleic acid introduced by the personnel, equipment or reagent used in collecting and / or processing the sample from the subject.WSGR Docket No.: 47697-751601
[0184] In some embodiments, a sample provided herein comprises circulating mcfDNA or mcfRNA. In some embodiments, mcfDNA or mcfRNA in human blood can originate from bacteria. In some embodiments, mcfDNA or mcfRNA can be detected in subjects that have been exposed to one or more infectious disease and / or one or more non-infectious diseases. In some embodiments, mcfDNA or mcfRNA can be detected in healthy individuals. An aspect of the present disclosure is the detection of mcfDNA or mcfRNA as a biomarker of infection. In some embodiments, mcfDNA can comprise fragments of double-stranded DNA. In some embodiments, the mcfDNA or mcfRNA provided herein or fragments thereof can be approximately less than about 20 bp, less than about 30 bp, less than about 40 bp, less than about 50 bp, less than about 60 bp, less than about 70 bp, less than about 80 bp, less than about 90 bp, less than about 100 bp, less than about 110 bp, less than about 120 bp, less than about 130 bp, less than about 140 bp, less than about 150 bp, less than about 160 bp, less than about 170 bp, less than about 180 bp, less than about 190 bp, less than about 200, less than about 210, less than about 220, less than about 230, less than about 240, or less than about 250 bp long. In some embodiments, the mcfDNA or mcfRNA provided herein or fragments thereof can be about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, or about 200 bp long. In some embodiments, the mcfDNA or mcfRNA provided herein or fragments thereof can be from about 10 bp to about 100 bp long. In some embodiments, the mcfDNA or mcfRNA provided herein or fragments thereof can be from about 30 bp to about 80 bp long.
[0185] In some embodiments, the mcfDNA or mcfRNA can be present in blood, plasma, serum, saliva, or other bodily fluids. In some embodiments, the mcfDNA or mcfRNA is not encapsulated by cells. In some embodiments, the mcfDNA or mcfRNA can be associated with communicable and / or non-communicable diseases. In some embodiments, the mcfDNA or mcfRNA can be associated with a range of diseases and conditions of the subject, including but not limited to an infection by a pathogen, inflammatory bowel disease (IBD), Kawasaki disease (KD), human immunodeficiency virus (HIV), cardiovascular diseases (CVD), cystic fibrosis (CF), and pneumonia, sepsis, cancer, gastric cancer (GC), hepatocellular carcinoma (HCC), and melanoma.Cell-free Nucleic acids
[0186] In some embodiments, the nucleic acids described herein comprise cell-free nucleic acids (cfNAs). Cell-free nucleic acids are generally fragments of nucleic acids that float freely outside of cells in any body fluid of a subject. In some embodiments, cfNA can comprise plasma cfNA, serum cfNA, whole-blood cfNA, cerebrospinal fluid (CSF) cfNA, saliva cfNA,WSGR Docket No.: 47697-751601 bronchoalveolar lavage (BAL) cfNA, urine cfNA, amniotic cfNA, fetal cfNA, synovial fluid cfNA, lymphatic cfNA, or any combination thereof. In some cases, cfNA comprise circulating cfNA in a subject’s bloodstream. In some embodiments, the nucleic acids comprise circulating cfDNA, circulating cfRNA, cfDNA, cfRNA, circulating DNA, circulating RNA, or any combination thereof. In some embodiments, a sample of a body fluid can comprise a cfNA.
[0187] The cfNA described herein are, in some embodiments, nucleic acids that, when floating within a body fluid of a subject, are not encapsulated by a cell. In some embodiments, a cfNA comprises nucleic acids that are not encapsulated by a human cell, not encapsulated by a microbial cell, or not encompassed by either a human cell or a microbial cell. The cfNAs can be free-floating, such as cfDNA fragments in plasma. In some embodiments, cfNA includes vesicle- associated cfNA, such as cfNA associated with exosomes, extracellular vesicles, microvesicles, apoptotic bodies, or any combination thereof. In some embodiments, cfNA do not comprise vesicle-associated cfNA, such as cfNA associated with exosomes, extracellular vesicles, microvesicles, apoptotic bodies, or any combination thereof. cfNAs can also be associated with proteins or other cellular constituents, outside of an intact cell. For example, cfNA can comprise free-floating nucleosome-associated cfNA. cfNAs can arise from various biological processes, including cell death (apoptosis, necrosis) or active secretion.
[0188] In some embodiments, cfNA can comprise viral nucleic acids that are not encapsulated by a capsid, that are fragmented, or a combination thereof. Of note, the methods provided herein can, in some embodiments, be practiced using viral nucleic acids derived from whole viruses floating in a bodily fluid.
[0189] In some embodiments, cfNA can be alternatively referred to as free-circulating nucleic acids. In some embodiments, a cfNA can originate from cell death and other processes that release fragments of nucleic acids into a bloodstream or other biological fluid. In some embodiments, a cfNA can be derived from any source of nucleic acids provided herein. In some embodiments, cfNA present in a raw biological sample can be isolated from genomic nucleic acid in the raw biological sample by processing the raw biological sample into an initial sample by removing intact cells. In some embodiments, removing intact cells can comprise centrifuging or filtering a raw biological sample to produce a cell-free fraction of a biological fluid comprising cfNA. In some cases, the removal of intact cells may use a technique targeting a specific cell type. For example, centrifugation at a relative low RPM can remove human or mammalian cells. In some cases, centrifugation at a higher RPM can be used to remove microbial cells such as bacterial cells. In some cases, centrifugation can be performed at an even higher speed (e.g, via ultracentrifugation) in order to remove viral particles. In some cases, afaster spin can be performed after an initial removal (e.g., by centrifugation) of mammalian cells. In someWSGR Docket No.: 47697-751601 embodiments, a sample can comprise non-host nucleic acids. In some embodiments, non-host nucleic acids can comprise microbial nucleic acids. In some embodiments, microbial nucleic acids can comprise microbial cell-free nucleic acid (mcfNA). In some embodiments, the phrase “target nucleic acids” as used herein can refer to target cfNA. In some embodiments, the phrase “target nucleic acids” as used herein can refer to target mcfNA. In some embodiments, an mcfNA can be derived from one or more kingdoms, divisions, classes, orders, families, genuses, species and / or strain of microbe. In some embodiments, a mcfNA can be derived from a prokaryotic or a eukaryotic microbe. In some embodiments, an mcfNA can comprise a bacterial cfNA, a fungal cfNA, a viral cfNA, a protozoan cfNA, an archaeal cfNA, an algal cfNA, or any combination thereof. In some embodiments, a sample can comprise a non-microbial nucleic acid (e.g., a non- microbial cell-free nucleic acid). In some embodiments, a sample can comprise mcfNAs from one or more species of microbes. In some embodiments, a sample can comprise mcfNAs from at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 species of microbes.
[0190] In some embodiments, a sample can comprise a mixture of nucleic acids. In some embodiments, a sample can comprise target nucleic acids (e.g., target cfNAs) and / or non-target nucleic acids. In some cases, the cfNAs comprise cfNA derived from a subject or microbial genome. In some cases, the cfNAs are derived from a housekeeping gene, e.g., a human or mammalian housekeeping gene, a microbial housekeeping gene, a microbe-specific housekeeping gene, and / or a microbe non-specific housekeeping gene.
[0191] In some embodiments, a sample can further comprise contaminant nucleic acids. In some embodiments, contaminant nucleic acids can comprise nucleic acids from a general environment (e.g., a sample collection site). In some embodiments, contaminant nucleic acids are introduced during any step of sample processing. In some embodiments, the contaminant nucleic acids comprise contaminating microbial nucleic acids, contaminating nucleic acids from a different sample, contaminating host nucleic acids, or any combination thereof.
[0192] In some embodiments, cell-free nucleic acids (cfNAs) can comprise a mixture of cfNAs. In some embodiments, a mixture of cfNAs can comprise cfNAs originated from one or more organisms. In some embodiments, a mixture of cfNAs can comprise microbial nucleic acids (e.g., mcfNAs) originated from one or more species of microbes. In some embodiments, an mcfNA can comprise a bacterial-derived cfNA, a fungal -derived cfNA, a viral-derived cfNA, a protozoan-derived cfNA, an archaeal-derived cfNA, an algal-derived cfNA, or any combination thereof.WSGR Docket No.: 47697-751601
[0193] In some embodiments, a cfNA can comprise a double-stranded nucleic acid (dsNA), a single-stranded nucleic acid (ssNA), nicked double-stranded, or a combination thereof. In some embodiments, a cfNA can comprise a cell-free DNA (cfDNA), a cell-free RNA (cfRNA), a cell-free DNA-RNA hybrid (cfDNA-RNA), or a combination thereof. In some embodiments, hcfNA can comprise host cell-free DNA (hcfDNA), host cell-free RNA (hcfRNA), host cell-free DNA-RNA hybrid (hcfDNA-RNA), or a combination thereof. In some embodiments, mcfNA can comprise microbial cell-free DNA (mcfDNA), microbial cell-free RNA (mcfRNA), microbial cell-free DNA-RNA hybrid (mcfDNA-RNA), or any combination thereof.
[0194] In some embodiments, the cfNA comprises a modified nucleotide base or a nucleotide analogue. In some embodiments, the nucleotide base or nucleotide analogue comprises modifications at one or more of a deoxyribose moiety, ribose moiety, phosphate moiety, nucleoside moiety, or a combination thereof. The modification can comprise an H, OR, R, halo, SH, SR, NH2, NHR, NR2, or CN, wherein R is an alkyl moiety. For example, the cfNA can comprise 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), or a derivative thereof, or any combination thereof.
[0195] As used herein, a cell-free samples is generally a sample devoid, or almost devoid, of cells. In some instances, the cell-free sample is devoid of human cells (e.g., intact human cells). In some instances, the cell-free sample is devoid of microbial cells (e.g., intact microbial cells). In some instances, the cell-free sample is devoid of microbial and human cells. In some instances, the cell-free sample is devoid of a particular type of microbial cell or virus, while comprising a different type of microbial cell or virus. For example, the cell-free sample can, in some embodiments, be devoid of intact bacterial cells while containing intact viruses. In some instances, the cell-free sample is devoid of all types of microbes, including eukaryotic or prokaryotic cells. The cell-free sample can be obtained from a biological sample provided herein. In some instances, the cell-free sample is a plasma, which has been processed in order to remove blood cells, subject cells, and / or intact microbes, or fragments thereof. In some instances, the cell-free sample can be obtained by a sample preparation process, such as centrifuging, ultracentrifuging, or filtering a biological sample.
[0196] In some embodiments, a cfNA can comprise a host nucleic acid, a non-host nucleic acid, a target nucleic acid, or a combination thereof. In some embodiments, a cfNA can be derived from a host (host cell free nucleic acids or “hcfNA”) or a non-host. In some embodiments, a hcfNA can be derived from nuclear nucleic acids, mitochondria nucleic acids, exosomal nucleic acids, fetal nucleic acids, or any combination thereof. In some embodiments, aWSGR Docket No.: 47697-751601 sample can comprise a host nucleic acid (e.g., a host cell-free nucleic acid). In some embodiments, a host is any subject provided herein.
[0197] In some embodiments, the nucleic acids from a sample can be extracted and / or enriched to generate enriched nucleic acids. In some embodiments, the enriched nucleic acids comprise degraded nucleic acids, ultra-short nucleic acids, single stranded nucleic acids, double stranded nucleic acids, nicked nucleic acids, microbial cell-free nucleic acids (mcfNA), subject’s cell-free nucleic acids, circulating tumor nucleic acids (ctNA), mitochondrial nucleic acids (mtNA), or any combination thereof. In some embodiments, the methods comprise enriching for degraded nucleic acids, ultra-short nucleic acids, single stranded nucleic acids, or nicked double stranded nucleic acids. As used herein, “degraded nucleic acid” refers to fragments of DNA and RNA that are released into circulation due to cell death, turnover, or pathological processes. Degraded nucleic acids can originate from various sources, including natural sources or artificial sources. In some embodiments, the degraded nucleic acids originate from normal cellular apoptosis, necrosis, or disease-related processes such as cancer or infections. In some embodiments, the degraded nucleic acids originate from factors during laboratory handling, including, but not limited to, sample degradation due to prolonged storage, or temperature fluctuations. In some embodiments, ultrashort nucleic acids comprise nucleic acids less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 60 bp, less than 50 bp, less than 40 bp, or less than 30 bp in length. In some embodiments, at least 70%, 75%, 80%, 85%, 90%, or 95% of the mcfNA are degraded nucleic acids.
[0198] In some embodiments, a cfNA can comprise any nucleic acid that is not encapsulated by a cell (e.g., a eukaryotic or microbial cell). In some embodiments, a cfNA can originate from any nucleic acids. In some embodiments, a cfNA can comprise a plurality of chemical forms of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or DNA / RNA hybrid. In some embodiments, nucleic acids can comprise a plurality of structural forms of DNA, RNA, or DNA / RNA hybrid. In some embodiments, a cfNA can comprise linear nucleic acids or circular nucleic acids. In some embodiments, a cfNA can comprise single stranded nucleic acids (ssNA), double strand nucleic acids (dsNA) or hybrid nucleic acids. In some embodiments, nucleic acids can be from a genome of an organism or an organelle of a cell (e.g., an exosome or a mitochondria). In some embodiments, a cfNA can comprise a mitochondrial DNA, an intercellular signal nucleic acid, an exogenous nucleic acid, a DNA enzyme, a RNA enzyme, a food-derived nucleic acid, any metabolic form of nucleic acid-based therapeutic, or any combination thereof. In some embodiments, a cfNA can be derived from a member selected from the group consisting of genomic DNA, cDNA, mRNA, cRNA, tRNA, ribosomal RNA, miRNA,WSGR Docket No.: 47697-751601 siRNA, nuclear DNA, nuclear RNA, mitochondrial DNA, mitochondrial RNA, exosomal DNA, exosomal RNA, fetal DNA, fetal RNA, plasmids, vectors, and any combination thereof.
[0199] In some embodiments, nucleic acids can comprise a mixture of nucleic acids from various sources. In some embodiments, nucleic acids can be derived from a plurality of biological fluids. In some embodiments, nucleic acids can be from a plurality of organisms. In some embodiments, nucleic acids can be from a subject. In some embodiments, nucleic acids can be from one or more species of microbes. In some embodiments, nucleic acids can comprise environmental nucleic acids. In some embodiments, environmental nucleic acids can comprise any nucleic acid at or near a sample collection site, or any nucleic acid introduced by personnel, equipment or a reagent used in collecting and / or processing a sample from a subject.Sizes of cfNAs
[0200] In some embodiments, cfNAs or fragments thereof are less than about 10 bases, less than about 15 bases, less than about 20 bases, less than about 25 bases, less than about 30 bases, less than about 35 bases, less than about 40 bases, less than about 45 bases, less than about 50 bases, less than about 55 bases, less than about 60 bases, less than about 65 bases, less than about 70 bases, less than about 75 bases, less than about 80 bases, less than about 85 bases, less than about 90 bases, less than about 95 bases, less than about 100 bases, less than about 105 bases, less than about 110 bases, less than about 115 bases, less than about 120 bases, less than about 125 bases, less than about 130 bases, less than about 135 bases, less than about 140 bases, less than about 145 bases, less than about 150 bases, less than about 155 bases, less than about 160 bases, less than about 165 bases, less than about 170 bases, less than about 175 bases, less than about 180 bases, less than about 185 bases, less than about 190 bases, less than about 195 bases, or less than about 200 bases long. In some embodiments, the cfNAs are less than about 55 bases long. In some embodiments, the cfNA is less than about 60 bases long. In some embodiments, the cfNA is less than about 80 bases long.
[0201] In some embodiments, cfNAs are about 10 bases, about 15 bases, about 20 bases, about 25 bases, about 30 bases, about 35 bases, about 40 bases, about 45 bases, about 50 bases, about 55 bases, about 60 bases, about 65 bases, about 70 bases, about 75 bases, about 80 bases, about 85 bases, about 90 bases, about 95 bases, about 100 bases, about 105 bases, about 110 bases, about 115 bases, about 120 bases, about 125 bases, about 130 bases, about 135 bases, about 140 bases, about 145 bases, about 150 bases, about 155 bases, about 160 bases, about 165 bases, about 170 bases, about 175 bases, about 180 bases, bout 185 bases, about 190 bases, about 195 bases, or about 200 bases long.
[0202] In some embodiments, the cfNAs comprise ultra short cfNAs. As used herein, “ultra short nucleic acids” refer to subnucleosomal nucleic acids or fragments shorter thanWSGR Docket No.: 47697-751601 subnucleosomal nucleic acids. In some embodiments, ultra short cfNAs can be from about 10 bases to about 100 bases long. In some embodiments, the ultra short cfNAs can be from about 30 bases to about 80 bases long. In some embodiments, the ultra short cfNAs can be from about 40 bases to about 60 bases long. In some embodiments, the ultra short cfNAs are about 50 bases long.
[0203] In some embodiments, microbial cfNA (mcfNA) can be present at higher concentrations relative to host cfNA (hcfNA) at lengths that fall outside a nucleosomal interval. In some embodiments, mcfNA can be enriched relative to hcfNA by enriching for cfNA of less than 180 bases, less than 170 bases, less than 160 bases, less than 150 bases, less than 140 bases, less than 130 bases, less than 120 bases, less than 110 bases, less than 100 bases, less than 90 bases, less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, less than 30 bases, or less than 20 bases. In some embodiments, ultra short hcfNA can be enriched by enriching for cfNA of less than 180 bases, less than 170 bases, less than 160 bases, less than 150 bases, less than 140 bases, less than 130 bases, less than 120 bases, less than 110 bases, less than 100 bases, less than 90 bases, less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, less than 30 bases, or less than 20 bases.Definitions
[0204] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this disclosure belongs.
[0205] “A,” “an,” and “the”, as used herein, can include plural references unless expressly and unequivocally limited to one reference.
[0206] As used herein, the term “or” is used to refer to a nonexclusive “or”; as such, “A or B” includes “A but not B,” “B but not A,” and “A and B,” unless otherwise indicated.
[0207] As used throughout the specification herein, the term “about” when referring to a number or a numerical range means that the number or numerical range referred to is an approximation within experimental variability (or within statistical experimental error), and the number or numerical range can vary from, for example, from 10% to 25% of the stated number or numerical range. In examples, the term “about” refers to ±20% of a stated number or value. In other examples, the departure from equimolarity in the case of mixes intended to be equimolar, such as but not limited to, some control molecules in the spike-in mixes, is no more than a tenfold disparity, an eight-fold disparity, a six-fold disparity, a four-fold disparity or a two-fold disparity.
[0208] As used herein, “abundance” refers to the quantity of something, such as, for example, the quantity or number of molecules, such as nucleic acids. As used herein, “relative abundance” is the abundance of a molecule or molecules of interest per abundance of a referenceWSGR Docket No.: 47697-751601 molecule or molecules of interest. For example, relative abundance of target nucleic acid molecules (e.g., microbial cell-free nucleic acids refers to abundance per reference nucleic acids (e.g., host nucleic acids, synthetic nucleic acid added to the sample, etc.). As used herein, “absolute abundance” is the concentration or abundance of molecules per a defined unit of initial sample or sample quantity. For example, the absolute abundance of target nucleic acid molecules (e.g., microbial cell-free nucleic acids refers to the abundance per defined unit of sample quantity (e.g., sample volume, sample mass etc.). In some cases, the absolute abundance or concentration is measured by molecules per milliliter (MPM).
[0209] As used herein, “adapter” or “portions of an adapter” refers to a chemically synthesized, single-stranded, or double-stranded oligonucleotide that can be attached, e.g., covalently (e.g., ligation) or non-covalently (e.g., hybridization), to the ends of nucleic acid molecules, such as DNA or RNA molecules. Adapter can refer to either a full-length adapter or a portion of the adapter, e.g., partial adapters can be attached in some embodiments before the full lengths are introduced by e.g., indexing primers in amplification steps. 3 '-end adapters and 5'-end adapters can be full-length or a portion of an adapter sequence that are attached to the opposite ends of a target nucleic acid, a copy of a target nucleic acid, or a target nucleic acid complement. 3'-end adapters and 5'-end adapters sequences end up being attached to the opposite ends of e.g., a template that can be sequenced that comprises target nucleic acid, a copy of a target nucleic acid, and / or a target nucleic acid complement. The 3'-end adapter and 5'-end adapter sequences can be the same or they can be different. Adapter sequences can be of any length.
[0210] As used herein, “control” refers to a standard of comparison. A “negative control” refers to a standard of comparison that is used to identify contaminants from samples or to identify the nature of a signal in the absence of a sample. A “positive control” refers to a standard of comparison that is used to identify normal substances from an initial sample or sample. Some embodiments of the disclosure can comprise a positive and / or negative control. Some embodiments of the disclosure can comprise an initial sample or samples without a positive and / or negative control. Some embodiments of the disclosure can comprise an initial sample or samples without a positive control. Some embodiments of the disclosure can comprise an initial sample or samples without a negative control.
[0211] As used herein, “denaturing” refers to a process in which biomolecules, such as proteins or nucleic acids, lose their native or higher order structure. Native and higher order structure can include, for example, without limitation, quaternary structure, tertiary structure, or secondary structure. For example, a double-stranded nucleic acid molecule can be denatured into two single-stranded molecules.WSGR Docket No.: 47697-751601
[0212] As used herein, “detect” refers to quantitative or qualitative detection, including, without limitation, detection by identifying the presence, absence, quantity, frequency, concentration, sequence, form, structure, origin, or amount of an analyte.
[0213] As used herein, “removal” or “extraction,” and their cognates, of nucleic acids refers to steps prior to the start of generating or preparing a nucleic acid library that separate nucleic acids from at least one component with which they are normally associated. Removal or extraction of nucleic acids can refer to the process of creating an initial sample from a raw biological sample. For example, without limitation, the fractionation of whole blood into its component parts, such as plasma, can be considered to involve removal or extraction. Similarly, purification or isolation of DNA from a sample (e.g., plasma sample) can be considered extraction.
[0214] As used herein, “host” refers to an organism that harbors another organism or microbe. For example, a living thing e.g., a mammal such as a human being can be a host that harbors a microbe, the microbe being the non-host.
[0215] As used herein, a “nucleic acid library” refers to a collection of nucleic acid fragments. The collection of nucleic acid fragments can be used, for example, for sequencing.
[0216] As used herein, “pathogen” refers to a microbe that can cause a disease, ailment, or an infection.
[0217] As used herein, “microbe,” or “microbial,” generally refers to archaea, bacteria, fungi, protists, parasites, viruses, or other entities that are usually detectable using a microscope. As used herein, the term “microorganism” refers to a uni- or multi- cellular organism, such as, for example, a microscopic organism or macroscopic organism including but not limited to bacteria, fungi, protists, and parasites. Microbes herein can be a prokaryote or a eukaryote. Microbes are often pathogens responsible for disease, but can also exist in a non-pathogenic, symbiotic, commensalistic, mutualistic, or amensalistic relationship with a host, such as a human.
[0218] As used herein, “plasma” or “blood plasma” refers to the liquid component or fraction of blood. Plasma is generally obtained by spinning a whole blood sample and removing the liquid component.
[0219] As used herein, the phrase “process control molecules” refers to molecules that are added to a sample before or during nucleic acid library generation to aid in the identification or quantification of nucleic acids in a sample. Process control molecules are separate from and not integrated in the target molecules, such as nucleic acids. Process control molecules can have special features such as specific sequences, lengths, GC content, degrees of degeneracy, degrees of diversity, different secondary, tertiary, or quaternary structures, and / or known starting concentrations. Process control molecules can be used for normalizing the signal in a sample inWSGR Docket No.: 47697-751601 order to account for variations in sample processing or to control process performance. Process control molecules can include, for example, without limitation, ID Spike(s), Spanks, and / or Sparks or GC Spike-in Panel molecules. Process control molecules can additionally include dephosphorylation control molecules, denaturation control molecules, and / or ligation control molecules. By “adapter attachment control molecule” is intended a control molecule that allows monitoring of the efficiency of an adapter attachment reaction be it ligation-based, TdT-based, template-switching-based, primer- extension-based, or amplification-based. By “degradation assessment molecules” is intended a control molecule used to evaluate sample and spiked sample integrity during processing.
[0220] The term “sequencing,” as used herein, generally refers to methods and technologies for determining the sequence of nucleotide bases in one or more polynucleotides. Sequencing can involve basic methods including Maxam-Gilbert sequencing and chaintermination methods, or de nova sequencing methods including shotgun sequencing and bridge PCR, or next-generation sequencing (NGS) methods (or massively-parallel sequencing method) including but not limited to polony sequencing, 454 pyrosequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent semiconductor sequencing, Heli Scope single molecule sequencing, SMRT® sequencing, nanopore sequencing and others. Sequencing can be performed by various systems currently available, such as, without limitation, a sequencing system by Illumina, Pacific Biosciences, Oxford Nanopore, Genia Technologies, or Life Technologies (Ion Torrent) and others. Such devices can provide a plurality of raw genetic data corresponding to the genetic information of a host (e.g., human), a non-host (e.g., a pathogen, an organ donor), a host-derived variant genetic sequence (e.g., a single nucleotide polymorphism), and / or combinations thereof as generated by the device from a sample provided by the subject.
[0221] As used herein, “Spanks” refers to degenerate pools, or pools of nucleic acids with diverse sequences, which degenerate pools can often be used for diversity assessment, abundance calculation, and / or determination of information transfer efficiency (See, for example, United States patent 9,976,181).
[0222] As used herein, “Sparks” “GC Spike-in Panel” or “GC dSPARKS” refers to nucleic acids that are size or length or GC-content markers, which can be used for abundance normalization, development, and / or analysis purposes and other purposes (See, for example, United States patent 9,976,181).
[0223] As used herein, “ID Spike(s)” refers to identification spikes that can be used, for example without limitation, for sample identification tracking, cross-contamination detection, reagent tracking, and / or reagent lot tracking (See, for example, United States patent 9,976,181).WSGR Docket No.: 47697-751601
[0224] As used herein, “sample” refers to any material comprising nucleic acids that has been derived from a subject described herein. A sample may comprise a raw biological sample, such as whole blood. A sample may have been processed or manipulated, such as plasma or serum. As used herein, the phrase “raw biological sample” refers to an unmanipulated sample obtained from a subject, e.g., host, containing or presumed to contain target nucleic acids. In other words, a raw biological sample, once obtained from the subject, has not been subjected to any extraction methods, e.g., alcohol-based extraction, size separation, etc., needed to generate an initial sample. The raw biological sample can be a sample for the sequencing assays described herein if no manipulation of the raw biological sample is needed, e.g., whole blood, to obtain the target nucleic acids. A raw biological sample can also be manipulated, such as, for example to create a fraction of whole blood (e.g., plasma, serum, etc.) to yield a sample for the sequencing assays described herein.
[0225] As used herein, the term “initial sample” refers to a sample comprising nucleic acids derived from a raw biological sample. An initial sample, for example, can comprise target or desired nucleic acids obtained or extracted from a raw biological sample. In some cases, the sample for sequencing assays described herein comprise the initial samples. In some cases, “sample” and “initial sample” can be used interchangeably.
[0226] As used herein, the phrase “spiked initial sample” refers to an initial sample to which process control molecules (or synthetic spike-ins) have been added prior to the start of generating a sequencing library.
[0227] The term “derived from” encompasses the terms “originated from,” “obtained from,” “obtainable from” and “created from,” and generally indicates that one specified material finds its origin in another specified material or has features that can be described with reference to the specified material. For example, a sample can be derived from a blood draw.
[0228] As used herein, the turn of phrase “uniformly distributed” refers to a distribution that is continuous or uniform between members of a family such that for each member of a family there is a predictable or symmetric interval between them. The term “non-uniformly distributed” refers to a distribution of members of a family that does not have a predictable or symmetric interval between them.Examples of Machine Learning Methodologies
[0229] In some embodiments, machine learning (ML) may be applied to the methods and the systems disclosed herein. For example, ML may be used in predicting a risk of a certain medical condition in a user. Further, for example, ML may be used in predicting symptoms in a user. Further, for example, ML may be used in predicting an effective treatment for a user. ToWSGR Docket No.: 47697-751601 accomplish these example predictions, ML may analyze one or more of: research data, genealogical data, medical data, demographic data, geographic data, assay data, etc.
[0230] In some cases, ML may generally involve identifying and recognizing patterns in existing data in order to facilitate making predictions for subsequent data. ML may include a ML model (which may include, for example, a ML algorithm). Machine learning, whether analytical or statistical in nature, may provide deductive or abductive inference based on real or simulated data. The ML model may be a trained model. ML techniques may comprise one or more supervised, semi-supervised, self-supervised, or unsupervised ML techniques. For example, an ML model may be a trained model that is trained through supervised learning (e.g., various parameters are determined as weights or scaling factors). ML may comprise one or more of regression analysis, regularization, classification, dimensionality reduction, ensemble learning, meta learning, association rule learning, cluster analysis, anomaly detection, deep learning, or ultra-deep learning. ML may comprise: k-means, k-means clustering, k-nearest neighbors, learning vector quantization, linear regression, non-linear regression, least squares regression, partial least squares regression, logistic regression, stepwise regression, multivariate adaptive regression splines, ridge regression, principal component regression, least absolute shrinkage and selection operation (LASSO), least angle regression, canonical correlation analysis, factor analysis, independent component analysis, linear discriminant analysis, multidimensional scaling, non-negative matrix factorization, principal components analysis, principal coordinates analysis, projection pursuit, Sammon mapping, t-distributed stochastic neighbor embedding, AdaBoosting, boosting, gradient boosting, bootstrap aggregation, ensemble averaging, decision trees, conditional decision trees, boosted decision trees, gradient boosted decision trees, random forests, stacked generalization, Bayesian networks, Bayesian belief networks, naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, hidden Markov models, hierarchical hidden Markov models, support vector machines, encoders, decoders, auto-encoders, stacked autoencoders, perceptrons, multi-layer perceptrons, artificial neural networks, feedforward neural networks, convolutional neural networks, recurrent neural networks, residual neural networks, physics-informed neural networks, long short-term memory, deep belief networks, deep Boltzmann machines, deep convolutional neural networks, deep recurrent neural networks, large language models, transformer models, vision transformers, or generative adversarial networks.
[0231] Training the ML model may include, in some cases, selecting one or more untrained data models to train using a training data set. The selected untrained data models may include any type of untrained ML models for supervised, semi-supervised, self-supervised, or unsupervised machine learning. The selected untrained data models may be specified based upon input (e.g., user input) specifying relevant parameters to use as predicted variables or otherWSGR Docket No.: 47697-751601 variables to use as potential explanatory variables. For example, the selected untrained data models may be specified to generate an output (e.g., a prediction) based upon the input. Conditions for training the ML model from the selected untrained data models may likewise be selected, such as limits on the ML model complexity or limits on the ML model refinement past a certain point. The ML model may be trained (e.g., via a computer system such as a server) using the training data set. In some cases, a first subset of the training data set may be selected to train the ML model. The selected untrained data models may then be trained on the first subset of training data set using appropriate ML techniques, based upon the type of ML model selected and any conditions specified for training the ML model. In some cases, due to the processing power requirements of training the ML model, the selected untrained data models may be trained using additional computing resources (e.g., cloud computing resources). Such training may continue, in some cases, until at least one aspect of the ML model is validated and meets selection criteria to be used as a predictive model.
[0232] In some cases, one or more aspects of the ML model may be validated using a second subset of the training data set (e.g., distinct from the first subset of the training data set) to determine accuracy and robustness of the ML model. Such validation may include applying the ML model to the second subset of the training data set to make predictions derived from the second subset of the training data. The ML model may then be evaluated to determine whether performance is sufficient based upon the derived predictions. The sufficiency criteria applied to the ML model may vary depending upon the size of the training data set available for training, the performance of previous iterations of trained models, or user-specified performance requirements. If the ML model does not achieve sufficient performance, additional training may be performed. Additional training may include refinement of the ML model or retraining on a different first subset of the training dataset, after which the new ML model may again be validated and assessed. When the ML model has achieved sufficient performance, in some cases, the ML may be stored for present or future use. The ML model may be stored as sets of parameter values or weights for analysis of further input (e.g., further relevant parameters to use as further predicted variables, further explanatory variables, further user interaction data, etc.), which may also include analysis logic or indications of model validity in some instances. In some cases, a plurality of ML models may be stored for generating predictions under different sets of input data conditions. In some cases, the ML model may be stored in a database (e.g., associated with a server).
[0233] In some embodiments, the methods and the systems disclosed herein improve predictive power of ML by using a combination of features to more robustly form predictions. For example, the first feature may be an expanded training set to train a neural network. ThisWSGR Docket No.: 47697-751601 expanded training set may be developed by applying mathematical transformation functions on an acquired set of training data. These transformations can include affine transformations, for example, rotating, shifting, or mirroring or filtering transformations, for example, smoothing or contrast reduction. The neural networks may then be trained with this expanded training set (e.g., using stochastic learning with backpropagation, which is a type of machine learning algorithm that uses the gradient of a mathematical loss function to adjust the weights of the network). The introduction of an expanded training set may however increase false positives when classifying data outside the training data. Accordingly, the second feature of the methods and the systems disclosed herein is the minimization of these false positives by performing an iterative training algorithm, in which the ML model may be retrained with an updated training set containing the false positives produced after prediction has been performed on a set of non-training data. This combination of features provides a robust prediction model with limited false positives. Accordingly, training the ML model may comprise: (A) obtaining a first set of training data; (B) training a machine learning model on the first set of training data; (C) obtaining a second set of training data; and (D) training the machine learning model on the second set of training data. For example, the second training data may comprise or correspond to at least some of the first training data. Further, for example, one or both of the first set of training data or the second set of training data may be transformed, such as by one or more operations: mirroring, rotating, smoothing, filtering, shifting, thresholding, contrast adjusting, etc.Data Management
[0234] In some embodiments, the methods and the systems disclosed herein may incorporate one or more data management techniques. For example, these techniques may include standardizing data across different patients, different studies, different data types, etc. Further, for example, these techniques may include sharing data (e.g., automatically) with different users in a network. These users may include, for example, the patient, a care provider (e.g., a doctor, a nurse, a parent, etc.), a pharmacy, a research study, etc. The data may, for example, be shared real-time (or near real-time) across these users.
[0235] Patients may often visit various care providers (e.g., doctors, nurses, etc.), pharmacies, researchers, etc. for diagnosis and treatment. It may be difficult for all these different individuals / entities to share updated information about a patient’s condition with each providers using patient management systems, due to, for example, issues with inconsistent data formatting, data types, data sizes, etc. This can lead to problems with managing treatments, research studies, prescriptions or having patients duplicate tests, for example.
[0236] Further, individuals / entities often continually monitor a patient’s medical records (e.g., biological, genealogical, phenological, demographic, etc.) for updated information, whichWSGR Docket No.: 47697-751601 is often-times incomplete as records across different individuals / entities are often not shared timely in useful data types, formats, sizes, etc.
[0237] To address these challenges, the methods and the systems disclosed herein provide a network-based patient management that collects, converts, and consolidates patient information from various care providers, research studies, pharmacies, etc. into a standardized format for network-based storage and sharing.
[0238] In some embodiments, the methods and the systems disclosed herein provide a graphical user interface (GUI) by a content server, which is hardware or a combination of both hardware and software. A user, such as a care provider, a pharmacy, a researcher, or a patient, is given remote access through the GUI to view or update information about a patient’s medical condition using the user’s own local device (e.g., a personal computer or wireless handheld device). When a user wants to update the records, the user can input the update in any format used by the user’ s local device. Whenever the patient information is updated, it may be converted into the standardized format and then stored in a collection of medical records on one or more of the network-based storage devices. After the updated information about the patient’s condition has been stored in the collection, the content server, which is connected to the network-based storage devices, a message may be generated containing the updated information about the patient’s condition. This message may be transmitted in a standardized format over the computer network to users (e.g., care providers) in a network that have access to the patient’s information (e.g., to a care provider the updated information about the patient’s medical condition) so that the users can quickly be notified of any changes (e.g., without having to manually look up or consolidate all of the updates of patient condition). This helps ensure each of a group of care providers is provided real-time notice and access to changes so they can readily adapt their own medical diagnostic and treatment strategy in accordance with other care providers’ actions. The message can be in the form of an email message, text message, a push notification, etc.
[0239] Accordingly, in some embodiments, data management may comprise: (A) obtaining, from a first source, first data corresponding to a patient condition for a patient; (B) obtaining, from a second source, second data corresponding to a patient condition for the patient; (C) standardizing the first data and the second data to generate standardized data; (D) generating (e.g., automatically) a message corresponding to the standardized data; and (E) transmitting the message to a plurality of users (e.g., care providers of the patient, pharmacists, researchers, the patient, etc.) over a computer network (e.g., in real time), so that each user of the plurality of users has access to the standardized data. For example, the first data and the second data may be of a different data format, different data size, different data type, etc.WSGR Docket No.: 47697-751601
[0240] In some embodiments, the methods and the systems disclosed herein help reduce data size. For example, this data size may correspond to data used in assays, research, medical, etc. contexts. In reducing data size, the methods and the systems disclosed herein may improve efficiency. This improved efficiency may help reduce the amounts of physical resources consumed, e.g., chemicals, devices, testing kits, etc. Further, this improved efficiency may help reduce computing resources, e.g., electricity consumption, heat loss, computing power etc.
[0241] Still further advantageously, improved efficiency from the methods and the systems disclosed herein may help reduce traffic over a network by reducing data size. Accordingly, the methods and the systems disclosed herein may be useful in optimizing network performance, resolving network issues, and improving network security through reducing network traffic and burden on the network.
[0242] In some embodiments, the methods and the systems disclosed herein help reduce network traffic. In some embodiments, the methods and the systems disclosed herein help reduce network traffic by varying (e.g., reducing) the amount of network data. For example, the methods and the systems disclosed herein may reduce data based at least in part on a threshold. This threshold may correspond too a risk threshold (e.g., a risk of a patient having a medical condition), an accuracy threshold (e.g., an accuracy threshold needed for a conclusion in a research study, assay, test, etc.), or a data threshold (e.g., a data size threshold, a network performance threshold, etc.). Accordingly, the methods and the systems disclosed herein may perform operations comprising: (A) collecting (e.g., over a network) data corresponding to a patient; (B) analyzing the data with respect to a threshold (e.g., risk threshold, accuracy threshold, data threshold, etc.); and (C) transmitting a subset of the data over the network, the subset of the data determined based at least in part on the threshold.EXAMPLES
[0243] The following examples are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the disclosure; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.
[0244] Described herein are methods and systems for bioinformatically identifying and / or eliminating contaminant nucleic acids in samples obtained from one or more blood draws from a subject (e.g., a pig or a monkey). The blood draws herein are taken from a same collection site of the subject. By comparing sequencing reads generated from nucleic acids that originated from the one or more blood draws, contaminant nucleic acids present at the sample collection site can be identified and discarded if desired. Moreover, identifying and / or eliminating the contaminant nucleic acids from the sequencing reads provides a higher signal-to-background ofWSGR Docket No.: 47697-751601 the sequencing assay, which provides more accurate methods for characterizing the target nucleic acids from the subject. Therefore, the methods provided herein may enhance detection of target nucleic acids (e.g., microbial cell-free nucleic acids).Example 1: Analyzing microbial cell-free nucleic acids (mcfNA) from samples collected from Yucatan porcine donors.
[0245] Yucatan miniature pigs served as donors for this experiment. At the time of the blood collection, at least one pig had high white blood cell counts and / or was Streptococcus suis (S. suis) positive, a pathogen in pigs that can cause severe systemic infection in humans. These porcine donors were excluded from a main xenotransplant study; however, blood samples were collected for the present study on mcfDNA sequencing.
[0246] The sample collection site on the pig can be sterilized with alcohol swabs prior to vein puncture by a phlebotomist. In some cases, the body sample collection site was shaved and sterilized. In some cases, the animals were sedated. Three whole blood samples were collected sequentially from the same site on each pig via one needle puncture in multiple BD Vacutainer tubes. After plasma had been separated from cells, the samples were shipped to the lab on dry ice overnight. Upon receipt, control molecules for the next sequencing steps were added to the samples and next generation sequencing assays were conducted. In some cases, controls for sequencing bias, metagenomic sequencing quality, and sample mix-ups were added to the sample prior to next generation sequencing. Porcine decoy Sequencing data were processed using a proprietary analytical pipeline and microbial reads were aligned to a database comprising curated assemblies from many species and taxa including bacteria, viruses, fungi, and parasites. To test the effectiveness of the porcine decoy, a commercial database of Yucatan miniature pigs was used for sequencing alignment and eliminating host cross-reactivity. Overall, the risk for crossreactivity (XR) is very low. The host filter removes approximately 92% of the Yucatan porcine reads on average in the 12 pilot samples. This is less than the 98% filtering sensitivity of the human filter in commercial samples. Simulated 28 bp and 56 bp reads from the Yucatan reference genome were used to determine the host XR levels. The host XR levels were approximately 0.2 effective data rate (EDR) / 1M porcine reads (not including the porcine type-C oncovirus). The Porcine decoy refers to genome assembly "Sscrofal l.l" deposited under accession GCF 000003025.6 at NCBI. Reads were assigned to taxa using a mixture model and EM algorithm and their abundance was estimated. Statistical significance values were computed for each taxon with its estimated abundance.
[0247] Candidate microbial detections refer to taxa determined to be at high significance relative to a statistical model designed to control for reagent contamination. Additional analytical filters were applied to account for read uniformity from the attributed genome, read percentWSGR Docket No.: 47697-751601 identity and cross-reactivity caused by pathogen genome homology. Candidate detections include species at high significance levels above background comprised after additional filtering was applied, which accounted for read location uniformity, read percent identity, and crossreactivity originating from higher abundance detections. The microorganism detections that passed these filters were reported along with abundances in molecules / pL (MPM) in plasma.
[0248] As shown by FIGs. 1A-1B, the pilot porcine samples were rich in a plurality of taxonomically distinct microbes. In pigs, the median report indicated 10 microbial detections, with the minimum being 2 detections. In contrast, for the healthy human cohort, approximately 80% of individuals had no microbial detections, and the majority of reports showed only one detection. FIG. 2 depicts a panel of observed taxa. Ubiquitous observations detected include E. coli (present in all samples), S. suis, and H. parasuis (approximately 50 MPM). Singleton observations detected include Treponema pedis (from subject ID 23013) and Mycoplasma hyosynoviae (from subject ID 21456). The preliminary sequencing results confirmed the clinical diagnosis of S. suis infections in the porcine donors, as depicted in FIG. 3.
[0249] As depicted in FIGs. 4A-4E, concentrations of total microbial cell-free nucleic acids (mcfDNA) in the samples decreased from the first blood draw to the third blood draw, with the exception of subject ID 22135, which appeared to be an outlier. FIG. 5A depicts normalized MPM of mcfDNA for different kingdoms. Abundance of mcfDNA (in MPM) in the third blood draw was normalized to abundance of mcfDNA in the first blood draw. FIG. 5B depicts MPM and normalized MPM of Staphlococcus epidermidis mcfDNA. Overall, the first blood draws showed higher numbers of significant observations, as depicted in FIGs. 6A-6B. Subject ID A4564 and A4580 had many more calls only in the first draw. The total MPM concentration dropped the most from first to third draw in these two subjects. Subject 22207 is the one with the third largest drop in MPM concentration. It has six calls made only in the first draw and 8 calls made in both draws.
[0250] Subject 22206 has only two calls that were made in both draws and seven calls that are made exclusively in the first and third draw. Subject 22129 has most calls shared by both draws and one unique call per draw. Subject 22135 has no unique calls in the first draw, three calls that are shared, and three calls that were made only in the third draw. This is also the subject for which the total MPM concentration increased from first to third draw.
[0251] As an approximation for porcine mcfDNA concentration between draws, the mcfDNA concentration of porcine type-c oncovirus was used as a control. FIGs. 7A-7B show that mcfDNA concentrations remained relatively stable between the blood draws in all subjects. Pairwise correlation between MPMs of microbes of all samples were computed as follows: 1) obtained abundance of all microbes for each pair of samples across all draws of a given subject;WSGR Docket No.: 47697-7516012) select microbes: remove microbes that fail uniformity in both samples; remove microbes that have higher p-value than a certain threshold in both samples and that are not called in any sample; remove microbes that have less than 3 EDR in both samples; 3) compute the Pearson correlation: highlight values with p-value <= 0.0003; starting from p-value 0.05, dividing value by 144 ( number of comparisons performed). Correlations for different p-values were shown in FIGs. 8A-8D (“*” in a cell means that p-value is <= 0.0003). Overall, first and third draw of same subject are more correlated than other draws.
[0252] Initial analysis revealed 116 species with reduced mcfDNA concentrations from the first to the third blood draw, and S. suis was the only taxon that was called. Bioinformatic analyses were performed to investigate the species of microbes that contributed to the decrease of the total mcfDNA across successive draws (FIGs. 9A-9C). The top hits driving the decreased mcfDNA concentration included Lactobacillus amylovorus, Corynebacterium xerosis, Acinetobacter gandensis, Acinetobacter Iwoffii, and Aerococcus urinaeequi (FIG. 9B, right), as well as Propionibacterium acnes and Raoultella planticola (FIG. 9C, right). Unintentional sample swapping between the first and the third blood draws likely explains why samples obtained from subject ID 22135 behaved differently from the rest (FIG. 10). Bioinformatic analysis was performed to track the changes in mcfDNA concentration between blood draws for each subject, as depicted in FIG. 11. The signature observed in Subject ID 22206 was unique, with unique calls from 7 bacterial species in the third blood draw that were not observed in the first draw. The calls from Subject ID 22206 were compared to over 15, 000 environment control (EC) samples but only 2 out of the 7 species were observed at higher concentration in the EC samples (FIG. 12A). Some taxa from 22206 were also present at high abundance in 15,852 EC samples sequencing with over 25,000 of synthetic spike-in molecules ddSPANK (FIG. 12B). FIG. 12C is a bar graph showing the number of taxa of interest per EC sample. 477 EC samples had sequencing reads from at least 4 of the 7 species from the third draw. In these 477 EC samples, higher concentrations of mcfDNA were recorded from the 7 species that were called only in the third draw of Subject ID 22206 (FIG. 12D). There were 9 EC samples that showed mcfDNA concentrations of Staphylococcus hominis higher than 400 MPM. MPM distribution of taxa in these 9 EC samples are shown in FIG. 12E. Based on the data (FIGs. 12A-12E), it is highly unlikely that the taxa of the unique calls from the third draw of the 22206 sample were introduced as environment contamination at the lab.
[0253] TABLE 1 shows the taxa or genera of microbes that were called only in the first draws of the six subjects. The “count” column shows the number of subjects in whose samples the calls were made. In the environment column, “x” means no environment was found and means the corresponding taxa was not looked up. Thirty-three environmental microbes wereWSGR Docket No.: 47697-751601 called. 25 of these (75.8%) might come from an GI, skin, or soil environment, suggesting that they could be environmental contamination originating from the site of collection (TABLE 1 and FIG. 13).TABLE 1 Microbes called only in the first draw of the six porcine subjectsWSGR Docket No.: 47697-751601
[0254] TABLE 2 shows the taxa or genera of microbes that were called in both the first and third draws. Burkholderia cepacia complex was called in 5 porcine subjects. Corynebacterium xerosis and Pseudomonas aeruginosa were called in 6 subjects.
[0255] Results from this experiment showed that it is possible to identify contaminant microbes from the environment by analyzing the sequencing data obtained from sequencing assays of mcfDNA in a sample, wherein the sample has been obtained from at least two blood draws in a series of blood draws from a single collection site on an animal. The method described herein can identify microbes from the environmental contaminants in a sequencing assay, including contaminants from soil, the skin or GI tract of the animal donor.Example 2: Analyzing microbial cell-free nucleic acids (mcfNA) from samples collected from cynomolgus monkeys.
[0256] All animal care, surgical procedures and postoperative care of animals were conducted in accordance with NIH Guidelines for the care and use of primates. Cynomolgus macaques (Macaca fascicularis) and baboons served as donors for this experiment. The sample collection site on each non-human primates (NHP) was sterilized with alcohol swabs prior toWSGR Docket No.: 47697-751601 vein puncture by a phlebotomist. Whole blood samples were collected using BD Vacutainer plasma preparation tubes. After plasma had been separated from cells, the samples were shipped to the lab on dry ice overnight. Upon receipt, controls molecules for the next sequencing steps were added to the samples and next generation sequencing assays were conducted. In some cases, controls for sequencing bias, metagenomic sequencing quality, and sample mix-ups were added to the sample prior to next generation sequencing. The blood samples collected from the NHP can be included in a single sequencing run with samples from other animals (e.g., the porcine samples).
[0257] Baboon genome Panubisl .l and cynomolgus macaques’ genome Mfascicularis_CE1976F were used to filter host reads and assess risk for monkey cross-reactivity (XR). Simulated reads were processed initially to estimate the risk for monkey XR. Based on simulations from nine cfDNA samples obtained from cynomolgus macaques, the pipeline had an average filter sensitivity of 92.4%. Based on simulated 56 bp reads from two baboon and two cynomolgus macaques’ genomes, it was estimated that the host-XR levels to be approximately 0.07 EDR / IM monkey reads. Based on simulated 28 bp reads from two baboon and two cynomolgus macaques’ genomes, it was estimated that the host-XR levels to be approximately 2.9 EDR / IM monkey reads. It was estimated that the real read length distributions will be dominated by longer reads, resulting in less host-XR potential.
[0258] While preferred embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the present disclosure may be employed in practicing the present disclosure. It is intended that the following claims define the scope of the present disclosure and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
WSGR Docket No.: 47697-751601CLAIMSWHAT IS CLAIMED IS:
1. A method comprising performing a sequencing assay on a sample from a subject to generate sequencing reads, wherein:(a) the sample from the subject:(i) comprises a nucleic acid, and(ii) is from a later blood draw in a series of blood draws; and(b) the series of blood draws is taken from a same sample collection site of the subject.
2. The method of claim 1, wherein the series of blood draws comprises at least two, at least three, or at least four blood draws.
3. The method of claim 1 or 2, wherein the series of blood draws comprises two blood draws.
4. The method of any one of the preceding claims, wherein the same sample collection site comprises a same area of skin on the subject, a same blood vessel of the subject, or an access to a same blood vessel of the subject.
5. The method of any one of the preceding claims, wherein the series of blood draws was taken using a vacuum blood collection system.
6. The method of any one of the preceding claims, wherein the series of blood draws was taken using one needle.
7. The method of any one of the preceding claims, wherein the series of blood draws was taken from a same vein of the subject.
8. The method of any one of the preceding claims, wherein the series of blood draws was taken from a single puncture to the vein of the subject.
9. The method of any one of the preceding claims, wherein the collection site of the subject comprises a vein or a venous access of the subject.
10. The method of claim 9, wherein the venous access of the subject comprises skin above the vein or a catheter.
11. The method of claim 9 or 10, wherein the vein comprises a jugular vein or an ear vein.
12. The method of any one of the preceding claims, further comprising discarding a sequencing read generated from the later blood draw.WSGR Docket No.: 47697-75160113. The method of claim 12, wherein the discarding the sequencing read is performed based in part on the following criteria being met:(a) the sequencing read is assigned to a single species of microbe, and(b) the abundance of sequencing reads generated from the later blood draw assigned to the single species of microbe is lower than the abundance of sequencing reads generated from an earlier blood draw assigned to the same species of microbe.
14. The method of claim 13, wherein the abundance of sequencing reads generated from the later blood draw assigned to the single species of microbe is at least 1 molecule per milliliter (MPM), at least 2 MPM, at least 3 MPM, at least 4 MPM, at least 5 MPM, at least 10 MPM, at least 15 MPM, or at least 20 MPM lower than the abundance of sequencing reads generated from the earlier blood draw assigned to the same species of microbe.
15. The method of claim 13 or 14, wherein the abundance of sequencing reads generated from the later blood draw assigned to the single species of microbe is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, or at least 60% lower than the abundance of sequencing reads generated from the earlier blood draw assigned to the same species of microbe.
16. The method of any one of the preceding claims, further comprising comparing an abundance of sequencing reads generated from the later blood draw to an abundance of sequencing reads generated from the earlier blood draw.
17. The method of any one of the preceding claims, further comprising calculating a variance of the abundance of the sequencing reads.
18. The method of claim 17, comprising comparing a variance of abundance of the sequencing reads from the later blood draw to a variance of the abundance of sequencing reads from the earlier blood draw.
19. The method of claim 17 or 18, wherein the variance of sequencing reads from the later blood draw is lower than the variance of sequencing reads from the earlier blood draw.
20. The method of any one of the preceding claims, further comprising identifying a contaminant nucleic acid sequence, at least in part based on a change in an abundance of the contaminant nucleic acid sequence from the earlier blood draw to the later blood draw.
21. The method of any one of the preceding claims, further comprising identifying a contaminant nucleic acid sequence, at least in part based on a reduced abundance of the contaminant nucleic acid sequence from the earlier blood draw to the later blood draw.
22. The method of any one of the preceding claims, further comprising identifying a contaminant nucleic acid sequence by comparing an abundance of the contaminant nucleic acid to an abundance of a control nucleic acid applied to the collection site.WSGR Docket No.: 47697-75160123. The method of any one of the preceding claims, further comprising identifying a contaminant microbe based on the contaminant nucleic acid sequence, wherein the contaminant microbe comprises a commensal microbe from the subject or a microbe of the subject’s general environment.
24. The method of any one of the preceding claims, further comprising identifying a contaminant nucleic acid sequence from a soil microbe, a gastrointestinal tract microbe, or a skin microbe.
25. The method of any one of the preceding claims, further comprising identifying a contaminant nucleic acid sequence from one or more Staphylococcus spp., Escherichia coli, Salmonella spp., Pseudomonas spp., Clostridium spp., Candida spp., Aspergillus spp., Penicillium spp., Rodent Parvovirus, Mouse Hepatitis Virus (MHV), Sendai Virus, helminths (e.g., Syphacia obvelata), protozoa (e.g., Giardia spp.), various soil microorganisms, or any combination thereof.
26. The method of any one of the preceding claims, wherein the series of blood draws comprises blood draws taken sequentially from the same collection site of the subject.
27. The method of any one of the preceding claims, wherein the later blood draw was taken at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at least about 45 seconds, or at least about 1 minute later than the earlier blood draw in the series of blood draws.
28. The method of any one of the preceding claims, wherein the later blood draw was taken at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at least about 45 seconds, at least about 1 minute, at least about 2 minutes, at least about 3 minutes, at least about 4 minutes, or at least about 5 minutes later than the first blood draw in the series of blood draws.
29. The method of any one of the preceding claims, wherein each blood draw in the series of blood draws was taken at intervals of at least about 1 second, at least about 5 seconds, at least about 10 seconds, at least about 30 seconds, at least about 45 seconds, or at least about 1 minute.
30. The method of any one of the preceding claims, wherein the series of blood draws was taken over the period of about 5 seconds, about 10 seconds, about 30 seconds, about 45 seconds, about 1 minute, about 1.5 minutes, about 2 minutes, about 2.5 minutes, about 2.5 minutes, about 3 minutes, about 3.5 minutes, about 4 minutes, about 4.5 minutes, or about 5 minutes.
31. The method of any one of the preceding claims, wherein the subject is an animal.WSGR Docket No.: 47697-75160132. The method of claim 31, wherein the animal comprises a mammal.
33. The method of claim 32, wherein the mammal comprises a canine, a swine, a feline, or a nonhuman primate.
34. The method of claim 31, wherein the subject is a human.
35. The method of any one of the preceding claims, further comprising preparing the collection site prior to taking the series of blood draws from the subject.
36. The method of any one of the preceding claims, wherein preparing the collection site comprises shaving, scrubbing, or cleaning.
37. The method of any one of the preceding claims, wherein the sequencing reads comprise sequencing reads generated from microbial cell free nucleic acids (mcfNA) in the sample.
38. The method of any one of the preceding claims, wherein the sequencing reads comprise at most 50, at most 100, at most 150, at most 200, at most 250, at most 300, at most 350, at most 400, at most 450, or at most 500 sequencing reads generated from the mcfNA in the sample.
39. The method of any one of the preceding claims, wherein the sequencing reads comprise at least 500, at least 1500, at least 2500, or at least 3500 sequencing reads generated from the mcfNA in the sample.
40. The method of any one of the preceding claims, further comprising applying a control nucleic acid to the same collection site of the subject prior to collecting the series of blood draw.
41. The method of claim 40, wherein the sequencing reads generated from the nucleic acids in the sample comprise a sequencing read generated from the control nucleic acid applied to the collection site.
42. The method of any one of the preceding claims, further comprising mapping a sequencing read to a reference sequence.
43. The method of claim 42, wherein the reference sequence comprises a region of a microbial genome.
44. The method of any one of the preceding claims, further comprising assigning a sequencing read to a microbial sequence.
45. The method of any one of the preceding claims, further comprising adding a known amount of synthetic spike-in molecules to the sample and generating sequencing reads from the synthetic spike-in molecules.
46. The method of any one of the preceding claims, further comprising calculating an abundance of a sequencing read generated from the nucleic acids in the sample.WSGR Docket No.: 47697-75160147. The method of any one of the preceding claims, wherein the abundance of the sequencing reads comprises an absolute abundance.
48. The method of any one of the preceding claims, wherein the abundance of the sequencing reads comprises a relative abundance or a normalized abundance.
49. The method of any one of the preceding claims, further comprising calculating the relative abundance of a sequencing read at least in part by comparing the abundance of the sequencing read to the abundance of another sequencing read.
50. The method of any one of claims 45 to 49, further comprising calculating the normalized abundance of a sequencing read at least in part by comparing the abundance of the sequencing read to the abundance of sequencing reads generated from the synthetic spike-in molecules in the sample.
51. The method of any one of the preceding claims, wherein the method increases the signal-to-background ratio of the sequencing assay.
52. The method of any one of the preceding claims, further comprising generating a report of the species of microbes identified in the sample.
53. The method of any one of the preceding claims, wherein the microbe comprises a virus, a bacterium, a protozoa, or a fungus.
54. The method of any one of the preceding claims, wherein the nucleic acid comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
55. The method of any one of the preceding claims, wherein the nucleic acid comprises a cell-free nucleic acid (cfNA).
56. The method of any one of the preceding claims, wherein the nucleic acid comprises a microbial cell-free nucleic acid (mcfNA).
57. The method of any one of the preceding claims, wherein the sample comprises mcfNA from at least one, at least two, or at least three blood draws in the series of blood draws.
58. The method of any one of the preceding claims, wherein the sample comprises mcfNA from the later blood draws in the series of blood draws.
59. The method of any one of the preceding claims, wherein the sample comprises mcfNA from at least one, at least two, or at least three microbial species.
60. The method of any one of the preceding claims, wherein the sample comprises whole blood or plasma.
61. The method of any one of the preceding claims, wherein the sequencing assay comprises a high throughput sequencing assay.
62. The method of any one of the preceding claims, wherein the sequencing assay comprises a next generation sequencing assay.WSGR Docket No.: 47697-75160163. The method of any one of the preceding claims, wherein the sequencing assay comprises a sequencing-by-synthesis assay.
64. The method of any one of the preceding claims, further comprising identifying a commensal microbe of the subject at least in part based on a result of the sequencing assay.
65. The method of any one of the preceding claims, further comprising identifying a pathogen infecting the subject at least in part based on a result of the sequencing assay.
66. The method of any one of the preceding claims, further comprising administering a treatment to the subject, wherein the subject is infected by the pathogen.
67. The method of any one of the preceding claims, wherein the treatment comprises an antimicrobial agent.
68. A method of distinguishing between a contaminant nucleic acid from a target nucleic acid in a sample using the method of any one of claims 1-67.
69. The method of claim 68, wherein the target nucleic acid comprises mcfDNA from a pathogen.
70. A method of determining an eligibility of a subject for transplantation, comprising analyzing a sample from the subject using the method of any one of claims 1-67.
71. The method of claim 70, wherein the subject is a xenotransplant animal donor.
72. The method of claim 70 or 71, wherein the subject comprises a pig or a nonhuman primate.