Detection and prediction of infectious diseases
By generating fragment length profiles of nucleic acid libraries and performing high-throughput sequencing, the problem of accurately distinguishing the infection stages of Helicobacter pylori was solved, enabling non-invasive and precise infection detection and treatment decisions.
Patent Information
- Application Number
- CN202511721869.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-17
- Filing Date
- 2019-11-21
- Publication Date
- 2026-02-27
AI Technical Summary
Current technologies struggle to accurately distinguish between the invisible and invasive stages of Helicobacter pylori infection, leading to misdiagnosis and overtreatment. Furthermore, existing nucleic acid sequencing methods suffer from bias and invasiveness issues.
By generating fragment length profiles of nucleic acid libraries, and using bias correction or reproducible bias methods, nucleic acid libraries are prepared from initial samples. Fragment length characteristics are analyzed, and combined with high-throughput sequencing and bioinformatics analysis, microorganisms and localization sites are identified.
It provides a non-invasive method to accurately distinguish the stages of infection, reduce misdiagnosis and overtreatment, and improve the accuracy and safety of testing.
Smart Images

Figure CN121575489A_ABST
Abstract
Description
[0001] This application is a divisional of the Chinese Patent Application with the application date of November 21, 2019, application number 201980083444.7, entitled “Detection and Prediction of Infectious Disease” which corresponds to the PCT Application with the application date of November 21, 2019, application number (PCT / US2019 / 062665). Cross Reference to Related Applications
[0002] This application claims priority to and benefit of U.S. Provisional Application No. 62 / 770,182, filed November 21, 2018, entitled “Detection and Prediction of Infectious Disease,” U.S. Provisional Application No. 62 / 770,181, filed November 21, 2018, entitled “Direct-to-Library Methods, Systems and Compositions,” and U.S. Provisional Application No. 62 / 849,618, filed May 17, 2019, entitled “Fragment Length Distributions and Methods of Using Such,” the entire contents of which are incorporated herein by reference for all purposes. TECHNICAL FIELD
[0003] The present invention relates to using fragment length distributions in nucleic acid libraries to identify microorganisms, identify the type of host-microorganism biological interaction, identify the site of infection or localization, select a therapy or treatment, monitor a treatment, monitor cytotoxicity, detect transplant rejection, monitor immune system response or activity, identify the stage of infection, monitor transplant rejection, and for cancer diagnosis. BACKGROUND
[0004] For many microbial infections, the first stage is colonization. In some cases, a microbial infection can progress to a persistent infection and can develop into an invasive disease stage. Examples of microorganisms that can develop into an invasive disease include cytomegalovirus (CMV) cytomegalovirus ), Epstein-Barr virus (EBV) Epstein-Barr virus ), Helicobacter pylori (H. pylori) heliobacter pylori ), Clostridium difficile (C. difficile) clostridium difficileCertain sexually transmitted infections, etc. For patients infected with these types of microorganisms, identifying the corrective, colonizing, or invasive stage of infection can be an important factor in making effective treatment decisions. The location of the infection can also affect its significance and available treatment options. Some microorganism-related diseases can occur in the absence of colonization that is considered typical of colonization. For example: Clostridium botulinum (… C . botulinum Intake may be sufficient to cause symptoms.
[0005] Furthermore, infections in the asymptomatic stage often exist without symptoms or with nonspecific symptoms that may resemble a variety of other diseases. Therefore, such infections are frequently undiagnosed, misdiagnosed, or treated symptomatically, allowing the microbe to persist and increasing the risk that the patient's infection will progress to an invasive disease.
[0006] Helicobacter pylori ( Helicobacter pylori (H. pylori Helicobacter pylori is the most common chronic bacterial infection in humans. It is estimated that 50% of the world’s population is infected. In the United States, about 30% of adults are infected before the age of 50, while most individuals are infected in childhood. Chen, Y. and MJ Blaser, Journal of Infectious Diseases, 2008. 198(4): p.553-60. Helicobacter pylori is strongly associated with gastrointestinal (GI) conditions including chronic gastritis, peptic ulcer disease, gastric adenocarcinoma, and lymphoma. Peptic ulcer disease (PUD) is the most common manifestation of Helicobacter pylori infection and the annual incidence of physician-diagnosed PUD is 0.1-0.19%. Sung, JJ, EJ Kuipers and HB El-Serag, Aliment Pharmacol Ther, 2009. 29(9): p.938-46. It is estimated that the lifetime risk of infected individuals developing peptic ulcer disease is 10%-20%. (Kuipers, EJ et al., Gastrointestinal Pharmacology and Therapeutics, 1995, Supplement 2: pp. 59-69.)
[0007] The primary phenomenon responsible for these disease manifestations is mucosal inflammation in response to the presence of Helicobacter pylori. However, only a small percentage of individuals with Helicobacter pylori will experience inflammation associated with invasive Helicobacter pylori infection.
[0008] Currently, it is challenging to distinguish between patients with asymptomatic stages of H. pylori infection and patients with symptomatic stages or at risk of progressing to symptomatic stages. Although most infections with H. pylori are asymptomatic, patients with invasive disease can begin to experience symptoms of persistent dyspepsia, such as abdominal pain, nausea or vomiting, and lack of appetite. However, these non-specific symptoms can also be caused by other conditions and are experienced by healthy people. Some physicians test all patients with unexplained persistent dyspepsia. Other physicians follow current guidelines, which recommend testing for H. pylori in individuals with active PUD, a documented history of peptic ulcer, or gastric MALT lymphoma. Chey, W.D., et al., Am J Gastroenterol, 2007. 102(8): p. 1808-25. Thus, physicians following guidelines will only test patients who are at risk of having H. pylori -related disease, which can result in undertreatment.
[0009] Currently, there are several methods for testing for H. pylori. Existing non-invasive testing methods for H. pylori include fecal antigen testing, urea breath testing, and H. pylori serology. However, these methods can only determine if H. pylori is present, not if H. pylori invasion or associated inflammation is present. Some practitioners will initiate primary therapy for eradication based on a positive result from one of the non-invasive tests, which can result in overtreatment.
[0010] The current gold standard for diagnosing H. pylori disease is to perform an upper endoscopy for documentation (by biopsy) of specific pathological changes due to H. pylori invasion, such as inflammation, atrophy, and intestinal metaplasia, in combination with detection of H. pylori in biopsy samples. Dixon, M.F., et al., Helicobacter, 1997. 2 Suppl 1 : p. S17-24. However, there are serious risks and potential complications from this procedure, including bleeding that can sometimes require blood transfusion, infection, and tearing of the GI tract.
[0011] Overall, about 75% of patients who comply with primary therapy for H. pylori infection are considered cured after the first treatment based on a negative H. pylori diagnostic assay for active infection, which was previously positive prior to initiating therapy. If the diagnostic test for active gastrointestinal H. pylori infection remains positive after completion of first-line therapy, there is a possibility of antibiotic-resistant H. pylori and additional therapy will be required before a negative diagnostic test result is obtained.
[0012] Next generation sequencing (NGS) can be used to aggregate large amounts of data about the nucleic acid content of a sample. The data can be particularly useful for analyzing nucleic acids in complex samples, such as clinical samples. However, prior to using NGS methods, the starting sample must often be processed, which can reduce nucleic acid recovery, delay sequencing, delay reporting of clinical calls, introduce errors, introduce bias, and generally result in chemical waste that requires controlled disposal. Errors and bias can affect results in many cases, such as when there are low abundance nucleic acids or target nucleic acids in a patient sample. Current NGS methods focus on the abundance or relative abundance of particular reads or sequences. Further, many sequencing library preparation methods and some next generation sequencing systems produce target nucleic acid fragment lengths and fragment length distributions that deviate from experimentally observed endogenous fragment lengths and fragment length distributions, particularly those methods and systems that utilize variable poly A tail tags, unaccounted for poly A tail tags, thermal inactivation of enzymes, use of biased extraction methods, or use of other metrics that introduce nucleic acid length, secondary structure, and / or GC bias across the entire range or partial range of target nucleic acid lengths and GC content. Some such methods and systems prevent successful correction for bias even in the presence of process control molecules, provided the bias is large such that insufficient target nucleic acids and / or process control molecules are recovered across the entire or certain segments of relevant lengths and GC content for final analysis.
[0013] Various methods including NGS have been used to identify microorganisms present in a host, but most of these methods focus on the abundance of microorganism reads rather than the physical properties of the molecules being read. For example, many extraction protocols, library generation protocols, and sequencing protocols contain steps or processes designed to remove short nucleic acid fragment lengths. Short nucleic acid fragment lengths are also often sacrificed to minimize undesirable or incomplete byproducts of extraction, library generation, or amplification, such as primer dimers or adapter dimers. Free nucleic acids of microorganisms are an example of target nucleic acids that are particularly susceptible to bias and depletion of short nucleic acids due to their fragment lengths being below about 100 bp.
[0014] Current methods for partitioning infections between the inapparent or latent phase of infection and other phases after identification of potential pathogens can sometimes require invasive biopsy procedures. Non-invasive tests, such as serology, can detect markers of exposure to a microorganism, but cannot indicate whether the infection is active or at risk of progressing to invasive disease. Thus, there is a need for precise non-invasive methods for determining whether an organ of a patient has been infected and distinguishing which patients will remain in the colonization phase, and which patients are at risk of developing secondary invasive disease. The present disclosure provides non-invasive methods, compositions, and kits for detecting infection in a subject and determining whether the infection is in the colonization or invasive disease phase. The present disclosure also provides non-invasive methods for determining a site of localization in a subject and / or a phase of infection in a subject. SUMMARY
[0015] Embodiments of the present application provide a fragment length profile from a nucleic acid library, wherein the nucleic acids used to prepare the nucleic acid library were obtained from a sample by an unbiased method, a method enabled for bias correction, or a method with reproducible bias. In various aspects, the nucleic acid library is generated from an initial sample, and the nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library or prior to initiating the library generation process. Aspects of the method can include nucleic acid sequencing as a step after nucleic acid preparation and prior to determining the fragment length profile of a target nucleic acid, a plurality of target nucleic acids, or a subset of nucleic acids within the nucleic acid library. In aspects of embodiments, the fragment length profile includes one or more properties selected from the group comprising: shape of the distribution, amplitude of the segments, fraction of segments, shape of the peaks, number of peaks, position of the largest peak in the peaks, ratio of segment counts of two or more segments, height of the helical phasing peak, ratio of segment counts at two different fragment lengths, ratio of segment counts within two different fragment length ranges, amount of fragments within a segment, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads, slope within a segment, peak width, rate of count decay or increase within a segment, number of peaks, scaling of count decay or increase within a segment.
[0016] Methods of generating a fragment length profile of a nucleic acid library are provided. Individual methods comprise the steps of preparing a nucleic acid library from an initial sample using a bias-corrected recovery method or a method with reproducible bias, determining read counts or normalized counts for a plurality of fragment lengths within the nucleic acid library, determining one or more fragment length characteristics of the nucleic acid library, and generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics. In aspects of the embodiments, the fragment length profile comprises one or more fragment length characteristics selected from the group comprising: shape of the distribution, segment amplitude, peak shape, ratio of fragment counts for two or more segments, height of a helical phasing peak, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes for two or more segments, position of the maximum peak in a peak, number of peaks, and fragment length distribution within a subset of reads. Methods of generating a fragment length profile of a nucleic acid library are provided. Individual methods comprise the step of preparing a nucleic acid library from an initial sample, the step comprising the steps of: optionally adding one or more process control molecules to the initial sample to provide a spiked initial sample, and generating a nucleic acid library from the spiked initial sample, wherein optionally nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library. Aspects of the methods can comprise nucleic acid sequencing as a step after nucleic acid preparation and prior to determining a fragment length profile. The methods of generating a fragment length profile of a target nucleic acid within the nucleic acid library further comprise the steps of: determining read counts for a plurality of fragment lengths within the nucleic acid library, determining one or more fragment length characteristics of the nucleic acid library, and generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics. In aspects of the embodiments, the fragment length profile comprises one or more fragment length characteristics selected from the group comprising: shape of the distribution, segment amplitude, peak shape, number of peaks, position of the maximum peak in a peak, ratio of fragment counts for two or more segments, height of a helical phasing peak, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes for two or more segments, and fragment length distribution within a subset of reads. In certain aspects, the step of generating the nucleic acid library from the initial sample further comprises, consists of, or consists essentially of the steps of: dephosphorylating nucleic acids from the initial sample to produce a set of dephosphorylated nucleic acids, denaturing the dephosphorylated nucleic acids to produce denatured nucleic acids, ligating 3' end adapters to the denatured nucleic acids to produce adapter-ligated nucleic acids, isolating adapter-ligated nucleic acids, taping primers to the adapter-ligated nucleic acids, and extending primers with a polymerase to generate complementary strands, ligating 5' end adapters, eluting the strands, and amplifying the complementary strands.Aspects of the method can include nucleic acid sequencing as a step after nucleic acid preparation and before determining a fragment length profile. In various embodiments, the number of reads is a normalized number of reads. In some embodiments, the fragment length profile is for at least one subset of reads in the nucleic acid library. In such embodiments, the method further includes the steps of identifying at least one subset of reads within the nucleic acid library, and determining a fragment length distribution within each selected subset of reads. In some embodiments, the step of generating the fragment length profile further includes using two or more fragment length characteristics.
[0017] Methods of identifying microorganisms present in a sample are provided. Methods of identifying or characterizing microorganisms present in a sample include the steps of generating a fragment length profile for sequencing reads from a nucleic acid library generated from the sample and aligned to a microorganism reference sequence, comparing the fragment length profile to a reference fragment length profile of one or more microorganisms, and identifying the microorganism as present in the sample if the fragment length profile from the sample is similar to a reference fragment length profile of a microorganism. Aspects of the methods include comparing fragment length profiles of target sequences from a nucleic acid library. In various embodiments, the fragment length profile can indicate that the microorganism is present in the form of a pathogen or a commensal microorganism. In aspects of the methods, generating a fragment length profile of the nucleic acid library includes the steps of preparing a nucleic acid library from an initial sample, quantifying the number of reads of a plurality of fragment lengths within the nucleic acid library; determining one or more fragment length characteristics of the nucleic acid library or reading at least a subset of the nucleic acid library, and generating a fragment length profile of the nucleic acid library or at least a subset of reads using the one or more fragment length characteristics. The step of preparing a nucleic acid library from an initial sample further includes the steps of adding one or more process control molecules to the initial sample to provide a spiked initial sample, and generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library. Aspects of the methods can include nucleic acid sequencing as a step after nucleic acid preparation and prior to determining a fragment length profile. In aspects of the embodiments, the fragment length profile includes one or more fragment length characteristics selected from the group comprising: shape of the distribution, amplitude of the segments, peak shape, number of peaks, position of the largest peak in the peaks, ratio of segment counts of two or more segments, height of helical phasing peak, ratio of segment counts at two different fragment lengths, ratio of segment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads. In various aspects of the methods, the fragment length profile includes at least one fragment length characteristic selected from the group comprising: ratio of segment counts of two or more segments, peak shape, peak width, rate of decay or increase of counts within a segment, number of peaks, scaling of decay or increase of counts within a segment, position of the largest peak in the peaks.
[0018] Methods of determining a localization site in a subject are provided. The methods include the steps of generating a fragment length profile of a target nucleic acid in a nucleic acid library generated from the sample or the entire nucleic acid library, comparing the fragment length profile to a reference fragment length profile of one or more source sites, and predicting a first site as a localization site if the fragment length profile in the sample is similar to a fragment length profile of the first source site and predicting a second site as a localization site if the fragment length profile in the sample is similar to a fragment length profile from the second source site. In embodiments of the methods, generating one or more fragment length profiles of the nucleic acid library includes the steps of preparing a nucleic acid library from an initial sample, quantifying the number of reads of a plurality of fragment lengths within the nucleic acid library, and generating a fragment length profile of a target nucleic acid in the nucleic acid library or the entire nucleic acid library using one or more fragment length characteristics. In embodiments of the methods, preparing a nucleic acid library from an initial sample further includes the steps of adding one or more process control molecules to the initial sample to provide a spiked initial sample, and generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library. Aspects of the methods can include nucleic acid sequencing as a step after nucleic acid preparation and prior to determining a fragment length profile. In aspects of embodiments, the fragment length profile includes one or more fragment length characteristics selected from the group comprising a shape of a distribution, a segment amplitude, a peak shape, a number of peaks, a position of the largest peak in the peaks, a ratio of segment counts of two or more segments, a height of a helical phasing peak, a ratio of segment counts at two different fragment lengths, a ratio of segment counts within two different fragment length ranges, a range of fragment lengths within a segment, a ratio of maximum amplitudes of two or more segments, a peak width, a rate of decay or increase of counts within a segment, a number of peaks, a scaling of decay or increase of counts within a segment, and a fragment length distribution within a subset of reads. In aspects of the methods, the localization site is selected from the group of source sites consisting of or consisting essentially of deep tissue, lung, liver, bone, kidney, brain, heart, sinus, GI tract, spleen, skin, joint, ear, nose, mouth, blood stream, and blood.
[0019] Methods of monitoring the status of a transplant in a subject are provided. The methods of monitoring the status of a transplant include the steps of generating a baseline fragment length profile from a nucleic acid library generated from a sample obtained from the subject; generating a second fragment length profile of a nucleic acid library generated from a second sample obtained from the subject; and comparing the second fragment length profile to the baseline fragment length profile. If the second fragment length profile is different from the baseline fragment length profile, then an increased amount of an anti-rejection therapy is administered internally to the subject, wherein the risk of rejection in the subject having a transplant is reduced after administration of the anti-rejection therapy. If the second fragment length profile is similar to the baseline fragment length profile, then the anti-rejection therapy is maintained or reduced, wherein the risk of side effects in the subject is lower than the risk of side effects in the subject receiving an increased amount of the anti-rejection therapy. Aspects of the methods include the steps of comparing the fragment length profile of a target nucleic acid in a nucleic acid library or whole library from a sample obtained from a subject having a transplant, and comparing the profile to a reference fragment length profile.
[0020] Methods of monitoring toxicity of a compound administered to a subject are provided. The methods include the steps of generating a fragment length profile of a nucleic acid library or a target nucleic acid in the nucleic acid library prepared from a sample obtained from the subject, and comparing the fragment length profile to one or more reference fragment length profiles. In aspects of the methods, the subject has, is at risk of having, or exhibits symptoms related to cancer. In aspects of the methods, the one or more reference fragment length profiles are generated from a nucleic acid library obtained from a subject or cells exposed to the compound. In aspects of the methods, the one or more reference fragment length profiles include a baseline fragment length profile. In aspects of the methods, the compound is a chemotherapeutic agent. In embodiments of the methods, the step of generating a fragment length profile of a nucleic acid library includes the steps of preparing a nucleic acid library from an initial sample using a bias-corrected recovery method; determining the number of reads of a plurality of fragment lengths within the nucleic acid library; determining one or more fragment length characteristics of the nucleic acid library; and generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics. Aspects of the methods can include nucleic acid sequencing as a step after nucleic acid preparation and before determining a fragment length profile. In aspects of the embodiments, the fragment length profile includes one or more fragment length characteristics selected from the group comprising: shape of the distribution, amplitude of the segments, shape of the peaks, ratio of fragment counts of two or more segments, height of the helical phasing peak, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads. In embodiments of the methods, generating a fragment length profile of the nucleic acid library includes the step of preparing a nucleic acid library from an initial sample, the step further comprising: adding one or more process control molecules to the initial sample to provide a spiked initial sample, and generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library; quantifying the number of reads of a plurality of fragment lengths within the nucleic acid library; determining one or more fragment length characteristics of the nucleic acid library; and generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics. In aspects of the embodiments, the fragment length profile includes one or more fragment length characteristics selected from the group comprising: shape of the distribution, amplitude of the segments, shape of the peaks, ratio of fragment counts of two or more segments, height of the helical phasing peak, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads.
[0021] The present invention relates to methods of predicting the risk of a biological (or plurality of biological) organism present in a host's body producing a local or systemic environmental change or invading an organ or anatomical system with a substantial negative outcome on health. An organism is invasive if it crosses a barrier or translocates from one organ or anatomical structure to another, invades structures beyond the tissue layer it occupies in a colonized state to produce a local invasion, it changes the environment of a structure such that it has a significant negative impact on the structure or causes DNA mutation or inflammation, or it otherwise overwhelms the host's immune system.
[0022] In certain embodiments, the risk level is based on the abundance of the organism in the host's body compared to an asymptomatic control or an infected control. In other embodiments, the abundance is a threshold or a range. In yet other embodiments, the risk level is calculated as a clinical decision score based on one or more of the following: abundance of the organism, patient's clinical history, chronicization of the disease, genetic biomarker factors and patient characteristics (such as age, gender, etc.), fragment length distribution profile, and fragment length distribution profile characteristics.
[0023] In one aspect, a method of determining the stage of infection of a subject suspected of having a microbial infection is provided, the method comprising: (a) performing high-throughput sequencing of nucleic acids from the biological sample; (b) performing bioinformatics analysis to identify microbial nucleic acid sequences present in the biological sample; and (c) calculating a measurement of the nucleic acids and comparing the measurement to a control, thereby determining the stage of infection of any microorganism identified in the biological sample.
[0024] In some embodiments, the method further comprises one or more steps selected from the group consisting of: (a) extracting nucleic acids from a portion of a biological sample obtained from the subject, and (b) adding synthetic nucleic acid spike-ins.
[0025] In one embodiment, the measurement of step (c) is selected from the group consisting of absolute abundance of free microbial nucleic acid sequences, distribution of fragment lengths of nucleic acid sequences, characteristics of nucleic acid fragment length distribution profile, or a combination thereof. In another embodiment, the measurement of step (c) is absolute abundance and distribution of fragment lengths if the target pathogen.
[0026] In a second embodiment, the subject has symptoms of infection or is at risk of infection.
[0027] In a third embodiment, the infection phase is an asymptomatic phase, a symptomatic infection phase, a treatment phase, or an eradication phase. In a fourth embodiment, the method further includes repeating the method over time to monitor infection, infection phase, the efficacy of treatment for the infection, or to detect the onset of infection. In various aspects, the method may further include changing the treatment regimen.
[0028] In a fifth embodiment, the method further includes administering a treatment regimen to the subject based on a determined stage of infection.
[0029] In the sixth embodiment, the high-throughput sequencing assay is next-generation sequencing, massively parallel sequencing, pyrosequencing, successive synthesis sequencing, single-molecule real-time sequencing, polymerase cloning sequencing, DNA nanosphere sequencing, helicopter single-molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.
[0030] In the seventh embodiment, the sample is blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, sputum, urine, feces, saliva, or nasal sample.
[0031] In the eighth embodiment, the method further includes identifying one or more antibiotic resistance genes of the target pathogen.
[0032] In a ninth embodiment, the method further includes identifying at least one risk factor in the subject's genomic DNA.
[0033] In the tenth embodiment, the nucleic acid is cell-free DNA and / or cell-free RNA. The nucleic acid may include cell-free pathogen DNA. The nucleic acid may include cell-free pathogen RNA. The nucleic acid may include cell-free microbial DNA. The nucleic acid may include cell-free microbial RNA.
[0034] In the eleventh embodiment, the target pathogens are Helicobacter pylori, Clostridium difficile, and Haemophilus influenzae (H. pylori). haemophilus influenza ),salmonella( salmonella Streptococcus pneumoniae () streptococcus pneumoniae ), Cytomegalovirus ( cytomegalovirus Hepatitis B virus, Hepatitis C virus, Human papillomavirus, Epstein-Barr virus, Human T-cell lymphoma virus, Merkel cell polyomavirus, Kaposi's sarcoma virus, Human herpesvirus, Chlamydia, Gonorrhea, Syphilis, or Trichomonas vaginalis.
[0035] In a twelfth embodiment, the subject has previously undergone another test or other clinical test. In one embodiment, the other clinical test is a fecal antigen test, a urea breath test, serology, urease test, histology, bacterial culture and susceptibility testing, biopsy, or endoscopy.
[0036] In a thirteenth embodiment, the target pathogen nucleic acid is DNA and / or RNA. The pathogen nucleic acid includes cell-free DNA. The nucleic acid includes pathogen cell-free RNA.
[0037] In a fourteenth embodiment, the synthetic nucleic acid spike-in includes at least 1,000 unique synthetic nucleic acids of the sample, wherein each of the 1,000 unique synthetic nucleic acids includes (i) an identifying tag; and (ii) a variable region comprising at least 5 degenerate bases. In further embodiments, the method further comprises (a) optionally extracting nucleic acids from the spiked sample; (b) generating a library of the spiked sample; (c) optionally enriching the library of the spiked sample; (d) performing a high-throughput sequencing assay to obtain sequence reads from the library of the spiked sample; (e) calculating a loss of diversity value for the 1,000 unique synthetic nucleic acids; and (f) calculating a measurement of the nucleic acids and comparing the measurement to a control, thereby determining the stage of infection of the subject.
[0038] In yet further embodiments, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in U.S. 9,976,181.
[0039] In another aspect, there is a method of determining a stage of infection of Helicobacter pylori in a subject, the method comprising: a) optionally, extracting cell-free nucleic acids from a biological sample obtained from the subject; b) adding synthetic nucleic acid spike-in to the sample; c) performing high-throughput sequencing of nucleic acids from the biological sample; d) performing bioinformatics analysis to identify Helicobacter pylori nucleic acid sequences present in the biological sample; and e) calculating a measurement of the Helicobacter pylori nucleic acids and comparing the measurement to a control, thereby determining the stage of infection of Helicobacter pylori in the subject.
[0040] In a first embodiment, the measurement is absolute abundance of H. pylori or distribution of fragment lengths or a combination thereof. In one embodiment, the measurement is absolute abundance of H. pylori. In another embodiment, the measurement is distribution of fragment lengths of H. pylori. In yet another embodiment, the measurement is absolute abundance and distribution of fragment lengths of H. pylori. In various embodiments, the steps of the method can be performed in varying order.
[0041] In a second embodiment, the subject has symptoms of or is at risk for H. pylori infection. In one embodiment, the stage of infection is asymptomatic, symptomatic infection, treatment, or eradication.
[0042] In a third embodiment, the method further comprises repeating the method over time to monitor the infection, efficacy of treatment of the infection.
[0043] In one aspect, there is a method of determining a stage of infection of H. pylori in a subject, the method comprising: (a) making a spiked sample by obtaining a sample comprising cell-free nucleic acid from the subject and adding one or more process control molecules; (b) optionally, extracting the nucleic acid from the spiked sample; (c) generating a spiked sample library, wherein the generating comprises (i) ligating adaptors to the nucleic acid; and (ii) amplifying; (d) optionally, enriching the spiked sample library; (e) performing a high-throughput sequencing assay to obtain sequence reads from the spiked sample library; (f) calculating a diversity loss value for 1,000 unique synthetic nucleic acids; and (g) calculating a measurement of the cell-free nucleic acid and comparing the measurement to a control, thereby determining a stage of infection of H. pylori in the subject.
[0044] In yet further embodiments, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in U.S. 9,976,181.
[0045] In a second embodiment, the high-throughput sequencing assay is next-generation sequencing, massively parallel sequencing, pyrosequencing, sequencing by synthesis, single molecule real-time sequencing, polymerase cloning sequencing, DNA nanoball sequencing, helicopter single molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.
[0046] In a third embodiment, the sample is blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, fecal, saliva, or nasal sample.
[0047] In a fourth embodiment, the method further comprises administering a treatment regimen to the subject, wherein the treatment can be administered at any stage of the infection cycle.
[0048] In a fifth embodiment, the method further comprises identifying one or more antibiotic resistance genes of the target pathogen.
[0049] In a sixth embodiment, the free nucleic acid is DNA and / or RNA. The nucleic acid comprises free pathogen DNA. The nucleic acid comprises free pathogen RNA.
[0050] In a twelfth embodiment, the subject has previously undergone another other clinical test. In one embodiment, the other clinical test is a fecal antigen test, a urea breath test, serology, urease test, histology, bacterial culture and susceptibility testing, biopsy, or endoscopy.
[0051] In an eighth embodiment, the target pathogen nucleic acid is DNA and / or RNA. The pathogen nucleic acid comprises free DNA. The nucleic acid comprises free pathogen RNA. The target pathogen nucleic acid comprises a mixture of free DNA and free RNA.
[0052] Another aspect provides a method of determining a localization site in a subject infected with a pathogen, the method comprising: (a) obtaining a sample comprising nucleic acid from the subject, and adding one or more process control molecules, thereby generating a spiked sample; (b) optionally, extracting the nucleic acid from the spiked sample; (c) generating a library from the spiked sample, wherein generating comprises ligating adaptors to the nucleic acid and amplifying; (d) optionally, enriching the spiked sample; (e) performing a high-throughput sequencing assay by comparing to a reference genome to obtain sequence reads from the spiked sample; (f) optionally, calculating a diversity loss value; and (g) calculating a measurement of the nucleic acid and comparing the measurement to a control, thereby determining a localization site in the subject.
[0053] In a first embodiment, the measurement is an absolute abundance of the target pathogen or a distribution of fragment lengths or a combination thereof. In one embodiment, the measurement is an absolute abundance of the target pathogen. In another embodiment, the measurement is a distribution of fragment lengths of the target pathogen. In yet another embodiment, the measurement is an absolute abundance and a distribution of fragment lengths of the target pathogen.
[0054] In a second embodiment, the localization site is a tissue. In a further embodiment, the localization site is a tissue type. In yet a further embodiment, the localization site is an organ. In another further embodiment, the localization site is a tissue type comprising an organ.
[0055] In a third embodiment, the subject has symptoms of or is at risk for an infection. In a further embodiment, the subject has previously been identified as infected with Helicobacter pylori, Clostridium difficile, Haemophilus influenzae, Salmonella, Streptococcus pneumoniae, Cytomegalovirus, Hepatitis B virus, Hepatitis C virus, Human Papillomavirus, Epstein-Barr virus, Human T-cell lymphotropic virus 1, Merkel cell polyomavirus, Kaposi's sarcoma-associated herpesvirus, Human herpesvirus 8, Chlamydia virus, Herpes simplex virus, Neisseria, Treponema, or Trichomonas.
[0056] In a fourth embodiment, the method is repeated over time to monitor an infection, efficacy of treatment of an infection.
[0057] In a fifth embodiment, the method further comprises administering a treatment regimen to the subject based on the determined stage of infection.
[0058] In a sixth embodiment, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in U.S. 9,976,181.
[0059] In a seventh embodiment, the high-throughput sequencing assay is next-generation sequencing, massively parallel sequencing, pyrosequencing, sequencing by synthesis, single molecule real-time sequencing, polymerase cloning sequencing, DNA nanoball sequencing, helicopter single molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.
[0060] In an eighth embodiment, the sample is blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, fecal, saliva, nasal, or a tissue sample.
[0061] In a ninth embodiment, the method further comprises identifying one or more antibiotic resistance genes of the pathogen.
[0062] In a tenth embodiment, the method further comprises identifying a risk factor in the genomic DNA of the subject.
[0063] In an eleventh embodiment, the target pathogen nucleic acid is DNA and / or RNA. The pathogen nucleic acid comprises cell-free DNA. The nucleic acid comprises pathogen cell-free RNA. The target pathogen nucleic acid comprises a mixture of cell-free DNA and cell-free RNA.
[0064] In a twelfth embodiment, the free nucleic acids are DNA and / or RNA. The nucleic acids include free pathogen DNA. The nucleic acids include free RNA. The nucleic acids include free pathogen RNA. The nucleic acids include free subject RNA. The nucleic acids include pathogen and subject free RNA.
[0065] In an aspect, a method of determining a stage of infection of a subject suspected of having a microbial infection is provided, the method comprising: (a) providing a sample comprising nucleic acids from the subject; (b) adding at least 1,000 unique synthetic nucleic acids to the sample, thereby generating a spiked sample; (c) generating a library from the spiked sample; (d) performing a high-throughput sequencing assay to obtain sequence reads from the spiked sample; (e) determining the stage of infection of the subject based on the sequence reads.
[0066] In one embodiment, the sample is selected from blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, stool, saliva, nasal, and tissue samples. The sample is blood, plasma, serum, cerebrospinal fluid, or synovial fluid.
[0067] In yet another embodiment, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in U.S. 9,976,181.
[0068] In a further embodiment, the high-throughput sequencing assay is next-generation sequencing, massively parallel sequencing, pyrosequencing, sequencing by synthesis, single molecule real-time sequencing, polymerase cloning sequencing, DNA nanoball sequencing, helicopter single molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.
[0069] In another further embodiment, the determination of the stage of infection is based on absolute abundance of the target pathogen or distribution of fragment lengths or a combination thereof. In one embodiment, the determination is based on absolute abundance of the target pathogen. In another embodiment, the determination is based on distribution of fragment lengths of the target pathogen. In yet another embodiment, the determination is based on absolute abundance and distribution of fragment lengths of the target pathogen.
[0070] One aspect of the application provides a method of determining an infection stage of a subject. The method comprises the steps of: generating a fragment length profile of a nucleic acid library generated from a sample obtained from the subject; comparing the fragment length profile to a reference fragment length profile; and if the fragment length profile from the sample is similar to a fragment length profile from a symptomatic subject, determining that the infection stage indicates an increased risk that the subject will exhibit a microbe-related symptom, and if the fragment length profile from the sample is similar to a fragment length profile from an asymptomatic subject, determining that the infection is in an invisible stage. In one aspect, the fragment length profile is a non-microbe host nucleic acid library fragment length profile. In various aspects, the method further comprises the steps of: determining an abundance of at least one significant microbe in the sample from the subject; comparing the abundance to a threshold; and comparing the fragment length profile to a reference fragment length profile. If the fragment length profile from the sample is similar to a fragment length profile from a symptomatic subject, and the abundance is at or above the threshold, then determining that the infection stage indicates an increased risk that the subject will exhibit a microbe-related symptom. If the fragment length profile from the sample is similar to a fragment length profile from an asymptomatic subject, then determining that the infection is in an invisible stage. In one aspect, the method further comprises the step of administering an antimicrobial agent to a subject determined to have an increased risk of exhibiting a microbe-related symptom.
[0071] A method of determining an infection stage of a subject suspected of having a microbial infection, the method comprising performing high-throughput sequencing of nucleic acids from a biological sample, performing bioinformatics analysis to identify nucleic acid sequences present in the biological sample, and calculating a measurement of the nucleic acids, and comparing the measurement to a control, thereby determining an infection stage of a microorganism identified in the biological sample. The method can further comprise one or more steps selected from the group consisting of: (i) extracting nucleic acids from a biological sample obtained from the subject, and (ii) adding synthetic nucleic acid spike-in to a biological sample obtained from the subject. In an aspect, the nucleic acids comprise microbial nucleic acids, host nucleic acids, or both microbial and host nucleic acids. In an aspect, the nucleic acids comprise free microbial nucleic acids, host nucleic acids, or both microbial and host nucleic acids. In an aspect, the measurement is selected from the group of measurements consisting of absolute abundance of nucleic acids, fragment length distribution profile of nucleic acids, and both absolute abundance and fragment length distribution profile. In an aspect, the infection stage is selected from the group consisting of an asymptomatic stage, a colonization stage, a symptomatic stage, an active stage, an invasive disease stage, a resolution stage, a treatment period, or an eradication stage. In an aspect, the method further comprises administering a treatment regimen to the subject based on the determined infection stage. The method can further comprise repeating the method over time to monitor the infection or efficacy of treatment of the infection. In some embodiments, the microorganism is selected from the group consisting of Helicobacter pylori, Clostridium difficile, Haemophilus influenzae, Salmonella, Streptococcus pneumoniae, Cytomegalovirus, Hepatitis B virus, Hepatitis C virus, Human Papillomavirus, Epstein-Barr virus, Human T-cell lymphotropic virus 1, Merkel cell polyomavirus, Kaposi's sarcoma-associated herpesvirus, Human herpesvirus 8, Chlamydia virus, Herpes simplex virus, Neisseria, Treponema, or Trichomonas. In aspects, adding synthetic nucleic acid spike-in further comprises making a spiked sample by obtaining a sample comprising free nucleic acids from a subject and adding one or more process control molecules; extracting nucleic acids from the spiked sample; generating a spiked sample library; enriching the spiked sample library; performing a high-throughput sequencing assay to obtain sequence reads from the spiked sample library; calculating a diversity loss value for 1,000 unique synthetic nucleic acids; and calculating a measurement of the free nucleic acids and comparing the measurement to a control, thereby determining an infection stage of the subject.
[0072] In one embodiment, the present application provides a method of determining the stage of infection of Helicobacter pylori in a subject, the method comprising extracting nucleic acids from a biological sample obtained from the subject, adding synthetic nucleic acid spike-in to the sample, performing high-throughput sequencing of the nucleic acids from the biological sample, performing bioinformatics analysis to identify free Helicobacter pylori nucleic acid sequences present in the biological sample, and calculating a measure of the free Helicobacter pylori nucleic acids, and comparing the measure to a control, thereby determining the stage of infection of Helicobacter pylori in the subject.
[0073] In one embodiment, the present application provides a method of determining the stage of infection of Helicobacter pylori in a subject, the method comprising: preparing a spiked sample by obtaining a sample comprising free nucleic acids from a subject and adding one or more process control molecules; extracting nucleic acids from the spiked sample; generating a spiked sample library, wherein the generating comprises (i) ligating adaptors to the nucleic acids; and (ii) amplifying; optionally, enriching the spiked sample library; performing a high-throughput sequencing assay to obtain sequence reads from the spiked sample library; calculating a diversity loss value for 1,000 unique synthetic nucleic acids; and calculating a measure of the free nucleic acids and comparing the measure to a control, thereby determining the stage of infection of Helicobacter pylori in the subject.
[0074] One embodiment provides a method of determining a localization site in a subject infected with a pathogen, the method comprising obtaining a sample comprising nucleic acids from a subject, adding one or more process control molecules to the initial sample to provide a spiked sample, optionally extracting the nucleic acids from the spiked sample, generating a library from the spiked sample, wherein the generating comprises ligating adaptors to the nucleic acids and amplifying; optionally, enriching the spiked sample, performing a high-throughput sequencing assay by comparing to a reference genome to obtain sequence reads from the spiked sample; determining one or more fragment length properties of the nucleic acid library, generating a fragment length profile of the nucleic acid library generated from the sample, comparing the fragment length profile to reference fragment length profiles of one or more source sites, and identifying a first site as a localization site if the fragment length profile from the sample is similar to a fragment length profile from the first source site; identifying a second site as a localization site if the fragment length profile from the sample is similar to a fragment length profile from the second source site.
[0075] In one aspect, a method of determining a localization site in a subject infected with a pathogen is provided, the method comprising obtaining a sample comprising cell-free nucleic acids from the subject, and adding one or more process control molecules, thereby generating a spiked sample; optionally extracting nucleic acids from the spiked sample; generating a library from the spiked sample, wherein generating comprises ligating adaptors to the nucleic acids and amplifying; optionally, enriching the spiked sample; performing a high-throughput sequencing assay by comparing to a reference genome to obtain sequencing reads from the spiked sample; calculating a diversity loss value for 1000 unique synthetic nucleic acids; and calculating a measurement of the cell-free nucleic acids and comparing the measurement to a control, thereby determining a localization site in the subject. INCORPORATION BY REFERENCE
[0076] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF DRAWINGS
[0077] The novel features of the application are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present application will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the application are utilized, and the accompanying drawings of which: Figure 1 Methods of the disclosure are depicted.
[0078] Figure 2 Cell-free methods of the disclosure are depicted.
[0079] Figure 3 A schematic of an exemplary infection is shown.
[0080] Figure 4 One of the methods of infection site detection of the disclosure is depicted.
[0081] Figure 5 Basic scheme of a method for determining a diversity loss value is depicted.
[0082] Figure 6 A diagnostic workflow terminating in treatment for a positive diagnosis for H. pylori is shown.
[0083] Figure 7 A computer control system programmed or otherwise configured to implement the methods provided herein is depicted.
[0084] Figure 8The distribution of fragment lengths from reads of three microorganisms detected in three different human plasma samples generated from nucleic acid libraries is depicted. The fragment length characteristic of interest in the figures is the distribution shape. Each figure provides an example of a different distribution shape. In each figure, the y-axis shows the normalized number of reads, and the x-axis indicates the fragment length. The left figure provides an example of a “50-base-pair peak” distribution shape. The middle figure provides an example of a short-range exponential distribution shape. The right figure provides an example of a complex distribution shape, where this particular complex distribution shape includes aspects of both an exponentially decaying distribution shape and a single-peak 50-base-pair distribution. It is generally accepted that each distribution shape depicted reflects the distribution of fragment lengths in nucleic acid libraries generated from the different human plasma samples and provides an example of the indicated distribution shape type. Other distribution shapes are described elsewhere in this document. Other distribution shapes are also possible.
[0085] Figure 9 Examples of fragment length characteristics involving distribution segment amplitude and segment amplitude ratio are provided. The figure depicts the same pathogen (Candida tropicalis) from three different clinical samples. Candida tropicalis The distribution of fragment lengths of reads. In each figure, the y-axis shows the normalized number of reads, and the x-axis indicates fragment length. For the purposes of this figure, clinical samples are numbered 1 to 3. Compared to Candida tropicalis in clinical sample 3, Candida tropicalis in clinical samples 1 and 2 showed a distribution with a higher long fraction (>65 bp) relative to the 50 bp peak, while all fragment length spectra had a clear peak of approximately 45–50 bp. The ratio of short reads (<40 bp) to the 50 bp peak also varied among the three samples. The distribution of segment amplitude and segment amplitude ratio (<40 bp to 50 bp peaks and >65 bp to 50 bp peaks) reflects results obtained from one experiment.
[0086] Figure 10 The fragment length distribution of WU polyomavirus from two clinical samples is depicted. The left panel shows the distribution of fragments with individual peaks of approximately 50 base pairs (bp) in length. The right panel shows a combined pattern including contributions from exponential distribution shape, peaks, and length fractions. Mechanistically independent, short exponential fractions may indicate viral incorporation into the human genome or degradation of microbial nucleic acids through processes different from those that generate fragments within the “50 bp peak”.
[0087] Figure 11Examples are provided that relate to the fragment length characteristic of the ratio of fragment counts in different distributions. The plots depict the ratio of fragment counts in the "50 bp peak" fraction to the short class index fraction (read density 40-55 bp / read density 23-35 bp, x-axis) versus normalized counts (y-axis). The same human and human mitochondrial fractions are added for reference. The ratio varies between kingdom types. The ratio of bacterial reads varies greatly, while the ratio of fungal reads shows a bimodal pattern. The ratio of viral reads is also shown.
[0088] Figure 12 A summary of the fragment length distribution of maternal (dashed line) and fetal (solid line) cell-free nucleic acids is provided. The "50 bp peak" appears narrower in the fetal distribution, indicating a smaller range of fragment lengths within the peak from fetal nucleic acids. Additionally, the ratio of fetal to maternal reads in the "50 bp peak" region is higher compared to the nucleosome length fragments (e.g., 150-200 bp region).
[0089] Figure 13 A summary of the fragment length distribution of microorganisms that exist as pathogens or as commensal microorganisms is provided. In end-repairable double-stranded DNA-based assays, the fragment length of pathogens tends to be longer than that of commensal microorganisms.
[0090] Figure 14 A summary of the fragment length distribution of pathogens in nucleic acid libraries generated from samples confirmed to be infected by urine or blood culture is provided. Pathogens detected in nucleic acid libraries from samples utilizing orthogonal blood culture tests show a higher long read rate compared to pathogens detected in nucleic acid libraries from samples utilizing orthogonal urine culture. Read length is shown on the x-axis; fraction of reads is shown on the y-axis. The mean of urine culture samples (light solid line) and the mean of blood culture samples (light dashed line) are shown in the plot, as well as the difference between urine and blood (thick dashed line).
[0091] Figures 15A-15F Data from asymptomatic samples (AP), diagnostic positive samples (DP), diagnostic positive samples confirmed by orthogonal methods (DP c ), diagnostic positive samples confirmed by orthogonal NGS methods (DP NGS ), and diagnostic positive samples confirmed by orthogonal non-NGS microbial methods (DP micro ) are summarized, as indicated. Figure 15A A plot of the abundance in molecules per microliter (MPM) of microorganisms present at a significant level in the indicated sample types is provided. Figure 15B A plot of the MPM abundance of microorganisms of the same species present in both types of samples in asymptomatic samples (AP) and diagnostic positive samples (DP) is provided.Figure 15C An example of a representative TapeStation electropherogram of a library obtained from a diagnostic positive sample included in this study is provided. Data was obtained on a TapeStation using an HS TapeStation Cassette D1000 with loading buffer and DNA ladder according to the manufacturer's instructions. The higher, lower DNA markers are indicated in the figure. The orientation of the subset of the fragment length range of interest is indicated in the plot (note that the fragment length in the electropherogram of the library reflects the length of the fully adaptered nucleic acid molecule, not the actual length of the endogenous original sequence). Library fragment length is shown on the x-axis; normalized intensity (FU) is shown on the y-axis. Figure 15D A plot of the molar fraction of sequencing reads that map to the human reference and are longer than 64 bp (i.e. the majority of these reads have nucleosome length) after the adaptor sequence trimming step for the asymptomatic samples (AP) and diagnostic positive samples (DP) included in this study is provided. Figure 15E A summary comparison of the maximum MPM abundance of microorganisms present at significant levels in each asymptomatic (AP) and diagnostic positive (DP) sample in this study with the fraction of long human reads as defined in the title of Figure 15D A summary comparison of the maximum MPM abundance of microorganisms present at significant levels in each asymptomatic (AP) and diagnostic positive (DP) sample in this study with the fraction of long human reads as defined in the title of Figure 15F A summary comparison of the maximum MPM abundance of microorganisms present at significant levels in each asymptomatic (AP) and diagnostic negative (DN) sample with the fraction of long human reads as defined in the title of Figure 15D A summary comparison of the maximum MPM abundance of microorganisms present at significant levels in each asymptomatic (AP) and diagnostic negative (DN) sample with the fraction of long human reads as defined in the title of
[0092] Figure 16A Results of training predictors of infection status based on human fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the human training model. The right panel depicts the regions of fragment length associated with each infection status used by the human training model. Figure 16B Results of training predictors of infection status based on human mitochondrial fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the human mitochondrial training model. The right panel depicts the regions of fragment length associated with each infection status used by the human mitochondrial training model. Figure 16CResults depicting the training of predictors of infection status based on all pathogen fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the model trained on all pathogen fragments. The right panel depicts the bins of fragment length associated with each infection status used by the model trained on all pathogen fragments. Figure 16D Results depicting the training of predictors of infection status based on significant pathogen fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the model trained on only reads derived from significant pathogens. The right panel depicts the bins of fragment length associated with each infection status identified by the model trained on significant pathogens. Figure 16E Results depicting the training of predictors of infection status based on bacterial fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the bacterial training model. The right panel depicts the bins of fragment length associated with each infection status identified by the bacterial training model. Figure 16F Results depicting the training of predictors of infection status based on eukaryotic microorganism fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the eukaryote training model. The right panel depicts the bins of fragment length associated with each infection status identified by the eukaryotic cell training model. Figure 16G Results depicting the training of predictors of infection status based on viral fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the viral training model. The right panel depicts the bins of fragment length associated with each infection status identified by the viral training model. Figure 16H Results depicting the training of predictors of infection status based on archaeal fragments recovered from asymptomatic and symptomatic patients for sequencing are depicted. The left panel shows the probability of asymptomatic samples based on the archaeal training model. The right panel depicts the bins of fragment length associated with each infection status identified by the archaeal training model.
[0093] Figure 17A1 - 17A10 shows the normalized fragment length distribution of microbes suspected to infect the lung, where each plot shows one distribution of microbes for the indicated species, and the sample ID is indicated at the top of each plot. Frequency is defined as the read count aligned to the reference of the indicated microbe for a particular read (fragment) length, normalized by the total count of reads aligned to the reference of the indicated microbe. Figure 17B1- 17B10 shows normalized fragment length distributions of suspected blood stream infections with microorganisms, where each plot shows one distribution of microorganisms of the indicated species, and the sample ID is indicated at the top of each plot. Frequency is defined as the count of reads that aligned to the reference of the indicated microorganism with a particular read (fragment) length, normalized by the total count of reads that aligned to the reference of the indicated microorganism.
[0094] Figure 18A1 - A2 depicts representative normalized fragment length distributions of two microorganisms detected in venous draws of two different donors. Shown in the left plot is Haemophilus influenzae (H. influenzae) - normalized fragment length distribution of reads of the microorganism detected in plasma obtained from a venous draw of donor 1. Shown in the right plot is Streptococcus thermophilus (S. thermophilus) - normalized fragment length distribution of reads of the microorganism detected in plasma obtained from a venous blood draw of donor 2. haemophilus influenzae streptococcus thermophilus Figure 18B1 - B4 depicts normalized fragment length distributions of microorganisms detected in biological samples obtained during capillary draw collection processes performed from the same two donors as in A2, and at the same sampling time as the venous draws in A2. The upper left plot shows the normalized fragment length distribution of H. influenzae detected in biological samples obtained during capillary draw collection processes performed from donor 1. The lower left plot shows the normalized fragment length distribution of additional microorganisms detected in biological samples obtained during capillary draw collection processes performed from donor 1. The average distribution pattern is shown in a thick black line. The upper right plot shows the normalized fragment length distribution of S. thermophilus detected in biological samples obtained during capillary draw collection processes performed from donor 2. The lower right plot shows the normalized fragment length distribution of additional microorganisms detected in biological samples obtained during capillary draw collection processes performed from donor 2. The average distribution pattern is shown in a thick black line. Figure 18A1 Figure 18C1 - C2 compares the abundance of microorganisms co-existing in two replicates of biological samples obtained during capillary draw collection processes of donor 1 (left plot) and donor 2 (right plot). Figure 18D1 - D2 depicts a comparison of the microorganism abundance (x-axis) of microorganisms detected in biological samples obtained with capillary blood draw procedures to the microorganism abundance in negative Microvette samples. The results obtained for donor 1 and donor 2 are shown in the left and right plots, respectively.
[0095] Figure 19A1 - A3 orthogonally confirms that blood stream of subject RD-02 is affected by Enterobacteriaceae species (e.g., Escherichia coli) and Staphylococcus species (e.g., S. aureus). nterobacter Infection with *Enterobacter cloacae* (a species) is depicted in the nucleic acid library generated from plasma samples collected at different collection times indicated in each of the above figures. enterobacter cloacae The normalized fragment length distribution of the aligned sequences. Figure 19B1 -B5 orthogonally confirmed that subject RD-11 had a Staphylococcus aureus infection ( staphylococcus aureus Endocarditis caused by infection. The figure depicts the normalized fragment length distribution of sequences aligned with Staphylococcus aureus in nucleic acid libraries generated from plasma samples collected at different collection times indicated in each of the above figures. Figure 19C1 -C4 orthogonally confirmed that subject RD-13 had an E. coli infection ( Escherichia coli Febrile neutropenia caused by infection. The figure depicts the normalized fragment length distribution of sequences aligned with E. coli sequences from nucleic acid libraries generated from plasma samples collected at different collection times indicated in each of the above figures.
[0096] Figure 20A The fractions of reads outside the “50 bp peak” region (< 30 bp and > 60 bp) depicting the fragment length distribution of all orthogonally confirmed microorganisms vary over time after admission. Time traces of only orthogonally confirmed microorganisms are shown, where more than 50 unique sequences aligned with the microorganism reference were detected. Figure 20B The abundance of orthogonally confirmed microorganisms detected by the method, in MPM, was depicted over time after admission.
[0097] Figure 21A1 -A4 shows paired orthogonally confirmed and orthogonally unconfirmed microorganisms in plasma samples collected at the admission time point (t = 0) of two subjects, RD-06 and RD-13. The upper left panel shows the orthogonally confirmed microorganism (Staphylococcus aureus) in RD-06. The lower left panel shows the unconfirmed microorganism (Haemophilus influenzae) in RD-06. The upper right panel shows the orthogonally confirmed microorganism (Escherichia coli) in RD-13. The lower right panel shows the unconfirmed microorganism (Prevotella melanogaster) in RD-13. prevotella melaninogenica )). Figure 21B1 -B2 Enterococcus quinquefolius — Normalized fragment length distribution of orthogonal unidentified microorganisms detected in plasma samples collected from subject RD-15 at several time points after admission. Time points are indicated above the figure.
[0098] Figure 22A- C depicts three major response patterns of human fragment length distribution during treatment of infected subjects. The left panel shows an example where the long human fraction (> 60 bp) decreases during treatment. The middle panel shows an example where the long human fraction (> 60 bp) floats during treatment. The right panel shows an example where the long human fraction (> 60 bp) increases during treatment.
[0099] Figure 23 A summary of fragment length information and GC content from a sample of Pasteurella trehalosi is provided. Relative frequency is shown on the y-axis; GC content is shown on the x-axis. Fragment length ranges of less than 45 base pairs, 45-54 base pairs, 55-64 base pairs, 65-74 base pairs, and longer than 74 base pairs are shown. The combination of fragment length distribution with GC content information indicates that the process induced a temperature shift of this microorganism. DETAILED DESCRIPTION
[0100] Next generation sequencing (NGS) can be used to aggregate a large amount of data about the nucleic acid content of a sample. The data can be particularly useful for analyzing nucleic acids in complex samples, such as clinical samples. To date, these NGS systems have focused on determining the abundance of individual reads. Prior to performing this work, the primary properties of interest are the sequence of each read and the abundance of reads associated with a particular source. This is particularly true for microbial nucleic acids and free microbial nucleic acids. This is due in part to the fact that the prior sample processing required by many NGS systems often results in errors and biases, which are particularly true for low abundance nucleic acids. Karius has developed methods for preparing nucleic acid libraries from initial samples that either reduce the bias in recovering nucleic acid libraries from initial samples or allow for correction of the bias. The reduced bias in nucleic acid libraries obtained from initial samples allows for the development of methods for fragment length profiling and generating a fragment length profile of a nucleic acid library or a target nucleic acid within a nucleic acid library. There is a need for effective and accurate methods for generating a fragment length profile of a nucleic acid library. This need can be seen, for example, in distinguishing between closely related microorganisms, determining whether a microorganism is present as a pathogen or a commensal microorganism, determining the biological relationship of a microorganism to a host, predicting the site of infection or colonization of a subject, monitoring the status of a transplant, monitoring fetal development and status, monitoring a tumor, monitoring the status and response of the immune system, and monitoring the toxicity of a compound administered to a subject.
[0101] A fragment length profile includes one or more fragment length characteristics of a nucleic acid library or a subset of reads from within a nucleic acid library. A fragment length profile can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more fragment length characteristics. One or more fragment length characteristics in a fragment length profile can be assigned a weighting value such that the one or more fragment length characteristics can have equal or different weights or values within the fragment length profile. Fragment length characteristics include, but are not limited to, shape of distribution, amplitude of segment, peak shape, ratio of segment counts of two or more segments, height of helical phasing peak, ratio of segment counts at two different fragment lengths, ratio of segment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, position of one or more peaks, and fragment length distribution within a subset of reads. It is intended that the ratio between “two or more segments” encompasses, but is not limited to, two or more segments from one nucleic acid library, two or more segments from two or more nucleic acid libraries, two or more segments of the same peak shape, two or more segments of different peak shape, two or more segments from similar or different nucleic acid library types, and two or more segments from similar or different subsets of reads out of a nucleic acid library.
[0102] Distribution types include, but are not limited to, single peak shape, multiple peak shape, exponential or quasi-exponential distribution, distribution of elongated or short fragments, flat or uniform distribution, complex distribution shape, and combinations thereof. A complex distribution can include aspects of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, or more peak shapes. A single peak shape can exist around any fragment length, including but not limited to around 50 base pair fragment length. Long fragments can include fragment lengths greater than about 60 base pairs, about 65 base pairs, about 70 base pairs, about 75 base pairs, about 80 base pairs, about 85 base pairs, about 90 base pairs, about 95 base pairs, about 100 base pairs, about 150 base pairs, about 175 base pairs, about 200 base pairs, about 250 base pairs, about 300 base pairs, about 350 base pairs, and about 400 base pairs. Short fragments can include fragment lengths shorter than about 500 bp, about 400 bp, about 300 bp, about 200 bp, about 100 bp, about 50 bp, about 40 bp, about 35 bp, about 30 bp, about 25 bp, about 20 bp. Aspects of peak shape include, but are not limited to, segment range, segment amplitude, and total number of reads within a segment, peak width, slope of peak, derivative of peak; aspects of peak shape can vary.
[0103] A single peak shape distribution can encompass a range of fragment lengths, including but not limited to a range of fragment lengths of at least about 5 base pairs, at least about 10 base pairs, at least about 15 base pairs, at least about 20 base pairs, at least about 30 base pairs, at least about 35 base pairs, at least about 40 base pairs, or greater than at least about 45 base pairs within a segment. The range of fragment lengths within a segment can vary. For example, a range of fragment lengths for a 50 base pair or so single peak distribution includes but is not limited to a range of fragment lengths of 30 to 60 base pairs, 35 to 60 base pairs, 40 to 60 base pairs, and 45 to 55 base pairs.
[0104] A distribution amplitude encompasses the abundance or relative abundance of reads of a defined fragment length within a defined segment. In some aspects, a distribution amplitude can be the highest abundance or relative abundance within a defined range of fragment lengths; a distribution amplitude can also encompass the average highest abundance or relative abundance within a defined range of fragment lengths. In some aspects of the application, a fragment length distribution or fragment length distribution profile is obtained for a subset of reads from a nucleic acid library. A subset of reads from a nucleic acid library is intended to encompass less than the complete collection of reads from a nucleic acid library. A subset can reflect reads determined to be from a particular microorganism type, from a particular microorganism species, host reads, maternal reads, fetal reads, organ donor reads, non-host reads, microorganism free nucleic acid reads, free nucleic acid reads, microorganism reads, or any other group; alternatively, a subset of reads can reflect the complete collection of reads minus those from a particular microorganism type, maternal reads, fetal reads, or any other group. In some aspects of the application, a fragment length distribution is obtained for a target nucleic acid. A "target nucleic acid" can be a nucleic acid fragment derived from a microorganism, a transplanted organ, a tumor cell, a cancer cell, host or non-host mitochondrial DNA, an antibiotic resistance gene sequence, a host genomic DNA, a microorganism sequence integrated into a host genome, or any other sequence of interest in a nucleic acid library. A target sequence can be migrated from another site, such as an infection site or a donated organ.
[0105] In some cases, the target nucleic acid can constitute only a very small fraction of the entire sample, e.g., less than 0.1%, less than 0.01%, less than 0.001%, less than 0.0001%, less than 0.00001%, less than 0.000001%, less than 0.0000001% of the total nucleic acid in the sample. Generally, the total nucleic acid in the original sample can vary. For example, the total free nucleic acid (e.g., DNA, mRNA, RNA) can be in the range of 0.01-10,000 ng / ml (e.g., about 0.01, 0.1, 1, 5, 10, 20, 30, 40, 50, 80, 100, 1000, 5000, 10000 ng / ml). In some cases, the total concentration of free nucleic acid in the sample is outside this range (e.g., less than 0.01 ng / ml; in other words, greater than 10,000 ng / ml). The same is true for free nucleic acid (e.g., DNA) samples that are primarily composed of human DNA and / or RNA. In such samples, the presence of pathogen target nucleic acids can be less than the human or host nucleic acid.
[0106] The length of the target nucleic acid can vary. In some particular embodiments, the target nucleic acid is relatively short; in other embodiments, the target is relatively long. In some particular embodiments, the target nucleic acid is shorter than 110 bp.
[0107] As used herein, “nucleic acid” refers to a polymer or oligomer of nucleotides, and is generally synonymous with the term “polynucleotide” or “oligonucleotide.” A nucleic acid can include, consist of, or consist essentially of, deoxyribonucleotides, ribonucleotides, deoxyribonucleotide analogs, chemically modified deoxyribonucleotides, ribonucleotides, and / or ribonucleotide analogs, nucleic acids having a modified backbone, or any combination thereof.
[0108] The nucleic acid can be any type of nucleic acid, including but not limited to: double-stranded (ds) nucleic acid, single-stranded (ss) nucleic acid, DNA, RNA, cDNA, mRNA, cRNA, tRNA, ribosomal RNA, dsDNA, ssDNA, miRNA, siRNA, short hairpin RNA, circulating nucleic acid, circulating cell-free nucleic acid, circulating DNA, circulating RNA, cell-free nucleic acid, cell-free DNA, cell-free RNA, circulating cell-free DNA, cell-free dsDNA, cell-free ssDNA, circulating cell-free RNA, genomic DNA, exosome, cell-free pathogen nucleic acid, circulating microbial or pathogen nucleic acid, mitochondrial nucleic acid, non-mitochondrial nucleic acid, nuclear DNA, nuclear RNA, chromosomal DNA, circulating tumor DNA, circulating tumor RNA, circular nucleic acid, circular DNA, circular RNA, circular single-stranded DNA, circular double-stranded DNA, plasmid, bacterial nucleic acid, fungal nucleic acid, parasitic nucleic acid, viral nucleic acid, cell-free bacterial nucleic acid, cell-free fungal nucleic acid, cell-free parasitic nucleic acid, viral particle-associated nucleic acid, mitochondrial DNA, host nucleic acid, host cell-free nucleic acid, intercellular signaling nucleic acid, exogenous nucleic acid, DNAse, RNAse, therapeutic nucleic acid, or any combination thereof. The nucleic acid can be a nucleic acid derived from a microbe or pathogen, including but not limited to viruses, bacteria, fungi, parasites, and any other microbe, particularly an infectious microbe or a potentially infectious microbe. The nucleic acid can be derived from an archaea, bacteria, fungi, mold, eukaryote, and / or virus. In some embodiments, the nucleic acid can be derived directly from the subject or host as opposed to a microbe or pathogen.
[0109] As used herein, a "nucleic acid library" refers to a collection of nucleic acid fragments. The collection of nucleic acid fragments can be used, for example, for sequencing. A nucleic acid library can be prepared from an initial sample using a bias-corrected recovery method that generates a sequencing library or using a bias recovery method that generates a sequencing library that achieves bias correction. As used herein, a "bias-corrected recovery" method is: a method that utilizes consistent fragment lengths that recovers sample nucleic acid fragments within a targeted length and GC range without appreciable length and GC bias; a method that achieves bias correction; a method that addresses bias of a sample; and a method that addresses bias introduced by the process of generating a nucleic acid library. Bias-corrected recovery methods can include, but are not limited to, adding process control molecules, extraction, library generation, sequencing, amplification, and any combination thereof. Bias-free recovery methods include, but are not limited to, those described in U.S. Provisional Nos. 62 / 770,181 and 62 / 644,357. Methods are provided for generating a nucleic acid library from an initial sample without extracting nucleic acids from the initial sample prior to starting the nucleic acid library generation process. In some embodiments, substances that can themselves reduce yield or inhibit the generation of a nucleic acid library can be extracted or removed, but nucleic acids are not extracted from the initial sample prior to nucleic acid library generation. Methods include, consist of, or consist essentially of adding one or more process control molecules to an initial sample, and generating a nucleic acid library from the spiked initial sample. Methods include, consist of, or consist essentially of generating a nucleic acid library from the spiked initial sample. The nucleic acid library can utilize single-stranded and / or double-stranded nucleic acids.
[0110] Methods for generating a nucleic acid library from a sample with extraction are also contemplated.
[0111] The process control molecules can be one or more of an ID spike, a SPANK, a Spark, or GC spike panel, a dephosphorylation control molecule, a denaturation control molecule, and / or a ligation control molecule. See, e.g., published U.S. Patent Application No. 2015-0133391 and published U.S. Patent Application No. 2017-0016048, the entire disclosure of each of which is incorporated by reference herein in its entirety for all purposes. In some embodiments, the initial sample includes, consists of, or consists essentially of circulating donor nucleic acids (see, e.g., US20150211070, which is incorporated by reference herein in its entirety, including any drawings).
[0112] As used herein, "denaturation" refers to a process in which a biological molecule, such as a protein or nucleic acid, loses its native or higher order structure. Native and higher order structures can include, for example, but are not limited to, quaternary structure, tertiary structure, or secondary structure. For example, a double-stranded nucleic acid molecule can be denatured into two single-stranded molecules.
[0113] As used herein, the term "dephosphorylation" or "dephosphorylating" refers to the removal of a phosphate, such as a 5' and / or 3' end phosphate, from a nucleic acid, such as DNA.
[0114] As used herein, "detecting" is a designated amount or qualitative detection, including but not limited to detection by identifying the presence, absence, amount, frequency, concentration, sequence, form, structure, source, or quantity of an analyte.
[0115] In some embodiments, ligating the 3' end adapter to the nucleic acid, e.g., denatured or dephosphorylated nucleic acid, and / or ligating the 5' end adapter comprises, consists of, or consists essentially of ligation to an enzyme, including a ligase, e.g., T4 DNA ligase, CircLigase II, consisting of, or consisting essentially of a ligase. In some embodiments, the ligase is a single-stranded ligase. In some embodiments, ligating the 3' end adapter to the nucleic acid, e.g., denatured or dephosphorylated nucleic acid, and / or ligating the 5' end adapter comprises, consists of, or consists essentially of a template switch reaction. In some embodiments, ligating the 3' end adapter to the nucleic acid, e.g., denatured or dephosphorylated nucleic acid, comprises, consists of, or consists essentially of enzymatic extension, including a polymerase, e.g., TdT polymerase, consisting of, or consisting essentially of a polymerase. In some embodiments, the method further comprises, consists of, or consists essentially of utilizing a DNA polymerase, e.g., Klenow fragment, Superscript IV reverse transcriptase, SMART MMLV reverse transcriptase, etc., to extend a primer hybridized to the nucleic acid or the adaptered nucleic acid and generate a complementary strand. In some embodiments, the target nucleic acid can be ligated to one or more adapters. In some embodiments, the target nucleic acid is ligated to the same adapter or different adapters at both ends.
[0116] As used herein, "GC bias" refers to differential performance, processing, or recovery of nucleic acids of different GC content but of the same length.
[0117] As used herein, "GC content" or "guanine-cytosine content" refers to the percentage of nitrogenous bases in a nucleic acid, such as a DNA or RNA molecule, that are guanine or cytosine or chemical modifications thereof.
[0118] As used herein, "host" refers to an organism that has another organism. The latter is defined as a "non-host" organism. For example, a human can be a host to a microorganism, pathogen, or fetus, which is a non-host. Host nucleic acids or materials are derived from the host. Non-host nucleic acids or materials can be derived from a non-host organism, from transplanted materials, or from a fetus or fetal material within the host.
[0119] As used herein, "microbe," "microbial," or "microorganism" refers to an organism that can exist as a single cell or as a colony of cells, capsids, spores, filaments, or multicellular organisms, such as microscopic or macroscopic organisms. Microorganisms include all single-celled organisms and some multicellular organisms, such as those from archaea, bacteria, protozoa, nematodes, viruses, and eukaryotes. Microorganisms are often the causative agents of disease, but can also exist in non-pathogenic, commensal relationships with a host, such as a human. "Commensal microorganism" is intended to include microorganisms that exist in non-pathogenic, commensal relationships with a host. A host organism can have multiple types of non-host organisms simultaneously. In a co-infection, a host organism has multiple types of non-host organisms. The multiple types of non-host organisms can include one or more pathogens, one or more commensal microorganisms, or at least one pathogen and at least one commensal microorganism. The methods of the instant application can be used to distinguish between closely related microorganisms, between microorganisms that exist as pathogens, commensal microorganisms, or as incidental but clinically unimportant microorganisms.
[0120] The microorganism or pathogen can comprise an archaea, bacteria, yeast, fungus, mold, protozoa, nematode, eukaryote, and / or virus. The microorganism or pathogen can also comprise a DNA virus, RNA virus, cultivable bacteria, additional fastidious and uncultivable bacteria, mycobacteria, and eukaryotic pathogens (see Bennett J.E., Mandell, D., Blaser, M.J., Principles and Practice of Infectious Diseases, 8thEd. Saunders, Philadelphia, PA, 2014; and Netter's Infectious Disease, 1stEd. Edited by Elaine C. Jong, MD and Dennis L. Stevens, MD, PhD, (2015)). The microorganism or pathogen can also comprise any of the microorganisms shown at https: / / www.ncbi.nlm.nih.gov / genome / microbes / or https: / / www.ncbi.nlm.nih.gov / biosample / .
[0121] Examples of microorganisms are one or more of the species or strains from one or more of the following genera: Coniosporium, Hantavirus, Talaromyces, Machlomovirus, Betatetravirus, Raoultella, Aeromonas, Ephemerovirus, Ephemerovirus, Loa, Macluravirus, Stenotrophomonas, Alfamovirus, Rosavirus, Emmonsia, Aggregatibacter, Orthopneumovirus, Weeksella, Nairovirus, Salivirus, Weissella, Mosavirus, Gammapartitivirus, Strongyloides, Passerivirus, Erysipelatoclostridium, Bacillarnavirus, Iotatorquevirus, Taenia, Trypanosoma, Olsenella, Cladosporium, Rhizobium, Prevotella, Leclercia, Paracoccus, Ilarvirus, Lagovirus, Rasamsonia, Plasmodium, Acremonium, Chlamydia, Clonorchis, Vibrio, Bartonella, Nakazawaea, Franconibacter, Anisakis, Norovirus, Nocardia,Solobacterium, Parechovirus, Avenavirus, Orthohepevirus, Aphthovirus, Hepandensovirus, Microbacterium, Lichtheimia, Lomentospora, Achromobacter, Ipomovirus, Tsukamurella, Elizabethkingia, Hepevirus, Seadornavirus, Alternaria, Trueperella, Gammatorquevirus, Bifidobacterium, Chrysosporium, Thogotovirus, Curtovirus, Deltatorquevirus, Balamuthia, Mastrevirus, Bdellomicrovirus, Mupapillomavirus, Pseudozyma, Wickerhamiella, Aquamavirus, Alloscardovia, Thielavia, Idaeovirus, Henipavirus, Coxiella, Haemophilus, Gammacoronavirus, Negevirus, Brevibacterium, Peptoniphilus, Alphacarmotetravirus, Nosema, Trichovirus, Arenavirus, Thermomyces, Necator, Waikavirus,Blosnavirus, Jonet, Tetraparvovirus, Emaravirus, Plectrovirus, Sclerodarnavirus, Toxocara, Umbravirus, Burkholderia, Chromobacterium, Paracoccidioides, Brugia, Eragrovirus, Macrococcus, Absidia, Colletotrichum, Inovirus, Phycomyces, Wickerhamomyces, Acidaminococcus, Moraxella, Rothia, Phlebovirus, Slackia, Purpureocillium, Betapapillomavirus, Tupavirus, Cryspovirus, Saksenaea, Erysipelothrix, Kobuvirus, Mimoreovirus, Echinococcus, Mannheimia, Bergeyella, Cyclospora, Xylanimonas, Leptospira, Finegoldia, Curvularia, Cryptosporidium, Babuvirus, Pecluvirus, Lambdatorquevirus, Pythium, Carlavirus, Entomobirnavirus, Kocuria, Anaplasma, Ampelovirus,Avihepatovirus, Nepovirus, Rhodococcus, Bordetella, Mischivirus, Scedosporium, Gardnerella, Maculavirus, Trichoderma, Aveparvovirus, Salmonella, Avastrovirus, Copiparvovirus, Trachipleistophora, Clostridioides, Nanovirus, Siccibacter, Leptotrichia, Citrivirus, Odoribacter, Sanguibacter, Novirhabdovirus, Acremonium, Hafnia, Chaetomium, Tenuivirus, Yokenella, Rubulavirus, Varicellovirus, Alphamesonivirus, Sicinivirus, Leuconostoc, Microvirus, Gallantivirus, Morbillivirus, Lolavirus, Pantoea, Hepatovirus, Nupapillomavirus, Metschnikowia, Barnavirus, Kytococcus, Tritimovirus, Tannerella, Respirovirus, Pneumocystis, Dirofilaria, Pediococcus, Lactococcus, Blastomyces, Dianthovirus,Actinobacillus, Teschovirus, Oscivirus, egomovirus, Potyvirus, Byssochlamys, lphacoronavirus, Molluscipoxvirus, Lymphocryptovirus, Sapelovirus, Parabacteroides, Pyrenochaeta, Listeria, Senecavirus, Brevidensovirus, Potexvirus, Parvimonas, Flavivirus, Recovirus, Toxoplasma, Yatapoxvirus, Opisthorchis, Trichuris, Cyphellophora, Morganella, Perhabdovirus, Micrococcus, Pequenovirus, Mastadenovirus, Anaeroglobus, Tropheryma, Dolosigranulum, Wolbachia, Lelliottia, Mycoplasma, Tobravirus, Shewanella, Paeniclostridium, Erythroparvovirus, Sutterella, Sporopachydermia, Narnavirus, Nyavirus, Francisella, Arthroderma, Epsilon torquevirus, Sigmavirus, Amdoparvovirus, Actinomyces,Alphapermutotetravirus, Cardiobacterium, Influenzavirus C, Orthopoxvirus, Poacevirus, Phialophora, Lactobacillus, Polyomavirus, Debaryomyces, Foveavirus, Bymovirus, Mycoflexivirus, Grimontia, Mucor, Rhytidhysteron, Quadrivirus, Thermoascus, Aureusvirus, Trichosporon, Myceliophthora, Dermacoccus, Dysgonomonas, Pseudoramibacter, Becurtovirus, Gordonia, Sapovirus, Orthobunyavirus, Spiromicrovirus, Pomovirus, Exophiala, Sneathia, Helicobacter, Photorhabdus, Mogibacterium, Betapartitivirus, Avibirnavirus, Ambidensovirus, Oleavirus, Orientia, Deltacoronavirus, Anulavirus, Trichomonasvirus, Budvicia, Geotrichum, Enamovirus, Lachnoclostridium, Schistosoma, Paecilomyces,Panicovirus, Rhizoctonia, Brevibacillus, Beauveria, Pestivirus, Tombusvirus, Cilevirus, Cokeromyces, Peptostreptococcus, Phanerochaete, Proteus, Invertebrate iridescent virus, Aspergillus, Pasteurella, Malassezia, Hanseniaspora, Endornavirus, Azospirillum, Velarivirus, Cystovirus, Avisivirus, Bacteroides, Picobirnavirus, Myroides, Circovirus, Arterivirus, Aquaparamyxovirus, Onchocerca, Cosavirus, Kluyveromyces, Fijivirus, Candida, Hepacivirus, Dermabacter, Ourmiavirus, Allexivirus, Enterobacter, Acidovorax, Bracorhabdovirus, Carmovirus, Pluralibacter, Coltivirus, Fonsecaea, Streptobacillus, Corynebacterium, Macrophomina, Marburgvirus, Comovirus, Fabavirus, Alphanodavirus, Cellulomonas,Enterobius, Catabacter, Moellerella, Nakaseomyces, Cucumovirus, Valsa, Deltapartitivirus, Plesiomonas, Pseudomonas, Torovirus, Cuevavirus, Hypovirus, Trichomonas, Influenzavirus D, Giardiavirus, Crinivirus, Tepovirus, Sakobuvirus, Cyberlindnera, Paenalcaligenes, Bafinivirus, Rymovirus, Pegivirus, Yarrowia, Treponema, Borreliella, Rubivirus, Aureobasidium, Angiostrongylus, Filobasidium, Photobacterium, Rhizopus, Orthoreovirus, Ustilago, Simplexvirus, Aquareovirus, Protoparvovirus, Propionibacterium, Sprivivirus, Hunnivirus, Apophysomyces, Meyerozyma, Alphapapillomavirus, Candida, Brucella, Gallivirus, Dinovernavirus, Anaerobiospirillum, Eubacterium, Tatlockia,Terrisporobacter, Quaranjavirus, Sobemovirus, Dicipivirus, Arcanobacterium, Macanavirus, Atopobium, Vesivirus, Lodderomyces, Dinornavirus, Betatorquevirus, Kerstersia, Aparavirus, Neisseria, Agrobacterium, Edwardsiella, Labyrnavirus, Totivirus, Actinomadura, Tobamovirus, Influenzavirus B, Mandarivirus, Anaerococcus, Kunsagivirus, Naegleria, Campylobacter, Veillonella, Yamadazyma, Filobasidiella, Oerskovia, Penicillium, Anncaliia, Leptosphaeria, Pneumovirus, Psychrobacter, Isavirus, Granicella, Torradovirus, Cladophialophora, Influenzavirus A, Ophiostoma, Aerococcus, Ureaplasma, Etatorquevirus, Bocaparvovirus, Megasphaera, Reptarenavirus, Comamonas, Capnocytophaga,Alphatorquevirus, Syncephalastrum, Wallemia, Betacoronavirus, Hyphopichia, Nocardiopsis, Legionella, Trichinella, Paraburkholderia, Mammarenavirus, Echinostoma, Sphingobacterium, Enterovirus, Methanobrevibacter, Ochroconis, Cheravirus, Pasivirus, Enterococcus, Mycoreovirus, Tospovirus, Betanodavirus, Phytoreovirus, Enterocytozoon, Ferlavirus, Stemphylium, Filifactor, Leishmaniavirus, Gemella, Bromovirus, Alloiococcus, Cunninghamella, Cronobacter, Oribacterium, Orbivirus, Chrysovirus, Cripavirus, Tatumella, Pandoraea, Ogataea, Dracunculus, Volvariella, Iflavirus, Benyvirus, Rhadinovirus, Histoplasma, Rahnella, Morococcus, Verticillium, Janibacter, Gyrovirus,Alphapartitivirus, Mycobacterium, Roseomonas, Varicosavirus, Chryseobacterium, Parapoxvirus, Rhizomucor, Aureimonas, Levivirus, Leishmania, Luteovirus, Cypovirus, Ochrobactrum, Microsporum, Piscihepevirus, Ceratocystis, Sporothrix, Vesiculovirus, Cupriavidus, Cryptococcus, Metapneumovirus, Alphanecrovirus, Eikenella, Brevundimonas, Escherichia, Leifsonia, Schizophyllum, Granulibacter, Gordonibacter, Lachancea, Madurella, Ophiovirus, Phellinus, Nebovirus, Acanthamoeba, Fusobacterium, Pichia, Verruconis, Ehrlichia, Tibrovirus, Higrevirus, Wohlfahrtiimonas, Rhinocladiella, Neorickettsia, Sadwavirus, Roseobacter, Sequivirus, Pannonibacter, Rotavirus, Turicella,Cardiovirus, Propionimicrobium, Furovirus, Naumovozyma, Closterovirus, Fluoribacter, Zeavirus, Clavispora, Megrivirus, Gammapapillomavirus, Rickettsia, Polemovirus, Corynespora, Encephalitozoon, Shimwellia, Fusarium, Yersinia, Capronia, Delftia, Victorivirus, Marafivirus, Kluyvera, Iteradensovirus, Isoptericola, Vitivirus, Roseolovirus, Conidiobolus, Abiotrophia, Babesia, Phoma, Sanguibacteroides, Staphylococcus, Rhodotorula, Zetatorquevirus, Hymenolepis, Fasciola, Cytorhabdovirus, Cardoreovirus, Memnoniella, Trichophyton, Mitovirus, Phaeoacremonium, Providencia, Lysinibacillus, Giardia, Oligella, Streptomyces, Paraclostridium, Ralstonia, Coccidioides,Brambyvirus, Biatriospora, Allolevivirus, Acinetobacter, Starmerella, Omegatetravirus, Porphyromonas, Avulavirus, Streptococcus, Arcobacter, Topocuvirus, Mamastrovirus, Ancylostoma, Bornavirus, Capillovirus, Alphavirus, Tymovirus, Nucleorhabdovirus, Diaporthe, Chlamydiamicrovirus, Turncurtovirus, Saccharomyces, Riemerella, Betanecrovirus, Clostridium, Mobiluncus, Cercospora, Marnavirus, Mortierella, Aquabirnavirus, Xanthomonas, Dependoparvovirus, Ebolavirus, Neofusicoccum, Borrelia, Leminorella, Klebsiella, Blastocystis, Alcaligenes, Citrobacter, Eggerthella, Cedecea, Serratia, Penstyldensovirus, Bacillus, Laribacter, Wuchereria, Hordeivirus, Cytomegalovirus,Actinomucor, Ascaris, Shigella, Vittaforma, Torulaspora, Kingella, Oryzavirus, Polerovirus, Tremovirus, Erbovirus, Entamoeba, Lyssavirus, Paenibacillus, Facklamia, Kappatorquevirus, Metarhizium, Stachybotrys, Okavirus, Botrexvirus, Thetatorquevirus, and Basidiobolus.
[0122] As used herein, infection stage or stage of infection refers to the pre-symptomatic infection period, symptomatic infection period, resolving infection period, treatment period, relapse period, reoccurrence period, acute period or infection, chronic period or infection, slow or latent period or infection, persistent infection, disseminated infection stage, first, second, or third period of infection. The pre-symptomatic infection period occurs before symptoms appear or before the subject or other person notices symptoms. Synonyms for the “pre-symptomatic period” would include “pre-symptomatic infection stage,” “first period of infection,” and “early infection stage.” The symbiont can persist during the pre-symptomatic stage of infection. The symptomatic infection period occurs when the subject or other person notices symptoms or clinical changes, such as fever, pain, rash, headache, malaise, respiratory problems, etc. The resolving infection stage occurs during the period when the infection resolves, either by itself or by administration of a treatment. The treatment period can be part of the resolving period of the administration of a treatment. The relapse period occurs when the subject experiences a relapse of the infection in any of the above stages. The reoccurrence period occurs when the infection comes back without proper or adequate treatment of the infection at the first time and the infection comes back. Chronic infection is a type of persistent infection that eventually gets cleared. The acute period or infection occurs suddenly, such as hepatitis. The slow or latent period or infection is an infection that persists for the rest of the host’s life. Persistent infection is an infection that persists for a long period of time; persistent infection occurs when the host does not clear the primary infection. Some microorganisms infect the host with first, second, and third periods of infection; examples are Treponema pallidum (syphilis), Plasmodium (malaria), and HIV (AIDS). treponema palliduminfection. Infection can remain at any of the above stages for an indefinite period of time and does not necessarily progress to a different stage. Symbiotic or commensal microorganisms can remain indefinitely in the non-pathogenic stage of infection or can not infect.
[0123] A variety of host-microbe biological relationships or interactions are known in the art. Host-microbe biological interactions include, but are not limited to, symbiosis, mutualism, facultative parasitism, parasitism, commensalism, and competition. It is recognized that a microbe can exhibit one type of interaction with a host when it is located at certain sites within the host, but can exhibit another type of interaction with the host when it is located at another site. For example, a microbe can exist in a symbiotic relationship with a host on the host’s skin, but can exist in a parasitic or competitive relationship inside the host. As used herein, “pathogen” refers to a microbe that causes or can cause or is suspected of being able to cause disease.
[0124] As used herein, the phrase “spiked initial sample” refers to an initial sample to which a process control molecule has been added prior to beginning generation of a sequencing library.
[0125] The term “derived from” encompasses the terms “originated from,” “obtained from,” “obtainable from,” and “produced from,” which generally indicate that one specified material is derived from another specified material or has a characteristic that can be described with reference to another specified material. For example, an initial sample can be derived from an original biological sample.
[0126] In some embodiments, the initial sample comprises, consists of, or consists essentially of a solid or a bodily fluid, such as blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, stool, saliva, peritoneal fluid, ascites fluid, peritoneal lavage fluid, gastric fluid, interstitial fluid, lymphatic fluid, bile, abscess fluid, tissue, amniotic fluid, meconium, sinus aspirate, lymph node, bone marrow, hair, fingernail, cheek swab, skin swab, urethral swab, cervical swab, nasopharyngeal swab, nasopharyngeal aspirate, vaginal swab, epithelial cells, semen, vaginal discharge, intercellular fluid, pericardial fluid, rectal swab, skeletal, skin tissue, soft tissue, tears, and / or nasal sample. In some embodiments, the initial sample comprises, consists of, or consists essentially of plasma. In some embodiments, the initial sample comprises, consists of, or consists essentially of urine. In some embodiments, the initial sample comprises, consists of, or consists essentially of cerebrospinal fluid. In some embodiments, the initial sample is from a human subject.
[0127] In some embodiments, the initial sample can consist, in whole or in part, of cells and / or tissues. The initial sample can be free or cell-depleted. The initial free sample can include, consist of, or consist essentially of nucleic acids derived from different sites in the body, such as the site of a pathogen infection. In the case of blood, serum, lymph, or plasma, the free sample or cell-depleted initial sample can contain "circulating" free nucleic acids originating at an anatomical location other than the site of bodily fluid collection for the fluid in question. In the case of urine, the free nucleic acids can be free nucleic acids derived from different sites within the body. The free sample or cell-depleted initial sample can be obtained by means of depletion or removal of cells, cell fragments, or exosomes by known techniques, such as by centrifugation or filtration.
[0128] As used herein, the term "invasive disease" refers to a disease based in part on the ability of a particular pathogen to severely compromise the health of certain infected subjects as opposed to merely colonizing other infected subjects in a commensal or mildly symptomatic infection. For example, certain microorganisms can colonize tissues locally in some hosts without causing any health problems, while in other hosts, it can invade tissues to the point where it causes severe inflammation, tissue or organ damage, sepsis, cancer, and other serious health problems. The microorganism can also colonize a subject who is asymptomatic at a certain point, but develop severe symptoms at a later point when the microorganism translocates and / or becomes "active."
[0129] As used herein, the term "free" refers to the condition of nucleic acids when they are outside of cells, viral particles, or virions at the point at which the sample is about to be obtained from the body. For example, circulating free nucleic acids in a sample can originate from free nucleic acids circulating in the bloodstream of a subject. In contrast, nucleic acids collected after extraction from intact microorganisms, such as blood-borne pathogens, or removal from intact virions in a plasma sample are not generally considered to be "free."
[0130] The present application provides methods of determining a localization site of a subject. Nucleic acids from a microbe or microbes from different sites within a subject can exhibit different fragment length profiles. If a microbial infection is circulating rather than localized to one or more localization sites, the fragment length profile of a nucleic acid library or subset of a nucleic acid library containing microbial nucleic acids is different. Thus, comparing the fragment length profile to a reference fragment length profile of one or more source sites can predict a localization site when the fragment length profile from the sample is similar to the reference fragment length profile from a source site. A "localization site" refers to any source site within a subject where a microbe is present, persists, survives, or proliferates. Source sites include, but are not limited to, the bloodstream, blood, deep tissue such as, but not limited to, the kidneys, liver, stomach, bladder, digestive organs, nerve cells, lungs, bones, brain, heart, heart lining, sinuses, GI tract, spleen, skin, joints, ear, nose, and mouth. It is contemplated that a subject can have more than one localization site for a particular microbe. It is further understood that some localization sites of a particular microbe can not contribute to a disease state or condition. Rather, some localization sites of a particular microbe can indicate a symbiotic relationship between the microbe and the host, while other localization sites of a particular microbe can indicate a parasitic or commensal relationship between the microbe and the host. It is further recognized that the presence of multiple localization sites of a particular microbe can indicate a systemic infection of the host. Additionally, it is recognized that a localization site of a particular microbe or pathogen of interest can impact the decision to treat or not treat, and can impact the selection of an appropriate treatment option. For example, and without being limited by mechanism, a fungal pathogen localized to the skin can be treated differently than a fungal pathogen localized to the lungs, and a bacterial microbe localized to heart tissue, including but not limited to heart lining, can be treated differently than a bacterial microbe localized to the blood or bloodstream.
[0131] In some embodiments, the initial sample comprises, consists of, or consists essentially of circulating tumor or fetal nucleic acids. (See, e.g., the analysis of serum or blood-derived nucleic acids, such as circulating tumor or fetal nucleic acids, described in U.S. Patents Nos. 8,877,442 and 9,353,414, or pathogen identification by, e.g., analysis of circulating microbial or viral nucleic acids, as described in published U.S. Patent Application No. 2015-0133391 and published U.S. Patent Application No. 2017-0016048, the entire disclosures of which are incorporated by reference herein in their entireties for all purposes). In some embodiments, the initial sample comprises, consists of, or consists essentially of circulating donor nucleic acids (see, e.g., US 20150211070, which is incorporated by reference herein in its entirety, including any drawings).
[0132] The initial sample can be derived from any subject (e.g., a human subject, a non-human subject, etc.). The subject can be healthy. In some embodiments, the subject is a human patient who has, is suspected of having, or is at risk of having a disease or infection. In some embodiments, the disease or infection is pathogen-associated.
[0133] The human subject can be male or female. In some embodiments, the sample can be from a human embryo or a human fetus. In some embodiments, the human can be an infant, a child, a young adult, an adult, or an elderly adult. In some embodiments, the subject is a female subject who is pregnant, suspected of being pregnant, or planning to become pregnant.
[0134] In some embodiments, the subject is a human subject who has undergone or is planning to undergo an organ transplant.
[0135] In some embodiments, the subject is a farm animal, a laboratory animal, or a domestic pet. In some embodiments, the animal can be an insect, a dog, a cat, a horse, a cow, a mouse, a rat, a pig, a fish, a bird, a chicken, or a monkey.
[0136] The subject can be an organism, such as a single-celled or multi-cellular organism. In some embodiments, the sample can be obtained from a plant, a fungus, a eubacterium, an archaebacterium, a protist, or any multi-cellular organism. The subject can be a cultured cell, which can be a primary cell or a cell from an established cell line.
[0137] In some embodiments, the subject has, is affected by, or is at risk of having a genetic disease or disorder. The genetic disease or disorder can be associated with a genetic variation, such as a mutation, an insertion, an addition, a deletion, a translocation, a point mutation, a trinucleotide repeat disorder, a single nucleotide polymorphism (SNP), or a combination of genetic variations.
[0138] In some aspects, the subject is healthy or asymptomatic, or exhibits mild or non-specific clinical symptoms. In some cases, the subject can be infected or suspected of being infected with a specific pathogen. In other cases, the subject is suspected of having an infection of unknown origin. In some cases, the subject has been exposed to a pathogen, or is suspected of having been exposed to a pathogen, such as through living conditions, through travel to a specific geographic region, or through interaction or sexual interaction with an infected individual.
[0139] The initial sample can be from a subject having a particular disease, condition, or infection, or suspected of having (or at risk of having) a particular disease, condition, or infection. For example, the initial sample can be from a cancer patient, a patient suspected of having cancer, or a patient at risk of having cancer. In some embodiments, the initial sample can be from a patient having an infection, a patient suspected of having an infection, or a patient at risk of having an infection. In some embodiments, the initial sample is from a subject who has undergone or will undergo an organ transplant.
[0140] The primer extension reaction can be performed with a DNA-dependent polymerase or an RNA-dependent polymerase or reverse transcriptase or a combination thereof. In some embodiments, the primer extension reaction can be performed by a DNA or RNA polymerase having strand displacement activity. In some embodiments, the primer extension reaction is performed by a DNA or RNA polymerase having non-templated activity. In some other embodiments, the primer extension reaction can be performed by a DNA or RNA polymerase having strand displacement activity and a DNA or RNA polymerase having non-templated activity. In some embodiments, the primer extension is performed with a Klenow fragment.
[0141] The reference fragment length profile is typically predetermined. One or more suitable reference fragment length profiles can vary depending on the method, type of comparison, or purpose of the method. One of skill in the art will select one or more appropriate reference fragment length profiles. The reference fragment length profile can be obtained from a subject or cell exposed to the compound of interest, a subject or cell exposed to a similar compound, obtained from a subject or cell similar to the subject, obtained from a subject or cell having a known microorganism, obtained from a subject or cell previously determined to have an infection at the source site, or a subject or cell in any other condition of interest as determined by one of skill in the art to be suitable for use.
[0142] Subjects with transplants are at risk for transplant rejection, even when provided with therapies to reduce the risk of rejection. Transplant rejection and conditions of transplant rejection are significant to subjects with transplants, often presenting a life-threatening risk. Many anti-rejection therapies suppress the immune system of the subject, thereby increasing the risk of infection or disease in the subject. Thus, there is a need to balance the use and dosage of anti-rejection therapies. The present application provides methods of monitoring the status of a transplant in a subject. The methods include the steps of generating a baseline fragment length profile of target nucleic acids within a nucleic acid library generated from a sample obtained from the subject or a donor. Target nucleic acids of particular interest in monitoring the status of a transplant include, but are not limited to, donor and recipient mitochondrial DNA (mtDNA). Methods of monitoring the status of a transplant can further include evaluating the abundance of mitochondrial DNA from the transplant. Monitoring the status of a transplant encompasses monitoring anything related to the status of a transplant, including but not limited to, host rejection of the transplant, immune response of the host to the transplant, response of the host to the transplant, deterioration of the transplant, health of the transplant, vascularization of the transplant, oxygenation of the transplant, and malfunction of the transplant. The baseline fragment length profile can be generated from a donor and / or recipient sample obtained prior to, at, or after transplantation. The methods further include the steps of generating a second fragment length profile from a sample obtained from the subject, and comparing the second fragment length profile to the baseline fragment length profile. If the second fragment length profile is different from the baseline fragment length profile, an increased amount of anti-rejection therapy can be administered internally to the subject.
[0143] Methods and systems of the present disclosure can be implemented by one or more algorithms. Algorithms can be implemented by way of software when executed using a central processing unit. Algorithms may, for example, facilitate enrichment, sequencing, and / or detection of a pathogen or microorganism or other target nucleic acid, or generation of a fragment length profile.
[0144] A compound can include, but is not limited to, a chemotherapeutic agent, an antiviral agent, an antibiotic agent, an antifungal agent, a pharmaceutical agent of interest, a small molecule, an experimental pharmaceutical agent, a clinical trial compound, a drug product, a drug, and an active ingredient.
[0145] Toxicity includes, but is not limited to, cytotoxicity. It is further recognized that toxicity can preferentially occur in specific classes of cells, including but not limited to, cancer cells and pathogens.
[0146] Fragment length profiles and methods of the present application can be used for non-invasive prenatal testing (NIPT). The methods allow for non-invasive monitoring, diagnosis, and tracking of fetal conditions.
[0147] In some embodiments, isolating the adaptered nucleic acids comprises, consists of, or consists essentially of immobilizing the adaptered nucleic acids. In some embodiments, the immobilization occurs on a magnetic bead or a functionalized magnetic bead. In some embodiments, the immobilization occurs on a modified glass, a modified capillary surface, and / or a modified column. In some embodiments, isolating the adaptered nucleic acids comprises, consists of, or consists essentially of purifying the adaptered nucleic acids. In some embodiments, isolating the adaptered nucleic acids comprises, consists of, or consists essentially of precipitating the adaptered nucleic acids. In some embodiments, isolating the adaptered nucleic acids comprises, consists of, or consists essentially of using a 3' end-protected 3' end adapter. In some embodiments, isolating the adaptered nucleic acids comprises, consists of, or consists essentially of separating the adaptered nucleic acids from non-adaptered nucleic acids by digesting the non-adaptered nucleic acids with a 3' exonuclease, the adaptered nucleic acids comprising, consisting of, or consisting essentially of a 3' end-protected 3' end adapter. Some embodiments further comprise, consist of, or consist essentially of enriching the nucleic acids for fragments of a certain length. In some embodiments, denaturation is used to further isolate the nucleic acids or target nucleic acids. In some embodiments, the denaturation comprises, consists of, or consists essentially of selective denaturation. In some embodiments, the selective denaturation comprises, consists of, or consists essentially of one or more denaturation steps effective for selecting fragments of a certain length and / or GC content. In some embodiments, isolating fragments of a certain length can occur by using a protease, a detergent, heparin, hemolysis, and plasma concentration.
[0148] The methods provided herein include various non-invasive methods for subjects suffering from infection, subjects at risk of infection, and / or subjects experiencing undefined symptoms mimicking a variety of other diseases. The methods provided herein can be used for a variety of purposes, such as diagnosing or detecting infection, determining the stage of infection, predicting the stage of infection of a microorganism, predicting whether an infection will progress to an invasive stage of disease, monitoring efficacy and / or response to a treatment or procedure, terminating a treatment, determining a site of infection, determining a site of colonization, or modifying or optimizing a therapy for better clinical response. Thus, the methods provided herein can reduce the adverse effects caused by misdiagnosis or by invasive procedures, such as biopsies, used to determine whether an organ of a subject is infected, which organs of a subject are infected, and how an organ of a subject is infected.
[0149] Figure 1 An overview of some of the methods provided herein is provided. Generally, the methods can include obtaining a clinical sample from an infected subject or a subject at risk of infection; making a "spiked sample" by adding a synthetic nucleic acid provided by the disclosure; optionally, extracting the nucleic acids from the spiked sample; generating a spiked sample library; optionally, enriching for target nucleic acids of interest; performing a detection assay, such as a sequencing assay, to obtain sequence reads from the spiked sample library; and determining a measurement from the detected nucleic acids, and comparing this measurement to a control or reference to determine the stage of infection of the subject, the biological relationship between the microorganism and the host, or the site of colonization (e.g., organ or tissue type). In some cases, a comparison of the absolute abundance of the target nucleic acids to the control or reference can indicate the stage of infection or site of colonization of the subject. In some cases, a comparison of the distribution of fragment lengths of the target nucleic acids to the control or reference can indicate the stage of infection or site of colonization of the subject. In some cases, a comparison of the absolute abundance and distribution of fragment lengths of the target nucleic acids to the control or reference can indicate the stage of infection or site of colonization of the subject.
[0150] The methods provided herein can be applied to any type of nucleic acid present in a clinical sample. Figure 2An overview of examples of cell-free methods is provided. FIG. 17 provides a schematic of an exemplary infection of a subject. The source of pathogen infection can be, for example, in the lungs or any other organ (e.g., brain, skin, heart tissue, stomach, liver, intestine). Free nucleic acids, such as free DNA, derived from the pathogen can travel through the bloodstream and can be collected in a plasma sample for analysis. Some of the cell-free methods provided herein can include: obtaining a clinical sample from an infected subject or a subject at risk of infection; making a “spiked sample” by adding synthetic nucleic acids provided by the disclosure; isolating the free nucleic acids, optionally, extracting the free nucleic acids from the spiked sample; generating a spiked sample library; optionally, enriching for target nucleic acids of interest; performing a detection assay, such as a sequencing assay, to obtain sequence reads from the spiked sample library; and determining a measurement from the detected free nucleic acids and comparing this measurement to a control or reference to determine the stage of infection or site of localization of the subject.
[0151] In some cases, the methods can be combined with sequencing methods to identify organs or tissues that can be infected or to rule out the possibility that an organ of a subject is infected (see Koh W. et al., “Noninvasive in vivo monitoring of tissue-specific global gene expression in humans”, PNAS 2014: 111 (7361-7366), which publication is incorporated by reference herein in its entirety for all purposes). Figure 4 Examples of organ-site methods using free RNA sequencing are provided. Organ-site detection assays can be used in cases where the methods of the disclosure or another clinical test determines that a subject has an infection at the stage of invasive disease. In this case, the methods can further include performing one of the organ-site methods provided herein to detect whether an organ has been infected.
[0152] The disclosure also provides methods for individualized treatment of infected subjects or subjects susceptible to or at risk of infection (e.g., immunosuppressed, immunocompromised, living conditions, or genetic variations that increase susceptibility to infection). Individualized treatment provided by the disclosure includes methods to predict whether an infection will progress to the stage of invasive disease, methods to monitor the efficacy of a therapy for a subject, methods to modify a treatment regimen depending on a subject’s response to a therapy, and methods to determine a pathogen’s resistance to a particular therapeutic agent or a subject’s genetic predisposition to respond to a given therapeutic agent.
[0153] The nucleic acids produced according to the methods of the application can be analyzed to obtain various types of information, including genomic, epigenetic (e.g., methylation), and RNA expression. Methylation analysis can be performed, for example, by converting methylated bases and then performing DNA sequencing. RNA expression analysis can be performed, for example, by polynucleotide array hybridization, by RNA sequencing techniques, or by sequencing cDNA produced from RNA.
[0154] Sequencing can be by any method known in the art. Sequencing methods include, but are not limited to, Maxam-Gilbert sequencing-based techniques, chain termination-based techniques, shotgun sequencing, bridge PCR sequencing, single molecule real-time sequencing, ion semiconductor sequencing (e.g., Ion Torrent sequencing), nanopore sequencing, pyrosequencing (454), sequencing by synthesis, sequencing by ligation (SOLiD sequencing), sequencing by electron microscopy, dideoxy sequencing reactions (Sanger method), massively parallel sequencing, polymerase clonal sequencing, and DNA nanoball sequencing. The term "next generation sequencing (NGS)" refers herein to sequencing methods that allow for massively parallel sequencing of nucleic acid molecules, during which multiple, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced simultaneously. Non-limiting examples of NGS include sequencing by synthesis, sequencing by ligation, real-time sequencing, and nanopore sequencing. In some embodiments, sequencing involves: hybridizing a primer to a template to form a template / primer duplex; contacting the duplex with a polymerase in the presence of detectably labeled or unlabeled nucleotides under conditions that allow the polymerase to add labeled or unlabeled nucleotides to the primer in a template-dependent manner; detecting a signal from the incorporated labeled nucleotides or detecting a signal produced by the process of incorporating labeled or unlabeled nucleotides (e.g., proton release); and repeating the contacting and / or detecting steps in sequence at least once, wherein the sequence of the nucleic acid is determined by the sequence of the incorporated labeled or unlabeled nucleotides.
[0155] Exemplary detectable labels include radioactive labels, fluorescent labels, protein labels, dye labels, enzyme labels, and the like. In some embodiments, the detectable label can be an optically detectable label, such as a fluorescent label. Exemplary fluorescent labels include cyanines, rhodamines, fluoresceins, coumarins, BODIPY, alexa, or conjugated polydyes.
[0156] In some embodiments, sequencing comprises, consists of, or consists essentially of obtaining paired-end reads. In some embodiments, sequencing comprises, consists of, or consists essentially of obtaining consensus reads.
[0157] The accuracy or average accuracy of the sequence information can be greater than about 80%, about 90%, about 95%, about 99%, about 99.98%, or about 99.99%. The sequence accuracy or average accuracy can be greater than about 95% or about 99%. The sequence coverage can be greater than about 0.00001 fold, 0.0001 fold, 0.001 fold, about 0.01 fold, about 0.1 fold, about 0.5 fold, about 0.7 fold, or about 0.9 fold. The sequence coverage can be less than about 200,000 fold, about 100,000 fold, about 10,000 fold, about 1,000 fold, or about 500 fold.
[0158] In some embodiments, sequence information is obtained per nucleic acid template of greater than about 10 base pairs, about 15 base pairs, about 20 base pairs, about 50 base pairs, about 100 base pairs, or about 200 base pairs. Sequence information can be obtained in less than 1 month, 2 weeks, 1 week, 2 days, 1 day, 14 hours, 10 hours, 3 hours, 1 hour, 30 minutes, 10 minutes, or 5 minutes.
[0159] While the examples (below) use specific sequences for certain sequencing systems, e.g., Illumina systems, it should be understood that reference to these sequences is for illustrative purposes only, and the methods described herein can be configured for use with other sequencing systems incorporating different primers, linkers, indices, and other operational sequences used in these systems, e.g., systems available from Ion Torrent, Oxford Nanopore, Genia Technologies, Pacific Biosciences, Complete Genomics, etc.
[0160] The methods provided herein can comprise use of a system, such as a system containing a nucleic acid sequencer (e.g., a DNA sequencer RNA sequencer) for generating DNA or RNA sequence information. The system can comprise a computer comprising software that performs bioinformatics analysis of the DNA or RNA sequence information. Bioinformatics analysis can include, but is not limited to, assembling sequence data, detecting and quantifying genetic variants in a sample, including germline variants and somatic variants (e.g., genetic variations associated with cancer or precancerous conditions, genetic variations associated with infection).
[0161] Sequencing data can be used to determine genetic sequence information, ploidy status, identification of one or more genetic variants, and quantitative measures of variants, including relative and absolute relative measures.
[0162] In some cases, sequencing of the genome involves whole genome sequencing or partial genome sequencing. Sequencing can be unbiased and can involve sequencing all or substantially all (e.g., greater than 70%, 80%, 90%) of the nucleic acids in the nucleic acids in the sample. Sequencing of the genome can be selective, e.g., directed to a portion of the genome of interest. Sequencing of a gene or portion of a gene can be sufficient for the desired analysis. Polynucleotides mapping to a particular locus in the genome of the subject of interest can be isolated for sequencing by, e.g., sequence capture or site-specific amplification.
[0163] Aligning sequence reads Following sequencing, the dataset of sequences can be uploaded to a data processor for bioinformatics analysis to subtract host or host-related sequences, e.g., human, cat, dog, etc., from the analysis; and to determine the presence and prevalence of pathogen or contaminant sequences (e.g., microbial sequences) by, e.g., comparing the coverage of sequences mapping to microbial reference sequences to the coverage of host reference sequences. Subtraction of host sequences can include the step of identifying reference host sequences, and masking microbial sequences or microbial mock sequences present in the reference host genome. Similarly, determining the presence of microbial sequences by comparison to microbial reference sequences can include the step of identifying reference microbial sequences, and masking host sequences or host mock sequences present in the reference microbial genome sequences.
[0164] The dataset can optionally be cleaned to check sequence quality, remove residuals of sequencer-specific nucleotides (e.g., adaptor sequences), and merge overlapping pairs of end reads to produce higher quality consensus sequences with fewer read errors. Duplicate value sequences can be identified as those with the same start site and length or identical or nearly identical sequences. Optionally, duplicates can be removed from the analysis.
[0165] In some aspects, host or host-related (e.g., human) sequences can be subtracted from the analysis. In some aspects, host sequences are retained in the analysis. In some aspects, the amplification / sequencing step can be unbiased, and the preponderance of sequences in the sample will be host sequences. The subtraction step can be optimized in several ways to improve the speed and accuracy of the process, e.g., by performing multiple subtractions at a coarse filter, e.g., with an initial alignment set by a fast aligner, and performing additional alignments with a fine filter, such as a sensitive aligner or an extended reference database.
[0166] The dataset of reads can initially be aligned to a host reference genome including but not limited to Genbank hg19 or Genbank hg38 reference sequence, thereby bioinformatically subtracting host DNA. Each sequence can be aligned to the best set of sequences in the host reference sequence. Sequences identified as host can be bioinformatically removed from analysis.
[0167] Subtraction of host or host-related sequences can also be optimized by adding contigs with high hit rates that include but are not limited to highly repetitive sequences present in the genome that are not well represented in the reference database. For example, it has been observed that a significant amount of reads that do not align to hg19 or hg38 are ultimately identified as human when using databases containing large sets of human sequences, such as the entire NCBI NT database, in later stages of the pipeline. These reads can be removed early in the analysis by constructing an expanded host or host-related reference. This reference can be created by identifying sequences in databases of sequences, such as the NCBI NT database, that have high coverage of host contigs after initial host read subtraction. These contigs can be added to the host reference to create a more comprehensive set of references. Additionally, newly assembled host-related contigs from cohort studies can be used as additional references for filtering host-derived reads.
[0168] Regions of the host genome reference sequence containing relevant non-host sequences, such as viral and bacterial sequences integrated into the genome of the reference sample, can be masked.
[0169] Optionally, host or host-related sequences can be identified and removed by non-alignment based methods, such as by sequence characteristic recognition sequences including frequency of certain motifs, sequence patterns, word frequencies, or nucleotide bias.
[0170] The sequence reads identified as non-human can then be aligned to a nucleotide database of microbial reference sequences. The dataset can be selected for those microbial sequences known to be associated with a collection of host, such as human commensal and pathogenic microorganisms.
[0171] The microbial database can be optimized to mask or remove contaminant sequences. For example, many public dataset entries contain artificial sequences that are not derived from microorganisms, e.g., primer sequences, host sequences, and other contaminants. It can be desirable to perform an initial alignment or multiple alignments on the database. Regions showing irregularities in read coverage when multiple samples are aligned can be masked or removed as artifacts. Detection of this irregular coverage can be done by various metrics, such as the ratio between the coverage of a particular nucleotide and the average coverage of the entire contig in which this nucleotide is present. Generally, sequences that are represented as greater than about 5X, about lOX, about 25X, about 50X, about lOOX of the average coverage of their reference sequence can be artificial. Alternatively, a binomial test can be applied to provide a per-library coverage likelihood given the overall coverage of a contig. Removal of contaminant sequences from the reference database allows accurate identification of microorganisms.
[0172] Each high-confidence read can be aligned to multiple organisms in a given microbial database. In order to correctly assign organismal abundances based on this possible mapping redundancy, algorithms can be used to calculate the most likely organism (see, e.g., Lindner et al., Nucl. Acids Res. (2013) 41 (1): e10). For example, the GRAMMy or GASiC algorithms can be used to calculate the most likely organism from which a given read came.
[0173] Alignment to and designation of host sequences or to and of non-host (e.g., microbial) sequences can be performed according to art-recognized methods. For example, a read can be designated as matching a given genome if the read length is present with no more than 1 mismatch, no more than 2 mismatches, no more than 3 mismatches, no more than 4 mismatches, no more than 5 mismatches, etc. Alignment and identification can be performed using publicly available algorithms. A non-limiting example of such an alignment algorithm is the bowtie2 program (Johns Hopkins University).
[0174] These assignments of reads to organisms (e.g., host organisms, non-host organisms, microorganisms, pathogens, etc.) can then be aggregated in determining the prevalence of organisms in a sample (e.g., a free nucleic acid sample), and used to calculate an estimated number of reads assigned to each organism in a given sample. This information can be used to determine the source of a pathogen or contaminant. The analysis can normalize counts of the size of microbial genomes to provide a calculation of coverage of microorganisms. The normalized coverage of each microorganism can be compared to the host sequence coverage in the same sample to resolve differences in sequencing depth between samples.
[0175] Further, the data sets of microorganism organisms represented by the sequence lists and the prevalence of these microorganisms in the sample can optionally be aggregated and presented for instant visualization, for example, in a report.
[0176] The present disclosure provides normalization methods. In some cases, the methods of the present disclosure can include one or more normalization methods. The normalization methods provided by the present disclosure allow for efficient and improved measurement or quantification of disease-specific, pathogen-specific, or organ-specific nucleic acids detected in a sample.
[0177] The normalization methods of the present disclosure generally use spiking synthetic nucleic acids. The spiking synthetic nucleic acids can be used to normalize the sample in a variety of different ways. The spiking nucleic acids can normalize across all samples and all methods of measuring disease-specific nucleic acids, pathogen-specific nucleic acids, or other target nucleic acids. In some cases, the use of spiking can increase the precision of the relative abundance calculations of pathogen nucleic acids (or disease-specific nucleic acids or target nucleic acids) in a sample compared to other pathogen nucleic acids in the sample.
[0178] Generally, one or more known concentrations of species of synthetic nucleic acids can be spiked into each sample. In many cases, the species of synthetic nucleic acids can be spiked at equimolar concentrations of each species. In some cases, the concentrations of the species of synthetic nucleic acids can be different.
[0179] The abundance of nucleic acid species can change due to inherent biases in sample handling, preparation, and measurement (e.g., detection). After measurement, the efficiency of recovering nucleic acids of each length can be determined by comparing the measured abundance of spiked nucleic acids of each "species" to the amount originally spiked. This can result in a "length-based recovery profile."
[0180] The "length-based recovery profile" can be used to normalize all (or most or some) of the disease-specific nucleic acids, pathogen nucleic acids, or other target nucleic acids by normalizing the disease-specific nucleic acid abundance (or pathogen nucleic acid or other target nucleic acid abundance) to the closest length spiked molecule or to a function that fits the different length spiked molecules.
[0181] This process can be applied to target nucleic acids, such as pathogen-specific nucleic acids, and can result in an estimate of the "original length distribution of all pathogen-specific nucleic acids" at the time of spiking the sample. The "original length distribution of all target nucleic acids" can show the length distribution profile of the target nucleic acids (e.g., pathogen-specific nucleic acids or organ-specific nucleic acids) at the time of spiking the sample. It is this length distribution that the spiked nucleic acids can seek to recapitulate to achieve perfect or near perfect abundance normalization. It is this length distribution that the spiked nucleic acids can seek to recapitulate to achieve determination of the endogenous fragment length distribution of the target nucleic acids.
[0182] Because it is not possible to spike a sample with a mixture of known nucleic acids that perfectly recapitulate the relative abundance profile of disease-specific nucleic acids, pathogen nucleic acids, or other target nucleic acids in the particular sample, in part because the sample can have been used up or time can have altered the relative abundance profile, one can weight each "species" of spike-in by its proportion of the relative abundance of all disease-specific nucleic acids of the original length distribution. The sum of all "weighting factors" can equal 1.0.
[0183] Normalization can involve a single step or a series of steps. In some cases, the abundance of a disease-specific nucleic acid (or pathogen nucleic acid or other target nucleic acid) can be normalized using the raw measurement of the abundance of the nearest size spike-in to yield a "normalized disease-specific nucleic acid (or pathogen nucleic acid or other target nucleic acid) abundance." The "normalized disease-specific nucleic acid abundance" (or pathogen nucleic acid or other target nucleic acid abundance) can then be multiplied by a "weighting factor" to adjust for the relative importance of recovering the length to yield a "weighted normalized disease-specific (or pathogen-specific or other target) nucleic acid abundance." One advantage of this approach to normalization can be that it allows for the comparable measurement of target nucleic acid (e.g., disease-specific nucleic acid, pathogen nucleic acid) abundance across all (or most) methods of measuring disease-specific nucleic acid abundance, regardless of the method.
[0184] The assay can involve measuring the amount of a target nucleic acid (e.g., disease-specific nucleic acid) in a biological sample (e.g., plasma) to detect the presence of a pathogen or to identify a disease state or to determine that the target nucleic acid is sample-based, reagent-based, or environment-based. The methods described herein can make these measurements comparable across samples, measurement times, methods of nucleic acid extraction, methods of nucleic acid manipulation, methods of nucleic acid measurement, and / or various sample handling conditions.
[0185] The present disclosure provides diversity loss value measurements. In some cases, the methods of the present disclosure can comprise determining a diversity loss value.
[0186] The number of de-duplicated (e.g., copy-removed) SPANK molecules detected in a particular library is a proxy for the minimum concentration of a SPANK molecule that can be detected in the library. This can be used to set a threshold based on the minimum concentration of a SPANK molecule that can be detected in the library. The threshold can be used to ensure sufficient sequencing depth for detection of a pathogen. The threshold can also be used to ensure that the pathogen signal is not due to cross-contamination from other samples. For example, the enrichment of a pathogen relative to the threshold set by the SPANK molecules can be compared between different samples. More generally, it is directly proportional to the efficiency with which the library converts DNA molecules in the original sample into reads in the DNA sequencing data.
[0187] The spiked SPANK molecules provided by the disclosure can be used to calculate a diversity loss value. The diversity loss value can be determined as shown in Figure 5 In some cases, if the diversity of the SPANK sequences is sufficiently high, the SPANK sequences spiked into the sample can be assumed to be substantially all unique. Thus, any duplicate value SPANK sequences sequenced can be due to PCR amplification, rather than multiple copies of the same SPANK sequence being added to the sample, and can be removed from the analysis. Additionally, if each SPANK sequence is unique, the total number of SPANK sequences initially added to the sample is known based on the concentration and volume of nucleic acid added to the sample, and the total number of unique SPANK sequencing reads after sequencing is known, these values together can be used to calculate the diversity loss value.
[0188] C: Absolute abundance (MPM) The disclosure provides for absolute abundance measurements (also referred to as “molecules per microliter” (MPM)).
[0189] In general, the absolute abundance of a target nucleic acid (e.g., DNA or RNA) in a sample can be determined by normalizing the number of sequence reads of the target nucleic acid by an empirically determined diversity loss value.
[0190] In some cases, absolute abundance measurements can include spiking a sample with nucleic acids at various lengths or a single length and at a known concentration. In some cases, a fraction of the information actually observed in the sequencing data from the sample can be observed for each spiked length (e.g., by comparing observed reads to reads associated with the spiked nucleic acids, or by separating observed reads by spiked reads). The original number of non-host or pathogen molecules at each length can also be reverse calculated (e.g., inferred in part from the number of spiked reads at each length). This load can be converted to a “molecules per microliter” measurement.
[0191] In many cases, methods for detecting molecules per microliter (as well as other methods provided herein) can involve removing or isolating low quality reads. Removing low quality reads can improve the accuracy and reliability of the methods provided herein. In some cases, the methods can include removing or isolating (in any combination): un-mappable reads, reads resulting from PCR duplicates, low quality reads, adapter dimer reads, sequencing adapter reads, non- uniquely mapped reads, and / or reads mapping to uninformative sequences.
[0192] In some cases, sequence reads can be mapped to a reference genome, and reads that do not map to such a reference genome can be mapped to one or more target or pathogen genomes. In some cases, reads can be mapped to a human reference genome (e.g., hgl9), while the remaining reads are mapped to a curated reference database of viral, bacterial, fungal, and other eukaryotic pathogen (e.g., fungal, protozoan, parasitic) genomes.
[0193] The present disclosure provides various controls and references that can be used to determine that a measurement provided by the present disclosure indicates that a subject has a certain stage of infection or infection at a certain site.
[0194] Generally, the method includes processing a reference or control using the methods of the present disclosure. In some cases, the control or reference value measurement can be measured as a concentration or number of sequencing reads. The level can be a qualitative or quantitative level. Based on the sequence reads from the control or reference sample, a baseline level of a target nucleic acid (e.g., pathogen species, genetic variant, contamination introduced from a laboratory environment or organ derived) can be determined.
[0195] In some cases, the control or reference value can be pathogen dependent. For example, a control value for H. pylori can be different than a control value for C. difficile. A database of levels or control values can be generated based on samples obtained from one or more subjects, one or more pathogens, and / or one or more time points. This database can be curated or proprietary.
[0196] In some cases, the control or reference value is a predetermined absolute value that indicates the presence or absence of free pathogen nucleic acid or free organ derived nucleic acid. The control or reference value can be a value obtained by analyzing free nucleic acid levels from a subject that is not infected. In some cases, the control or reference value can be a positive control value and can be obtained by analyzing free nucleic acid from a subject that has a particular known infection or a particular known infection of a specific organ.
[0197] In some cases, the control can include identifying a set of commensal microorganisms or natural microflora that cause infection or do not cause infection using control samples from healthy individuals. A threshold can be set based on the set of commensal microorganisms in the control samples.
[0198] A Poisson model or other statistical model can be used to determine whether a determined baseline level of a clinical sample is significantly higher than a reference control. In cases where sequence reads from a clinical sample are significantly higher than a reference control, this indicates that the reads are informative. In some cases, such informative reads can be selected to determine a threshold for two different clinical groups.
[0199] Depending on the target nucleic acid and the level of background observed across samples, it can be desirable to use one or more references to subtract or filter out sequence reads. Filtering can be combined with selection, and done before or after selection. In some embodiments, at least one reference value is based on the level of pathogen nucleic acid detected in one or more samples selected from the group consisting of: water sample, blood sample, plasma sample, serum sample, urine sample, bodily fluid sample, reagent sample, sample from a healthy subject, or any combination thereof.
[0200] The control value can be the level of free pathogen or free organ-specific nucleic acid obtained from the subject at a different time point.
[0201] In some cases, a sample can be extracted at a time point prior to a later test time point (e.g., after a therapeutic intervention or after some time has elapsed for watchful waiting). In such cases, comparison of the levels at different time points can indicate the presence of infection, the presence of infection in a particular organ, improvement in infection, or worsening of infection. For example, an increase in pathogen or organ-specific free nucleic acid over time by an amount can indicate the presence of infection or worsening of infection, e.g., an increase of at least 5%, 10%, 20%, 25%, 30%, 50%, 75%, 100%, 200%, 300%, or 400% compared to an original value can indicate the presence of infection or worsening of infection. In other examples, a decrease in pathogen or organ-specific free nucleic acid of at least 5%, 10%, 20%, 25%, 30%, 50%, 75%, 100%, 200%, 300%, or 400% compared to an original value can indicate the absence of infection or improvement in infection (e.g., eradication of infection).
[0202] A sample can be extracted over a particular time period, such as daily, every other day, weekly, every other week, monthly, or every other month. For example, an increase in pathogen or organ free nucleic acid of at least 50% over a week can indicate the presence of infection.
[0203] The method can include determining a threshold or range of values. The threshold can be used to identify samples in a certain clinical group (colonization stage vs. invasive disease stage or no organ infection vs. infected organ). The threshold can be used to identify or select informative sequence reads from a clinical sample. In general, the desired threshold will be one that maximizes the number of true positives while minimizing the number of false positives. In some cases, the threshold can be selected using ROC curve analysis. In some cases, the threshold can be selected based on a performance metric.
[0204] Threshold selection Thresholds can be selected based on their performance using various statistical methods, such as receiver operating characteristic (ROC) curve analysis. ROC analysis can be used to assess the performance of a classifier across its entire range of operation before a cutoff threshold is selected. To use the ROC curve to determine which threshold cutoffs should perform best, the threshold can be moved stepwise across a range (e.g., 0 to 1.0) to find the results of the cutoffs in reducing the number of false positives as well as increasing the number of true negatives.
[0205] ROC analysis can be performed by plotting the data obtained from the methods of the present disclosure as TP (sensitivity) versus FP (1 - specificity). Using the ROC plot, a perfect or near-perfect classifier will typically travel straight along the Y-axis and then along the X-axis, while a classifier with no ability to classify samples in different clinical groups will typically sit on the diagonal. Most classifiers will be somewhere between these two extremes, and the user can pick a threshold based on its most likely or desired performance.
[0206] Thresholds can be selected using performance metrics such as accuracy, sensitivity, specificity, positive predictive value, or negative predictive value. In some cases, a performance metric can be used to select a threshold. In some cases, multiple performance metrics can be used to select a threshold.
[0207] Any threshold applied to a dataset (where PP is the positive population and NP is the negative population) will result in true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN).
[0208] In some cases, an accuracy performance metric can be used to determine the probability of correct classification. Accuracy can be calculated by applying the following equation: (TP + TN) / (PP + NP). In some cases, accuracy is calculated using a trained algorithm.
[0209] In some cases, a sensitivity performance metric can be used to determine the ability of a test to detect disease in a population of individuals with the disease. The sensitivity percentage can be calculated by applying the following equation: TP / (TP + FN).
[0210] In some cases, a specificity performance metric can be used to determine the ability of a test to correctly rule out disease in a population without the disease. Specificity can be calculated by applying the following equation: TN / (TN + FP).
[0211] In classifying a sample for diagnosis of infection, there are generally four possible outcomes from a binary classifier. If the outcome from the prediction is p, and the actual value is also p, it is called a true positive (TP); however, if the actual value is n, it is called a false positive (FP). Conversely, a true negative occurs when both the predicted outcome and the actual value are n, and a false negative is when the predicted outcome is n while the actual value is p. For a test that detects a disease or condition such as an infection, a false positive can occur in this case when a subject tests positive but is not actually infected. On the other hand, a false negative can occur when a subject is actually infected but tests negative for the infection.
[0212] The positive predictive value (PPV) or precision rate or post-test probability of disease is the proportion of patients with a positive test result who are correctly diagnosed. It can be calculated by applying the following equation: PPV = TP / (TP + FP) x 100. The PPV can reflect the probability that a positive test reflects the underlying condition being tested for. However, its value can indeed depend on the prevalence of the disease, which can vary.
[0213] The negative predictive value (NPV) can be calculated by the following equation: TN / (TN + FN) x 100. The negative predictive value can be the proportion of patients with a negative test result who are correctly diagnosed. PPV and NPV measures can be estimated using an appropriate prevalence of disease.
[0214] The threshold can be set based on the user's desired performance in terms of specificity and sensitivity to distinguish between two clinical groups. In some cases, the specificity of the methods provided by the present disclosure can be greater than 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%, and the sensitivity can be greater than 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or more.
[0215] Applications The methods provided by the present disclosure can be used for a variety of purposes, such as diagnosing or detecting an infection, determining the biological relationship between a microorganism and a host, the stage of infection, predicting whether an infection will progress to an invasive disease stage, monitoring the efficacy and response of a therapy for an infection, modifying or optimizing a therapy for better clinical response, stopping a treatment or therapy. Thus, using the methods provided by the present disclosure, individualized treatment can be provided to a subject based on the data obtained by the methods.
[0216] The pathogen that is expected to cause infection in the subject has several properties, such as, but not limited to, elevated absolute abundance levels compared to asymptomatic reference or controls, abnormal nucleic acid length distribution profile, or it can have both. Likewise, the pathogen that is expected to infect the organ of the subject has elevated absolute abundance levels compared to asymptomatic reference or controls, abnormal nucleic acid length distribution profile, or it can have both. The pathogen that causes infection in the subject can have several properties, such as, but not limited to, nucleic acid length distribution profile that can be compared to symptomatic reference or controls.
[0217] A: Infection stages The methods provided by the present disclosure can be used to detect, diagnose, treat, monitor, predict, or prognosticate the stage of infection in a subject. The pathogen that causes infection can be a bacterium, virus, fungus, parasite, yeast, or other microorganism, especially an infectious microorganism. In some cases, the methods can be used to determine whether the subject is in a colonization or invasive disease stage. In some cases, the methods can be used to detect whether the subject is in an incubation stage, prodromal stage, disease stage, decline stage, convalescence stage, eradication stage, chronic stage, or invasive stage. In some cases, the methods can determine that the infection is in an active stage or a latent stage.
[0218] The methods of the present disclosure can be used in conjunction with other medical tests. For example, the methods can be used before or after performing a fecal antigen test, urea breath test, serology, urease test, histology, bacterial culture and susceptibility testing, biopsy, endoscopy from the subject. In some cases, the methods described herein are performed without performing a fecal antigen test, urea breath test, serology, urease test, histology, bacterial culture and susceptibility testing, biopsy, or endoscopy on the subject.
[0219] In some cases of the methods described herein, the methods reduce the risk of progression of the infection to an invasive disease stage by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In some cases of the methods described herein, the methods reduce the mortality rate and / or mortality rate associated with complications of the invasive disease stage by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%.
[0220] The methods described herein can further comprise RNA sequencing (RNA-Seq) of the cell-free nucleic acids derived from the organ of the subject. Tissue damage caused by infection can result in the release of cell-free nucleic acids from the infected organ or tissue into the blood. Figure 3Examples of release of cell-free DNA are depicted. An increase in organ-derived, e.g., cell-free RNA, in a sample can indicate that an organ of the subject has been infected with a pathogen.
[0221] For example, a method can include analyzing circulating cell-free pathogen nucleic acid from a pathogen associated with one or more clinical symptoms. The method can further include performing RNA-Seq to detect an increase in organ-derived cell-free RNA in the subject's blood. The combination of these test results can indicate that a pathogen has infected the subject, and which organ of the subject is infected.
[0222] The RNA-Seq test can be performed simultaneously with another clinical method for detecting infection, after a clinical method for detecting infection, or before a detecting and infection clinical method. In other cases, RNA-Seq can be used independently to study organ health, or can provide increased confidence that an infection detected by another clinical method described herein is an infection of a particular organ.
[0223] In some cases, the RNA-Seq test can be able to determine whether an infection is in an invasive disease stage. In some cases, the RNA sequencing test can be repeated over time to determine whether an infection of a particular organ or tissue is worsening or improving, or whether it will spread to a different organ or tissue of the subject. Likewise, the pathogen detection assays provided herein in combination with the organ infection assays can also be repeated over time.
[0224] The RNA-Seq test (or series of RNA-Seq tests) can sometimes be performed after a method described herein produces a positive test result (e.g., detects a pathogen infection). The RNA-Seq test can be particularly useful for confirming an infection or for identifying the site of infection. For example, the method can detect the presence of a pathogen in the subject by analyzing circulating cell-free nucleic acid, but the site of infection can not be clear. In this case, the method can further include sequencing cell-free RNA from the subject to confirm that the infection is within an organ.
[0225] Absolute abundance of organ-specific RNA In some cases, the absolute abundance level of an organ-specific RNA sequence can be used as an indicator that an organ of the subject is infected with a pathogen. Detection of an organ infection can involve comparing the level of an organ-specific nucleic acid to a control or reference value to determine the presence or absence of the organ nucleic acid and / or the amount of the organ-specific nucleic acid. The level can be a qualitative or quantitative level.
[0226] In some cases, the control or reference value is a predetermined absolute value that is indicative of the presence or absence of organ-derived nucleic acids in circulation. For example, detection of a level of circulating pathogen nucleic acids above the control value can be indicative of the presence of an infection in the organ, while a level below the control value can be indicative of the absence of an infection in the organ.
[0227] The control value can be a value obtained by analyzing the level of circulating nucleic acids from a subject without infection (e.g., a healthy control). In some cases, the control value can be a positive control value obtained by analyzing circulating nucleic acids from a subject with a particular infection or a particular infection in a particular organ.
[0228] The control or reference value measurement can be measured as a concentration or number of sequencing reads. The control or reference value can be pathogen-dependent, organ-dependent, or both pathogen-dependent and organ-dependent. A database of levels or control values can be generated based on samples obtained from one or more subjects, one or more pathogens, and / or one or more time points. This database can be curated or proprietary.
[0229] In some embodiments, the control or reference absolute abundance value is indicative of the presence or absence of a localization site in the subject. For example, detection of an absolute abundance level of circulating pathogen nucleic acids above the control or reference value can be indicative of an infection in the organ, while an absolute abundance value below the control or reference value can be indicative of an infection not in the organ. In some cases, detection of an absolute abundance level of circulating pathogen nucleic acids above the control or reference value can be indicative of an infection in the organ, while an absolute abundance value below the control or reference value can be indicative of an infection not in the organ.
[0230] Distribution of fragment length of organ-specific RNA In some cases, the distribution of fragment lengths of organ-specific RNA sequences is indicative of an infection of an organ of the subject by a pathogen.
[0231] For example, detection of an abnormal distribution of circulating organ-specific nucleic acids can be indicative of an infection of the organ, while a normal distribution of circulating organ-specific nucleic acids can be indicative of an infection not in the organ.
[0232] The control fragment length distribution can be predetermined by analyzing the level of circulating nucleic acids from a subject without infection in the organ (e.g., a healthy control). The control fragment length distribution can be obtained in parallel by analyzing the level of circulating nucleic acids not associated with an infection from a subject with an infection in the organ.
[0233] In some embodiments, the control or reference distribution of fragment lengths indicates the presence or absence of a localization site. For example, detection of an abnormal distribution of free pathogen nucleic acids can indicate that an infection is in an organ, while a normal distribution of free pathogen nucleic acids can indicate that an infection is not in an organ. In some cases, detection of an abnormal distribution of free pathogen nucleic acids can indicate that an infection is in an organ, while a normal distribution of free pathogen nucleic acids can indicate that an infection is not in an organ.
[0234] Range of thresholds or values for organ-specific RNA In some cases, a threshold cutoff value can be used as an indicator of infection of an organ of a subject by a pathogen provided herein. The threshold cutoff value can be determined by using organ-specific RNA sequences from subjects infected with a pathogen, and comparing these to controls or references, as provided herein.
[0235] In some cases, the sample is identified as an infected organ with greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or more accuracy. In some cases, the sample is identified as an infected organ with greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or more sensitivity. In some cases, the sample is identified as an infected organ with greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or more than 95% specificity.
[0236] In some cases, the sample is identified as an infected organ with a positive predictive value of at least 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or more. In some cases, the sample is identified as an infected organ with a negative predictive value of at least 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or more of an infected organ.
[0237] In some cases, the sample is identified with greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or more sensitivity and greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or more 95% specificity for the infected organ.
[0238] B: Individualized treatment and monitoring The present disclosure also provides methods for individualized treatment of an infected subject or a subject susceptible to or at risk of infection (e.g., immunosuppressed, immunocompromised, living conditions, or genetic variations that increase susceptibility to infection). Individualized treatment can include predicting whether an infection will progress to an invasive stage of disease, monitoring efficacy of a therapy for a subject, modifying a treatment regimen depending on a subject's response to a therapy, and determining pathogen resistance to a particular therapeutic agent.
[0239] In some cases, the methods can be used to detect, diagnose, predict, or prognose pathogen resistance to a particular therapeutic agent. In some cases, the methods can further include sequencing the subject's DNA for genetic variations associated with resistance to a therapeutic agent or to a particular therapeutic agent.
[0240] In some cases, samples can be collected continuously at various times before or during the course of infection to determine the pathogen and the subject's response to treatment, thereby providing individualized regimens. In some cases, the continuously collected samples are compared to one another to determine whether the subject's infection is improving or worsening.
[0241] Treatment can involve administration of a drug or other therapy to reduce or eliminate colonization or invasive disease associated with infection. In some cases, a subject can be treated prophylactically to prevent infection from developing. Any medical procedure or therapy that can be used involving administration of a drug can be used to ameliorate or reduce symptoms of infection. Some non-limiting exemplary drugs that can be used are antibiotics (e.g., ampicillin, sulbactam, penicillin, vancomycin, gentamycin, aminoglycoside, clindamycin, cephalosporin, metronidazole, timentin, ticarcillin, clavulanic acid, cefoxitin), antiretroviral drugs (e.g., highly active antiretroviral therapy (HAART), reverse transcriptase inhibitors, nucleoside / nucleotide reverse transcriptase inhibitors (NRTIs), non-nucleoside RT inhibitors, and / or protease inhibitors), or immunoglobulins.
[0242] The present disclosure also provides methods of adjusting a treatment regimen. For example, a subject can have been administered a drug for treating an infection. The efficacy of the drug treatment can be tracked or monitored using the methods provided herein. In some cases, the treatment regimen can be adjusted depending on the rising or falling course of the infection. For example, if the methods provided herein indicate that the infection is not improving with drug treatment, the treatment regimen can be adjusted by changing the type of drug or treatment, discontinuing use of the drug, continuing use of the drug, increasing dosing of the drug, or adding a new drug or therapy to the treatment regimen of the subject.
[0243] In some cases, the treatment regimen can involve a particular procedure. For example, in some cases, the methods can indicate that a surgical procedure or an invasive diagnostic procedure is needed, such as removal of a tumor or performance of a biopsy to determine whether an organ is infected. Likewise, if the methods indicate that the infection is improving or has resolved with therapeutic intervention, adjusting the treatment regimen can involve reducing or discontinuing treatment. In other cases, no treatment regimen can be given, and instead a “watchful waiting” or “watch and wait” approach can be used to observe whether the infection clears without any additional medical intervention.
[0244] Methods of the disclosure can comprise detecting a pathogen in a subject. In some cases, the methods can comprise whole genome sequencing of a sample. In some cases, the methods can comprise targeted sequencing of a sample, wherein specific primers are used to detect specific pathogens of interest. Typically, a pathogen can have a recommended treatment cycle. For example, a treatment cycle for H. pylori is shown in Figure 6 Methods provided by the disclosure can be used at any stage of the treatment cycle.
[0245] Methods of the disclosure can be applied to any pathogen having various stages of infection. The methods can be particularly useful for pathogens having a colonization stage and an invasive disease stage. In some cases, the invasive disease stage can be caused by infection with the pathogen. In some cases, the invasive disease stage can be associated with infection by the pathogen.
[0246] The disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, monitoring, predicting, or preventing colonization by H. pylori (H. pylori). H. pylori colonization can be asymptomatic. In some cases, colonization can manifest as acute gastritis, with abdominal pain (stomach pain) or nausea. The disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive H. pylori disease. Subjects with invasive H. pylori disease can develop complications, such as chronic gastritis, peptic ulcer disease, gastric adenocarcinoma, gastric cancer, and / or lymphoma.
[0247] The disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by C. difficile (CDI). CDI can exist in asymptomatic or symptomatic forms. The clinical spectrum of CDI infection can range from mild to moderate, severe, or complicated disease. Subjects with mild to moderate CDI can present with diarrhea, colitis, including fever, leukocytosis, and / or cramping. The severity of abdominal and systemic symptoms of CDI can increase with the severity of infection. The methods can be used to detect, monitor, diagnose, prognose, treat, or prevent invasive CDI disease. Subjects with complicated or invasive CDI disease can develop pseudomembranous colitis, toxic megacolon, colonic perforation, and / or sepsis.
[0248] The disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by H. influenzae. Typically, H. influenzae colonizes the upper respiratory tract of a subject. The disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive H. influenzae disease. Subjects with invasive H. influenzae disease can develop complications, such as sepsis and / or meningitis.
[0249] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization by Salmonella. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive Salmonella disease. Some non-limiting examples of Salmonella serotypes associated with invasive disease include, but are not limited to, murine typhus, typhoid fever, enteric fever, Heidelberg, Dublin, paratyphoid A, swine cholera, and Swartzwelder. Subjects with invasive Salmonella disease can develop bacteremia, meningitis, enteric fever, and / or invasive non-typhoidal Salmonella (iNTS) disease.
[0250] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by Streptococcus pneumoniae. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive Streptococcus pneumoniae disease. Subjects with invasive pneumococcal disease can develop bacteremia and / or meningitis.
[0251] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization by Cytomegalovirus (CMV). Subjects infected with CMV can be asymptomatic, as the virus can cycle into a dormant phase. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive CMV disease. Subjects with invasive CMV disease can develop complications in their eyes, lungs, and / or digestive system.
[0252] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by Human Papillomavirus (HPV). Subjects with HPV colonization can exhibit non-invasive cervical intraepithelial neoplasia and / or genital warts. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive HPV disease. Subjects with invasive HPV disease can develop cervical cancer, anal squamous cell carcinoma, and / or anal carcinoma in situ.
[0253] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by Epstein-Barr virus (EBV). Subjects with EBV colonization can be asymptomatic or exhibit fatigue, fever, sore throat, swollen neck lymph nodes, enlarged spleen, enlarged liver, and / or rash. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive EBV disease. Subjects with invasive EBV disease can develop infectious mononucleosis (e.g., glandular fever), can be at higher risk of having certain autoimmune diseases, can develop cancer, such as Hodgkin's lymphoma, Burkitt's lymphoma, stomach cancer, nasopharyngeal carcinoma, hairy leukoplakia, and / or central nervous system lymphoma.
[0254] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by hepatitis B virus (HBV). HBV infection can be transient or chronic. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive disease associated with HBV infection. Subjects with invasive HBV disease can develop cirrhosis, hepatocellular carcinoma, liver infection, and / or liver failure.
[0255] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by hepatitis C virus (HCV). HCV infection can be acute or chronic. Generally, HCV colonization can be asymptomatic. When signs and symptoms are present, they can include jaundice, along with fatigue, nausea, fever, and muscle aches. Some subjects can have spontaneous viral clearance, while other subjects can progress to the chronic stage. However, in cases where HCV infection becomes chronic, it can lead to invasive HCV disease. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive HCV disease. Subjects with invasive HCV disease can develop cirrhosis, hepatocellular carcinoma, liver infection, and / or liver failure.
[0256] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by human T-cell lymphotrophic virus 1 (HTLV-1). HTLV-1 infects the T cells of a subject. Subjects infected with HTLV-1 can be asymptomatic for many years. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive HTLV-1 disease. Subjects with invasive HTLV-1 disease can develop T-cell (ATL) leukemia, HTLV-1 -associated myelopathy / tropical spastic paraparesis (HAM / TSP), or other cancerous conditions.
[0257] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization by gonorrhea. Subjects with colonizing infection can be asymptomatic, while other subjects can exhibit symptoms such as burning urination, testicular or pelvic pain, and / or discharge from the genitals. The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive gonorrhea disease. Subjects with invasive gonorrhea disease can develop skin lesions, joint infection (e.g., joint pain and swelling), endocarditis, and / or meningitis.
[0258] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization by syphilis. Syphilis infection can be classified into first, second, latent, and third stages. Subjects in the first stage can exhibit soreness. Subjects in the second stage can exhibit skin rashes, swollen lymph nodes, and / or fever. In the latent or invisible stage of syphilis, subjects are typically asymptomatic. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive syphilis disease. Subjects with the third stage or invasive disease can develop complications in other organ systems, including but not limited to the heart, blood vessels, brain, and / or nervous system.
[0259] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization by trichomoniasis. Subjects with colonizing infection can be asymptomatic or can develop inflammation in their genital area. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive trichomoniasis disease. Subjects with invasive trichomoniasis disease can develop cervical cancer and / or prostate cancer.
[0260] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization by human herpesvirus 8 (HHV-8), also known as Kaposi sarcoma-associated herpesvirus or KSHV. Healthy subjects with colonizing infection are typically asymptomatic. However, subjects with weakened immune systems can develop invasive HHV-8 disease. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive HHV-8 disease. Subjects with invasive HHV-8 disease can develop Kaposi sarcoma and / or several lymphoproliferative disorders, such as primary effusion lymphoma, multicentric Castleman disease, or B-cell lymphoma.
[0261] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization by Merkel cell polyomavirus. Subjects with colonizing infection can be asymptomatic. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive Merkel cell polyomavirus disease. Subjects with invasive Merkel cell polyomavirus disease can develop Merkel cell carcinoma (MCC) tumors, a rare but aggressive form of skin cancer.
[0262] The present disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization by Chlamydia. Subjects with colonizing infection can be asymptomatic or can exhibit a burning sensation when urinating or discharging from the genitalia. The present disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive Chlamydia disease. Untreated Chlamydia can progress to the invasive disease stage, spreading to the uterus and / or fallopian tubes of a female subject. Subjects with invasive Chlamydia disease can develop pelvic inflammatory disease (PID), which can lead to long-term pelvic pain, inability to conceive, and ectopic pregnancy.
[0263] In some cases, a subject is infected or at risk of infection by a pathogen at different stages of infection, such as the colonization stage and the invasive disease stage. A colonized subject can have no clinical signs or symptoms. In other cases, a colonized subject can have clinical signs or symptoms. A subject with invasive disease can exhibit clinical signs or symptoms. In other cases, a subject with invasive disease can exhibit no clinical signs or symptoms.
[0264] The subject can have or be at risk of another disease or condition. For example, the subject can have, be at risk of, or be suspected of having cancer (e.g., breast cancer, lung cancer, gastric cancer, hematological cancer).
[0265] In some cases, a subject has a risk factor that can increase the risk of infection or progression to the invasive disease stage. In some cases, the risk factor is associated with living conditions. Some non-limiting examples of risk factors associated with living conditions include, but are not limited to, crowded living conditions, lack of access to clean water, living in or visiting a developing country, and / or cohabitation with an infected individual.
[0266] In some cases, a risk factor for infection or progression to the invasive disease stage is a genetic variant of the genomic DNA of the subject. Genetic variants that can be a risk factor for infection include, but are not limited to, single nucleotide polymorphisms, deletions, insertions, and the like. In some other cases, a subject can have a family history of a disease, such as gastric cancer, lymphocytic gastritis, hyperplastic gastric polyps, or a family history of pregnancy-induced vomiting.
[0267] The subject can have or be co-infected with more than one pathogen or be at risk of having or being co-infected with more than one pathogen. In some cases, a subject is immunosuppressed (e.g., an organ transplant patient). In some cases, a subject is immunocompromised (e.g., by chemotherapy treatment, immunodeficiency caused by AIDS or general diseases such as diabetes or lymphoma).
[0268] In some cases, the subject can exhibit one or more clinical symptoms. Non-limiting examples of clinical symptoms can include abdominal pain or burning, abdominal pain that worsens when the tail is emptied, nausea, lack of appetite, frequent belching, bloating in the stomach area, weight loss, severe or persistent abdominal pain, difficulty swallowing, blood or black tarry stool, and / or blood or black vomit. Other clinical symptoms are known in the art.
[0269] In some cases, the subject can exhibit a clinical pathology such as atrophic gastritis, acute or chronic gastritis, gastric acid hypersecretion, antigenic stimulation, active peptic ulcer disease, a history of PUD, low-grade gastric mucosa-associated lymphoid tissue lymphoma, a history of endoscopic resection of early gastric cancer, dyspepsia, Barrett's esophagus, functional dyspepsia, unexplained iron deficiency, or idiopathic thrombocytopenic purpura (ITP).
[0270] The subject can be infected with any type of pathogen or microorganism, including bacteria, viruses, fungi, parasites, prokaryotes, eukaryotes, and the like. In some cases, the pathogen is known, while in other cases it can be a known commensal.
[0271] In some cases, the subject can have an active or latent infection. In some cases, the subject is infected, but the infection is below the level of diagnostic sensitivity of other tests previously performed on the subject. In some cases, the subject is infected but asymptomatic, or the infection is at a subclinical level.
[0272] In some cases, the subject can have been previously treated or can be treated with a drug or medical procedure such as an antimicrobial, antibacterial, antiviral, and / or antiparasitic drug, or the like. In some cases, the subject can not have had a biopsy, endoscopy, colonoscopy, blood culture, or other such procedure prior to using the methods herein. In some cases, the subject can have had or can have had a fecal antigen test, urea breath test, serology, urease test, histology, bacterial culture and sensitivity testing, biopsy, or endoscopy prior to using the methods herein.
[0273] The present disclosure provides methods for determining the stage or site of infection of a subject using nucleic acids obtained from a clinical sample (e.g., blood, serum, cells, or tissue). In some embodiments, the methods include making a spiked sample by adding a synthetic nucleic acid provided by the present disclosure; extracting nucleic acids from the spiked sample; generating a spiked sample library; enriching the spiked sample library for target nucleic acids of interest; performing a sequencing assay to obtain sequence reads from the spiked sample library; and determining a measurement from the detected nucleic acids (e.g., DNA, RNA, cell-free DNA, or cell-free RNA) and comparing this measurement to a control or reference to determine the stage or site of infection of a subject (e.g., organ or type of tissue).
[0274] Embodiments of the methods can include extracting nucleic acids or target nucleic acids from a sample or purifying nucleic acids or target nucleic acids from unwanted components in a reaction mixture (e.g., ligation, amplification, restriction enzymes, end repair, etc.). Any means of extracting nucleic acids known in the art can be used in the methods of the present application.
[0275] Extraction can include separating nucleic acids from other cellular components and contaminants that can be present in a sample. Nucleic acids can be extracted from a sample using liquid extraction (e.g., Trizol, DNAzol) techniques. In some cases, extraction is performed with the aid of a phenol chloroform extraction or precipitation by organic solvents (e.g., ethanol or isopropanol). In some cases, extraction is performed using nucleic acid binding columns.
[0276] In some cases, extraction is performed using commercially available kits, such as Qiagen Qiamp Circulating Nucleic Acid Kit, Qiagen Qubit dsDNA HS Assay Kit, Agilent™ DNA 1000 Kit, TruSeq™ Sequencing Library Preparation, QIAamp Circulating Nucleic Acid Kit, Qiagen DNeasy Kit, QIAamp Kit, Qiagen Midi Kit, QIAprep spin Kit) or nucleic acid binding spin columns (e.g., Qiagen DNA mini-prep kit). In some cases, extraction of cell-free nucleic acids can involve filtration or ultrafiltration.
[0277] Nucleic acids can be extracted or purified by using magnetic beads. For example, magnetic beads having an iron oxide core and surface coated with molecules containing free carboxylic acids or synthetic polymers can be used. Salt concentration or polyalkylene glycols can be adjusted to control the strength of the bond between the functional group and the nucleic acid, allowing for controlled and reversible binding. Finally, the nucleic acid can be released from the magnetic particles with an elution buffer. In some cases, extraction or purification is performed using a commercially available kit, such as Omega Biotek Mag-Bind® magnetic bead kits, Agencourt®, RNAClean® and / or XP magnetic beads.
[0278] The methods can include purifying the target nucleic acid. Purification can be performed where a user desires to separate the target nucleic acid from unwanted components in the reaction mixture. Non-limiting exemplary purification methods include ethanol precipitation, isopropanol precipitation, phenol chloroform purification, and column purification (e.g., affinity-based column purification), dialysis, filtration, or ultrafiltration.
[0279] Methods of generating nucleic acid libraries are known in the art.
[0280] Computer control system The present disclosure provides a computer control system programmed to implement the methods of the present disclosure. Figure 7 A computer system 201 programmed or otherwise configured to implement the methods of the present disclosure is shown.
[0281] The computer system 201 includes a central processing unit (CPU, also "processor" and "computer processor" herein) 205, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 201 also includes memory or memory location 210 (e.g., random access memory, read only memory, flash memory), electronic storage unit 215 (e.g., hard disk), communication interface 220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 225, such as cache, other memory, data storage, and / or electronic display adapters. The memory 210, storage unit 215, interface 220, and peripheral devices 225 are in communication with the CPU 205 through a communication bus (solid lines), such as a motherboard. The storage unit 215 can be a data storage unit (or data repository) for storing data. The computer system 201 can be operatively coupled to a computer network ("network") 230, by way of the communication interface 220. The network 230 can be the Internet, an internet and / or an extranet, or an intranet and / or extranet that is in communication with the Internet. The network 230 in some cases is a telecommunication and / or data network. The network 230 can include one or more computer servers that can implement a distributed computing methodology, such as cloud computing. The network 230 in some cases operatively couples the computer system 201 to a peer computer system, which can act as a server or a client in terms of the process discussed herein.
[0282] The CPU 205 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory 210. The instructions can be directed to the CPU 205, which can subsequently program or otherwise configure the CPU 205 to implement methods of the present disclosure. Examples of actions performed by the CPU 205 can include fetch, decode, execute, and writeback.
[0283] The CPU 205 can be part of a circuit, such as an integrated circuit. One or more other components of the system 201 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0284] The storage unit 215 can store files, such as drivers, libraries and saved programs. The storage unit 215 can store user data, e.g., user preferences and user programs. The computer system 201 in some cases can include one or more additional data storage units that are external to the computer system 201, such as located on a remote server that is in communication with the computer system 201 through an intranet or the Internet.
[0285] The computer system 201 can communicate over the network 230 with one or more remote computer systems. For instance, the computer system 201 can communicate with a remote computer system of a user (e.g., a healthcare provider). Examples of remote computer systems include a personal computer (e.g., a laptop), a tablet or tablet PC (e.g., Apple® iPad, Samsung® Galaxy Tab), a telephone, a smart phone (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or a personal digital assistant. The user can access the computer system 201 via the network 230.
[0286] Methods as described herein can be implemented through the use of instructions that are stored in storage memory locations, such as the memory 210 or the electronic storage unit 215, on the computer system 201. The instructions, when executed by the processor 205, can cause the processor 205 to perform one or more features described herein. The instructions can be software written in any suitable computer programming language to accomplish the desired results. The software can be compiled or interpreted.
[0287] The software can be supplied in a memory, such as the memory 210 or the electronic storage unit 215. The software, when executed by the processor 205, causes the processing system to perform one or more features described herein. The software can be executable, i.e., locatable by the processor 205, from the electronic storage unit 215 or from another storage medium or device (e.g., a memory).
[0288] Aspects of the systems and methods provided herein, such as server 201, can be embodied in programming. Various aspects of the technology can be thought of as "products" or "articles of manufacture" typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, hard disk drives or the like which can provide non-transitory storage at any time for the software programming. All or portions of the software can at times be communicated by an associated physiological entity, such as by the networks described herein that can carry the software. Such communications, for example, can enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that can bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces to communicate software from one part to another, through a physical medium such as over physical wires and / or wireless interfaces, optical interfaces, etc. The physical elements that carry the waves, such as wires or physical interfaces, optical interfaces, etc. can also be considered to be media as described herein. As used in this document, the term "machine-readable medium" includes any medium that can carry or store software for execution by a processor, unless the medium is specifically designated to be non-transitory, tangible "storage" medium.
[0289] Accordingly, a machine readable medium, such as computer-executable code, can take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as can be used to implement databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include, for example: a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, CDTV, DVD or DVD-ROM, any other optical medium, punch cards, paper tape, any other physical storage medium that can be used to store or transfer data or information, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer can read or write a program code and / or data. Many of these forms of computer readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0290] The computer system 201 can include or can be in communication with an electronic display 235, which comprises a user interface (UI) 240 for providing report outputs, which can include a diagnosis for a subject or a therapeutic intervention for a subject. Examples of UIs include, without limitation, a graphical user interface (GUI) and a web-based user interface. The analysis can be provided in the form of a report. The report can be provided to a subject, a healthcare professional, a laboratory worker, or other individual.
[0291] The methods and systems of this disclosure can be implemented by one or more algorithms. The algorithms can be implemented by way of software upon execution by the central processing unit 205. The algorithms may, for example, facilitate enrichment, sequencing, and / or detection of pathogen nucleic acids.
[0292] Information about 'a' can be entered into a computer system, such as patient identifiers, including information about the stage or risk of infection, patient background, patient medical history, previous infections, or ultrasound scans. Patient identifiers can be separated from clinical samples to obtain deidentified samples, for example, through the sample sender or recipient. Patient identifiers can be replaced with accession numbers or other non-personally identifiable codes. Clinical samples can be sequenced using a high-throughput sequencer. The deidentified sample sequence data generated by the sequencer can be uploaded to a server, such as a cloud server. Using the methods disclosed herein, pathogen nucleic acids within the deidentified sample can be detected to obtain deidentified result data. Deidentified result data can be downloaded from the server. The deidentified result data can be associated with patient identifiers, for example, through the sample sender or recipient.
[0293] Electronic reports can be generated to indicate the stage of infection by a pathogen. Electronic reports can be generated to indicate prognosis. Electronic reports can be generated to indicate diagnosis. If the electronic report indicates the presence of a treatable infection, an electronic report can be generated to specify a treatment plan or protocol. Computer systems can be used to analyze the results from the methods described herein, reporting the results to the patient or physician or proposing a treatment plan.
[0294] Kit Reagents and kits thereof for carrying out one or more of the methods described herein are also provided. The reagents and kits thereof of the present invention may vary considerably. The reagents of interest comprise those specifically designed for the identification, detection, and / or quantification of one or more pathogen nucleic acids in samples obtained from subjects infected with or at risk of infection with pathogens.
[0295] The kit may include reagents necessary for performing nucleic acid extraction and / or nucleic acid detection using the methods described herein, such as PCR and sequencing. The kit may further include software for data analysis, which may contain a reference spectrum for comparison with test profiles from clinical samples, and specifically may contain a reference database. The kit may include reagents such as buffers and water.
[0296] Such kits can also include information, such as scientific literature references, package insert materials, clinical trial results and / or summaries of these, and / or other information useful to a health care professional in connection with the activity and / or benefits of the compositions, dosages, administration, side effects, drug interactions, etc. Such kits can also include instructions for accessing a database. Such information can be based upon the results of various studies, e.g., studies using experimental animals involving in vivo models and studies based upon human clinical trials. The kits described herein can be provided, sold, and / or promoted to health care providers, including physicians, nurses, pharmacists, formulary officials, etc. In some embodiments, the kits can also be sold directly to consumers.
[0297] It will be understood that the following examples are for illustrative purposes only and do not limit the scope of the claims.
[0298] The present invention provides, including but not limited to the following embodiments: 1. A fragment length profile of a nucleic acid library, wherein the nucleic acid library is generated from an initial sample, and wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library, wherein the fragment length profile comprises one or more properties selected from the group comprising: shape of the distribution, amplitude of the bins, shape of the peaks, ratio of fragment counts of two or more bins, height of the helical phasing peak, ratio of fragment counts at two different fragment lengths, ratio of fragment counts in two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads.
[0299] 2. A method of generating a fragment length profile of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample using a bias-corrected recovery method; (b) determining the number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length properties of the nucleic acid library, wherein the one or more fragment length properties are selected from the group comprising: shape of the distribution, amplitude of the bins, shape of the peaks, ratio of fragment counts of two or more bins, height of the helical phasing peak, ratio of fragment counts at two different fragment lengths, ratio of fragment counts in two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length properties.
[0300] 3. A method of generating a fragment length profile of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample, the preparing a nucleic acid library from an initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library; (b) determining a number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of the segments, shape of the peaks, ratio of segment counts of two or more segments, height of the helical phasing peak, ratio of segment counts at two different fragment lengths, ratio of segment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.
[0301] 4. The method of embodiment 3, wherein generating the nucleic acid library from the initial sample comprises, consists of, or consists essentially of: (a) dephosphorylating nucleic acids from the initial sample to produce a set of dephosphorylated nucleic acids; (b) denaturing the dephosphorylated nucleic acids to produce denatured nucleic acids; (c) ligating 3’ end adaptors to the denatured nucleic acids to produce adapted nucleic acids; (d) isolating adapted nucleic acids; (e) priming the adapted nucleic acids and extending the primers with a polymerase to generate complementary strands; (f) ligating 5’ end adaptors; (g) eluting the strands; and (h) amplifying the complementary strands.
[0302] 5. The method of embodiment 2, wherein the number of reads is a normalized number of reads.
[0303] 6. The method of embodiment 2, wherein the fragment length profile is for at least one subset of reads, and the method further comprises: (a) identifying at least one subset of reads within the nucleic acid library; and (b) determining the fragment length profile within the at least one subset of reads.
[0304] 7. The method of embodiment 2, wherein the step of generating at least one fragment length profile further comprises using two or more fragment length characteristics.
[0305] 8. A method of identifying a microorganism present in a sample, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from the sample; (b) comparing the fragment length profile to a reference fragment length profile of one or more microorganisms; and (c) identifying the microorganism as present in the sample if the fragment length profile from the sample is similar to a reference fragment length profile of a microorganism.
[0306] 9. The method of embodiment 8, wherein generating a fragment length profile of the nucleic acid library comprises the steps of: (a) preparing a nucleic acid library from an initial sample, the preparing a nucleic acid library from an initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library; (b) quantifying the number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of the segments, shape of the peaks, ratio of segment counts of two or more segments, height of the helical phasing peak, ratio of segment counts at two different fragment lengths, ratio of segment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.
[0307] 10. The method of embodiment 8, wherein the fragment length profile indicates that the microorganism is present in the form of a pathogen or a commensal microorganism.
[0308] 11. The method of embodiment 8, wherein the fragment length profile comprises at least one fragment length characteristic selected from the group comprising: ratio of segment counts of two or more peaks and shape of the fragment length distribution.
[0309] 12. A method of identifying a localization site in a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample; (b) comparing the fragment length profile to a reference fragment length profile of one or more source sites; and (c) identifying a first site as a localization site if the fragment length profile from the sample is similar to a fragment length profile from the first source site and a second site as a localization site if the fragment length profile from the sample is similar to a fragment length profile from the second source site.
[0310] 13. The method of embodiment 12, wherein generating a fragment length profile of the nucleic acid library comprises the steps of: (a) preparing a nucleic acid library from an initial sample, the preparing a nucleic acid library from an initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library; (b) quantifying a number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of the segments, shape of the peaks, ratio of segment counts of two or more segments, height of the helical phasing peak, ratio of segment counts at two different fragment lengths, ratio of segment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.
[0311] 14. The method of embodiment 12, wherein the localization site is selected from the group of source sites comprising: deep tissue, blood flow, skin, lung, heart, brain, and blood.
[0312] 15. A method of monitoring a status of a transplant in a subject having a transplant, the method comprising the steps of: (a) generating a baseline fragment length profile of a nucleic acid library generated from a sample obtained from the subject; (b) generating a second fragment length profile of a nucleic acid library generated from a second sample obtained from the subject; (c) comparing the second fragment length profile to the baseline fragment length profile; if the second fragment length profile is different from the baseline fragment length profile, internally administering an increased amount of an anti-rejection therapy, wherein after administering the anti-rejection therapy, the subject having a transplant has a reduced risk of rejection; if the second fragment length profile is similar to the baseline fragment length profile, maintaining or reducing the anti-rejection therapy, wherein the anti-rejection therapy has a lower risk of side effects in the patient than the patient receiving the increased amount of the anti-rejection therapy.
[0313] 16. A method of monitoring toxicity of a compound administered to a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample; and (b) comparing the fragment length profile to one or more reference fragment length profiles.
[0314] 17. The method of embodiment 16, wherein the one or more reference fragment length profiles are generated from a nucleic acid library obtained from a subject or cell exposed to the compound.
[0315] 18. The method of embodiment 16, wherein the subject has, is at risk of having, or exhibits symptoms related to cancer.
[0316] 19. The method of embodiment 16, wherein the compound is a chemotherapeutic agent.
[0317] 20. The method of embodiment 16, wherein generating a fragment length profile of the nucleic acid library comprises the steps of: (a) preparing a nucleic acid library from an initial sample using a bias-corrected recovery method; (b) determining a number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of the segments, shape of the peaks, ratio of fragment counts of two or more segments, height of the helical phasing peak, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitudes of two or more segments, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.
[0318] 21. The method of embodiment 16, wherein generating a fragment length profile of the nucleic acid library comprises the steps of: (a) preparing a nucleic acid library from an initial sample, the preparing a nucleic acid library from an initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparing the nucleic acid library; (b) quantifying a number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: a shape of a distribution, a segment amplitude, a peak shape, a ratio of segment counts of two or more segments, a height of a helical phasing peak, a ratio of segment counts at two different fragment lengths, a ratio of segment counts within two different fragment length ranges, a fragment length range within a segment, a ratio of maximum amplitudes of two or more segments, and a fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.
[0319] 22. A method of determining a stage of infection of a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample obtained from the subject; (b) comparing the fragment length profile to a reference fragment length profile; and (c) determining that the stage of infection indicates an increased risk that the subject exhibits a symptom associated with the microbe if the fragment length profile from the sample is similar to a fragment length profile from a symptomatic subject, and determining that the infection is in an asymptomatic stage if the fragment length profile from the sample is similar to a fragment length profile from an asymptomatic subject.
[0320] 23. The method of embodiment 22, wherein the fragment length profile is a non-microbiome host or microbiome subset of a nucleic acid library fragment length profile.
[0321] 24. The method of embodiment 22, further comprising the steps of: (a) determining an abundance of at least one significant microorganism in a sample from the subject; (b) comparing the abundance to a threshold value and comparing the fragment length profile to a reference fragment length profile; and (c) determining that the infection stage indicates an increased risk that the subject exhibits a microorganism-related symptom if the fragment length profile from the sample is similar to a fragment length profile from a symptomatic subject and the abundance is at or above a threshold value; and determining that the infection is in an asymptomatic stage if the fragment length profile from the sample is similar to a fragment length profile from an asymptomatic subject.
[0322] 25. The method of embodiment 22, further comprising administering an antimicrobial agent to a subject determined to have an increased risk of exhibiting a microorganism-related symptom.
[0323] 26. A method for determining an infection stage of a subject suspected of having a microorganism infection, the method comprising: a) performing high-throughput sequencing of nucleic acids from a biological sample; b) performing bioinformatic analysis to identify cell-free nucleic acid sequences present in the biological sample; and c) obtaining a measurement of the cell-free nucleic acids and comparing the measurement to a control, thereby determining an infection stage of a microorganism identified in the biological sample.
[0324] 27. The method of embodiment 26, further comprising one or more steps selected from the group consisting of: (a) extracting cell-free nucleic acids from a biological sample obtained from the subject; and (b) adding synthetic nucleic acid spike-ins to the cell-free fraction.
[0325] 28. The method of embodiment 26, wherein the nucleic acids comprise microorganism nucleic acids, host nucleic acids, or both microorganism nucleic acids and host nucleic acids.
[0326] 29. The method of embodiment 26, wherein the measurement is selected from the group of measurements consisting of: an absolute abundance of the cell-free nucleic acids, a distribution of fragment lengths of the cell-free nucleic acids, and both an absolute abundance and a distribution of fragment lengths of a target microorganism.
[0327] 30. The method of embodiment 26, wherein the infection stage is selected from an asymptomatic stage, a symptomatic infection stage, a treatment stage, or an eradication stage.
[0328] 31. The method of embodiment 26, further comprising administering a treatment regimen to the subject, wherein the treatment regimen is appropriate for the determined infection stage.
[0329] 32. The method of embodiment 26, further comprising repeating the method on a sample obtained from a subject at a plurality of time points to monitor an infection or the efficacy of a treatment for an infection.
[0330] 33. The method of embodiment 26, wherein the microorganism is selected from the group comprising Helicobacter pylori (H. pylori) heliobacter pylori ), Clostridium difficile (C. difficile) clostridium difficile ), Haemophilus influenzae (H. influenzae) haemophilus influenza ), Salmonella (Salmonella) salmonella ), Streptococcus pneumoniae (S. pneumoniae) streptococcus pneumoniae ), Cytomegalovirus (CMV) cytomegalovirus ), Hepatitis B virus, Hepatitis C virus, Human Papillomavirus, Epstein-Barr virus, Human T-cell lymphotropic virus 1, Merkel cell polyomavirus, Kaposi's sarcoma virus, human Herpesvirus 8, chlamydia, gonorrhea, Syphilis, or trichomoniasis.
[0331] 34. The method of embodiment 27, wherein adding synthetic nucleic acid spikert further comprises: (a) preparing a spiked sample by obtaining a sample comprising cell-free nucleic acids from a subject and adding at least 1000 unique synthetic nucleic acids to the sample, wherein each of the 1000 unique synthetic nucleic acids comprises: (i) an identifying tag; and (ii) a variable region comprising at least 5 degenerate bases; (b) extracting nucleic acids from the spiked sample; (c) generating a spiked sample library; (d) enriching the spiked sample library; (e) performing a high-throughput sequencing assay to obtain sequence reads from the spiked sample library; (f) calculating a diversity loss value for the 1,000 unique synthetic nucleic acids; and (g) calculating a measurement of the cell-free nucleic acids and comparing the measurement to a control, thereby determining a stage of infection in the subject.
[0332] 35. A method of determining the stage of infection of Helicobacter pylori in a subject, the method comprising: (b) extracting cell-free nucleic acids from a biological sample obtained from the subject; (c) adding synthetic nucleic acid spikins to the cell-free fraction; (d) performing high-throughput sequencing on nucleic acids from the biological sample; (e) performing bioinformatic analysis to identify cell-free Helicobacter pylori nucleic acid sequences present in the biological sample; and (f) calculating a measurement of the cell-free Helicobacter pylori nucleic acids and comparing the measurement to a control, thereby determining the stage of infection of Helicobacter pylori in the subject.
[0333] 36. A method of determining the stage of infection of Helicobacter pylori in a subject, the method comprising: a) preparing a spiked sample by obtaining a sample comprising cell-free nucleic acids from a subject and adding at least 1000 unique synthetic nucleic acids to the sample, wherein each of the 1000 unique synthetic nucleic acids comprises: (i) an identifying tag; and (ii) a variable region comprising at least 5 degenerate bases; b) extracting nucleic acids from the spiked sample; c) generating a spiked sample library, wherein the generating comprises (i) ligating adaptors to end-repaired spiked sample; and (ii) amplification; d) enriching the spiked sample library; e) performing a high-throughput sequencing assay to obtain sequence reads from the spiked sample library; f) calculating a diversity loss value for the 1,000 unique synthetic nucleic acids; and g) calculating a measurement of the cell-free nucleic acids and comparing the measurement to a control, thereby determining the stage of infection of Helicobacter pylori in the subject.
[0334] 37. A method of determining a host-microbe biological interaction in a subject, the method comprising: (a) generating a fragment length profile of a nucleic acid library generated from a sample from the subject; (b) optionally, determining the abundance of a target nucleic acid and comparing the abundance to a threshold; (c) comparing the fragment length profile to one or more reference fragment length profiles of host-microbe biological interactions; and (d) identifying a host-microbe biological interaction if the fragment length profile is similar to a reference fragment length profile of a host-microbe biological interaction.
[0335] 38. The method of embodiment 37, wherein if the fragment length profile and abundance of the target nucleic acid is similar to a reference fragment length profile and threshold of a host-microorganism biological interaction, the host-microorganism biological interaction is identified.
[0336] 39. The method of embodiment 32, further comprising altering a treatment regimen.
[0337] 40. A method of identifying the presence of a viral infection in a subject suspected of having a microbial infection, the method comprising: a) generating a fragment length profile of a nucleic acid library generated from a sample from the subject; b) comparing the fragment length profile to a viral reference fragment length profile; c) optionally, quantifying the abundance of a target nucleic acid and comparing the abundance to a threshold; d) identifying the presence of a viral infection in the subject if the fragment length profile is similar to the reference profile.
[0338] Examples Example 1 : Distribution shape and microbial state Processing biological samples with methods that lack bias in the region of fragment lengths of interest or that correct for bias allows for the measurement of endogenous fragment length distributions and the potential for using endogenous fragment length profile to inform diagnostics and therapeutic aspects of treatment. Thus several different clinical samples processed show the diversity of fragment length profile. A direct to library method that lacks detectable length and GC bias in the range of fragment lengths investigated was applied to obtain the shape of the endogenous fragment length distribution.
[0339] Clinical plasma samples: 36 cases of diagnostic positivity (i.e. presence of a microorganism confirmed with orthogonal tests, e.g. blood culture, targeted PCR, Karius test) were collected from 36 human subjects. A single centrifugation step plasma extraction process from whole blood was performed on each sample within 24 hours of sample collection, as previously described (see first centrifugation step in Fan HC et al., Proc Natl Acad Sci U S A. 2008; 105(42): 16266-16271, which is incorporated by reference herein in its entirety, including any drawings), and stored at -80°C prior to use. Samples were then thawed and 2 µL of a spike master mix (see below) was added to 200 µL of each plasma.
[0340] Positive assay control samples: For each set of 18 samples, two positive controls, referred to as assay control samples (ACs), were processed. The AC samples were prepared from human asymptomatic plasma spiked with enzymatically sheared genomes of human pathogens purchased in purified form from ATCC (American Type Culture Collection). The human pathogens selected were Aspergillus fumigatus, Escherichia coli, Pseudomonas aeruginosa, and Staphylococcus epidermidis. 10 pL of a spike master mix (see below) was added to each 1 mL AC sample.
[0341] Negative control samples: Four 500 pL negative control samples (ECs) were made from aqueous buffer (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) with 5 pL of a spike master mix (see below) per 18 samples and used as a control for environmental contamination (e.g., microbial and pathogen nucleic acid contamination introduced through reagents, instruments, consumables, operators, and / or air during processing). These synthetic nucleic acids were used to normalize the signal in the samples to account for variation in sample processing.
[0342] Spike master mix: A set of process controls were pre-mixed together in individual spike master mixes, each containing a unique “ID spike” process control molecule, see, e.g., U.S. Patent 9,976,181. The spike master mix contains three classes of molecules: ID spike molecules, SPANK molecules, and SPARK molecules. The latter set of molecules consists of two classes of SPARKs: GC dSPARKs and long SPARKs. The molar concentration of ID spike, SPANK molecules, and long SPARK molecules in the spike master mix is 500 pM per molecule, while the GC dSPARK molecules are present at 50 pM per molecule.
[0343] “ID spike” molecules: Each sample received a unique ID spike single-stranded DNA molecule characterized by a 50 base pair long unique sequence that does not exist in any reference genome available in public databases at the time of processing.
[0344] SPANK molecules: The SPANK molecules used were a pool of single-stranded DNA molecules, each 50 base pairs long, with identical 3' and 5' end sequences that were not present in any reference genome available in public databases at the time of processing. In addition, there were two stretches of 8 base pairs nested between the constant 3' and 5' end sequences and fully degenerated in the pool. The SPANK molecule pool contained 416 unique SPANK molecules. The two degenerate stretches were separated by stretches of four non-degenerate bases.
[0345] SPARK molecules: The GC spike set was a collection of 32, 42, 52, and 75 nt long molecules with 7 different sequences for each length containing GC content of 20%, 30%, 40%, 50%, 60%, 70%, and 80%. Like some of the other molecules provided above, the GC dSPARK sequences did not appear in available reference genomes. The long SPARK sequence set was a group of 4 non-natural sequences with a GC content of 50% and lengths of 100 nt, 125 nt, 150 nt, and 175 nt. The complete set of SPARK molecules contained 32 different sequences.
[0346] Library generation: Direct to library generation was performed as described in U.S. provisional application 62 / 770,181, filed November 21, 2018, the entirety of which is incorporated by reference herein in its entirety. Here, a template switch based method with Proteinase K was utilized. Briefly, 50.0 pL of each spiked sample was mixed with 20.0 pL of 10x terminal transferase reaction buffer (New England Biolabs, Ipswich, MA), 5.0 pL of Proteinase K (Sigma), 2.0 pL of 10% Tween-20 (Thermo-Fisher Scientific, Waltham, MA), 2.0 pL of 10% Triton X100 (Thermo-Fisher Scientific, Waltham, MA), and 121.0 pL of nuclease-free water. The mixture was heated to 60°C for 20 minutes and to 95°C for 10 minutes, and it was placed on ice until cooled. 2.0 pL of 10 mM dATP, 2.0 pL of terminal transferase (20 u / pL, New England Biolabs, Ipswich, MA), and 6.0 pL of nuclease-free water were added to make an A-tail reaction, which was then incubated at 37°C for 40 minutes. 300.0 pL of lysis / binding buffer (Thermo-Fisher Scientific, Waltham, MA) was added to the reaction. The entire volume was then added to 50.0 pL of Dynabeads oligo(dT)25 (Thermo-Fisher Scientific, Waltham, MA), which was then washed once with lysis / binding buffer (Thermo-Fisher Scientific, Waltham, MA). The mixture was incubated at 25°C and 600 RPM. The beads were then washed twice with 600.0 pL of wash buffer A (Thermo-Fisher Scientific, Waltham, MA) and twice with 300.0 pL of wash buffer B (Thermo-Fisher Scientific, Waltham, MA), then eluted in 24.0 pF of elution buffer (Thermo-Fisher Scientific, Waltham, MA) at 80°C and 600 RPM for 3 minutes. The entire eluate was transferred to a new plate. To the eluate was added 2.0 pL of 1 pM Poly dT primer (IDT) and 6 pL of SMARTScribe 1st strand buffer (5x) (Takara, Kusatsu, Japan), and the resulting mixture was incubated at 95°C for 1 minute, then placed on ice.An extension and template switch mixture was prepared by combining 4.5 pL SMARTScribe 1st strand buffer (5x) (Takara, Utsu, Japan), 0.5 pL dNTP mix (25 mM per nucleotide, Thermo Fisher Scientific, Waltham, MA), 2.0 pL SMARTScribe Reverse Transcriptase (100 u / pL, Takara, Utsu, Japan), 2.0 pL 5 pM template switch oligo (TS oligo) (IDT), 5.0 pL DTT (20 M, Takara, Utsu, Japan), and 4 pL nuclease-free water. The resulting reaction mixture was incubated at 42 °C for 90 minutes and the reaction was heat denatured at 70 °C for 15 minutes. Next, 50.0 pL NEBNext Ultra II Q5 (New England BioLabs, Ipswich, MA) and 8.0 pL indexing primer mix (New England BioLabs, Ipswich, MA) were added to the reaction from the previous step. Amplification of the nucleic acids was then performed using the following temperature cycling program: 98 °C for 30 seconds, 98 °C for 8 cycles for 10 seconds, 65 °C for 75 seconds, and 65 °C for a final extension for 5 minutes. The final nucleic acid library was then pooled into groups of four EC, two AC, and eighteen clinical samples, and the pools were then purified using RNAclean™ Ampure beads as described above. After purification, the concentration of nucleic acids in the library pools was measured with a TapeStation as described above, and loaded on a sequencer according to the manufacturer’s recommendations.
[0347] Sequencing: Samples were sequenced to obtain sequence reads using an Illumina NextSeq™ 500 sequencer. Sequencing was performed using 76 cycles according to the manufacturer’s instructions.
[0348] Sequencing data analysis: Primary sequencing output was demultiplexed by bcl2fastq v2.17.1.14 (with default parameters), followed by removal of template switch oligos using Cutadapt. Poly A tails were removed, and reads were quality trimmed, and subsequently filtered by Trimmomatic v 0.32 if shorter than 20 bases. Reads that passed these filters were aligned to human and synthetic (containing process control molecules and sequencing adapters) references using Bowtie v2.2.4. Reads that aligned to either sequence were set aside. Reads potentially representing human satellite DNA were also filtered by a k-mer based approach. Remaining reads were aligned to a microbial reference database using BLAST v2.2.30. Reads that aligned exhibited both a high percent identity and a high query coverage were retained, with the exception of reads that aligned to any mitochondrial or plasmid reference sequences. PCR duplicates were removed based on their alignment. Relative abundances were assigned to each taxonomic unit in a sample based on sequencing reads and their alignment. For each combination of reads and taxonomic units, a read sequence probability was defined that resolved the ambiguity between the microorganism present in the sample and the reference assembly in the database. A mixture model was used to assign a likelihood to the complete set of sequencing reads that contained the read sequence probabilities and the (unknown) abundances of each taxonomic unit in the sample. An expectation-maximization algorithm was applied to compute the maximum likelihood estimate of each taxonomic unit abundance. From these abundances, the number of reads produced by each taxonomic unit was aggregated into a taxonomic tree. A set of libraries can be prepared from the respective negative control buffer and processed and sequenced within each batch. The estimated taxonomic unit abundances of the batch's internal negative control samples can be combined to parameterize that their variation by the environment is driven by a count noise. A statistical significance value can be computed for each estimated taxonomic unit abundance, and those at a high significance level within the CRR include candidate calls (i.e., significant calls). Final calls (i.e., reportable calls) are made after additional filtering that resolves read position consistency, read percent identity, and cross-reactivity from higher abundance calls. The number of reads of multiple fragment lengths for each reportable microorganism within each processed nucleic acid library is determined, the fragment length distribution is evaluated, and the fragment length characteristics of the distribution shape are determined. Figure 8Examples of some of the different fragment length distribution shapes observed in the detected microorganisms within the tested clinical samples are shown. Due to the minimum mapping length and the 68 bp set by combining the maximum read length in the described sequencing experiment and adaptor trimming algorithm, the range of fragment lengths shown is limited on the shorter end to 22 bp. Thus, fragments longer than 68 bp contribute to the count in the 68 bp length bin. The three microorganisms detected in the three examples shown are Candida tropicalis, Aspergillus oryzae, and W U polyomavirus 。 The fragment length distribution shapes vary greatly between these microorganisms and are not related to the specific species or superkingdom as shown by the rest of the data (not shown).
[0349] Candida tropicalis was detected in three different clinical samples processed here. A subset of reads from each sample that aligned to the Candida tropicalis reference genome was identified and its fragment length distribution determined. Figure 9 The results with the Candida tropicalis fragment length distribution from each of the three samples in the separate group are shown in the middle. In comparison to the right plot, the left and middle plots show distributions with a higher fraction of short (< 40 bp) and long (> 65 bp) fragments relative to the 50 bp peak, while it has a clear peak at around 45-50 bp. The left 2 plot is from a patient with disseminated Candida tropicalis infection, not limited by mechanism, which can explain the increased amount of long and short fragments relative to the peak. Different fragment length distributions can be indicative of different states of disease or condition. WU polyomavirus is another example of a microorganism detected in multiple clinical samples processed in this study and exhibiting different fragment length distributions in each sample (see Figure 10 ). In one subject, W WU polyomavirus shows only the "50 bp peak". The second subject shows a considerable contribution of short-like index fraction as well as a higher fraction of reads longer than 68 bp. While not limited by mechanism, the WU polyomavirus can have incorporated into the human genome in this sample, or its genome was released into the body fluids, which resulted in a different fragmentation pattern. Out of the total of 36 clinical samples (see above), 60, 24, and 13 bacterial, fungal, and viral microorganisms were detected, respectively. The fragment length distributions of these microorganisms vary greatly as demonstrated by the above examples. Next, the ratio of the count of reads detected in the "50 bp peak" peak to the short-like index region of the distribution of all detected microorganisms or pathogens was determined. The obtained ratios were grouped by their superkingdom and a histogram of the ratio characteristics of each superkingdom was generated. Figure 11Results from one such analysis are presented in FIG. 1. The same analysis was performed for human DNA (i.e., host DNA) and human mitochondrial DNA (i.e., host mitochondrial DNA) as controls Figure 11 The behavior of microorganisms depends on the superkingdom, an aspect that must be addressed when using fragment length distribution shapes and properties for diagnostic purposes.
[0350] Example 2: Analysis of plasma samples from pregnant subjects Many types of non-host nucleic acids can be found in samples obtained from a host. Fetal cell-free nucleic acids can be detected in maternal blood. In this sample, plasma samples were obtained from 15 consenting pregnant women and have been de-identified. The samples were processed and sequenced according to the direct-to-library method based on ligation described in Example 1 of U.S. Provisional Application 62 / 770,181, filed November 21, 2018, which is hereby incorporated by reference in its entirety. In this analysis, only samples from subjects who were pregnant with a male fetus were considered. Only reads that aligned to the Y chromosome were considered to be fetal. Reads were aligned to the human genome using bowtie2. Reads that mapped to chromosome Y were then aligned to an index created from all human chromosomes except Y using bowtie2. Any reads that aligned to this index were discarded, such that only reads specific to chromosome Y were retained.
[0351] Figure 12 The fragment length distribution of maternal (dashed line) and fetal (solid line) cell-free nucleic acids from one individual is presented in FIG. 1. In this example, the ratio of fetal to maternal reads in the “50 bp peak” region is higher compared to the nucleosome fragment region (e.g., 150-200 bp region). On average, the concentration of fetal fragments observed within the “50 bp peak” region was 4-fold higher than the nucleosome length fragment region. The methods employed here can be used to enrich the fetal fraction.
[0352] Example 3: Analysis of microorganisms using fragment length profiles Nucleic acid libraries were prepared from over 4000 cell-free plasma samples and sequenced using a validated Karius test that is based on an extraction method that recovers double-stranded DNA fragments in an unbiased manner with respect to their length and GC content within the fragment length range relevant for cell-free nucleic acids. The fragment length profile of detected microorganisms was generated and 33 taxonomic units that were called 10 or more times within the sample set under investigation were evaluated. More specifically, the ratio of the fraction of short reads in low probability calls and high probability calls was evaluated. Figure 13Results from one such experiment are presented. In this experiment, the plots indicate that there are more short reads in low probability calls compared to high probability calls. While not being bound by mechanism, these results can suggest that, given the end-repairable double-stranded cell-free DNA, clinical infections have a longer fragment length distribution than translocated colonizers or non-pathogenic organisms in the bloodstream.
[0353] Example 4: Analysis of colonization sites using fragment length profiles Nineteen clinical samples were obtained from subjects confirmed to be infected as determined by positive urine (n = 19) and / or blood culture tests (n = 11). Nucleic acid libraries were prepared from these samples and sequenced using a validated Karius test that is based on an extraction method that recovers double-stranded DNA fragments in an unbiased manner with respect to their length and GC content in the fragment length range associated with cell-free nucleic acids. In all nineteen subjects, blood and urine cultures were identified as 19 and 11 microorganisms, respectively. The fragment length distribution profile shape of the microorganisms detected by blood and urine cultures was evaluated. The results are shown in Table 14. While not being bound by mechanism, pathogen DNA from deep tissue infections (lungs, brain, etc.) can undergo different degradation mechanisms, affecting the fragment length observed as the DNA from the pathogen infects the blood.
[0354] Example 5: Length distribution profiles of host nucleic acids and infection state The fragment length distribution of host nucleic acids can help inform non-host nucleic acid signals within the host, such as microbial nucleic acid signals or the stage of infection of the host (e.g., asymptomatic versus symptomatic). For example, the abundance of microbial nucleic acids within a sample from a human host can vary by several orders of magnitude (Blauwkamp et al., (2016)). While samples obtained from asymptomatic individuals tend to exhibit lower abundance of microbial nucleic acids compared to infected individuals, the abundance measured in some asymptomatic samples can exceed the lowest abundance in infected individuals (Blauwkamp et al., (2016)). Additional properties of the pool of nucleic acids obtained from a sample can help distinguish between different stages of infection or biological relationship of a microbe to a host (e.g., symbiotic versus pathogenic). Here, the utility of the length distribution of host nucleic acids in predicting the state of infection of a microbe within the plasma from a host was tested. The methods enable access to the endogenous fragment length profile with fragment lengths that are not typically accessed in an unbiased manner with prior methods. The methods enable access to the endogenous fragment length profile with fragment lengths that are typically discarded, ignored, or deemed unimportant with prior methods.
[0355] Clinical plasma samples: Plasma samples were collected from human subjects, 100 asymptomatic (collection criteria: no active health tissue associated with infection and passed normal blood screening tests), 85 diagnostic positive (i.e., presence of a microorganism confirmed with orthogonal tests, e.g., blood culture, targeted PCR, Karius test), and 45 diagnostic negative. A single centrifugation plasma extraction process was performed on each sample from whole blood within 24 hours of sample collection, as previously described (see Fan HC et al., Proc Natl Acad Sci U S A. 2008; 105(42): 16266-16271, incorporated by reference in its entirety, including any figures), and stored at -80°C prior to use. Samples were then thawed and 5 pL of spike master mix (see above) was added to 500 pL of each plasma. If a smaller volume was obtained, a smaller volume of spike master mix was added in proportion to maintain a constant concentration of process control molecules in all initial samples and control samples.
[0356] Positive and negative control samples were prepared as described above.
[0357] Nucleic acid library generation and sequencing directly from plasma: Generation of direct-to-library is described in U.S. Provisional Application 62 / 770,181, filed November 21, 2018, which is incorporated by reference in its entirety. Libraries were prepared and sequenced as described in Example 1 above.
[0358] Results: The abundance of significant microorganisms present in each sample was determined as described above and given in units of concentration of the plasma sample, molecules per microliter (MPM), which is a normalized number that gives an estimate of the number of unique nucleic acid fragments of an organism in 1 microliter of a plasma sample. This calculation is derived from the number of sequences of unique or de-duplicated values of presence of each organism normalized to the known number of unique synthetic spike-in added to the plasma sample prior to the start of the process (see U.S. Patent No. 9,976,181). Figure 9 A shows the distribution of MPM values in asymptomatic (AP) and diagnostic positive (DP) sample types. The lower abundance values in the DP sample type overlap with the range of MPMs observed in the AP samples, even those containing only orthogonally confirmed microorganisms. (DP NGS Contain microorganisms confirmed by the Karius test, and DP microcontaining microorganisms confirmed by culture or PCR-based methods). In addition, if the analysis is limited to the microbial species present in the set of AP samples and also in the set of DP samples (in this dataset, the following species fit this description: Bacillus coagulans, Enterococcus, Enterococcus faecalis, Haemophilus influenzae, Haemophilus parainfluenzae, Human mammalian adenovirus D, Neisseria mucosa, Pediococcus acidilactici, Prevotella intermedia, Prevotella melaninogenica, Saccharomyces cerevisiae, Streptococcus agalactiae, Streptococcus salivarius, Streptococcus thermophilus), the abundance in the diagnostic positive group is still not always higher ( Figure 15B ). Thus, abundance is not sufficient to discriminate the infection status of non-microbial hosts.
[0359] A combination of several measurable parameters can then be used to distinguish between asymptomatic / healthy patients and patients experiencing an infection. To this end, a combination of MPM microbial abundance and nucleic acid fragment length distribution mapped to a host reference (i.e. the human reference in this sample cohort) was investigated as a potential classifier.
[0360] Figure 15C An example of a typical distribution of nucleic acid fragments as the library generation process is completed and as measured by a TapeStation instrument is shown. Two major peaks of fragment length can be observed: (1) the “nucleosomal” peak (in the range of 300-450 bp in the electropherogram) and (2) the “sub-nucleosomal” peak (in the range of 180-280 bp in the electropherogram). This signal is determined by the nature of the human (i.e. host) nucleic acids, as microbial (i.e. non-host) nucleic acids represent a tiny fraction of the total nucleic acid population in these samples, including the DP sample type. The molar and mass ratios of human fragments contributing to the two peaks vary between samples and are different between the AP sample type and the DP sample type ( Figure 15D ). The vast majority of AP samples (92%) show a “nucleosomal” peak molar fraction below 0.4, while the same value is equally distributed over a wider range for DP samples (< 0.7).
[0361] MPM microbial abundance and the nature of the human fragment length distribution show overlap between values between AP samples and DP samples. A combination of two independent measurements can help to distinguish between an asymptomatic call and an infection call in unknown samples for which the infection stage is unknown. Figure 15ELong human read fraction as measured from sequencing data (all reads mapping to human reference longer than 65 bp after adaptor trimming) and maximum MPM value measured in the same sample of all AP and DP samples are shown. The region encompassed by the coordinates [(0, 3000), (0, 0.4)] is exclusively populated by AP samples. Three of the 100 AP samples fall outside this space (arrow in 15E). The microorganisms detected in these three samples were Helicobacter pylori, human mammalian adenovirus, and Neisseria gonorrhoeae. All three microorganisms are known human pathogens, but it is not known whether they are pathogenic in these individuals.
[0362] Comparison between the properties of microbial MPM and human fragment length distribution in AP and DN sample types Figure 15F reveals that no DN samples fall within the typical asymptomatic range, even if they are negative according to the orthogonal test.
[0363] Non-microbial signals, such as properties of the fragment length distribution of non-microbial host nucleic acids, can be utilized to identify asymptomatic or non-infected status of a subject.
[0364] The data also indicate that asymptomatic individuals can be identified by combining abundance (e.g. maximum MPM) and fragment length distribution parameters, as presented herein, even if the MPM value of the microorganism overlaps with the range that can be observed in diagnostic positive samples. This also suggests that early detection of infection is possible in the absence of standard symptoms. The regions on this two-dimensional plane that can help distinguish between different infection states of an individual can be further optimized for MPM or kingdom of e.g. specific microbial species, as well as microbial fragment length to improve the performance of the test.
[0365] Finally, the normalized size distribution of fragments aligned to the human genome (nuclear genome dominant), human mitochondrial genome, all pathogens, significant pathogens, and bacteria, eukaryotes, viruses, and archaea were computed for all samples. To distinguish AP from DP / DN samples, a classifier was trained on the fragment size distribution (feature), in this case by using logistic regression with L2 regularization. Logistic regression is a linear model for classification that multiplies the features by a set of weights before transforming with a logistic function. The weights are determined using standard numerical optimization techniques with L2 regularization, providing additional constraints to minimize the sum of the squares of the weights. This has the effect of reducing overfitting as well as multiple collinearity in the features. The accuracy of this model was evaluated by using the trained model to predict the probability of each sample being asymptomatic or symptomatic. A value > 0.5 indicates that the sample was predicted to be asymptomatic, and a value < 0.5 indicates that the sample was predicted to be symptomatic. In addition, the trained model provides the weights (coefficients). Positive coefficients indicate an association with asymptomatic individuals, and negative coefficients indicate an association with symptomatic individuals. FIG. 16 shows the accuracy of the training to predict asymptomatic and symptomatic infection status using the normalized size distribution of fragments aligned to: human genome (nuclear genome dominant), human mitochondrial genome, all pathogens, significant pathogens; and bacteria, eukaryotes, viruses, and archaea. The subset of nucleic acids from the library used to train the model impacts the accuracy of the model. In addition, the subset of nucleic acids from the library impacts the region of the fragment length distribution that has a positive predictive value for asymptomatic or symptomatic status. For example, the presence of long human fragments (> 60 bp) predicts a symptomatic status (FIG. 16, right panel), as do short (< 30 bp) pathogen fragments (FIG. 16, right panel). On the other hand, a high concentration of fragments around 50 bp predicts an asymptomatic status (FIG. 16, right panel), as do long (> 65 bp) pathogen fragments (FIG. 16, right panel). Figure 16A Figure 16C Figure 16A Figure 16C
[0366] Example 6: Distinguishing asymptomatic patients colonized with H. pylori from patients with active H. pylori -associated inflammation Plasma processing and DNA extraction: Plasma was extracted from whole blood samples within 24 hours of sample collection as previously described (Fan HC et al., Proc Natl Acad Sci U S A. 2008; 105(42): 16266-16271) and stored at -80 °C. When analysis was required, the plasma sample was thawed and immediately 0.5-1 ml of plasma was used to extract circulating DNA.
[0367] Sequencing library preparation and sequencing: Sequencing libraries were prepared from purified patient plasma DNA using NEBNext DNA library preparation master mix with standard Illumina indexes (purchased from IDT) and end-repair purification (e.g., MagBind beads, NEBNext end-repair module) for Illumina setup or using a microfluidics-based automated library preparation platform (Mondrian ST, Ovation SP ultra-low library system). Libraries were characterized using an Agilent 2100 Bioanalyzer (High Sensitivity DNA Kit) and quantified by qPCR.
[0368] qPCR validation of sequencing results for selected bacterial targets. Standard qPCR kits for quantification of selected bacterial targets (e.g., H. pylori) were used to validate sequencing results for a subset of free DNA samples. The qPCR assay was run on cfDNA extracted from approximately 1 ml of plasma and eluted in 100 ml Tris buffer (50 mM [pH 8.1-8.2]). Plasma extraction and PCR experiments were performed in different facilities. No template controls were run to validate PCR reagents included in each experiment.
[0369] After removal of low quality reads, reads were mapped to the human reference genome. Remaining reads, assumed to be microbiome-derived, were mapped to a reference database of target microbial genomes. A proprietary algorithm was used to calculate the relative abundance of each microbe. The algorithm reported organisms that were present in a statistically significant amount compared to controls. Organisms with overrepresented sequences were reported as positive.
[0370] Quality control (QC) metrics include the addition of ID spike synthetic nucleic acids that are one type of spike added as unique to each sample in a sequencing batch and other synthetic nucleic acid spike-in molecules ("SPANK molecules") spiked in at a constant concentration across all libraries. Thus, the number of SPANK molecules detected in a particular library in de-duplicated values is a proxy for the minimum concentration of the SPANK molecule that can be detected in the library. This can be used to set a threshold based on the minimum concentration of the SPANK molecule that can be detected in the library. The threshold can be used to ensure adequate sequencing depth for detection of a pathogen. The threshold can also be used to ensure that the pathogen signal is not due to cross-contamination from other samples. For example, the enrichment of a pathogen relative to the threshold set by the SPANK molecule can be compared between different samples. More generally, it is directly proportional to the efficiency with which the library converts DNA molecules in the original sample into reads in the DNA sequencing data. The purpose of the SPANK molecules is to help establish the relative abundance of pathogen molecules within the mixture represented in the sample, reported in "molecules per ml" (MPM). The MPM data is used to construct heat maps and correlation plots. The sample purity ratio (SPR) is intended to capture how the number of reads associated with a taxon gives an estimate of the degree of cross-contamination in the sample. In the event of a de-duplicated value of SPANK and / or SPR failure, the sample is re-queued and re-run once. If QC fails twice on the same sample, it is reported as "no result."
[0371] Results. The method was able to detect H. pylori free DNA in plasma obtained from patients with H. pylori -associated peptic ulcer disease. The method was able to distinguish between patients with asymptomatic H. pylori from patients with H. pylori disease. For the latter case, samples were obtained from healthy (i.e., asymptomatic) and infected subjects and analyzed using cell-free plasma next-generation sequencing to detect pathogen DNA (Karius Test™ by Karius, Redwood City, CA) to detect pathogen DNA. In healthy volunteers, the test detected H. pylori in 8 / 106 samples assayed. Some patients were identified in the dataset as having asymptomatic colonization with H. pylori (C) (n = 1) or symptomatic chronic infection with H. pylori (CI) (n = 7) (see Table 1 below). H. pylori positive samples were associated with African-Weekly or Spanish ethnicity, which is consistent with the epidemiology of H. pylori infection.
[0372] Table 1: Detection of H. pylori in plasma. H. pylori infection was a likely, possible, or unlikely cause of sepsis
[0373] Without being bound by a mechanism, free nucleic acids can be derived from dead and dying pathogens. Thus, the present method is uniquely suited to detect organisms that are actively being cleared by the immune system. Indeed, the assay is able to distinguish between H. pylori in the context of active inflammation, rather than asymptomatic colonization.
[0374] Example 7: Method for detecting H. pylori GI tract infection among high risk patients The objective of this study was to evaluate the clinical utility of the present method (i) to detect active H. pylori infection in symptomatic patients with peptic ulcer disease (H. pylori PUD) compared to conventional diagnostic testing; (ii) to confirm eradication of active H. pylori gastrointestinal infection after first-line therapy compared to conventional diagnostic testing; and (iii) to evaluate the optimal MPM threshold to distinguish patients with active H. pylori PUD from those without (asymptomatic). Using this non-invasive method allows physicians to make effective treatment decisions without resorting to traditional invasive diagnostic methods.
[0375] Study Design. The percent positive agreement (PPA) and percent negative agreement (PA) of the present method compared to non-serological conventional H. pylori diagnostic tests were determined in two well-described adult study populations under specific test conditions as described below. At study entry, patients with symptomatic H. pylori PUD met clinical criteria and had at least one positive protocol-approved, non-serological conventional H. pylori diagnostic test prior to any administration of first eradication therapy. Plasma testing was performed on all documented symptomatic H. pylori PUD patients. Thereafter, these PUD patients received a standard eradication regimen (per standard of care) for 2-4 weeks followed by a 1 -month drug holiday. Within 30 days (+ / - 3 days) after completion of first therapy, all PUD patients who ended study participation underwent repeat plasma testing evaluation and at least one of the original non-serological, conventional H. pylori diagnostic tests performed prior to therapy.
[0376] At study entry, negative control patients who underwent colonoscopy for any reason had no evidence of active H. pylori gastrointestinal disease during screening based on clinical criteria and at least one negative protocol-approved non-serological conventional H. pylori diagnostic test. Thereafter, negative control colonoscopy patients were plasma tested to complete all protocol requirements.
[0377] Data from these diagnostic test comparisons provide insight into the utility of the present method for detecting active H. pylori disease and confirming eradication after first therapy compared to non-serological conventional H. pylori diagnostic tests.
[0378] Methods and Materials A quantitative test method is used to detect microorganisms by analyzing non-human DNA in blood plasma. The analyte of this method is microbial free nucleic acid that is very short (average length less than 100 nucleotides) compared to human cfDNA.
[0379] Whole blood is centrifuged twice to present free (cf) plasma. To address potential environmental contaminants, non-volatile buffers can be heated to a temperature exceeding 85°C and cooled prior to use. After the first centrifugation, internal control molecules are added to each sample using the methods set forth in PCT-US2017-024176. Plasma is extracted and purified free DNA (cfDNA) is used to prepare sequencing libraries using NEBNext DNA Library Preparation Master Mix with standard Illumina indexes (purchased from IDT) and end-repair purification (e.g., MagBind beads, NEBNext End Repair Module) set up for Illumina or using a microfluidics-based automated library preparation platform (Mondrian ST, Ovation SP Ultra-Low Library System). Adapters are ligated and purification is performed using AMPure beads without heating, followed by amplification by qPCR. Libraries are characterized using an Agilent HS TapeStation and total concentration of nucleic acids is measured to control loading volume by integrating the signal (e.g., between 50 bp and 1000 bp) for a small selection step.
[0380] Sequenced cfDNA fragments are mapped to a reference database of microbial sequences to determine the identity of non-human, non-internal control material present in the sample at significant levels (above assay background). Sequencing data is first converted to reads representing DNA sequences and then demultiplexed into sets of reads (read sets) derived from each library loaded into the sequencer based on index sequences. Reads aligned to human sequences are filtered and the remaining reads aligned to sequences of internal control molecules are set aside for additional analysis. Next, reads that do not match the human reference or internal control reference are aligned to known microbial genomes. Reads with one or more alignments to this database (pathogen reads) are the basis for subsequent analysis.
[0381] The relative abundance of each taxon associated with a reference sequence is inferred using the alignment of each pathogen read to a database of microorganism genomes. These abundances are aggregated into a taxonomic tree to give the abundance of all taxon ranks. Finally, on the same sequencing run, the abundances in the clinical sample are compared to the abundances in the negative control library to determine if they are elevated above the expected background level due to environmental DNA contamination. Taxa that meet this criterion are reported in molecules per milliliter (MPM) based on the ratio of microorganism reads to the abundance of certain internal control reads obtained. Prior to calling results, the pipeline applies a set of filters to limit the organisms that can be reported to be greater than, for example, 3-10% of the microorganism with the highest number of reads and greater than, for example, 25-50% of any other taxonomic family related organisms. The filters are applied to all patient samples and assay controls.
[0382] Potential sources of performance bias with sample-specific or microorganism-specific properties include: the class of microorganism-specific properties including the class of microorganism (e.g., bacteria, viruses, eukaryotes, prokaryotes, fungi, etc.), GC content, genome size, abundance of endogenous microbiome, level of environmental contamination (EC), and number of reference assemblies and quality of data. To address these sources of bias, the method includes using a representative panel of 10-100 microorganisms that capture the full spectrum of potential performance bias along GC content, genome size, and strain. These representative organisms should span kingdoms, range within GC content (e.g., 10-80%), and have genomes ranging from kilobases to megabases. The representative panel should include a mix of types, such as commensals and non-commensals, microorganisms that are commonly found as environmental contaminants, and closely related strains. The method additionally incorporates standard quality control measures, such as reference intervals for the levels of microorganisms in a healthy population and EC negative controls.
[0383] If the test shows that H. pylori is significant compared to the negative background control, the test will be considered positive. Note, however, that after the determination of the quantitative MPM cutoff, the negative percent identity (NPA) is unlikely to reflect the NPA of the test.
[0384] In addition to evaluating the PPA and NPA within each of the study cohorts, PUDs, and colonoscopies using the threshold for positivity and negativity in MPM as determined by the laboratory, other thresholds in MPM should be considered. First, the mean, standard deviation, median, and range of MPM will be used for each study cohort to aggregate MPM. Second, the receiver operating characteristic (ROC) curve will be used to identify the optimal cutpoint in MPM for maximizing the PPA and NPA in the samples.
[0385] Finally, to assess the ability of the method to identify eradication at 30 days, the proportion and 95% confidence interval within each of the study cohorts will be estimated for successful eradication.
[0386] Example 8: Fragment length distribution profiles and colonization sites The characteristics of the fragment length distribution of microbial sequencing reads obtained from clinical samples from patients infected in the bloodstream and lungs were compared as an example of deep tissue infection. The fragment length distribution characteristics varied depending on the site of localization. Without being limited by mechanism, different host responses to different infection sites can contribute to altering the fragment length distribution characteristics. Again, without being limited by mechanism, different infection sites can exhibit different non-host nucleic acid fragmentation mechanisms.
[0387] Clinical plasma samples: Ten de-identified clinical samples from patients confirmed to have bloodstream infections and ten de-identified clinical samples from patients confirmed to have lung infections were collected. A single centrifugation plasma extraction process was performed on each sample within 24 hours of sample collection from whole blood as previously described (see Fan HC et al., Proc Natl Acad Sci U S A. 2008; 105(42): 16266-16271, which is incorporated by reference herein in its entirety, including any figures), and stored at -80°C prior to use. The samples were then thawed and 1.5 μΐ of spike master mix (see below) was spiked into 150 μΐ of each plasma. If a different volume was obtained, a smaller or higher volume of spike master mix was added proportionally to maintain a constant concentration of process control molecules in all initial samples and control samples.
[0388] Negative control samples: Four 500 pF negative control samples (EC) were made from an aqueous buffer (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) with 5 μΐ of spike master mix (see below) and were used as controls for environmental contamination (e.g., microbial and pathogen nucleic acid contamination introduced through reagents, instruments, consumables, operators, and / or air during processing). These synthetic nucleic acids were then used to normalize the signal in the samples to account for variation in sample processing.
[0389] The spike master mix was prepared as described above with ID spike molecules, SPANK molecules, and SPARK molecules.
[0390] A ligation-based direct-to-library method as described in Example 1 of U.S. Provisional Application 62 / 770,181 was used to prepare sequencing libraries from 5 μΐ of spiked asymptomatic plasma. Sequencing and sequencing data analysis were performed as described in Example 8.
[0391] Results: Table 2 lists all 20 clinical samples as part of this example along with the site of infection and the species of infecting microorganism for each subject that donated a clinical sample. The fragment length distribution of the infecting microorganism in all of the samples tested is shown in Figure 17. The normalized fragment length distribution of reads mapping to the reference of the infecting microorganism was analyzed for the presence of the following fragment length distribution profile characteristics (e.g., short exponentially decaying fragments, peaks, long fragments): (1) short exponential-like distribution fraction (Table 2 as "short"), (2) peak fraction (Table 2 as "peak"), and (3) fraction of reads longer than the read length of the experiment (75 bp; Table 2 as "long"). Also, the fraction of reads of typical length ranges in the microbial fragment length distribution. Comparison of the fragment length distribution profile types revealed that bloodstream infections disproportionately exhibit a fragment length distribution profile characterized by: (1) a high fraction of fragments of short pseudo-exponential distribution, (2) the absence of peaks between 20 bp and 75 bp read length, and (3) a fraction of long reads (> 64 bp) greater than 10% in. In contrast, pulmonary infections disproportionately exhibit a fragment length distribution profile characterized by: (1) the presence of fragments of short pseudo-exponential distribution, (2) the presence of peaks between 20 bp and 75 bp read length, and (3) a fraction of long reads less than 10%. This indicates that the characteristics of the microbial fragment length distribution can be used to determine if the infection is present in the blood or in deep tissue.
[0392] Table 2: List of clinical samples and site of infection, species of infecting microorganism, and properties of the fragment length distribution of sequencing reads mapping to the reference of the species of infecting microorganism. For each property, its quantitative assessment (presence / absence) is indicated and the fraction of total reads present in the segment is given in parenthesis. Here, short fragment segments contain reads of 22 bp up to and including 29 bp; peak fragment length range contains reads of 30 bp up to and including 59 bp; and long fragment range contains reads longer than 59 bp.
[0393]
[0394] Example 9: Fragment length distribution profiles and colonization sites 2 The properties of the fragment length distribution of microbial sequencing reads obtained from clinical samples from patients with infections located in the bloodstream (plasma from venous blood draws) and from capillary blood that was in contact with the skin on the tip of the finger prior to its collection in a capillary draw collection system were compared as an example of skin infections.
[0395] Clinical plasma samples: Blood from 20 healthy adult donors was collected into PPT tubes with K2EDTA as the anticoagulant (Becton Dickinson, Franklin Lakes, NJ) according to the manufacturer’s instructions. Immediately after venous blood draw, capillary blood draws were performed on the same group of 20 healthy donors using a Microvette CB300 blood sampling device using K2EDTA as the anticoagulant (Sarstedt Inc, Sparks, NV). During the capillary draw, the following steps were performed: (1) the donor’s finger was held in an upward position and a lancet of appropriate size was used to prick the palmar surface of the finger, (2) the finger was not pressed down hard during the prick to prevent hemolysis of the drawn blood, and (3) a drop of blood that spread over the tip of the finger was collected into a clean Microvette CB300 blood sampling device. Isolation of plasma from whole blood was performed on each sample within 12 hours of sample collection according to the manufacturer’s instructions and the plasma was stored at -80°C until use. The samples were then thawed and a volume equivalent to 1% of the plasma volume of the spike master mix (see below) was added to each plasma.
[0396] Negative control samples: Four 500 µL negative control samples (EC) were made from aqueous buffer (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) with 5 µL of spike master mix (see below) and were used as controls for environmental contamination (e.g., microbial and pathogen nucleic acid contamination introduced by reagents, instruments, consumables, operators, and / or air during processing). These synthetic nucleic acids were later used to normalize the signal in the samples to account for variability in sample processing.
[0397] Negative Microvette samples: Four 300 µL of aqueous buffer (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) were added to four clean and unused Microvette CB300 blood sampling devices and incubated at room temperature for 6 hours before quantitatively collecting the contents and spiking with 3 µL of spike master mix (see below).
[0398] The spike master mix was prepared with the ID spike molecules, Spank molecules, and Spark molecules as described above.
[0399] Direct from plasma nucleic acid library generation: Twenty-five 0 μΐ, of each spiked sample was mixed with 10.0 μΐ, of 10x terminal transferase reaction buffer (New England BioLabs, Ipswich, MA), 2.5 μΐ, of proteinase K (Sigma), 1.0 μΐ, of 10% Tween-20 (Thermo Fisher Scientific, Waltham, MA), 1.0 μΐ, of 10% Triton X100 (Thermo Fisher Scientific, Waltham, MA), and 60.5.0 μΐ, of nuclease-free water. The mixture was heated to 60°C for 20 minutes and to 95°C for 10 minutes, and it was placed on ice until cooled. 1.0 μΐ, of 10 mM dATP, 1.0 μΐ, of terminal transferase (20 u / μL, New England BioLabs, Ipswich, MA), and 3.0 μΐ, of nuclease-free water were added to make an A-tail reaction, which was then incubated at 37°C for 40 minutes. 150.0 μΐ, of lysis / binding buffer (Thermo Fisher Scientific, Waltham, MA) was added to the reaction. The entire volume was then added to 25.0 μΐ, of Dynabeads oligo(dT) 25 (Thermo Fisher Scientific, Waltham, MA), then washed once with lysis / binding buffer (Thermo Fisher Scientific, Waltham, MA). The mixture was incubated at 25°C and 600 RPM. The remainder of the process followed the steps of the protocol outlined in Example 1.
[0400] Sequencing: The samples were sequenced to obtain sequence reads using an Illumina NextSeq™ 500 sequencer. Sequencing was performed according to the manufacturer's instructions. Sequencing analysis was performed as described in Example 1 above.
[0401] Results: Figure 18A shows the normalized fragment length distribution of microorganisms detected in venous draws in two of the donors in this study, and Figure 18B shows the normalized fragment length distribution of microorganisms detected in one of the replica capillary draws in the same two donors. The two microorganisms detected in the venous draws (e.g., H. influenzae in vivo in Donor 1 and S. thermophilus in vivo in Donor 2) were also detected in biological samples obtained during the capillary draw collection process, and the two microorganisms showed similar fragment length distributions in both collection types, i.e., the peaked fragment length distributions (Figure 18A and Figure 18B). The additional microorganisms detected in samples obtained with the methods applied during capillary draws comprised a more diverse set of microorganisms (Table 3). Most of these additional microorganisms co-existed in both replicas / each donor (Figure 18C). To confirm that these additional microorganisms were not caused by contaminants present in the Microvette CB300 blood sampling device used to collect the procedures applied during capillary draws or derived from process contamination, the sequencing data obtained from the negative Microvette samples were analyzed (see above). Figure 18D shows a comparison of the abundance in MPMs of the additional microorganisms in biological samples obtained from the process applied during capillary blood draws (x-axis) and the abundance in MPMs of the same microorganisms in the negative Microvette samples. The vast majority of the signal of the additional microorganisms in the data obtained from capillary draws was not caused by tube contamination profiles, and it can be concluded that the vast majority of the signal was derived from biological samples obtained by collecting blood droplets from fingertips. Since no signal of these microorganisms was detected in venous draws, it must have originated from the skin surface onto which blood spread after pricking the skin of the fingertips, indicating that microorganism nucleic acids derived from the skin showed different properties of their fragment length distributions, e.g., the absence of a peak between 20 bp and 75 bp, and an exponential-like decay of the fragment frequency with fragment length. The same trend was observed in other sample donors (data not shown).
[0402] Table 3: List of microorganism species detected in biological samples obtained during the process applied from capillary blood draws in donor 1 and donor 2. Example 10: Post-transplant infection
[0403] Example 11: Colonization site evaluation Ten transplant patients were monitored for possible infections after transplant surgery, and the fragment length distribution of pathogens detected at the presymptomatic stage was monitored for changes to correlate the infection stage with observed fragment lengths. Specifically, the presence of peaks between 20 and 75 bp and the fraction of fragments not associated with such peaks were tracked as the infection progressed through different stages. In addition to the ten transplant patients, ten de-identified, serial sampling sets from Karius production were selected to track the same behavior.
[0404] Example 12: Infection stage determination from microbial fragment length distribution A direct-to-library approach based on template switching was used to spike and process 1000 de-identified samples from Karius production with Proteinase K as described in U.S. provisional 62 / 770,181, along with assay controls and environmental controls. The 1000 de-identified samples contained plasma samples from patients with pneumonia, immunocompromised states, endocarditis, sepsis, or invasive fungal infections, and the analysis of umbilical cord abundance as well as microbial and host fragment length distributions was used to correlate features of the fragment length distribution (e.g. presence or absence of peaks between 20 and 75 bp, fraction of reads longer than 65 bp, fraction of reads shorter than 40 bp) to infection sites, especially to the presence of peaks in deep tissue infections or in symbiosis.
[0405] Figure 20A To determine the predictive value of fragment length profile diagnostics for measuring infection stage, a set of clinical plasma samples was collected from 16 different consenting subjects suspected of having an infection by drawing blood into PPT tubes and extracting plasma through a single centrifugation step according to the manufacturer’s instructions. The plasma samples were either frozen or shipped overnight at ambient temperature to the Karius lab in Redwood City, CA. For each subject, a first sample was obtained upon hospital admission, at which time orthogonal tests (e.g. blood cultures) were also performed to confirm possible microbial species responsible or partially responsible for the infection. Subsequently, additional samples were drawn from the subject at various time points during treatment to monitor the progression of the infection and the effect of treatment. In total, samples were collected at least at two time points per subject (including the time point of hospital admission). The maximum number of time points per subject was 7. The plasma samples and negative control samples were processed into nucleic acid libraries and sequenced as described above.
[0406] The group of subjects of this study contained 3 patients diagnosed orthogonally with bloodstream infection, 8 patients diagnosed orthogonally with endocarditis, and 5 patients diagnosed orthogonally as febrile agranulocytotic patients. Figures 19A, 19B, and 19C show the change in fragment length distribution in representative examples of bloodstream infection, endocarditis, and febrile agranulocytosis, respectively. The example fragment length distributions in Figure 19 indicate a high probability of short exponentially distributed fragments (range < 40 bp), and an increased probability of peaked distribution around 50 bp after treatment has begun. Thus, the fraction of short exponentially distributed or near exponentially distributed fragments in all treated samples was investigated. Figure 20B The kinetics of this short read fraction change were plotted. This shows that invasive infections can be diagnosed based on the presence of a short and exponentially distributed read fraction, especially in the case of bloodstream infection or bacteremia. Within a single subject, there is a high read fraction of > 64 bp, which can indicate saturation of the mechanism giving short exponentially distributed fragments (data not shown). Simultaneous measurement of microbial abundance elicobacter pylori ) enables determination of the stage of infection by combining the use of abundance and fragment length profile measurements.
[0407] The sequencing data also indicated the presence of microorganisms not confirmed by other microbial tests performed orthogonally. In the case of these microorganisms, the fragment length distribution can also be investigated. For example, Haemophilus influenzae and Prevotella melaninigenica were detected in the admission sample from subjects RD-06 and RD-13, respectively, by the disclosed methods (Figure 21A) 。 While the microorganisms were detected orthogonally, the presumed cause of the infection shows a high short read fraction in both cases, the additional microorganisms show variable trends; the Haemophilus influenzae fragment length distribution is consistent with an invasive or bacteremic infection, while the Prevotella melaninigenica shows only the presence of a peaked distribution, which is consistent with an asymptomatic patient’s infection in an amorphous stage or commensal behavior (see, e.g., “Helicobacter pylori fragment length distribution (H distribution”) in U.S. Provisional Application No. 62 / 770,181, filed November 21, 2018, entitled “Methods, systems, and compositions direct to library” or a managed infection footprint. Additionally, new microorganisms can emerge during the course of treatment, and fragment length analysis can also assist in diagnosing the infection status of these. For example, Figure 21B shows the fragment length distribution of reads aligned to Enterococcus gallinarum, which shows a detectable fraction of short exponentially distributed reads with a string of peak fractions. The decision to treat this infection can be based on the magnitude of the short read fraction. Examination of the clinical record confirms that the subject did undergo treatment of this infection.
[0408] Finally, changes in human fragment length distributions were analyzed as the subjects under study moved from the symptomatic phase of infection at admission and diagnosis, through the infectious cycle, and to the treatment of infection phase during therapy. FIG. 22 depicts the three main behavioral patterns of human fragment distributions for infected patients in this study: (1) the fraction of long (mainly nucleosomal) human fragments decreased during therapy (left panel of FIG. 22, 37.5% of total subjects in this study); (2) the fraction of long human reads floated during therapy (middle panel of FIG. 22, 37.5% of total subjects in this study); and (3) the fraction of long (mainly nucleosomal) human reads increased during therapy (right panel of FIG. 22, 37.5% of total subjects in this study). As shown above, the human fragment length distribution shape and properties can predict the infection phase of a subject. Parameters derived from the human distribution can then be used in combination with the fragment lengths of the infectious or other microorganisms detected in the sample to predict the recovery trajectory of the subject, e.g., whether the subject is recovering, whether another microorganism will infect the subject during treatment of the initial infection, or to recognize an invisible infection or symbiosis.
Claims
1. A fragment length profile from a nucleic acid library, wherein the nucleic acid library is generated from an initial sample, and wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparation of the nucleic acid library, wherein the fragment length profile comprises one or more characteristics selected from the group comprising: shape of the distribution, amplitude of bins, peak shape, ratio of fragment counts of two or more bins, height of helical phasing peaks, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads.
2. A method of generating a fragment length profile of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample using a bias- corrected recovery method; (b) determining the number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of bins, peak shape, ratio of fragment counts of two or more bins, height of helical phasing peaks, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.
3. A method of generating a fragment length profile of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample, the preparing a nucleic acid library from an initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparation of the nucleic acid library; (b) determining the number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of bins, peak shape, ratio of fragment counts of two or more bins, height of helical phasing peaks, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.
4. A method of identifying a localization site in a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample; (b) comparing the fragment length profile to a reference fragment length profile of one or more source sites; and (c) if the fragment length profile from the sample is similar to a fragment length profile from a first source site, then the first site is identified as a localization site; if the fragment length profile from the sample is similar to a fragment length profile from a second source site, then the second site is identified as a localization site.
5. A method of monitoring toxicity of a compound administered to a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample; and (b) comparing the fragment length profile to one or more reference fragment length profiles.
6. A method of determining a stage of infection in a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample obtained from the subject; (b) comparing the fragment length profile to a reference fragment length profile; and (c) if the fragment length profile from the sample is similar to a fragment length profile from a symptomatic subject, then the stage of infection is determined to indicate an increased risk that the subject exhibits a symptom associated with the microbe; if the fragment length profile from the sample is similar to a fragment length profile from an asymptomatic subject, then the infection is determined to be in an asymptomatic stage.
7. A method for determining a stage of infection in a subject suspected of having a microbial infection, the method comprising: a) performing high-throughput sequencing of nucleic acids from a biological sample; b) performing bioinformatic analysis to identify free nucleic acid sequences present in the biological sample; and c) obtaining a measurement of the free nucleic acids and comparing the measurement to a control, whereby a stage of infection of the microbe identified in the biological sample is determined.
8. A method of determining a stage of infection of Helicobacter pylori in a subject, the method comprising: (b) extracting free nucleic acids from a biological sample obtained from the subject; (c) adding synthetic nucleic acid spike-in to the free fraction; (d) performing high-throughput sequencing of nucleic acids from the biological sample; (e) performing bioinformatic analysis to identify free Helicobacter pylori nucleic acid sequences present in the biological sample; and (f) calculating a measurement of the free Helicobacter pylori nucleic acids and comparing the measurement to a control, whereby a stage of infection of Helicobacter pylori in the subject is determined.
9. A method of determining a host-microbe biological interaction in a subject, the method comprising: (a) generating a fragment length profile of a nucleic acid library generated from a sample from the subject; (b) optionally, determining an abundance of a target nucleic acid and comparing the abundance to a threshold; (c) comparing the fragment length profile to one or more reference fragment length profiles of host-microbe biological interactions; and (d) if the fragment length profile is similar to a reference fragment length profile of a host-microbe biological interaction, then the host-microbe biological interaction is identified.
10. A method of identifying a presence of a viral infection in a subject suspected of having a microbial infection, the method comprising: a) generating a fragment length profile of a nucleic acid library generated from a sample from the subject; b) comparing the fragment length profile to a viral reference fragment length profile; c) optionally, quantifying the abundance of the target nucleic acid and comparing the abundance to a threshold value; d) identifying the presence of a viral infection in the subject if the fragment length profile is similar to the reference profile.
Citation Information
Patent Citations
Cell-free nucleic acids for the analysis of the human microbiome and components thereof
US20150133391A1
Compositions and methods for analyzing heterogeneous samples
US20150211070A1
Compositions and methods for enriching populations of nucleic acids
US20170016048A1
Image forming apparatus and method of cloning using mobile device
US20170024176A9
Non-invasive determination of fetal inheritance of parental haplotypes at the genome-wide scale
US8877442B2