Detection and prediction of infectious disease

A non-invasive method using fragment length spectra of nucleic acid libraries addresses the limitations of current diagnostics by accurately differentiating infection stages and localization sites, enhancing diagnostic precision and reducing treatment errors.

HK40135052APending Publication Date: 2026-07-17KARIUS INC

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
KARIUS INC
Filing Date
2026-05-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Current methods for diagnosing microbial infections, particularly those caused by pathogens like Helicobacter pylori, are invasive, inaccurate, or lead to overtreatment/undertreatment due to inability to differentiate between asymptomatic and symptomatic stages, and existing next-generation sequencing (NGS) methods introduce biases and reduce nucleic acid recovery rates.

Method used

A non-invasive method using fragment length spectra of nucleic acid libraries generated without prior extraction, incorporating bias correction, to identify microorganisms, infection stages, and localization sites, by analyzing characteristics such as fragment length distribution and ratios.

Benefits of technology

Enables precise differentiation between asymptomatic and symptomatic infection stages and localization of infections without invasive procedures, improving diagnostic accuracy and reducing treatment errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention provides detection and prediction of infectious diseases. In particular, the present application provides fragment length profiles of nucleic acid libraries, methods of generating fragment length profiles of nucleic acid libraries, and methods of using fragment length profiles for diagnosis and / or prognosis. The present application further provides methods, compositions, and kits for determining the stage of infection or localization in a subject.
Need to check novelty before this filing date? Find Prior Art

Description

(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202511721869.X (22) Application Date 2019.11.21 (30) Priority Data 62 / 770,181 2018.11.21 US 62 / 770,182 2018.11.21 US 62 / 849,618 2019.05.17 US (62) Divisional Application Data 201980083444.7 2019.11.21 (71) Applicant: Karius Corporation Address: California, USA (72) Inventors: S. Belkovich, L. Blair, T.A. Broucamp, P.J. Eugster, D. Hollimon, D.K. Hon, T. Kali, M.A. Kowaski, M.M.S. Lindner, M.J. Rosen, D. Spike, I.D. Verfan (74) Patent Agency: Beijing Anxin Fangda Intellectual Property Agency Co., Ltd. 11262 Patent Attorneys: Xu Aiwen, Wu Jingjing (51) Int.Cl. C40B 40 / 06 (2006.01) C40B 30 / 06 (2006.01) C40B 50 / 06 (2006.01) C12Q 1 / 6869 (2018.01) (54) Invention Title: Detection and Prediction of Infectious Diseases (57) Abstract: This application provides for the detection and prediction of infectious diseases. Specifically, this application provides fragment length spectra of nucleic acid libraries, methods for generating fragment length spectra of nucleic acid libraries, and methods for diagnosis and / or prognosis using fragment length spectra. This application further provides methods, compositions, and kits for determining the stage or location of infection in a subject. Claims 2 pages, Description 68 pages, Drawings 47 pages. CN 121575489 A 2026.02.27 CN 1 21 57 54 89 A 1. A fragment length spectra from a nucleic acid library, wherein the nucleic acid library is generated from an initial sample, and wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library, wherein the fragment length spectra include one or more characteristics selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads. 2. A method for generating a fragment length spectrum of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample using a bias-corrected recovery method; and (b) determining the number of reads of multiple fragment lengths within the nucleic acid library;(c) Determine one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads; and (d) Use the one or more fragment length characteristics to generate a fragment length spectrum of the nucleic acid library. 3. A method for generating a fragment length spectrum of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample, the preparation of the nucleic acid library from the initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library; (b) determining the number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads; and (d) generating a fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics. 4. A method for identifying a localization site in a subject, the method comprising the steps of: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample; (b) comparing the fragment length spectrum with reference fragment length spectra of one or more source sites; and (c) identifying the first site as a localization site if the fragment length spectrum from the sample is similar to a fragment length spectrum from a first source site; and identifying the second site as a localization site if the fragment length spectrum from the sample is similar to a fragment length spectrum from a second source site. 5. A method for monitoring the toxicity of a compound administered to a subject, the method comprising the steps of: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample; and (b) comparing the fragment length spectrum with one or more reference fragment length spectra. 6. A method for determining the stage of infection in a subject, the method comprising the steps of: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample obtained from the subject; (b) comparing the fragment length spectrum with reference fragment length spectra; and (c) [Claims 1 / 2 page 2 CN]121575489 A (c) If the fragment length spectrum from the sample is similar to that from a symptomatic subject, the infection stage is determined to indicate an increased risk of the subject exhibiting microbial-related symptoms; if the fragment length spectrum from the sample is similar to that from an asymptomatic subject, the infection is determined to be in an asymptomatic stage. 7. A method for determining the infection stage of a subject suspected of having a microbial infection, the method comprising: a) performing high-throughput sequencing on nucleic acids from a biological sample; b) performing bioinformatics analysis to identify free nucleic acid sequences present in the biological sample; and c) obtaining a measurement of the free nucleic acid and comparing the measurement with a control, thereby determining the infection stage of the microorganism identified in the biological sample. 8. A method for determining the stage of Helicobacter pylori infection in a subject, the method comprising: (b) extracting free nucleic acid from a biological sample obtained from the subject; (c) adding a synthetic nucleic acid spiking compound to the free fraction; (d) performing high-throughput sequencing on the nucleic acid from the biological sample; (e) performing bioinformatics analysis to identify the free Helicobacter pylori nucleic acid sequence present in the biological sample; and (f) calculating the measurement result of the free Helicobacter pylori nucleic acid and comparing the measurement result with a control, thereby determining the stage of Helicobacter pylori infection in the subject. 9. A method for determining host-microbe biological interactions in a subject, the method comprising: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample from the subject; (b) optionally, determining the abundance of a target nucleic acid and comparing the abundance to a threshold; (c) comparing the fragment length spectrum to a reference fragment length spectrum of one or more host-microbe biological interactions; and (d) identifying the host-microbe biological interaction if the fragment length spectrum is similar to the reference fragment length spectrum of the host-microbe biological interaction. 10. A method for identifying the presence of a viral infection in a subject suspected of having a microbial infection, the method comprising: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample from the subject; (b) comparing the fragment length spectrum to a viral reference fragment length spectrum; (c) optionally, quantifying the abundance of a target nucleic acid and comparing the abundance to a threshold; and (d) identifying the presence of a viral infection in the subject if the fragment length spectrum is similar to the reference spectrum. Claims 2 / 2 Page 3 CN 121575489 A Detection and Prediction of Infectious Diseases

[0001] This application is filed on November 21, 2019, with application number 201980083444.7, and entitled "Infection..."This application claims a divisional application of the Chinese patent application entitled "Detection and Prediction of Infectious Disease" (the corresponding PCT application was filed on November 21, 2019, with application number PCT / US2019 / 062665). Cross-Reference to Related Applications

[0002] This application claims U.S. Provisional Application No. 62 / 770,182, filed November 21, 2018, entitled "Detection and Prediction of Infectious Disease"; U.S. Provisional Application No. 62 / 770,181, filed November 21, 2018, entitled "Direct-to-Library Methods, Systems and Compositions"; and U.S. Provisional Application No. 62 / 770,181, filed May 17, 2019, entitled "Fragment Length Distributions and Methods of Using...". Priority and interest in U.S. Provisional Application No. 62 / 849,618, “Such,” the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0003] This invention relates to the use of fragment length distributions in nucleic acid libraries to identify microorganisms, identify the type of host-microbe biological interaction, identify infection sites or localization sites, select therapies or treatments, monitor treatments, monitor cytotoxicity, detect transplant rejection, monitor immune system response or activity, identify infection stages, monitor transplant rejection, and for cancer diagnosis. Background Art

[0004] For many microbial infections, the first stage is colonization. In some cases, microbial infections may progress to persistent infection and may develop into an invasive disease stage. Examples of microorganisms that may develop into invasive diseases include cytomegalovirus, Epstein-Barr virus, Helicobacter pylori, and Clostridium difficile. (e.g., difficile infections), certain sexually transmitted infections, etc. For patients infected with these types of microorganisms, identifying the corrective, colonizing, or invasive stage of infection can be an important factor in making effective treatment decisions. The location of the infection can also affect its significance and available treatment options. Some microorganism-related diseases can occur in the absence of colonization that is considered typical of colonization. For example, ingestion of Clostridium botulinum may be sufficient to cause symptoms.

[0005] Furthermore, infections in the asymptomatic stage are often asymptomatic or may present with nonspecific symptoms that resemble a variety of other diseases.Such infections often go undiagnosed, are misdiagnosed, or are treated symptomatically, allowing the microbe to persist and increasing the risk that the infection will progress to an invasive disease.

[0006] Helicobacter pylori (H. pylori) is the most common chronic bacterial infection in humans. It is estimated that 50 percent of the world’s population is infected. In the United States, about 30 percent of adults are infected before the age of 50, while most individuals are infected in childhood. Chen, Y. and M.J. Blaser, Journal of Infectious Diseases, 2008. 198(4): p.553-60. Helicobacter pylori is strongly associated with gastrointestinal (GI) conditions including chronic gastritis, peptic ulcer disease, gastric adenocarcinoma, and lymphoma. Peptic ulcer disease (PUD) is the most common manifestation of Helicobacter pylori infection and the annual incidence of physician-diagnosed PUD is 0.1-0.19 percent. Sung, JJ, EJ Kuipers and instruction manual 1 / 68 page 4 CN 121575489 A HB El-Serag, Aliment Pharmacol Ther, 2009. 29(9): p.938-46. It is estimated that infected individuals have a lifetime risk of developing peptic ulcer disease of 10%-20%. Kuipers, EJ et al, Aliment Pharmacol Ther, 1995. 9 Supplement 2: p.59-69.

[0007] The main phenomenon responsible for these disease manifestations is mucosal inflammation in response to the presence of Helicobacter pylori. However, only a small percentage of individuals with Helicobacter pylori will have inflammation associated with invasive Helicobacter pylori.

[0008] Currently, it is challenging to distinguish between patients with the asymptomatic stage of Helicobacter pylori infection and those with the symptomatic stage or at risk of progressing to the symptomatic stage. Although most Helicobacter pylori infections are asymptomatic, patients with invasive disease may begin to experience persistent dyspepsia symptoms such as abdominal pain, nausea or vomiting, and loss of appetite. However, these nonspecific symptoms can also be caused by other conditions and are experienced by healthy individuals. Some physicians test all patients with unexplained persistent dyspepsia. Other physicians follow current guidelines that recommend testing individuals with active pyorrhea, a documented history of peptic ulcers, or gastric MALT lymphoma. Chey, W.D. et al., *American Journal of Gastroenterology*, 2007, 102(8):pp. 1808-25. Therefore, physicians following guidelines will only test patients with a probability of having Helicobacter pylori-related disease, which may lead to undertreatment.

[0009] Several methods exist for testing for Helicobacter pylori. Existing non-invasive testing methods for Helicobacter pylori include fecal antigen testing, urea breath testing, and Helicobacter pylori serology. However, these methods may only determine the presence of Helicobacter pylori, but not whether Helicobacter pylori invasion or associated inflammation is present. Some practitioners will initiate primary treatment for eradication based on a positive result from one of the non-invasive tests, which may lead to overtreatment.

[0010] The current gold standard for diagnosing Helicobacter pylori disease is to perform upper endoscopy to document (by biopsy) specific pathological changes caused by Helicobacter pylori invasion, such as inflammation, atrophy, and intestinal metaplasia, combined with detection of Helicobacter pylori in the biopsy sample. Dixon, MF et al., Helicobacter pylori, 1997. 2 Supplement 1: pp. S17-24. However, there are serious risks and potential complications from this procedure, including bleeding that may sometimes require blood transfusions, infection, and GI tract tears.

[0011] Overall, approximately 75% of patients following primary treatment for Helicobacter pylori infection are considered cured after the first treatment based on a negative diagnostic test for active Helicobacter pylori infection that was previously positive before treatment initiation. If a diagnostic test for active gastrointestinal Helicobacter pylori infection remains positive after completion of first-line therapy, there is a possibility of antibiotic-resistant Helicobacter pylori, and further treatment will be required until a negative diagnostic test result is obtained.

[0012] Next-generation sequencing (NGS) can be used to aggregate large amounts of data about the nucleic acid content of a sample. This data can be particularly useful for analyzing nucleic acids in complex samples, such as clinical samples. However, before using NGS methods, the starting sample must typically be processed, which can reduce nucleic acid recovery rates, delay sequencing, delay reporting of clinical calls, introduce errors, introduce biases, and often result in chemical waste requiring controlled disposal. Errors and biases can affect results in many cases, such as when low abundance of nucleic acids or target nucleic acids are present in patient samples. Current NGS methods focus on the abundance or relative abundance of specific reads or sequences. Furthermore, many sequencing library preparation methods and some next-generation sequencing systems have produced experimentally observed deviations in target nucleic acid fragment length and fragment length distribution from the endogenous fragment length and fragment length distribution, particularly through methods such as using variable polyA tail tags, unspecified polyA tail tags, thermal inactivation of enzymes, the use of biased extraction methods, or the introduction of nucleic acid length, secondary structure, and / or GC biases across the entire or partial range of target nucleic acid length and GC content.Other methods and systems for measurement. Some of these methods and systems prevent successful correction of bias even in the presence of process control molecules, provided the bias is large enough that insufficient target nucleic acids and / or process control molecules are recovered across the entire or certain segments of the relevant length and GC content for final analysis.

[0013] Various methods incorporating NGS have been used to identify microorganisms present in a host, but most of these methods focus on the abundance of microbial reads rather than the physical properties of the molecules being read. For example, many extraction protocols, library generation protocols, and sequencing protocols include steps or processes designed to remove short nucleic acid fragment lengths. Short nucleic acid fragment lengths are often also sacrificed to minimize unwanted or incomplete byproducts of extraction, library generation, or amplification, such as primer dimers or adaptor dimers. Free nucleic acids of microorganisms are examples of target nucleic acids that are particularly susceptible to bias and depletion due to their fragment lengths being less than about 100 bp.

[0014] Current methods for differentiating between infections in the asymptomatic or latent stage and infections in other stages after the identification of the underlying pathogen sometimes require invasive biopsy procedures. Non-invasive tests, such as serology, can detect markers of exposure to microorganisms but cannot indicate whether an infection is active or at risk of progressing to an invasive disease. Therefore, there is a need for precise non-invasive methods to determine whether a patient's organs are infected and to differentiate which patients will remain in the colonization stage and which patients are at risk of developing secondary invasive diseases. This disclosure provides non-invasive methods, compositions, and kits for detecting infection in a subject and determining whether the infection is in a colonization or invasive disease stage. This disclosure also provides non-invasive methods for determining the location of the infection site in a subject and / or the stage of infection in the subject. Summary of the Invention

[0015] Embodiments of this application provide a fragment length profile from a nucleic acid library, wherein nucleic acids for preparing the nucleic acid library are obtained from a sample by an unbiased method, a method that enables bias correction, or a method with reproducible bias. In various aspects, the nucleic acid library is generated from the initial sample, and the nucleic acids used to generate the nucleic acid library are not extracted from the initial sample before preparing the nucleic acid library or before initiating the library generation process. Aspects of the method may include nucleic acid sequencing as a step after nucleic acid preparation and before determining the fragment length profile of the target nucleic acid, multiple target nucleic acids, or a subset of nucleic acids within the nucleic acid library. In various aspects of the embodiments, the fragment length profile includes one or more characteristics selected from the group consisting of: shape of distribution, segment amplitude, segment fraction, peak shape, number of peaks, position of the largest peak, fragment count ratio of two or more segments, height of helical phasing peaks, and two different segments.The fragment count ratio at a segment length, the fragment count ratio within two different fragment length ranges, the number of fragments within a segment, the fragment length range within a segment, the ratio of the maximum amplitude of two or more segments, and the fragment length distribution within a subset of reads, the slope within a segment, the peak width, the rate of count decay or increase within a segment, the number of peaks, and the scaling factor of count decay or increase within a segment.

[0016] A method for generating a fragment length spectrum of a nucleic acid library is provided. Various methods include the following steps: preparing a nucleic acid library from an initial sample using a bias-corrected recovery method or a method with reproducible bias; determining the number of reads or normalized counts of multiple fragment lengths within the nucleic acid library; determining one or more fragment length characteristics of the nucleic acid library; and generating a fragment length spectrum of the nucleic acid library using one or more fragment length characteristics. In various aspects of the embodiments, the fragment length spectrum includes one or more fragment length characteristics selected from the group comprising: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, position of the largest peak among peaks, number of peaks, and fragment length distribution within a subset of reads. Methods for generating fragment length spectra of nucleic acid libraries are provided. Various methods include the step of preparing a nucleic acid library from an initial sample, the step comprising the steps of: optionally adding one or more process control molecules to the initial sample to provide a spiked initial sample, and generating a nucleic acid library from the spiked initial sample, wherein optionally, nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library. Aspects of the methods may include nucleic acid sequencing as a step after nucleic acid preparation and before determining the fragment length spectrum. The method for generating a fragment length spectrum of a target nucleic acid within the nucleic acid library further includes the following steps: determining the number of reads of multiple fragment lengths within the nucleic acid library, determining one or more fragment length characteristics of the nucleic acid library, and generating a fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics. In various aspects of the embodiments, the fragment length spectrum includes one or more fragment length characteristics selected from the group consisting of: shape of distribution, segment amplitude, peak shape, number of peaks, position of the largest peak, fragment count ratio of two or more segments, height of helical phasing peaks, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of the largest amplitude of two or more segments, and fragment length distribution within a subset of reads. In some aspects,The step of generating the nucleic acid library from the initial sample further includes, constitutes, or substantially comprises the following steps: dephosphorylating the nucleic acid from the initial sample to produce a set of dephosphorylated nucleic acids; denaturing the dephosphorylated nucleic acids to produce denatured nucleic acids; ligating a 3' end adaptor to the denatured nucleic acids to produce adapted nucleic acids; isolating the adapted nucleic acids; attaching primers to the adapted nucleic acids; and extending the primers with polymerase to generate complementary strands; ligating a 5' end adaptor; eluting the strands; and amplifying the complementary strands. Aspects of the method may include nucleic acid sequencing as a step after nucleic acid preparation and before determining the fragment length profile. In various embodiments, the number of reads is a normalized number of reads. In some embodiments, the fragment length profile is for at least one subset of reads in the nucleic acid library. In such embodiments, the method further includes the steps of: identifying at least one subset of reads within the nucleic acid library and determining the fragment length distribution within each selected subset of reads. In some embodiments, the step of generating the fragment length profile further includes using two or more fragment length characteristics.

[0017] A method for identifying microorganisms present in a sample is provided. The method for identifying or characterizing microorganisms present in a sample includes the steps of: generating a fragment length spectrum for sequencing reads from a nucleic acid library generated from the sample and aligned to a microbial reference sequence; comparing the fragment length spectrum with a reference fragment length spectrum of one or more microorganisms; and identifying the microorganism as present in the sample if the fragment length spectrum from the sample is similar to the reference fragment length spectrum of the microorganism. Aspects of the method include comparing fragment length spectra of target sequences from the nucleic acid library. In various embodiments, the fragment length spectrum may indicate that the microorganism is present in the form of a pathogen or a symbiotic microorganism. In aspects of the method, generating the fragment length spectrum of the nucleic acid library includes the steps of: preparing a nucleic acid library from an initial sample; quantifying the number of reads of multiple fragment lengths within the nucleic acid library; determining one or more fragment length characteristics of the nucleic acid library or at least a subset of the nucleic acid library; and using the one or more fragment length characteristics to generate a fragment length spectrum of at least a subset of the nucleic acid library or reads. The step of preparing a nucleic acid library from an initial sample further includes the following steps: adding one or more process control molecules to the initial sample to provide a spiked initial sample, and generating a nucleic acid library from the spiked initial sample, wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library. Aspects of the method may include nucleic acid sequencing as a step after nucleic acid preparation and before determining fragment length profiles. In various aspects of the embodiments, the...The fragment length spectrum includes one or more fragment length characteristics selected from the group consisting of: the shape of the distribution, segment amplitude, peak shape, number of peaks, position of the largest peak among the peaks, fragment count ratio of two or more segments, height of the spiral phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of the largest amplitude of two or more segments, and fragment length distribution within a subset of reads. In various aspects of the method, the fragment length spectrum includes at least one fragment length characteristic selected from the group consisting of: fragment count ratio of two or more segments, peak shape, peak width, rate of count decay or increase within a segment, number of peaks, scaling factor of count decay or increase within a segment, and position of the largest peak among the peaks.

[0018] A method for determining localization sites within a subject is provided. The method includes the following steps: generating a fragment length spectrum of the target nucleic acid in a nucleic acid library generated from the sample or the entire nucleic acid library; comparing the fragment length spectrum with a reference fragment length spectrum of one or more source sites; and predicting the first site as a localization site if the fragment length spectrum in the sample is similar to the fragment length spectrum of a first source site, and predicting the second site as a localization site if the fragment length spectrum in the sample is similar to the fragment length spectrum from a second source site. In an embodiment of the method, generating one or more fragment length spectra of the nucleic acid library includes the following steps: preparing a nucleic acid library from an initial sample; quantifying the number of reads of multiple fragment lengths within the nucleic acid library; and generating a fragment length spectrum of the target nucleic acid in the nucleic acid library or the entire nucleic acid library using one or more fragment length characteristics. In an embodiment of the method, preparing a nucleic acid library from an initial sample further includes the following steps: adding one or more process control molecules to the initial sample to provide a spiked initial sample; and generating a nucleic acid library from the spiked initial sample, wherein the nucleic acid used to generate the nucleic acid library is not extracted from the initial sample before preparing the nucleic acid library. Aspects of the method may include nucleic acid sequencing as a step after nucleic acid preparation and before determining the fragment length profile. In aspects of the embodiments, the fragment length profile includes one or more fragment length characteristics selected from the group consisting of: shape of distribution, segment amplitude, peak shape, number of peaks, position of the largest peak among the peaks, fragment count ratio of two or more segments, height of the helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, and two or more fragment length characteristics.The ratio of the maximum amplitude of more segments, peak width, rate of count decay or increase within segments, number of peaks, scaling factor of count decay or increase within segments, and fragment length distribution within subsets of reads. In various aspects of the method, the location site is selected from the group consisting of, or substantially consisting of, source sites including: deep tissues, lungs, liver, bones, kidneys, brain, heart, sinuses, GI tract, spleen, skin, joints, ears, nose, mouth, blood flow, and blood.

[0019] A method for monitoring graft status in a subject is provided. The method for monitoring graft status includes the steps of: generating a baseline fragment length spectrum from a nucleic acid library from a sample obtained from the subject; generating a second fragment length spectrum from a second sample obtained from the subject; and comparing the second fragment length spectrum with the baseline fragment length spectrum. If the second fragment length spectrum differs from the baseline fragment length spectrum, an increased amount of anti-rejection therapy is administered to the subject, wherein the risk of rejection in the subject with the graft is reduced after the administration of the anti-rejection therapy. If the second fragment length spectrum is similar to the baseline fragment length spectrum, anti-rejection therapy is maintained or reduced, wherein the risk of side effects from the anti-rejection therapy in the subject is lower than the risk of side effects from the subject receiving an increased dose of the anti-rejection therapy. Aspects of the method include the steps of: comparing fragment length spectra of target nucleic acids in a nucleic acid library or the entire library obtained from a sample of a subject having a graft, and comparing the spectra with a reference fragment length spectrum.

[0020] A method for monitoring the toxicity of a compound administered to a subject is provided. The method includes the steps of: generating fragment length spectra of a nucleic acid library or target nucleic acid in a nucleic acid library prepared from a sample obtained from the subject, and comparing the fragment length spectra with one or more reference fragment length spectra. In aspects of the method, the subject has cancer, is at risk of developing cancer, or exhibits cancer-related symptoms. In aspects of the method, the one or more reference fragment length spectra are generated from a nucleic acid library obtained from a subject or cells exposed to the compound. In aspects of the method, the one or more reference fragment length spectra include a baseline fragment length spectrum. In all aspects of the method described on page 5 / 68 of CN 121575489 A, the compound is a chemotherapeutic agent. In embodiments of the method, the step of generating a fragment length spectrum of a nucleic acid library includes the following steps: preparing a nucleic acid library from an initial sample using a bias-corrected recovery method; determining the number of reads of multiple fragment lengths within the nucleic acid library; determining one or more fragment length characteristics of the nucleic acid library; and generating a fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics. All aspects of the method...The method may include nucleic acid sequencing as a step after nucleic acid preparation and before determining the fragment length profile. In various aspects of the embodiments, the fragment length profile includes one or more fragment length characteristics selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads. In embodiments of the method, generating the fragment length profile of the nucleic acid library includes the step of preparing a nucleic acid library from an initial sample, the step further including: adding one or more process control molecules to the initial sample to provide a spiked initial sample, and generating a nucleic acid library from the spiked initial sample, wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample before preparing the nucleic acid library; quantifying the number of reads of multiple fragment lengths within the nucleic acid library; determining one or more fragment length characteristics of the nucleic acid library; and generating the fragment length profile of the nucleic acid library using the one or more fragment length characteristics. In various aspects of the embodiments, the fragment length spectrum includes one or more fragment length characteristics selected from the group consisting of: the shape of the distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of the helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of the maximum amplitude of two or more segments, and fragment length distribution within a subset of reads.

[0021] The present invention relates to a method for predicting the risk of an organism (or multiple organisms) present in a host causing local or systemic environmental alterations or invading an organ or anatomical system with substantial negative health consequences. An organism is considered invasive if it crosses a barrier or translocates from one organ or anatomical structure to another, invading a structure beyond the tissue layer it occupies in its colonizing state to produce local invasion, altering the environment of the structure to have a significant negative impact on the structure or causing DNA mutations or inflammation, or otherwise overwhelming the host's immune system.

[0022] In some embodiments, the risk level is based on the abundance of the organism in the host compared to an asymptomatic or infected control. In other embodiments, abundance is a threshold or range. In other embodiments, the risk level is calculated as a clinical decision score based on one or more of the following: bioavailability, patient's clinical history, disease chronicity, genetic biomarker factors and patient characteristics (such as age, sex, etc.), fragment length distribution profile, and fragment length distribution profile characteristics.

[0023] In one aspect, a method is provided for determining the infection stage of a subject suspected of having a microbial infection, the method comprising:(a) Performing high-throughput sequencing on nucleic acids from the biological sample; (b) Performing bioinformatics analysis to identify microbial nucleic acid sequences present in the biological sample; and (c) Calculating the measurement results of the nucleic acids and comparing the measurement results with a control, thereby determining the infection stage of any microorganism identified in the biological sample.

[0024] In some embodiments, the method further includes one or more steps selected from the group consisting of: (a) extracting nucleic acids from a portion of the biological sample obtained from the subject, and (b) adding a synthetic nucleic acid spike.

[0025] In one embodiment, the measurement results of step (c) are selected from the absolute abundance of free microbial nucleic acid sequences, the distribution of fragment lengths of nucleic acid sequences, the characteristics of nucleic acid fragment length distribution profiles, or a combination thereof. In another embodiment, the measurement of step (c) is, if it is, the absolute abundance and fragment length distribution of a target pathogen.

[0026] In a second embodiment, the subject has infection symptoms or is at risk of infection.

[0027] In a third embodiment, the infection stage is an asymptomatic period, a symptomatic infection stage, a treatment stage, or an eradication stage. In a fourth embodiment, the method further includes repeating the method over time to monitor infection, infection stage, the efficacy of treatment for the infection, or to detect the onset of infection. In various aspects, the method may further include changing the treatment regimen.

[0028] In a fifth embodiment, the method further includes administering a treatment regimen to the subject based on a determined infection stage.

[0029] In a sixth embodiment, the high-throughput sequencing assay is next-generation sequencing, massively parallel sequencing, pyrosequencing, successive synthesis sequencing, single-molecule real-time sequencing, polymerase chain cloning sequencing, DNA nanosphere sequencing, helicopter single-molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.

[0030] In a seventh embodiment, the sample is blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, sputum, urine, feces, saliva, or a nasal sample.

[0031] In an eighth embodiment, the method further includes identifying one or more antibiotic resistance genes of the target pathogen.

[0032] In a ninth embodiment, the method further includes identifying at least one risk factor in the subject's genomic DNA.

[0033] In a tenth embodiment, the nucleic acid is cell-free DNA and / or cell-free RNA. The nucleic acid may include cell-free pathogen DNA. The nucleic acid may include cell-free pathogen RNA. The nucleic acid may include cell-free microbial DNA. The nucleic acid may include cell-free microbial RNA.

[0034] In an eleventh embodiment, the target pathogen is Helicobacter pylori, Clostridium difficile, or Haemophilus influenzae.(haemophilus influenza), Salmonella, Streptococcus pneumoniae, cytomegalovirus, hepatitis B virus, hepatitis C virus, human papillomavirus, Epstein-Barr virus, human T-cell lymphoma virus I, Merkel cell polyomavirus, Kaposi's sarcoma virus, human herpesvirus 8, chlamydia, gonorrhea, syphilis, or trichomoniasis.

[0035] In the twelfth embodiment, the subject had previously undergone another test or other clinical test. In one embodiment, other clinical tests are fecal antigen test, urea breath test, serology, urease test, histology, bacterial culture and susceptibility test, biopsy, or endoscopy.

[0036] In the thirteenth embodiment, the target pathogen nucleic acid is DNA and / or RNA. The pathogen nucleic acid includes cell-free DNA. Nucleic acids include pathogen-free RNA.

[0037] In a fourteenth embodiment, the synthetic nucleic acid spike comprises at least 1,000 unique synthetic nucleic acids of a sample, each of the 1,000 unique synthetic nucleic acids comprising (i) an identification tag; and (ii) a variable region containing at least 5 degenerated bases. In another embodiment, the method further comprises (a) optionally extracting nucleic acids from the spiked sample; (b) generating a spiked sample library; (c) optionally enriching the spiked sample library; (d) performing high-throughput sequencing assays to obtain sequence reads from the spiked sample library; (e) calculating the diversity loss values ​​of the 1,000 unique synthetic nucleic acids; and (f) calculating the measurement results of the nucleic acids and comparing the measurement results with a control, thereby determining the infection stage of the subject. In another embodiment, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in US 9,976,181.

[0039] In another aspect, there is a method for determining the stage of Helicobacter pylori infection in a subject, the method comprising: a) optionally, extracting free nucleic acids from a biological sample obtained from the subject; b) adding a synthetic nucleic acid spike to the sample; c) performing high-throughput sequencing on the nucleic acids from the biological sample; d) performing bioinformatics analysis to identify Helicobacter pylori nucleic acid sequences present in the biological sample; ande) Calculate the measurement results of the Helicobacter pylori nucleic acid and compare the measurement results with a control to determine the Helicobacter pylori infection stage of the subject.

[0040] In a first embodiment, the measurement results are the absolute abundance of Helicobacter pylori or the distribution of fragment lengths, or a combination thereof. In one embodiment, the measurement results are the absolute abundance of Helicobacter pylori. In another embodiment, the measurement results are the distribution of fragment lengths of Helicobacter pylori. In yet another embodiment, the measurement results are the distribution of both the absolute abundance and fragment lengths of Helicobacter pylori. In various embodiments, the steps of the method may be performed in a varying order.

[0041] In a second embodiment, the subject has symptoms of Helicobacter pylori infection or is at risk of Helicobacter pylori infection. In one embodiment, the infection stage is an asymptomatic period, a symptomatic infection period, a treatment period, or an eradication period.

[0042] In a third embodiment, the method further includes repeating the method over time to monitor infection and the efficacy of treatment for the infection.

[0043] In one aspect, there is a method for determining the stage of Helicobacter pylori infection in a subject, the method comprising: (a) preparing a spiked sample by obtaining a sample comprising free nucleic acid from the subject and adding one or more process control molecules; (b) optionally, extracting the nucleic acid from the spiked sample; (c) generating a spiked sample library, wherein the generation comprises (i) linking an adaptor to the nucleic acid; and (ii) amplification; (d) optionally, enriching the spiked sample library; (e) performing high-throughput sequencing assays to obtain sequence reads from the spiked sample library; (f) calculating the diversity loss value of 1,000 unique synthetic nucleic acids; and (g) calculating the measurement result of the free nucleic acid and comparing the measurement result with a control, thereby determining the stage of Helicobacter pylori infection in the subject.

[0044] In yet another embodiment, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in US 9,976,181.

[0045] In the second embodiment, the high-throughput sequencing assay is next-generation sequencing, massively parallel sequencing, pyrosequencing, successive synthesis sequencing, single-molecule real-time sequencing, polymerase cloning sequencing, DNA nanosphere sequencing, helicopter single-molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.

[0046] In the third embodiment, the sample is blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, feces, saliva, or a nasal sample.

[0047] In the fourth embodiment, the method further includes administering a treatment regimen to the subject, wherein the treatment can be administered at any stage of the infection cycle.

[0048] In a fifth embodiment, the method further includes identifying one or more antibiotic resistance genes of the target pathogen.

[0049] In a sixth embodiment, the free nucleic acid is DNA and / or RNA. The nucleic acid includes free pathogen DNA. The nucleic acid includes free pathogen RNA.

[0050] In a twelfth embodiment, the subject has previously undergone another clinical test. In one embodiment, the other clinical test is a fecal antigen test, a urea breath test, serological tests, a urease test, histological tests, bacterial culture and susceptibility tests, biopsy, or endoscopy.

[0051] In an eighth embodiment, the target pathogen nucleic acid is DNA and / or RNA. The pathogen nucleic acid includes free DNA. The nucleic acid includes free pathogen RNA. The target pathogen nucleic acid includes a mixture of free DNA and free RNA.

[0052] On the other hand, a method is provided for determining localization sites in a subject infected with a pathogen, the method comprising: (a) obtaining a sample comprising nucleic acid from the subject, and adding one or more process control molecules to generate a spiked sample; (b) optionally, extracting the nucleic acid from the spiked sample; (c) generating a library from the spiked sample, wherein the generation includes linking an adaptor to the nucleic acid and amplifying it; (d) optionally, enriching the spiked sample; (e) performing high-throughput sequencing assays by comparing a reference genome to obtain sequence reads from the spiked sample; (f) optionally, calculating a diversity loss value; and (g) calculating a measurement result of the nucleic acid and comparing the measurement result with a control, thereby determining the localization site of the subject.

[0053] In a first embodiment, the measurement result is the absolute abundance of the target pathogen or the distribution of fragment lengths, or a combination thereof. In one embodiment, the measurement result is the absolute abundance of the target pathogen. In another embodiment, the measurement result is the distribution of fragment lengths of the target pathogen. In yet another embodiment, the measurement results are the distribution of the absolute abundance and fragment length of the target pathogen.

[0054] In a second embodiment, the localization site is a tissue. In another embodiment, the localization site is a tissue type. In yet another embodiment, the localization site is an organ. In yet another embodiment, the localization site is a tissue type including an organ.

[0055] In a third embodiment, the subject has symptoms of infection or is at risk of infection. In another embodiment, the subject has previously been identified as infected with Helicobacter pylori, Clostridium difficile, Haemophilus influenzae, Salmonella, Streptococcus pneumoniae, cytomegalovirus, hepatitis virus B, hepatitis virus C, human papillomavirus, Epstein-Barr virus, human T-cell lymphoma virus I, Merkel cell polyomavirus, Kaposi's sarcoma virus, human herpesvirus 8, Chlamydia virus, herpes simplex virus, Neisseria, Treponema, or Trichomonas.

[0056] In a fourth embodiment, the method is repeated over time to monitor infection and the efficacy of treatments administered to the infection.

[0057] In a fifth embodiment, the method further includes administering a treatment regimen to the subject based on a determined stage of infection.

[0058] In a sixth embodiment, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in US 9,976,181, page 9 / 68, CN 121575489 A.

[0059] In a seventh embodiment, the high-throughput sequencing assay is next-generation sequencing, massively parallel sequencing, pyrosequencing, successive synthesis sequencing, single-molecule real-time sequencing, polymerase cloning sequencing, DNA nanosphere sequencing, helicopter single-molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.

[0060] In an eighth embodiment, the sample is blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, feces, saliva, nasal, or tissue samples.

[0061] In a ninth embodiment, the method further includes identifying one or more antibiotic resistance genes of the pathogen.

[0062] In a tenth embodiment, the method further includes identifying risk factors in the subject's genomic DNA.

[0063] In an eleventh embodiment, the target pathogen nucleic acid is DNA and / or RNA. The pathogen nucleic acid includes cell-free DNA. The nucleic acid includes cell-free pathogen RNA. The target pathogen nucleic acid includes a mixture of cell-free DNA and cell-free RNA.

[0064] In a twelfth embodiment, the free nucleic acid is DNA and / or RNA. The nucleic acid includes cell-free pathogen DNA. The nucleic acid includes cell-free RNA. The nucleic acid includes cell-free pathogen RNA. The nucleic acid includes cell-free subject RNA. The nucleic acid includes pathogen and subject cell-free RNA.

[0065] In one aspect, a method is provided for determining the stage of infection in a subject suspected of having a microbial infection, the method comprising: (a) providing a sample comprising nucleic acids from the subject; (b) adding at least 1,000 unique synthetic nucleic acids to the sample to generate a spiked sample; (c) generating a library from the spiked sample; (d) performing high-throughput sequencing to obtain sequence reads from the spiked sample; and (e) determining the stage of infection in the subject based on the sequence reads.

[0066] In one embodiment, the sample is selected from blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, feces, saliva, nasal, and tissue samples. The sample is blood, plasma, serum, cerebrospinal fluid, or synovial fluid.

[0067] In yet another embodiment, the at least 1,000 unique synthetic nucleic acids are synthetic nucleic acids as described in US 9,976,181.

[0068] In another embodiment, high-throughput sequencing is determined by next-generation sequencing, massively parallel sequencing, and pyrosequencing.Sequencing, sequential synthesis sequencing, single-molecule real-time sequencing, polymerase cloning sequencing, DNA nanosphere sequencing, helicopter single-molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, or Gilbert sequencing.

[0069] In another embodiment, the determination of the infection stage is based on the absolute abundance of the target pathogen or a fragment length distribution spectrum, or a combination thereof. In one embodiment, the determination is based on the absolute abundance of the target pathogen. In another embodiment, the determination is based on the distribution of fragment lengths of the target pathogen. In yet another embodiment, the determination is based on both the absolute abundance of the target pathogen and the distribution of fragment lengths.

[0070] One aspect of this application provides a method for determining the infection stage of a subject. The method includes the steps of: generating a fragment length spectrum of a nucleic acid library obtained from a sample of the subject; comparing the fragment length spectrum with a reference fragment length spectrum; and determining that if the fragment length spectrum from the sample is similar to a fragment length spectrum from a symptomatic subject, the infection stage indicates an increased risk of the subject exhibiting microbiome-related symptoms, and if the fragment length spectrum from the sample is similar to a fragment length spectrum from an asymptomatic subject, the infection is in an asymptomatic stage. In one aspect, the fragment length spectrum is a fragment length spectrum of a non-microbial host nucleic acid library. In another aspect, the method further includes the steps of: determining the abundance of at least one significant microorganism in a sample from the subject; comparing the abundance to a threshold; and comparing the fragment length spectrum to a reference fragment length spectrum (see specification 10 / 68, page 13, CN 121575489 A). If the fragment length spectrum from the sample is similar to the fragment length spectrum from a symptomatic subject, and the abundance is equal to or higher than the threshold, an infection stage is determined indicating an increased risk of the subject exhibiting microbial-related symptoms. If the fragment length spectrum from the sample is similar to the fragment length spectrum from an asymptomatic subject, the infection is determined to be in an asymptomatic stage. In another aspect, the method further includes the step of administering an antimicrobial agent to a subject determined to have an increased risk of exhibiting microbial-related symptoms.

[0071] A method for determining the stage of infection in a subject suspected of having a microbial infection, the method comprising performing high-throughput sequencing on nucleic acids from a biological sample, performing bioinformatics analysis to identify nucleic acid sequences present in the biological sample, calculating a measurement result of the nucleic acids, and comparing the measurement result with a control to determine the stage of infection of the microorganism identified in the biological sample. The method may further include one or more steps selected from the group consisting of: (i) extracting nucleic acids from a biological sample obtained from the subject, and (ii) adding a synthetic nucleic acid spike to the biological sample obtained from the subject. In one aspect, the nucleic acids include microbial nucleic acids, host nuclei, etc.The nucleic acid may be either a microbial nucleic acid or a host nucleic acid. In one aspect, the nucleic acid includes free microbial nucleic acid, host nucleic acid, or both microbial and host nucleic acids. In another aspect, the measurement result is selected from the group consisting of the absolute abundance of nucleic acid, the fragment length distribution spectrum of nucleic acid, and a combination of both absolute abundance and fragment length distribution spectrum. In one aspect, the infection stage is selected from the asymptomatic stage, colonization stage, symptomatic stage, active stage, invasive disease stage, regression stage, treatment period, or eradication stage of infection. In one aspect, the method further includes administering a treatment regimen to the subject based on the determined infection stage. The method may further include repeating the method over time to monitor the infection or the efficacy of treatment administered to the infection. In some embodiments, the microorganism is selected from the group consisting of: Helicobacter pylori, Clostridium difficile, Haemophilus influenzae, Salmonella, Streptococcus pneumoniae, cytomegalovirus, hepatitis virus B, hepatitis virus C, human papillomavirus, Epstein-Barr virus, human T-cell lymphoma virus 1, Merkel cell polyomavirus, Kaposi's sarcoma virus, human herpesvirus 8, Chlamydia virus, herpes simplex virus, Neisseria, Treponema, or Trichomonas. In various aspects, adding a synthetic nucleic acid spike further includes preparing a spiked sample by obtaining a sample comprising free nucleic acid from the subject and adding one or more process control molecules; extracting nucleic acid from the spiked sample; generating a spiked sample library; enriching the spiked sample library; performing high-throughput sequencing to obtain sequence reads from the spiked sample library; calculating the diversity loss value of 1,000 unique synthetic nucleic acids; and calculating the measurement results of the free nucleic acid and comparing the measurement results with a control, thereby determining the infection stage of the subject.

[0072] In one embodiment, the present application provides a method for determining the stage of Helicobacter pylori infection in a subject, the method comprising extracting nucleic acid from a biological sample obtained from the subject, adding a synthetic nucleic acid spike to the sample, performing high-throughput sequencing on the nucleic acid from the biological sample, performing bioinformatics analysis to identify free Helicobacter pylori nucleic acid sequences present in the biological sample, calculating a measurement result of the free Helicobacter pylori nucleic acid, and comparing the measurement result with a control, thereby determining the stage of Helicobacter pylori infection in the subject.

[0073] In one embodiment, the present application provides a method for determining the stage of Helicobacter pylori infection in a subject, the method comprising: preparing a spiked sample by obtaining a sample comprising free nucleic acid from the subject and adding one or more process control molecules; extracting nucleic acid from the spiked sample; generating a spiked sample library, wherein the generation includes (i) linking an adaptor to the nucleic acid; and (ii) amplification; optionally, enriching the spiked sample library; performing high-throughput sequencing on the nucleic acid from the biological sample; performing bioinformatics analysis to identify free Helicobacter pylori nucleic acid sequences present in the biological sample, and calculating a measurement result of the free Helicobacter pylori nucleic acid; and comparing the measurement result with a control to determine the stage of Helicobacter pylori infection in the subject.Throughput sequencing was performed to obtain sequence reads from the spiked sample library; the diversity loss value of 1,000 unique synthetic nucleic acids was calculated; and the measurement results of the free nucleic acids were calculated and compared with the measurement results to a control, thereby determining the stage of Helicobacter pylori infection in the subject.

[0074] One embodiment provides a method for determining localization sites in a subject infected with a pathogen, the method specification 11 / 68 pages 14 CN 121575489 A comprising obtaining a sample comprising nucleic acid from the subject, adding one or more process control molecules to an initial sample to provide a spiked sample, optionally extracting the nucleic acid from the spiked sample, generating a library from the spiked sample, wherein the generation includes linking an adaptor to the nucleic acid and amplifying it; optionally, enriching the spiked sample, performing high-throughput sequencing assays by comparison with a reference genome to obtain sequence reads from the spiked sample; determining one or more fragment length characteristics of the nucleic acid library, generating a fragment length spectrum of the nucleic acid library generated from the sample, comparing the fragment length spectrum with a reference fragment length spectrum of one or more source sites, and identifying the first site as a localization site if the fragment length spectrum from the sample is similar to the fragment length spectrum from a first source site; and identifying the second site as a localization site if the fragment length spectrum from the sample is similar to the fragment length spectrum from a second source site.

[0075] One aspect provides a method for determining a localization site in a subject infected with a pathogen, the method comprising obtaining a sample comprising cell-free nucleic acid from the subject, and adding one or more process control molecules to generate a spiked sample; optionally extracting nucleic acid from the spiked sample; generating a library from the spiked sample, wherein the library includes linking an adaptor to the nucleic acid and amplifying it; optionally, enriching the spiked sample; performing high-throughput sequencing assays by comparison with a reference genome to obtain sequencing reads from the spiked sample; calculating a diversity loss value for 1000 unique synthetic nucleic acids; and calculating a measurement result of the cell-free nucleic acid and comparing the measurement result with a control, thereby determining the localization site of the subject.

[0076] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference in their entirety, as if each individual publication, patent or patent application were explicitly and individually designated as incorporated by reference.

[0077] The novel features of the invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of illustrative embodiments that utilize the principles of the invention, and in the accompanying drawings: FIG1 depicts the method of the present disclosure.

[0078] FIG2 depicts the cell-free method of the present disclosure.

[0079] Figure 3 illustrates a schematic diagram of an exemplary infection.

[0080] Figure 4 depicts one of the infection site detection methods of this disclosure.

[0081] Figure 5 depicts a basic scheme for a method for determining the diversity loss value.

[0082] Figure 6 illustrates a diagnostic workflow terminating in treatment for a positive diagnosis of Helicobacter pylori.

[0083] Figure 7 depicts a computer control system programmed or otherwise configured to implement the methods provided herein.

[0084] Figure 8 depicts the distribution of fragment lengths from reads of three microorganisms detected in three different human plasma samples generated from nucleic acid libraries. The fragment length characteristic of interest in the figures is the distribution shape. Each figure provides an example of a different distribution shape. In each figure, the normalized number of reads is shown on the y-axis, and the x-axis indicates the fragment length. The left figure provides an example of a “50-base-pair peak” distribution shape. The middle figure provides an example of a short class exponential distribution shape. The right figure provides an example of a complex distribution shape, where this particular complex distribution shape includes aspects of an exponentially decaying distribution shape and a single peak of 50 base pairs. It is generally accepted that each distribution shape depicted reflects the distribution of fragment lengths in nucleic acid libraries generated from different human plasma samples and provides an example of the indicated distribution shape type. Other distribution shapes are described elsewhere in this document. Other distribution shapes are also possible. Specification 12 / 68 pages 15 CN 121575489 A

[0085] Figure 9 provides an example of fragment length characteristics involving distribution segment amplitude and segment amplitude ratio. The figure depicts the distribution of fragment lengths of the same pathogen (Candida tropicalis) from three different clinical samples. In each figure, the normalized number of reads is shown on the y-axis, and the x-axis indicates the fragment length. For the purposes of this figure, the clinical samples are numbered 1 to 3. Compared to *Candida tropicalis* in clinical sample 3, *Candida tropicalis* in clinical samples 1 and 2 showed a distribution with a higher long fraction (>65 bp) relative to the 50 bp peak, while all fragment length spectra had clear peaks of approximately 45–50 bp. The ratio of short reads (<40 bp) to the 50 bp peak also varied among the three samples. The distribution segment amplitude and segment amplitude ratio (<40 bp to 50 bp peaks and >65 bp to 50 bp peaks) reflect results obtained from one experiment.

[0086] Figure 10 depicts the fragment length distribution of WU polyomavirus from two clinical samples. The left panel shows the distribution of fragments with individual peaks of approximately 50 base pairs (bp). The right panel shows a combined pattern including contributions from exponential distribution shape, peaks, and long fractions. Mechanistically unrestricted, short exponential fractions may indicate the presence of fragments within the “50 bp peak”.Different processes have led to the incorporation of viruses into the human genome or the degradation of microbial nucleic acids.

[0087] Figure 11 provides examples of fragment length characteristics involving fragment count ratios with different distributions. The figure depicts the ratio of fragment counts in the “50 bp peak” fraction to short class index fractions (read densities of 40–55 bp / read densities of 23–35 bp, x-axis) relative to normalized counts (y-axis). The same human and human mitochondrial fractions have been added for reference. The ratios vary between kingdom types. The ratios of bacterial reads vary considerably, while the ratios of fungal reads show a bimodal pattern. The ratios of viral reads are also shown.

[0088] Figure 12 provides a summary of the fragment length distributions of cell-free nucleic acids from the mother (dashed line) and fetus (solid line). The “50 bp peak” appears narrower in the fetal distribution, indicating a smaller range of fragment lengths within the peaks from fetal nucleic acids. Additionally, the ratio of fetal to maternal reads in the "50 bp peak" region is higher compared to nucleosome length fragments (e.g., the 150–200 bp region).

[0089] Figure 13 provides a summary of fragment length distributions for microorganisms present as pathogens or symbiotic microorganisms. In assays based on end-repairable double-stranded DNA, pathogen fragment lengths tend to be longer than those of symbiotic microorganisms.

[0090] Figure 14 provides a summary of fragment length distributions for pathogens in nucleic acid libraries generated from samples confirmed to be infected by urine or blood cultures. Pathogens detected in nucleic acid libraries from samples tested using orthogonal blood cultures show a higher read rate compared to pathogens detected in nucleic acid libraries from samples using orthogonal urine cultures. Read lengths are shown on the x-axis; the fraction of reads is shown on the y-axis. The mean values ​​(light solid line) for urine culture samples and the mean values ​​(light dashed line) for blood culture samples are shown in the figure, as well as the difference between urine and blood (thick dashed line).

[0091] Figures 15A-15F summarize data from asymptomatic samples (AP), diagnostic positive samples (DP), diagnostic positive samples confirmed using orthogonal methods (DPc), diagnostic positive samples confirmed using orthogonal NGS methods (DPNGS), and diagnostic positive samples confirmed using orthogonal non-NGS microbiological methods (DPmicro), as indicated. Figure 15A provides a plot of the abundance of microorganisms present at significant levels in the indicated sample types, in molecules per microliter (MPM). Figure 15B provides a plot of the MPM abundance of microorganisms of the same species present in both types of samples, in asymptomatic samples (AP) and diagnostic positive samples (DP). Figure 15C provides an example of a representative TapeStation electrophoresis plot of a library obtained from diagnostic positive samples included in this study. The data were obtained using a TapeStation with an HS TapeStation D1000, utilizing...Loading buffer and DNA ladder were obtained according to the manufacturer's instructions. Higher and lower DNA markers are indicated in the figure. The orientation of subsets of regions of interest within the fragment length range is indicated in the plot (note that fragment lengths in the library electrophoresis plot reflect the length of the fully adapted nucleic acid molecule, not the actual length of the endogenous original sequence). Library fragment lengths are shown on the x-axis; normalized intensity (FU) is shown on the y-axis. Figure 15D provides a plot of the molar fractions of sequencing reads mapped to human references and longer than 64 bp (i.e., most of these reads are nucleosome length) after the adaptor sequence trimming steps in the asymptomatic (AP) and diagnostic positive (DP) samples included in this study (page 13 / 68, CN 121575489 A). Figure 15E provides a summary comparison of the maximum MPM abundance of microorganisms present at significant levels in each asymptomatic (AP) and diagnostic positive (DP) sample in this study with the summed fractions of long human reads present in the same samples as defined in the title of Figure 15D. This analysis included only AP and DP samples where significant levels of microorganisms were detected. Arrows indicate AP samples showing maximum MPM and long human read scores above 3000 and 0.4, respectively. Figure 15F provides a summary comparison of the maximum MPM abundance of microorganisms present at significant levels in asymptomatic (AP) and diagnostic negative (DN) samples with the scores of long human reads present in the same samples as defined in Figure 15D. This analysis included only AP and DN samples where significant levels of microorganisms were detected.

[0092] Figure 16A depicts the results of training a predictor of infection status based on human fragments recovered from asymptomatic and symptomatic patients for sequencing. The left panel shows the probability of asymptomatic samples based on the human training model. The right panel depicts the regions of fragment lengths associated with each infection status used in the human training model. Figure 16B depicts the results of training a predictor of infection status based on human mitochondrial fragments recovered from asymptomatic and symptomatic patients for sequencing. The left panel shows the probability of asymptomatic samples based on the human mitochondrial training model. The right panel depicts the regions of fragment length associated with each infection state used in the human mitochondrial training model. Figure 16C depicts the results of training a predictor of infection state based on all pathogen fragments recovered from asymptomatic and symptomatic patients for sequencing. The left panel shows the probability of asymptomatic samples trained on the model based on all pathogen fragments. The right panel depicts the regions of fragment length associated with each infection state used in the model trained on all pathogen fragments. Figure 16D depicts the results of training a predictor of infection state based on significant pathogen fragments recovered from asymptomatic and symptomatic patients for sequencing. LeftFigure 16E depicts the results of training a predictor of infection status based on a model trained only on reads derived from significant pathogens. The right figure depicts the region of fragment length associated with each infection status identified by the model trained on significant pathogens. Figure 16F depicts the results of training a predictor of infection status based on eukaryotic microbial fragments recovered from asymptomatic and symptomatic patients for sequencing. The left figure shows the probability of asymptomatic samples based on a eukaryotic training model. The right figure depicts the region of fragment length associated with each infection status identified by the eukaryotic cell training model. Figure 16G depicts the results of training a predictor of infection status based on viral fragments recovered from asymptomatic and symptomatic patients for sequencing. The left figure shows the probability of asymptomatic samples based on a viral training model. The right figure depicts the regions of fragment length associated with each infection state identified by the virus training model. Figure 16H depicts the results of training a predictor of infection state based on archaea fragments recovered from asymptomatic and symptomatic patients for sequencing. The left figure shows the probability of asymptomatic samples based on the archaea training model. The right figure depicts the regions of fragment length associated with each infection state identified by the archaea training model.

[0093] Figures 17A1-17A10 show the normalized fragment length distribution of microorganisms suspected of infecting the lungs, where each figure shows a distribution of microorganisms of the indicated species, and the sample ID is indicated at the top of each figure. The frequency is defined as the read count relative to a reference alignment of the indicated microorganism for a specific read (fragment) length, which is normalized by the total count of reads relative to the reference alignment of the indicated microorganism. Figures 17B1-17B10 show the normalized fragment length distribution of microorganisms suspected of infecting the bloodstream, where each figure shows a distribution of microorganisms of the indicated species, and the sample ID is indicated at the top of each figure. Frequency is defined as the number of reads compared to a reference for the indicated microorganism of a specific read (fragment) length, said count being normalized by the total count of reads compared to the reference for the indicated microorganism. Specification 14 / 68 pages 17 CN 121575489 A

[0094] Figures 18A1-A2 depict representative normalized fragment length distributions of two microorganisms detected in venous aspirations from two different donors. The left figure shows the normalized fragment length distribution of reads mapped to Haemophilus influenzae—the microorganism detected in venous plasma obtained from donor 1. The right figure shows...The diagram shows the normalized fragment length distribution of reads mapped to *Streptococcus thermophilus*—a microorganism detected in plasma obtained from venous blood drawn from donor 2. Figures 18B1-B4 depict the normalized fragment length distribution of microorganisms detected in biological samples obtained during capillary collection processes from the same two donors and taken at the same sampling time as the venous collection in Figures 18A1-A2. The upper left panel shows the normalized fragment length distribution of *Haemophilus influenzae* detected in biological samples obtained during capillary collection processes from donor 1. The lower left panel shows the normalized fragment length distribution of other microorganisms detected in biological samples obtained during capillary collection processes from donor 1. Their average distribution patterns are shown by thick black lines. The upper right panel shows the normalized fragment length distribution of *Streptococcus thermophilus* detected in biological samples obtained during capillary collection processes from donor 2. The lower right figure shows the normalized fragment length distribution of additional microorganisms detected in the biological sample obtained during the capillary collection process from donor 2. The average distribution pattern is shown by a thick black line. Figures 18C1-C2 compare the abundance of co-occurring microorganisms in two replicas of the biological samples obtained during the capillary collection process from donor 1 (left) and donor 2 (right). Figures 18D1-D2 depict a comparison of the microbial abundance (x-axis) of microorganisms detected in the biological sample obtained using the capillary blood extraction procedure with the microbial abundance in the negative Microvette sample. The results for donor 1 and donor 2 are shown in the left and right figures, respectively.

[0095] Figures 19A1-A3 orthogonally confirm that the bloodstream of subject RD-02 was infected with Enterobacter species. The figure depicts the normalized fragment length distribution of sequences aligned with *Enterobacter cloacae* in nucleic acid libraries generated from plasma samples collected at different collection times indicated in each of the above figures. Figures 19B1-B5 orthogonally confirm that subject RD-11 suffered from endocarditis caused by *Staphylococcus aureus* infection. The figure depicts the normalized fragment length distribution of sequences aligned with *Staphylococcus aureus* in nucleic acid libraries generated from plasma samples collected at different collection times indicated in each of the above figures. Figures 19C1-C4 orthogonally confirm that subject RD-13 suffered from febrile neutropenia caused by *Escherichia coli* infection. The figure depicts the normalized fragment length distribution of sequences aligned with sequences from plasma samples collected at different collection times indicated in each of the above figures.Normalized fragment length distribution of E. coli sequences aligned to nucleic acid libraries generated from plasma samples collected over time.

[0096] Figure 20A depicts the fraction of reads outside the “50 bp peak” region (< 30 bp and > 60 bp) of the fragment length distribution of all orthogonally confirmed microorganisms as of post-admission time. Time traces of orthogonally confirmed microorganisms are shown where more than 50 unique sequences aligned to the microorganism reference were detected. Figure 20B depicts the abundance of orthogonally confirmed microorganisms detected by the method, in MPM, as of post-admission time.

[0097] Figures 21A1-A4 show paired orthogonally confirmed and orthogonally unconfirmed microorganisms in plasma samples collected at the admission time point (t = 0) of two subjects, RD-06 and RD-13. The orthogonally confirmed microorganism (Staphylococcus aureus) in RD-06 is shown in the upper left figure. The lower left figure shows unidentified microorganisms (Haemophilus influenzae) in RD-06. The upper right figure shows orthogonally confirmed microorganisms (Escherichia coli) in RD-13. The lower right figure shows unidentified microorganisms (Prevotella melaninogenica) in RD-13. Figure 21B1-B2 Enterococcus quati – Normalized fragment length distribution of orthogonally confirmed unidentified microorganisms detected in plasma samples collected from subject RD-15 at several post-admission time points. Time points are indicated at the top of the figure.

[0098] Figures 22A-C depict three main response patterns of human fragment length distribution during treatment of infected subjects. The left figure shows an example in which the long human fraction (>60 bp) decreased during treatment. The middle figure shows an example in which the long human fraction (>60 bp) fluctuated during treatment instructions on page 15 / 68 of CN 121575489 A. The right figure shows an example where the fraction of long human cells (>60 bp) increased during treatment.

[0099] Figure 23 provides a summary of fragment length information and GC content for samples from *Pasteuranius*. Relative frequencies are shown on the y-axis; GC content is shown on the x-axis. Fragment length ranges of less than 45 base pairs, 45–54 base pairs, 55–64 base pairs, 65–74 base pairs, and longer than 74 base pairs are shown. The combination of fragment length distribution and GC content information indicates that the process induced a temperature deviation in this microorganism. Detailed Description

[0100] Next-generation sequencing (NGS) can be used to aggregate large amounts of data about the nucleic acid content of a sample. This data can be particularly used to analyze nucleic acids in complex samples, such as clinical samples. To date, these NGS systems have focused on identifying individual reads.Segment abundance. Prior to this work, the primary property of interest was the sequence of each read and the abundance of reads associated with a specific source. This is especially true for microbial nucleic acids and free microbial nucleic acids. This is partly due to the fact that prior sample processing required by many NGS systems often introduces errors and biases, particularly for low-abundance nucleic acids. Karius developed a method for preparing nucleic acid libraries from initial samples that reduces the bias in recovering nucleic acid libraries from initial samples, or allows for the correction of biases. The reduced bias in nucleic acid libraries obtained from initial samples allows for the development of fragment length profiles and methods for generating fragment length profiles of nucleic acid libraries or target nucleic acids within nucleic acid libraries. There is a need for efficient and accurate methods for generating fragment length profiles of nucleic acid libraries. For example, this need can be seen in distinguishing between closely related microorganisms, determining whether a microorganism exists as a pathogen or symbiotic microorganism, determining the biological relationship between a microorganism and its host, predicting infection or colonization sites in subjects, monitoring graft status, monitoring fetal development and status, monitoring tumors, monitoring the status and response of the immune system, and monitoring the toxicity of compounds administered to subjects.

[0101] The fragment length spectrum includes one or more fragment length characteristics of a nucleic acid library or a subset of reads from the nucleic acid library. The fragment length spectrum may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more fragment length characteristics. Weights may be assigned to one or more fragment length characteristics in the fragment length spectrum, such that one or more fragment length characteristics may have equal or different weights or values ​​within the fragment length spectrum. Fragment length characteristics include, but are not limited to, the shape of the distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of the helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of the maximum amplitude of two or more segments, position of one or more peaks, and fragment length distribution within a subset of reads. The intent is that the ratio "between two or more segments" encompasses, but is not limited to, two or more segments from one nucleic acid library, two or more segments from two or more nucleic acid libraries, two or more segments with the same peak shape, two or more segments with different peak shapes, two or more segments from similar or different nucleic acid library types, and two or more segments from similar or different subsets of reads from nucleic acid libraries.

[0102] Distribution types include, but are not limited to, single peak shape, multiple peak shapes, exponential or quasi-exponential distribution, dilated distribution of long or short segments, flat or uniform distribution, complex distribution shapes, and combinations thereof. Complex distributions may include at least two, at leastAspects of 3, at least 4, at least 5, at least 6, at least 7, at least 8 or more peak shapes. A single peak shape can exist at any fragment length, including but not limited to fragment lengths of about 50 base pairs. Long fragments can contain the following fragment lengths: greater than about 60 base pairs, about 65 base pairs, about 70 base pairs, about 75 base pairs, about 80 base pairs, about 85 base pairs, about 90 base pairs, about 95 base pairs, about 100 base pairs, about 150 base pairs, about 175 base pairs, about 200 base pairs, about 250 base pairs, about 300 base pairs, about 350 base pairs and about 400 base pairs. (See specification page 16 / 68, 19 CN 121575489 A) Short fragments may include the following fragment lengths: less than approximately 500 bp, approximately 400 bp, approximately 300 bp, approximately 200 bp, approximately 100 bp, approximately 50 bp, approximately 40 bp, approximately 35 bp, approximately 30 bp, approximately 25 bp, and approximately 20 bp. Aspects of peak shape include, but are not limited to, segment range, segment amplitude and the total number of reads within a segment, peak width, peak slope, and peak derivative; aspects of peak shape may vary.

[0103] The distribution of a single peak shape may cover a certain range of fragment lengths, including, but not limited to, fragment lengths of at least approximately 5 base pairs, at least approximately 10 base pairs, at least approximately 15 base pairs, at least approximately 20 base pairs, at least approximately 30 base pairs, at least approximately 35 base pairs, at least approximately 40 base pairs, or greater than at least approximately 45 base pairs within a segment. The range of fragment lengths within a segment may vary. For example, the range of fragment lengths with a unimodal distribution of about 50 base pairs includes, but is not limited to, fragment lengths of 30 to 60 base pairs, 35 to 60 base pairs, 40 to 60 base pairs, and 45 to 55 base pairs.

[0104] The segment amplitude covers the abundance or relative abundance of reads within a defined segment length. In some aspects, the distribution amplitude may be the highest abundance or relative abundance within a defined fragment length range; the distribution amplitude may also cover the average highest abundance or relative abundance within a defined fragment length range. In some aspects of this application, a fragment length distribution or fragment length distribution spectrum is obtained from a subset of reads from a nucleic acid library. The subset of reads from the nucleic acid library is intended to cover less than the complete set of reads from the nucleic acid library. A subset may reflect reads identified as: from a specific microbial type, from a specific microbial species, host reads, maternal reads, fetal reads, organ donor reads, non-host reads, cell-free nucleic acid reads, cell-free nucleic acid reads, microbial reads, or any other group; alternatively, a subset of reads may reflect the complete set of reads minus those from a specific microbial type, maternal reads, fetal reads, or any other group.In some aspects of this application, the fragment length distribution of the target nucleic acid is obtained. The “target nucleic acid” can be a nucleic acid fragment derived from: microorganisms, transplanted organs, tumor cells, cancer cells, host or non-host mitochondrial DNA, antibiotic resistance gene sequences, host genomic DNA, microbial sequences integrated into the host genome, or one or more other sequences of interest in a nucleic acid library. The target sequence may migrate from another site, such as an infection site or a donated organ.

[0105] In some cases, the target nucleic acid may constitute only a very small portion of the entire sample, for example, less than 0.1%, less than 0.01%, less than 0.001%, less than 0.0001%, less than 0.00001%, less than 0.000001%, less than 0.000001% of the total nucleic acids in the sample. Typically, the total nucleic acids in the original sample can vary. For example, total free nucleic acids (e.g., DNA, mRNA, RNA) can range from 0.01 to 10,000 ng / ml (e.g., about 0.01, 0.1, 1, 5, 10, 20, 30, 40, 50, 80, 100, 1000, 5000, 10000 ng / ml). In some cases, the total concentration of free nucleic acids in a sample is outside this range (e.g., less than 0.01 ng / ml; in other words, a total concentration greater than 10,000 ng / ml). The same is true for free nucleic acid (e.g., DNA) samples that are mainly composed of human DNA and / or RNA. In such samples, the presence of pathogen target nucleic acids can be less than that of human or host nucleic acids.

[0106] The length of the target nucleic acid can vary. In some specific embodiments, the target nucleic acid is relatively short; in other embodiments, the target is relatively long. In some specific embodiments, the target nucleic acid is shorter than 110 bp.

[0107] As used herein, “nucleic acid” refers to a polymer or oligomer of nucleotides and is generally synonymous with the terms “polynucleotide” or “oligonucleotide”. Nucleic acids may include, consist of, or consist substantially of: deoxyribonucleotides, ribonucleotides, deoxyribonucleotide analogs, chemically modified typical deoxyribonucleotides, ribonucleotides and / or ribonucleotide analogs, nucleic acids having a modified backbone, or any combination thereof.

[0108] Nucleic acids can be any type of nucleic acid, including but not limited to: double-stranded (ds) nucleic acids, single-stranded (ss) nucleic acids, DNA, RNA, cDNA, mRNA, cRNA, tRNA, ribosomal RNA, dsDNA, ssDNA, miRNA, siRNA, short hairpin RNA, circulating nucleic acids, circulating cell-free nucleic acids, circulating DNA, circulating RNA, cell-free nucleic acids, cell-free DNA, cell-free RNA, circulating cell-free DNA, cell-free dsDNA, specification 17 / 68 pages 20 CN 121575489 ACellular ssDNA, circulating cell-free RNA, genomic DNA, exosomes, free pathogen nucleic acid, circulating microbial or pathogen nucleic acid, mitochondrial nucleic acid, non-mitochondrial nucleic acid, nuclear DNA, nuclear RNA, chromosomal DNA, circulating tumor DNA, circulating tumor RNA, circular nucleic acid, circular DNA, circular RNA, circular single-stranded DNA, circular double-stranded DNA, plasmids, bacterial nucleic acid, fungal nucleic acid, parasitic nucleic acid, viral nucleic acid, free bacterial nucleic acid, free fungal nucleic acid, free parasitic nucleic acid, virus particle-associated nucleic acid, mitochondrial DNA, host nucleic acid, host free nucleic acid, intercellular signaling nucleic acid, exogenous nucleic acid, DNase, RNase, therapeutic nucleic acid, or any combination thereof. Nucleic acid can be derived from microorganisms or pathogens, including but not limited to viruses, bacteria, fungi, parasites, and any other microorganisms, particularly infectious or potentially infectious microorganisms. Nucleic acid can be derived from archaea, bacteria, fungi, molds, eukaryotes, and / or viruses. In some embodiments, unlike microorganisms or pathogens, nucleic acid can be directly derived from the subject or host.

[0109] As used herein, a “nucleic acid library” refers to a collection of nucleic acid fragments. The collection of nucleic acid fragments can be used, for example, for sequencing. Nucleic acid libraries can be prepared from an initial sample using a bias-corrected recovery method for generating a sequencing library or a bias-corrected recovery method for generating a sequencing library. As used herein, a “bias-corrected recovery” method is: a method utilizing consistent fragment length generation, which typically recovers sample nucleic acid fragments within a targeted length and GC range without perceptible length and GC bias; a method for achieving bias correction; a method capable of resolving biases with respect to the sample; and a method capable of resolving biases introduced by the process of generating the nucleic acid library. Bias-corrected recovery methods may include, but are not limited to, adding process control molecules, extraction, library generation, sequencing, amplification, and any combination thereof. Bias-free recovery methods include, but are not limited to, those described in U.S. Provisional Nos. 62 / 770,181 and 62 / 644,357. Methods are provided for generating nucleic acid libraries from an initial sample without extracting nucleic acids from the initial sample prior to initiating the nucleic acid library generation process. In some embodiments, substances that may reduce yield or inhibit the generation of nucleic acid libraries may be extracted or removed, but nucleic acids are not extracted from the initial sample before the generation of the nucleic acid library. The method comprises, consists of, or substantially comprises: adding one or more process control molecules to an initial sample, and generating a nucleic acid library from the spiked initial sample. The method comprises, consists of, or substantially comprises: generating a nucleic acid library from a spiked initial sample. The nucleic acid library may utilize single-stranded and / or double-stranded nucleic acids.

[0110] Methods for generating nucleic acid libraries from samples using extraction are also covered.

[0111] Process control molecules may be one or more of ID spiking, SPANK, Spark, or GC spiking groups, dephosphorylation control molecules, denaturation control molecules, and / or linkage control molecules. See, for example, published U.S. Patent Application No. 2015-0133391 and published U.S. Patent Application No. 2017-0016048, the entire disclosure of each of which is incorporated herein by reference for all purposes. In some embodiments, the initial sample comprises, consists of, or is substantially composed of circulating donor nucleic acids (see, for example, US 20150211070, which is incorporated herein by reference in its entirety, including any figures).

[0112] As used herein, “denaturation” refers to the process in which a biomolecule, such as a protein or nucleic acid, loses its native or higher-order structure. Native and higher-order structures may include, for example, but not limited to, quaternary, tertiary, or secondary structures. For example, a double-stranded nucleic acid molecule may denature into two single-stranded molecules.

[0113] As used herein, the terms “dephosphorylation” or “dephosphorylating” refer to the removal of phosphates, such as 5' and / or 3' phosphates, from nucleic acids, such as DNA.

[0114] As used herein, “detection” refers to quantitative or qualitative detection, including but not limited to detection by identifying the presence, absence, quantity, frequency, concentration, sequence, form, structure, source, or amount of an analyte.

[0115] In some embodiments, the 3' end adapter is ligated to a nucleic acid, such as a denatured or dephosphorylated nucleic acid, and / or the 5' end adapter is ligated to an enzyme, consists of an enzyme ligation, or is substantially composed of an enzyme ligation, including ligases such as T4 DNA ligase, CircLigase II, or consist of an enzyme ligase. In some embodiments, the ligase is a single-stranded ligase. In some embodiments, linking the 3' end adaptor to a nucleic acid, such as a denatured or dephosphorylated nucleic acid, and / or linking the 5' end adaptor comprises using a template switching reaction, or consists of or substantially consists of using a template switching reaction. In some embodiments, linking the 3' end adaptor to a nucleic acid, such as a denatured or dephosphorylated nucleic acid, comprises using an enzyme extension, or consists of or substantially consists of using an enzyme extension, wherein the enzyme comprises a polymerase, such as TdT polymerase, or consists of or substantially consists of a polymerase. In some embodiments, the method further comprises, consists of, or substantially consists of: using a DNA polymerase, such as...Klenow fragments, SuperScript IV reverse transcriptase, SMART MMLV reverse transcriptase, etc., are used to extend primers that hybridize with nucleic acids or linked nucleic acids and generate complementary strands. In some embodiments, the target nucleic acid may be linked to one or more adaptors. In some embodiments, the target nucleic acid is linked to the same adaptor or different adaptors at both ends.

[0116] As used herein, “GC bias” refers to the difference in performance, treatment, or recovery of nucleic acids with different GC contents but the same length.

[0117] As used herein, “GC content” or “guanine-cytosine content” refers to the percentage of nitrogenous bases in a nucleic acid, such as a DNA or RNA molecule, where the nitrogenous bases are guanine or cytosine or chemical modifications thereof.

[0118] As used herein, “host” refers to an organism that has another organism. The latter is defined as a “non-host” organism. For example, a human can be a host having a microorganism, pathogen, or fetus that is a non-host. The host nucleic acid or material is derived from the host. Non-host nucleic acids or materials can be derived from non-host organisms, from transplanted materials, or from fetuses or fetal materials within a host.

[0119] As used herein, “microbe,” “microbial,” or “microorganism” means an organism that can exist as a single cell or as a colony of cells, capsids, spores, filaments, or multicellular organisms, such as microscopic or macroscopic organisms. Microorganisms include all single-celled organisms and some multicellular organisms, such as those from archaea, bacteria, protozoa, nematodes, viruses, and eukaryotes. Microorganisms are typically pathogens responsible for diseases, but can also exist in a non-pathogenic, symbiotic relationship with a host, such as a human. “Symbiotic microorganism” is intended to include microorganisms that exist in a non-pathogenic, symbiotic relationship with a host. A host organism can simultaneously have multiple types of non-host organisms. In co-infection, the host organism has multiple types of non-host organisms. Multiple types of non-host organisms can include one or more pathogens, one or more symbiotic microorganisms, or at least one pathogen and at least one symbiotic microorganism. The methods of the present application can be used to distinguish between closely related microorganisms, or between microorganisms that exist as pathogens, symbiotic microorganisms, or incidental but clinically insignificant microorganisms.

[0120] Microorganisms or pathogens may include archaea, bacteria, yeasts, fungi, molds, protozoa, nematodes, eukaryotes, and / or viruses. Microorganisms or pathogens may also include DNA viruses, RNA viruses, culturable bacteria, other refractory and unculturable bacteria, mycobacteria, and eukaryotic pathogens (see Bennett J.E., D., R., Blaser, M.J.).Mandell, Douglas, and Bennett, Principles and Practice of Infectious Diseases; Saunders, Philadelphia, PA, 2014; and Netter's Infectious Disease, 1st Edition, edited by Elaine C. Jong, MD, and Dennis L. Stevens, MD, S.D., (2015). Microbes or pathogens may also include any of the microbes listed below: https: / / www.ncbi.nlm.nih.gov / genome / microbes / or https: / / www.ncbi.nlm.nih.gov / biosample / .

[0121] Examples of microorganisms are one or more species or strains from one or more of the following genera: *coniosporium*, *Hantavirus*, *Talaromyces*, *Machlomovirus*, *Betatetravirus*, and *Raoultella*. Aeromonas, Ephemerovirus, Ephemerovirus, Loa, Macluravirus, Stenotrophomonas, Alfamovirus, Rosavirus, Emmonsia, Aggregatibacter, Orthopneumovirus, Weeksella, Nairovirus, Salivirus, Weissella, Mosavirus, Gammapartitivirus, Strongyloides, Passerivirus, Erysipelatoclostridium, BasilanavirusBacillarnavirus, Iotatorquevirus, Taenia, Trypanosoma, Olsenella, Cladosporium, Rhizobium, Prevotella, Leclercia, Paracoccus, Ilarvirus, Lagovirus, Rasamsonia The genera *Plasmodium*, *Acremonium*, *Chlamydia*, *Clonorchis*, *Vibrio*, *Bartonella*, *Nakazawaea*, *Franconibacter*, *Anisakis*, *Norovirus*, *Nocardia*, *Solobacterium*, *Parechovirus*, and others are mentioned. Genus Avenavirus, Genus Orthohepevirus, Genus Aphthovirus, Genus Hepandensovirus, Genus Microbacterium, Genus Lichtheimia, Genus Lomentospora, Genus Achromobacter, Genus Ipomovirus, Genus Tsukamurella, Genus Elizabethkingia, Genus Hepatitis E Genus: Hepevirus, Seadornavirus, Alternaria, Trueperella, Gammatorquevirus, Bifidobacterium, Chrysosporium, Thogotovirus, Curtovirus, Deltatorquevirus, Balamuthia.Mastrevirus, Bdellomicrovirus, Mupapillomavirus, Pseudozyma, Wickerhamiella, Aquamavirus, Alloscardovia, Thievaria, Idaeovirus, Henipavirus, Coxiella, Haemophilus, Gammacoronavirus, Negevirus, B revibacterium, Peptone eptoniphilus), Alphacarmotetravirus, Nosema, Trichovirus, Arenavirus, Thermomyces, Necator, Waikavirus, Blosnavirus, Jonesella, Tetraparvovirus, Emaravirus, Plectrovirus. Sclerodarnavirus, Toxocara, Umbravirus, Burkholderia, Chromobacterium, Paraacoccidioides, Brugia, Eragrovirus, Macrococcus, Absidia, Colletotrichum, Inovirus, Phycomyces, Wickerhamomyces, Acidaminococcus (Instruction manual, page 20 / 68, CN 121575489 A), Moraxella, Rothia, Sandfly virusPhlebovirus, Slackia, Purpureocillium, Beta-pa pillomavirus, Tupavirus, Cryptosporidium, Saksenaea, Erysipelothrix, Kobuvirus, Mimoreovirus, Echinococcus, Mannheimia, Bergeyella, Cyclospora, Xylanmonella nimonas, Leptospira, Finegoldia, Curvularia, Cryptosporidium, Babuvirus, Pecluvirus, Lambdatorquevirus, Pythium, Carlavirus, Entomobirnavirus, Coxella Genus: Kocuria, Anaplasma, Ampelovirus, Avihepatovirus, Nepovirus, Rhodococcus, Bordetella, Mischivirus, Scedosporium, Gardnerella, Maculavirus, Trichoderma, Aveparvovirus, Salmonella, Avastrovirus, Copiparvovirus, Trachipleistophora, Clostridium, Nanovirus, Siccibacter, Leptotrichia, Citrivirus, Odor-causing bacteriaOdoribacter, Sanguibacter, Novirhabdovirus, Acremonium, Hafnia, Chaetomium, Tenuivirus, Yokenella, Rubulavirus, Varicellovirus, Alphamesonivirus, Sicinivirus, Leuconostoc, Microvirus, Gallantivirus, Morbillivirus, Lolavirus, Pantoea, Hepatovirus, Nupapillomavirus, Metschnikowia, Double RNA The genera *Barnavirus*, *Kytococcus*, *Tritimovirus*, *Tannerella*, *Respirovirus*, *Pneumocystis*, *Dirofilaria*, *Pediococcus*, *Lactococcus*, *Blastomyces*, *Dianthovirus*, *Actinobacillus*, *Teschovirus*, *Oscivirus*, *egomovirus*, and *Potato Y* Genus of viruses: Potyvirus, Byssochlamys, Iphacoronavirus, Molluscum cipoxvirus, Lymphocryptovirus, Sapelovirus, Parabacteroides, Pyrenochaeta, Listeria, Senecavirus, Brevidensovirus, Potato X disease*Potexvirus*, *Parvimonas*, *Flavavivirus*, *Recovirus*, *Toxoplasma*, *Yatapoxvirus*, *Opisthorchis*, *Richuris*, *Cyphellophora*, *Morganella*, *Perhabdovirus*, *Micrococcus*, *Pequenovirus*, *Mastad enovirus*, *Anaeroglobus*, *Tropheryma*, *Dolosigranulum*, *Wolbachia*. (Instruction manual, page 21 / 68, 24, CN 121575489 A) Wolbachia, Lelliottia, Mycoplasma, Tobravirus, Shewanella, Paeniclostridium, Erythroparvovirus, Sutterella, Sporopachydermia, Narnavirus, Nyavirus, Francisella, Arthroderma, Epsilontorquevirus, Sigmavirus, Amdoparvovirus, Actinomyces, Alphapermutotetravirus, Cardiobacterium, Influenza virus C) Genus Orthopoxvirus, Genus Poacevirus, Genus Phialophora, Genus Lactobacillus, Genus Polyomavirus, Genus Debaryomyces, Genus Foveavirus, Genus Bymovirus, Genus MicoccusMycoflexivirus, Grimonia, Mucor, Rhytidhysteron, Quadrivirus, Thermoascus, Aureusvirus, Trichosporon, Myceliophthora, Dermacoccus, Dysgonomonas, Pseudoramibacter, Becurtovirus, Gordonia, Sapovirus, Orthobunyavirus, Spiromicrovirus, Pomovirus, Exophiala, Sneathia, Helicobacter, Photorhabdella *As*, *Mogibatterium*, *Betapartitivirus*, *Avian BiRNA Virus* The genera *Ambidensovirus*, *Oleavirus*, *Orientia*, *Deltacoronavirus*, *Anulavirus*, *Trichomonasvirus*, *Budvicia*, *Geotrichum*, *Enamovirus*, *Lachnoclostridium*, *Schistosoma*, *Paecilomyces*, *Panicovirus*, *Rhizoctonia*, *Brevibacillus*, *Beauveria*, *Pestivirus*, *Tombusvirus*, *Cilevirus*, *Cokeromyces*, *Peptostreptococcus*, and *Pleurotus* are mentioned.(Phanerochaete), Proteus, Idnoreovirus, Aspergillus, Pasteurella, Malassezia, Hanseniaspora, Endornavirus, Azospirillum, Velarivirus, Cystovirus, Avisivirus, Bacteroides, Picobirnavirus, Myroides, Circovirus, Arterivirus, Aquaparamyxovirus, Onchocerca, Cosavirus, Kluyveromyces, Fijivirus, Candida IDA), Hepatovir, Dermabacter, Ourmiavirus, Allexivirus, Enterobacter, Acidovorax, Bracorhabdovirus, Carmovirus, Pluralibacter, Coltivirus, Fonsecaea, Streptobacillus, Corynebacterium, Macrophomina, Marburgvirus, Comovirus, Leguminosae Virus. (Instruction manual, page 22 / 68, 25 CN 121575489 A) (Fabavirus), Alphanodavirus, Cellulomonas, Enterobius, Catabacter, Moellerella, Nakaseomyces, Cucumovirus, Valsa virus, and type D double disease*Deltapartitivirus*, *Plesiomonas*, *Pseudomonas*, *Torovirus*, *Cuevavirus*, *Hypovirus*, *Trichomonas*, *Influenzavirus D*, *Giardiavirus*, *Crinivirus*, *Tepovirus*, *Sakobuvirus*, *Cyberlindnera*, *Paenalcaligenes*, *Bafinivirus*. Rymovirus, Pegivirus, Yarrowia, Treponema, Borreliella, Rubivirus, Aureobasidium, Angiostrongylus, Filobasidium, Photobacterium, Rhizopus, Orthoreovirus, Ustilago, Simplex virus, Aquatic reovirus Genus: Aquareovirus, Protoparvovirus, Propionibacterium, Sprivivirus, Hunnivirus, Apophysomyces, Meyerozyma, Alphapapillomavirus, Candida, Brucella, Gallivirus, Dinovernavirus, Anaerobiospirillum, Eubacterium, Tatlockia, Terrisporobacter, Quaranjavirus, Sobemovirus, Dipsisvirus(Dicipivirus), *Arcanobacterium*, *Macanavirus*, *Atopobium*, *Vesivirus*, *Lodderomyces*, *Dinornavirus*, *Betatorquevirus*, *Kerstersia*, *Aparavirus*, *Neisseria*, *Agrobacterium*, *Edwardsiella*, *Labyrnavirus*, *Totivirus*, *Actinomadura*, *Tobamovirus*, *Influenzavirus* B), Mandarivirus, Anaerococcus, Kunsagivirus, Naegleria, Campylobacter, Veillonella, Yamadazyma, Filobasidiella, Oerskovia, Penicillium, Anncaliia, Leptosphaeria, Pneumovirus, Psychrobacter, Isavirus, Streptococcus, Torradovirus, Cladophialophora, Influenzavirus A, Ophiostoma, Aerococcus, Urea plasma, E. heptavirus rq ue vi rus), Bocaparvovirus, Megasphaera, Reptarenavirus, Commonas, Capnocytophaga, and Clavivirus A(Alphatorquevirus), Syncephalastrum, Wallemia, Betacoronavirus, Hyphopichia, Nocardiopsis, Legionella, Trichinella, Paraburkholderia, Mammarenavirus, Echinostoma, Sphingomyelinum, 23 / 68 pages, CN 121575489, Sphingobacterium, Enterovirus us), Methanobrevibacter, Ochroconis, Cheravirus, Pasivirus, Enterococcus, Mycoreovirus, Tospovirus, β-Nodamuravirus etanodavirus, Phytoreovirus, Enterocytozoon, Ferlavirus, Stemphylium, Filifactor, Leishmaniavirus, Gemella, Bromovirus, Alloiococcus, Cunninghamella, Cronobacter, Oribacterium, Orbivirus, Chrysovirus, Cripavirus, Tatumella, Pandoraea, Ogataea, Dracunculus, Volvariella, Iflavirus, Benyvirus.Rhadinovirus, Histoplasma, Rahnella, Morococcus, Verticillium, Janibacter, Gyrovirus, Alphapartitivirus, Mycobacterium, Roseomonas, Varicosavirus, Chryseobacter ium), Parapoxvirus, Rhizomucor, Aureimonas, Levivirus, eishmania, Luteovirus, Cypovirus, Ochrobactrum, Microsporum, Piscihepevirus, Ceratocystis, Sporothrix, vesicular disease *Vesiculovirus*, *Cupriavidus*, *Cryptococcus*, *Metapneumovirus*, *Alphanecrovirus*, *Eikenella*, *Brevundimonas*, *Escherichia*, *Leifsonia*, *Schizophyllum*, *Granulibacter*, *Gordon* ibacter, Lachancea, Madurella, Ophiovirus, Phellinus, Nebovirus, Acanthamoeba, Fusobacterium, Pichia, Verruconis, Ehrlichia, Tibrovirus, Higrevirus, W ohlf ahrtiimo na s, R hinoceros*Ladiel la*, *Neorickettsia*, *Sadwavirus*, *Roseobacter*, *Sequivirus*, *Pannonibacter*, *Rotavirus*, *Turicella*, *Cardiovirus*, *Propionimicrobium*, *Furuvirus*, *Naumovozyme* ma), Closterovirus, Fluoribacter, Zeavirus, Clavispora, Megrivirus, Gammapapillomavirus, Rickettsia, Polemovirus, Corynespora, Encephalitozoon, Shimwellia, Fusarium, Yersinia, Capronia, Delftia, Victoriavirus, Marafivirus, Kluyvera, Iteradensovirus, Isoptericola, Vitivirus, (Instructions 24 / 68, page 27) CN 121575489 A Genus: Roseolovirus, Conidiobolus, Abiotrophia, Babesia, Phona, Sang uibacteroides, Staphylococcus, Rhodotorula, Zetatorquevirus, Hymenolepis, Fasciola, Cytorhabdovirus, CardiorespiratoryOrphanvirus (Cardoreovirus), *Memnoniella*, *Trichophyton*, *Mitovirus*, *Phaeoacremonium*, *Providencia*, *Lysinibacillus*, *Giardia*, *Oligella*, *Streptomyces*, *Paraclostridium*, *Ralstonia*, and others. Coccidioides, Brambyvirus, Biatriospora, Allolevivirus, Acinetobacter, Starmerella, Omegatetravirus, Porphyromonas, Avulavirus, Streptococcus, Arcobacter, Tomato pseudotopevirus uvirus, Mamastrovirus, Ancylostoma, Bornavirus, Capillovirus, Alphavirus, Tymovirus, Nucleorhabdovirus, Diaporthe, Chlamydiamicrovirus, Turncurtovirus, Saccharovirus myces, Riemerella, Betanecrovirus, Clostridium, Mobiluncus, Cercospora, Marnavirus, Mortierella, Aquabirnavirus, Xanthomonas, Dependoparvovirus, Ebolavirus, Clostridium neoformans(Neofusicoccum), Borrelia, Leminorella, Klebsiella, Blastocystis, Alcaligenes, Citrobacter, Eggerthella, Cedeceea, Serratia, Penstyldensovirus, Bacillus, Laribacter, Wuchereria, Hordeivirus, Cytomegalovirus, Actinomucor, Ascaris, Shigella, Vittaforma, Torulaspora, Kingella. The genera *Oryzavirus*, *Polerovirus*, *Tremovirus*, *Erbovirus*, *Entamoeba*, *Lyssavirus*, *Paenibacillus*, *Facklamia*, *Kappatorquevirus*, *Metarhizium*, *Stachybotrys*, *Okavirus*, *Botrexvirus*, *Thetatorquevirus*, and *Basidiobolus*.

[0122] As used herein, infection stage or stage of infection refers to the asymptomatic infection period, symptomatic infection period, remission infection period, treatment period, relapse period, recurrence period, acute phase or infection, chronic phase or infection, slow or latent phase or infection, persistent infection, disseminated infection stage, initial, secondary or tertiary infection. The asymptomatic infection period occurs before the onset of symptoms or before the subject or other person notices symptoms. Synonyms for “asymptomatic period” will include “presymptomatic infection stage,” “initial infection stage,” and “early infection stage.” Symbionts may persist during the asymptomatic phase of infection. The symptomatic infection period occurs when the subject or other person notices symptoms or clinical changes, such as fever, pain, etc.Symptoms include pain, rash, headache, aches, and breathing problems. The remission phase occurs during the period when the infection resolves on its own or through the application of treatment (Instructions for Use, page 25 / 68, CN 121575489 A). The treatment period can be part of the remission phase after the application of treatment. The relapse phase occurs when the subject experiences a recurrence of infection at any of the above phases. The relapse phase occurs when the infection is not properly or adequately treated in the first instance and the infection returns. Chronic infection is a type of persistent infection that will eventually be cleared. Acute phase or infection occurs suddenly, such as hepatitis. Slow or latent phase or infection persists for the remainder of the host's life. Persistent infection is an infection that lasts for a long period; persistent infection occurs when the host has not cleared the primary infection. Some microorganisms infect the host with primary, secondary, and tertiary infections; an example is infection via Treponema pallidum. Infection can remain at any of the above phases for an indeterminate period of time and may not necessarily progress to different phases. Symbiotic or commensal microorganisms may remain in an invisible phase of infection indefinitely or may not infect at all.

[0123] Various host-microbe biological relationships or interactions are known in the art. Host-microbe biological interactions include, but are not limited to, symbiosis, mutualism, parasitism, commensalism, and competition. It is generally accepted that when a microorganism is located at certain sites within a host, it may exhibit one type of interaction with the host, but when it is located at another site, it may exhibit another type of interaction with the host. For example, a microorganism may exist in a symbiotic relationship with a host on the host's skin, but may exist in a parasitic or competitive relationship within the host. As used herein, “pathogen” means a microorganism that causes, can cause, or is suspected of causing disease.

[0124] As used herein, the phrase “spiked initial sample” refers to an initial sample to which process control molecules have been added before the generation of a sequencing library begins.

[0125] The term “derived from” encompasses the terms “originating from,” “obtained from,” “available from,” and “produced by,” which generally indicates that a specified material is derived from another specified material or has features that can be described with reference to another specified material. For example, an initial sample may be derived from a primitive biological sample.

[0126] In some embodiments, the initial sample comprises, consists of, or is substantially composed of: solids or body fluids, such as blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, feces, saliva, peritoneal fluid, peritoneal lavage fluid, gastric juice, interstitial fluid, lymph, bile, abscess fluid, tissue, amniotic fluid, meconium, sinus aspirate, lymph nodes, bone marrow, hair, nails, cheek swabs, skin swabs, urethral swabs, cervical swabs, nasopharyngeal swabs, nasopharyngeal aspirate, vaginal swabsSamples include sperm, epithelial cells, semen, vaginal discharge, intercellular fluid, pericardial fluid, rectal swabs, bone, skin tissue, soft tissue, tears, and / or nasal samples. In some embodiments, the initial sample includes plasma, is composed of plasma, or is substantially composed of plasma. In some embodiments, the initial sample includes urine, is composed of urine, or is substantially composed of urine. In some embodiments, the initial sample includes cerebrospinal fluid, is composed of cerebrospinal fluid, or is substantially composed of cerebrospinal fluid. In some embodiments, the initial sample is derived from a human subject.

[0127] In some embodiments, the initial sample may consist wholly or partially of cells and / or tissues. The initial sample may be free or cell-depleted. The initial free sample may include nucleic acids derived from different sites in the body, such as pathogen infection sites, composed of said nucleic acids, or substantially composed of said nucleic acids. In the case of blood, serum, lymph, or plasma, the free sample or the cell-depleted initial sample may contain “circulating” free nucleic acids derived from anatomical locations other than the fluid collection sites of the fluids discussed. In the case of urine, the free nucleic acids may be free nucleic acids derived from different sites in the body. Free samples or cell-depleted initial samples can be obtained by means of depletion or removal of cells, cell debris, or foreign bodies using known techniques such as centrifugation or filtration.

[0128] As used herein, the term “invasive disease” refers in part to a disease based on the ability of a particular pathogen to severely impair the health of certain infected subjects, in contrast to colonizing other infected subjects only in a symbiotic or uninfected or mildly symptomatic form. For example, some microorganisms may locally colonize tissues in some hosts without causing any health problems, while in other hosts they may invade tissues to the point where they cause serious inflammation, tissue or organ damage, sepsis, cancer, and other serious health problems. Microorganisms may also colonize asymptomatic subjects at one point in time, but develop severe symptoms at a later point, when the microorganism translocates and / or becomes “active.”

[0129] As used herein, the term “free” refers to the state in which nucleic acids are outside of cells, viral particles, or virions when they are present in the body, prior to obtaining a sample from the body. For example, circulating cell-free nucleic acids in a sample may originate from cell-free nucleic acids circulating in the bloodstream of a subject. Conversely, nucleic acids extracted from intact microorganisms, such as bloodborne pathogens, or removed from intact virions in plasma samples are generally not considered "free".

[0130] This application provides a method for determining the localization site of a subject. Nucleic acids from microorganisms or from different sites within a subject can exhibit different fragment length profiles. If the microbial infection is circulating rather than located at one or more localization sites, the fragment length profile of a nucleic acid library containing microbial nucleic acids or a subset of the nucleic acid library will not be...Therefore, comparing the fragment length spectrum with a reference fragment length spectrum from one or more source sites can predict localization sites when the fragment length spectrum from the sample is similar to the reference fragment length spectrum from the source site. A “localization site” refers to any source site in the subject’s body where the microorganism is present, persists, survives, or proliferates. Source sites include, but are not limited to, blood, blood, and deep tissues such as, but not limited to, the kidneys, liver, stomach, bladder, digestive organs, nerve cells, lungs, bones, brain, heart, heart lining, sinuses, GI tract, spleen, skin, joints, ears, nose, and mouth. It is envisioned that a subject may have more than one localization site for a particular microorganism. It should be further understood that some localization sites of a particular microorganism may not contribute to a disease state or condition. Rather, some localization sites of a particular microorganism may indicate a symbiotic relationship between the microorganism and the host, while other localization sites of a particular microorganism may indicate a parasitic or symbiotic relationship between the microorganism and the host. It is further recognized that the presence of multiple localization sites of a particular microorganism may indicate a systemic infection in the host. Additionally, it is recognized that the localization site of a particular microorganism or pathogen of interest may influence the decision to treat or not treat, and may influence the selection of appropriate treatment options. For example, and not limited by mechanism, fungal pathogens localized to the skin may be treated differently than fungal pathogens localized to the lungs, and bacterial microorganisms localized to cardiac tissue, including but not limited to the cardiac lining, may be treated differently than bacterial microorganisms localized to the blood or bloodstream.

[0131] In some embodiments, the initial sample comprises, consists of, or is substantially composed of circulating tumor or fetal nucleic acids. (See, for example, the analysis of serum or blood-derived nucleic acids, such as circulating tumor or fetal nucleic acids, as described in U.S. Patent Nos. 8,877,442 and 9,353,414, or pathogen identification by, for example, analysis of circulating microbial or viral nucleic acids, as described in U.S. Patent Application No. 2015-0133391 and U.S. Patent Application No. 2017-0016048, the entire disclosure of which is incorporated herein by reference for all purposes). In some embodiments, the initial sample comprises, consists of, or is substantially composed of circulating donor nucleic acids (see, for example, US 20150211070, which is incorporated herein by reference in its entirety, including any figures).

[0132] The initial sample may be derived from any subject (e.g., human subject, non-human subject, etc.). The subject may be healthy. In some embodiments, the subject is a human patient who has a disease or infection, is suspected of having a disease or infection, or is at risk of having a disease or infection. In some embodiments, the disease or infection is pathogen-related.

[0133] Human subjects may be male or female. In some embodiments, the sample may be derived from a human embryo or human fetus. In some embodiments, a human may be an infant, child, youth, adult, or elderly person. In some embodiments, the subject is a female subject who is pregnant, suspected of being pregnant, or planning to become pregnant.

[0134] In some embodiments, the subject is a human subject who has undergone organ transplantation or plans to undergo organ transplantation.

[0135] In some embodiments, the subject is a farm animal, laboratory animal, or domestic pet. In some embodiments, the animal may be an insect, dog, cat, horse, cow, mouse, rat, pig, fish, bird, chicken, or monkey.

[0136] The subject may be an organism, such as a single-celled or multicellular organism. In some embodiments, the sample may be obtained from a plant, fungus, eubacteria, archaea, protist, or any multicellular organism. The subject may be a cultured cell, which may be a primary cell or a cell from an established cell line.

[0137] In some embodiments, the subject has a hereditary disease or condition, is affected by a hereditary disease or condition, or is at risk of having a hereditary disease or condition. Hereditary diseases or conditions can be associated with genetic variations such as mutations, insertions, additions, deletions, translocations, point mutations, trinucleotide repeat disorders, single nucleotide polymorphisms (SNPs), or combinations of hereditary variations.

[0138] In some aspects, the subject is healthy or asymptomatic, or exhibits mild or nonspecific clinical symptoms. In some cases, the subject may be infected with or suspected of being infected with a specific pathogen. In other cases, the subject is suspected of having an infection of unknown origin. In some cases, the subject has been exposed to, or is suspected of being exposed to, a pathogen, such as through living conditions, through travel to a specific geographic area, or through interaction with an infected individual or sexual interaction.

[0139] The initial sample may come from a subject who has a specific disease, condition, or infection, or is suspected of having, or is at risk of having, a specific disease, condition, or infection. For example, the initial sample may come from a cancer patient, a suspected cancer patient, or a patient at risk of having cancer. In some embodiments, the initial sample may come from a patient with an infection, a suspected infection patient, or a patient at risk of infection. In some embodiments, the initial sample comes from a subject who has undergone or will undergo an organ transplant.

[0140] Primer extension reactions can be performed using DNA-dependent polymerases, RNA-dependent polymerases, or reverse transcriptases, or combinations thereof. In some embodiments, primer extension reactions can be performed using DNA or RNA polymerases with strand displacement activity. In some embodiments, primer extension reactions can be performed using DNA or RNA polymerases with non-templatement activity. In some other...In the embodiments, primer extension reactions can be performed by DNA or RNA polymerases with strand displacement activity and DNA or RNA polymerases with non-templatement activity. In some embodiments, primer extension is performed using Klenow fragments.

[0141] Reference fragment length profiles are generally predetermined. One or more suitable reference fragment length profiles may vary depending on the method, the type of comparison, or the purpose of the method. Those skilled in the art will select one or more suitable reference fragment length profiles. Reference fragment length profiles may be obtained from subjects or cells exposed to the compound of interest, subjects or cells exposed to similar compounds, subjects or cells similar to said subjects, subjects or cells with known microorganisms, subjects or cells previously identified as having an infection at the source site, or subjects or cells in any other condition of interest suitable for use as determined by those skilled in the art.

[0142] Subjects with grafts are at risk of transplant rejection, even when therapies for reducing the risk of rejection are provided. Transplant rejection and transplant rejection syndrome are significant, often life-threatening risks to subjects with grafts. Many anti-rejection therapies suppress the subject's immune system, thereby increasing the subject's risk of infection or disease. Therefore, a balance needs to be struck between the use and dosage of anti-rejection therapy. This application provides a method for monitoring the graft status of a subject with a graft. The method includes the following steps: generating a baseline fragment length profile of target nucleic acids within a nucleic acid library or the entire nucleic acid library obtained from a sample obtained from the subject or donor. Target nucleic acids of particular interest in monitoring graft status include, but are not limited to, donor and recipient mitochondrial DNA (mtDNA). The method for monitoring graft status may further include evaluating the abundance of mitochondrial DNA from the graft. Monitoring graft status encompasses monitoring anything related to the graft's status, including but not limited to host rejection of the graft, host immune response to the graft, host response to the graft, graft deterioration, graft health, graft vascularization, graft oxygenation, and graft failure. The baseline fragment length profile may be generated from a donor and / or recipient sample obtained before, during, or after transplantation. The method further includes the following step: generating a second fragment length profile from a sample obtained from the subject, and comparing the second fragment length profile with the baseline fragment length profile. If the second fragment length spectrum differs from the baseline fragment length spectrum, an increased amount of anti-rejection therapy can be administered internally to the subject.

[0143] The methods and systems of this disclosure can be implemented by one or more algorithms. The algorithms can be implemented in software when executed by a central processing unit. The algorithms can, for example, facilitate the enrichment, sequencing and / or detection of pathogens or microorganisms or other target nucleic acids, or the generation of fragment length spectra.

[0144] The compound may include, but is not limited to, chemotherapeutic agents, antiviral agents, antibiotics, antifungals, agents of interest, small molecules, experimental agents, clinical trial compounds, pharmaceuticals, drugs, and active ingredients.

[0145] Toxicity includes, but is not limited to, cytotoxicity. It is further recognized that toxicity may preferentially occur in specific classes of cells, including, but not limited to, cancer cells and pathogens.

[0146] The fragment length profiles and methods of this application can be used for noninvasive prenatal testing (NIPT). The methods allow for noninvasive monitoring, diagnosis, and tracking of fetal condition.

[0147] In some embodiments, isolating the coupled nucleic acid includes immobilizing the coupled nucleic acid, consisting of, or substantially consisting of the immobilized coupled nucleic acid. In some embodiments, immobilization occurs on magnetic beads or functionalized magnetic beads. In some embodiments, immobilization occurs on modified glass, modified capillary surfaces, and / or modified columns. In some embodiments, isolating the coupled nucleic acid includes purifying the coupled nucleic acid, consisting of, or substantially consisting of the purified coupled nucleic acid. In some embodiments, isolating ligated nucleic acids includes precipitating ligated nucleic acids, consisting of precipitated ligated nucleic acids, or consisting substantially of precipitated ligated nucleic acids. In some embodiments, isolating ligated nucleic acids includes using a 3'-end protected 3'-end adaptor, consisting of a 3'-end protected 3'-end adaptor, or consisting substantially of a 3'-end protected 3'-end adaptor. In some embodiments, isolating ligated nucleic acids includes, consisting of, or consisting substantially of: separating ligated nucleic acids from unligated nucleic acids by digesting unligated nucleic acids with a 3'-end exonuclease, wherein the ligated nucleic acids include a 3'-end protected 3'-end adaptor, consisting of a 3'-end protected 3'-end adaptor, or consisting substantially of a 3'-end protected 3'-end adaptor. Some embodiments further include enriching nucleic acids of a certain length, consisting of nucleic acids enriched of a certain length, or consisting substantially of nucleic acids enriched of a certain length. In some embodiments, denaturation is used to further isolate nucleic acids or target nucleic acids. In some embodiments, denaturation includes selective denaturation, consisting of selective denaturation, or consisting substantially of selective denaturation. In some embodiments, selective denaturation includes one or more denaturation steps effective for selecting fragments of a certain length and / or GC content, consisting of one or more denaturation steps effective for selecting fragments of a certain length and / or GC content, or substantially consisting of one or more denaturation steps effective for selecting fragments of a certain length and / or GC content. In some embodiments, the separation of fragments of a certain length can occur using proteases, detergents, heparin, hemolysis, and plasma concentration.

[0148] The methods provided herein include those for subjects who have been infected, subjects at risk of infection, and / or those who have undergone [treatment / treatment].Various non-invasive methods are used to mimic undefined symptoms of a variety of other diseases in subjects. The methods provided herein can be used for a variety of purposes, such as diagnosing or detecting infection, determining the stage of infection, predicting the stage of microbial infection, predicting whether an infection will progress to an invasive disease stage, monitoring the efficacy and / or response to treatment or procedures, discontinuing treatment, identifying infection sites, identifying colonization sites, or modifying or optimizing therapy to achieve better clinical response. Therefore, the methods provided herein can reduce the adverse effects of misdiagnosis or invasive procedures, such as biopsies, used to determine whether a subject's organs are infected, which organs of the subject are infected, and how the subject's organs are infected.

[0149] Figure 1 provides a general overview of some of the methods provided herein. Typically, the method may include: obtaining a clinical sample from an infected subject or a subject at risk of infection; preparing a “spiked sample” by adding synthetic nucleic acids provided in this disclosure (page 29 / 68, 32 CN 121575489 A); optionally, extracting the nucleic acid from the spiked sample; generating a spiked sample library; optionally, enriching the target nucleic acid of interest; performing a detection assay, such as sequencing, to obtain sequence reads from the spiked sample library; and determining the measurement result based on the detected nucleic acid, and comparing this measurement result with a control or reference to determine the subject’s infection stage, the biological relationship between the microorganism and the host, or the localization site (e.g., organ or tissue type). In some cases, a comparison of the absolute abundance of the target nucleic acid with a control or reference may indicate the subject’s infection stage or localization site. In some cases, a comparison of the distribution of fragment lengths of the target nucleic acid with a control or reference may indicate the subject’s infection stage or localization site. In some cases, a comparison of the absolute abundance and distribution of fragment lengths of the target nucleic acid with a control or reference may indicate the subject’s infection stage or localization site.

[0150] The methods provided herein can be applied to any type of nucleic acid present in clinical samples. Figure 2 provides an overview of examples of cell-free methods. Figure 17 provides a schematic diagram of an exemplary infection in a subject. The source of pathogen infection can be, for example, the lungs or any other organ (e.g., brain, skin, heart tissue, stomach, liver, intestine). Cell-free nucleic acids derived from pathogens, such as cell-free DNA, can travel through the bloodstream and can be collected in plasma samples for analysis. Some methods of the cell-free methods provided herein may include: obtaining a clinical sample from an infected subject or a subject at risk of infection; preparing a “spiked sample” by adding synthetic nucleic acids provided herein; isolating the cell-free nucleic acid, optionally, extracting the cell-free nucleic acid from the spiked sample; generating a spiked sample library; optionally, enriching the target nucleic acid of interest; performing a detection assay, such as sequencing assay, to obtain sequence reads from the spiked sample library; and according to the detection assay...The detected cell-free nucleic acid confirms the measurement result, and this measurement result is compared with a control or reference to determine the stage or location of infection in the subject.

[0151] In some cases, the method can be combined with sequencing methods to identify potentially infected organs or tissues, or to rule out the possibility of an infected organ in the subject (see Koh W. et al., “Noninvasive in vivo monitoring of tissue-specific global gene expression in humans”, Proceedings of the National Academy of Sciences (PNAS) 2014: 111 (7361-7366), which is incorporated herein by reference in its entirety for all purposes). Figure 4 provides an example of an organ-site method using cell-free RNA sequencing. Organ-site detection assays can be used when the method of this disclosure or another clinical test has determined that the subject has an infection at a stage of invasive disease. In this case, the method may further include performing one of the organ-site methods provided herein to detect whether the organ has been infected.

[0152] This disclosure also provides methods for individualized treatment of infected subjects or subjects who are susceptible to or at risk of infection (e.g., immunosuppressed, immunocompromised, living under certain conditions, or with genetic variations that increase susceptibility to infection). The individualized treatment provided by this disclosure includes methods for predicting whether an infection will progress to an invasive disease stage, methods for monitoring the efficacy of a therapy for a subject, methods for modifying a treatment regimen based on the subject's response to a therapy, and methods for determining the pathogen's resistance to a particular therapeutic agent or the subject's genetic susceptibility to a given therapeutic agent.

[0153] Nucleic acids produced according to the methods of the present invention can be analyzed to obtain various types of information, including genomics, epigenetics (e.g., methylation), and RNA expression. Methylation analysis can be performed, for example, by converting methylated bases and then sequencing the DNA. RNA expression analysis can be performed, for example, by multinucleotide array hybridization, by RNA sequencing technology, or by sequencing cDNA generated from RNA.

[0154] Sequencing can be performed by any method known in the art. Sequencing methods include, but are not limited to, Maxam-Gilbert-based sequencing, chain termination-based sequencing, shotgun sequencing, bridged PCR sequencing, single-molecule real-time sequencing, ion semiconductor sequencing (e.g., ion torrent sequencing), nanopore sequencing, pyrosequencing (454), sequencing by synthesis, sequencing by ligation (SOLiD sequencing), sequencing by electron microscopy, and dideoxy sequencing reaction. (Instructions for use 30 / 68, page 33, CN 121575489 A)(Sanger method), massively parallel sequencing, polymerase cloning sequencing, and DNA nanosphere sequencing. The term "next-generation sequencing (NGS)" herein refers to a sequencing method that allows for massively parallel sequencing of nucleic acid molecules, during which multiple, for example, millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced simultaneously. Non-limiting examples of NGS include sequencing by synthesis, sequencing by ligation, real-time sequencing, and nanopore sequencing. In some embodiments, sequencing involves: hybridizing primers to a template to form a template / primer duplex; contacting the duplex in a template-dependent manner with a polymerase, in the presence of detectably labeled or unlabeled nucleotides, and allowing the polymerase to add labeled or unlabeled nucleotides to the primers; detecting a signal from the incorporated labeled nucleotides or a signal generated by a process (e.g., proton release) of incorporating labeled or unlabeled nucleotides; and sequentially repeating the contact and / or detection steps at least once, wherein the sequential detection of the incorporated labeled or unlabeled nucleotides determines the sequence of the nucleic acid.

[0155] Exemplary detectable markers include radioactive markers, fluorescent markers, protein markers, dye markers, enzyme markers, etc. In some embodiments, the detectable marker may be an optically detectable marker, such as a fluorescent marker. Exemplary fluorescent markers include cyanin, rhodamine, fluorescein, coumarin, BODIPY, Alexa, or conjugated multi-dyes.

[0156] In some embodiments, sequencing includes obtaining paired end reads, consisting of obtained paired end reads, or substantially consisting of obtained paired end reads. In some embodiments, sequencing includes obtaining consensus reads, consisting of obtained consensus reads, or substantially consisting of obtained consensus reads.

[0157] The accuracy or average accuracy of the sequence information may be greater than about 80%, about 90%, about 95%, about 99%, about 99.98%, or about 99.99%. Sequence accuracy or average accuracy may be greater than about 95% or about 99%. The sequence coverage can be greater than about 0.00001, 0.0001, 0.001, about 0.01, about 0.1, about 0.5, about 0.7, or about 0.9. The sequence coverage can be less than about 200,000, about 100,000, about 10,000, about 1,000, or about 500.

[0158] In some embodiments, the sequence information obtained per nucleic acid template is more than about 10 base pairs, about 15 base pairs, about 20 base pairs, about 50 base pairs, about 100 base pairs, or about 200 base pairs. The sequence information can be obtained in less than 1 month, 2 weeks, 1 week, 2 days, 1 day, 14 hours, 10 hours, 3 hours, 1 hour, 30 minutes, 10 minutes, or 5 minutes.

[0159] Although the examples (hereinafter) use specific sequences for certain sequencing systems, such as the Illumina system, it should be understood that references to these sequences are for illustrative purposes only, and the methods described herein can be configured for use with other sequencing systems incorporating specific initiators, ligations, indexes, and other operational sequences available from Ion Torrent, Oxford Nanopore, Genia Technologies, Pacific Biosciences, Complete Genomics, etc.

[0160] The methods provided herein may include the use of systems, such as systems containing nucleic acid sequencers (e.g., DNA sequencers, RNA sequencers) for generating DNA or RNA sequence information. The system may include a computer comprising software for performing bioinformatics analysis on the DNA or RNA sequence information. Bioinformatics analysis may include, but is not limited to, assembling sequence data, detecting and quantifying genetic variants in a sample, including germline variants and somatic variants (e.g., genetic variants associated with cancer or precancerous conditions, genetic variants associated with infection).

[0161] Sequencing data can be used to determine genetic sequence information, ploidy status, identification of one or more genetic variants, and quantitative measures of variants, including relative and fully relative measures.

[0162] In some cases, genome sequencing involves whole genome sequencing or partial genome sequencing. Sequencing can be unbiased and can involve sequencing all or substantially all nucleic acids (e.g., greater than 70%, 80%, 90%) in the nucleic acids of a sample. Genome sequencing can be selective, e.g., directed to a portion of the genome of interest. Selective sequencing of genes or portions of genes may be sufficient for the desired analysis. Polynucleotides mapped to specific loci in the genome of the subject of interest can be isolated for sequencing by, for example, sequence capture or site-specific amplification (see page 34 of specification 31 / 68, CN 121575489 A).

[0163] After sequencing the aligned reads, the sequence dataset can be uploaded to a data processor for bioinformatics analysis to subtract host or host-related sequences, such as those from humans, cats, dogs, etc., from the analysis; and to determine the presence and prevalence of pathogen or contaminant sequences (e.g., microbial sequences), for example, by comparing the coverage of sequences mapped to microbial reference sequences with the coverage of host reference sequences. Subtraction of host sequences may include the steps of identifying reference host sequences and masking microbial sequences or microbial mimicry sequences present in the reference host genome. Similarly, determining the presence of microbial sequences by comparison with microbial reference sequences may include identifying reference microbial sequences and masking reference microbial sequences.The steps include the presence of host sequences or host-mimicking sequences in the genome sequence.

[0164] The dataset may optionally be cleaned to check sequence quality, remove remnants of sequencer-specific nucleotides (e.g., adaptor sequences), and merge overlapping paired end reads to produce higher quality shared sequences with smaller read errors. Repetitive value sequences may be identified as those with the same start site and length or identical or nearly identical sequences. Optionally, repetitive values ​​may be removed from the analysis.

[0165] In some aspects, host or host-related (e.g., human) sequences may be subtracted from the analysis. In some aspects, host sequences are retained in the analysis. In some aspects, the amplification / sequencing steps may be unbiased, and the advantage of the sequence in the sample will be the host sequence. The subtraction steps may be optimized in several ways to improve the speed and accuracy of the process, for example by performing multiple subtractions at a coarse filter, for example, by setting the initial alignment using a fast aligner, and performing additional alignments using fine filters, such as sensitive aligners or extended reference databases.

[0166] Initially, the dataset of reads can be aligned with a host reference genome containing, but not limited to, Genbank hg19 or Genbank hg38 reference sequences to subtract host DNA in a bioinformatics manner. Each sequence can be aligned with the best set of sequences in the host reference sequences. Sequences identified as host can be removed from the analysis in a bioinformatics manner.

[0167] Removal of host or host-associated sequences can also be optimized by adding contigs with high hit rates, which include, but are not limited to, highly repetitive sequences present in the genome that are not well represented in the reference database. For example, it has been observed that in later stages of the pipeline, when using a database containing a large set of human sequences, such as the entire NCBI NT database, those reads that are significantly incomparable to hg19 or hg38 are ultimately identified as human. Removal of these reads early in the analysis can be performed by constructing an expanded host or host-associated reference. This reference can be created by identifying host contigs with high coverage after the initial host read subtraction from a sequence database, rather than the sequences themselves, such as the NCBI NT database. These contigs can be added to the host reference to create a more comprehensive reference set. Additionally, novel assembled host-related contigs from cohort studies can be used as an additional reference for filtering host-derived reads.

[0168] Regions in the host genome reference sequence containing relevant non-host sequences, such as viral and bacterial sequences integrated into the genome of the reference sample, can be masked.

[0169] Optionally, host or host-related sequences can be identified and removed using non-alignment-based methods, such as sequence identification by sequence characteristics including the frequency of certain motifs, sequence patterns, word frequencies, or nucleotide deviations.

[0170] The sequence reads identified as non-human can then be compared with a nucleotide database of microbial reference sequences. Data sets of microbial sequences known to be associated with a host, such as a set of human symbiotic and pathogenic microorganisms can be selected.

[0171] Microbial databases can be optimized to mask or remove contaminating sequences. For example, many public dataset entries contain artificial sequences not derived from microorganisms, such as primer sequences, host sequences, and other contaminants. It may be desirable to perform an initial alignment or multiple alignments on the database. Regions showing irregularities in read coverage when aligning multiple samples can be masked or removed as artifacts. This irregularity in coverage can be detected by various metrics, such as the ratio between the coverage of a particular nucleotide and the average coverage of the entire contiguous group where that nucleotide is present. Typically, sequences expressed as having an average coverage greater than that of their reference sequence (approximately 5X, 10X, 25X, 50X, or 100X) are likely artificial. Alternatively, a binomial test can be applied to provide per-library coverage likelihood given the overall coverage of the contiguous group. Removing contaminant sequences from the reference database allows for accurate identification of microorganisms.

[0172] Each high-confidence read can be aligned with multiple organisms in a given microbial database. To correctly assign organism abundance based on this possible mapping redundancy, an algorithm can be used to calculate the most probable organism (e.g., see Lindner et al., Nucleic Acids Res. (2013) 41(1): e10). For example, the GRAMMy or GASiC algorithm can be used to calculate the most probable organism from which a given read originates.

[0173] Alignment and assignment with host sequences or with non-host (e.g., microbial) sequences can be performed according to methods generally accepted in the art. For example, a read of 50 nt can be assigned as a match for a given genome if there are no more than 1 mismatch, no more than 2 mismatches, no more than 3 mismatches, no more than 4 mismatches, no more than 5 mismatches, etc., in the read length. Publicly available algorithms can be used for alignment and identification. A non-limiting example of such an alignment algorithm is the bowtie2 program (Johns Hopkins University).

[0174] Then, when determining the incidence of an organism in a sample (e.g., a free nucleic acid sample), these assignments of reads to organisms (e.g., host organisms, non-host organisms, microorganisms, pathogens, etc.) can be aggregated and used to calculate an estimated number of reads assigned to each organism in a given sample. This information can be used to determine the source of a pathogen or contaminant. The analysis can normalize the count of the size of the microbial genome to provide a count of the coverage of the microorganisms.The normalized coverage of each microorganism can be compared with the host sequence coverage in the same sample to address differences in sequencing depth between samples.

[0175] Further, a dataset of sequence-represented microbial organisms in a sample and the incidence of these microorganisms can optionally be aggregated and displayed for immediate visualization, e.g., in the form of a report.

[0176] This disclosure provides normalization methods. In some cases, the methods of this disclosure may include one or more normalization methods. The normalization methods provided by this disclosure allow for efficient and improved measurement of disease-specific, pathogen-specific, or organ-specific nucleic acids detected in a sample.

[0177] The normalization methods of this disclosure typically use spiked synthetic nucleic acids. Spiked synthetic nucleic acids can be used to normalize samples in a variety of different ways. Spiked nucleic acids can be normalized across all samples and all methods for measuring disease-specific nucleic acids, pathogen-specific nucleic acids, or other target nucleic acids. In some cases, the use of spikes can increase the accuracy of calculations of the relative abundance of pathogen nucleic acids (or disease-specific nucleic acids or target nucleic acids) in a sample relative to other pathogen nucleic acids in the sample.

[0178] Typically, one or more species of synthetic nucleic acids at known concentrations can be spiked into each sample. In many cases, the species of synthetic nucleic acids can be spiked at an equimolar concentration for each species. In some cases, the concentrations of the species of synthetic nucleic acids can be different.

[0179] The abundance of nucleic acid species can vary due to inherent biases in sample handling, preparation, and measurement (e.g., detection). After measurement, the efficiency of recovering each length of nucleic acid can be determined by comparing the measured abundance of spiked nucleic acids of each “species” with the amount initially spiked, see specification 33 / 68 pages 36 CN 121575489 A. This yields a “length-based recovery profile.”

[0180] A “length-based recovery profile” can be used to normalize all (or most or some) disease-specific nucleic acids, pathogen nucleic acids, or other target nucleic acids by normalizing the abundance of disease-specific nucleic acids (or pathogen nucleic acids or other target nucleic acids) according to the molecule spiked to the closest length or according to a function fitted to molecules spiked to different lengths.

[0181] This process can be applied to target nucleic acids, such as pathogen-specific nucleic acids, and may result in an estimation of the "raw length distribution of all pathogen-specific nucleic acids" when a sample is spiked. The "raw length distribution of all target nucleic acids" can show the length distribution profile of the target nucleic acids (e.g., pathogen-specific or organ-specific nucleic acids) when the sample is spiked. It is this length distribution that allows the spiked nucleic acids to be generalized to achieve perfect or near-perfect abundance normalization. It is this length distribution that allows the spiked nucleic acids to be generalized to determine the endogenous fragment length distribution of the target nucleic acids.

[0182] Since it is impossible to spike a sample with a mixture of relative abundance profiles of disease-specific nucleic acids, pathogen nucleic acids, or other target nucleic acids in the particular sample that accurately summarizes the profiles of known nucleic acids, partly because the sample may have been used up or the relative abundance profiles may have changed over time, each “species” of the spike can be weighted in proportion to its relative abundance within the “original length distribution of all disease-specific nucleic acids”. The sum of all “weighting factors” can be equal to 1.0.

[0183] Normalization can involve a single step or a series of steps. In some cases, the abundance of disease-specific nucleic acids (or pathogen nucleic acids or other target nucleic acids) can be normalized using the original measurement of the abundance of the spiked nucleic acids that is closest in size to obtain a “normalized disease-specific nucleic acid (or pathogen nucleic acid or other target nucleic acid) abundance”. The “normalized disease-specific nucleic acid abundance” (or pathogen nucleic acid or other target nucleic acid abundance) can then be multiplied by a “weighting factor” to adjust for the relative importance of restoring the length, thereby obtaining a “weighted normalized disease-specific (or pathogen-specific or other target) nucleic acid abundance”. One advantage of this normalization method is that it allows for the comparable measurement of target nucleic acid (e.g., disease-specific nucleic acid, pathogen nucleic acid) abundance across all (or most) methods for measuring disease-specific nucleic acid abundance, regardless of the method.

[0184] This assay may involve measuring the amount of target nucleic acid (e.g., disease-specific nucleic acid) in a biological sample (e.g., plasma) to detect the presence of a pathogen or identify a disease state or determine the target nucleic acid based on the sample, reagent, or environment. The methods described herein can make these measurements comparable across samples, measurement times, methods of nucleic acid extraction, methods of nucleic acid manipulation, methods of nucleic acid measurement, and / or various sample handling conditions.

[0185] This disclosure provides for the measurement of diversity loss values. In some cases, the methods of this disclosure may include determining a diversity loss value.

[0186] The number of duplicate-removed (e.g., copy-removed) SPANK molecules detected in a particular library is replaced by the minimum detectable concentration in said library. This can be used to set a threshold based on the minimum detectable concentration of SPANK molecules in said library. The threshold can be used to ensure sufficient sequencing depth for pathogen detection. The threshold can also be used to ensure that pathogen signals are not due to cross-contamination from other samples. For example, the enrichment of pathogens relative to a threshold set by the SPANK molecule can be compared between different samples. More generally, it is proportional to the efficiency with which the library converts DNA molecules in the original sample into reads in DNA sequencing data.

[0187] The spiked SPANK molecules provided in this disclosure can be used to calculate the diversity loss value. The diversity loss value can be...As shown in Figure 5. In some cases, if the diversity of SPANK sequences is high enough, the SPANK sequences spiked into the sample can be assumed to be substantially unique. Therefore, any repetitive SPANK sequences in the sequencing may be due to PCR amplification rather than to multiple copies of the same SPANK sequence being added to the sample, and can be removed from the analysis manual, page 34 / 68, 37 CN 121575489 A. Additionally, if each SPANK sequence is unique, the total number of SPANK sequences initially added to the sample is known based on the concentration and volume of nucleic acids added to the sample, and the total number of unique SPANK sequencing reads after sequencing is known; these values ​​together can be used to calculate the diversity loss value.

[0188] C: Absolute Abundance (MPM) This disclosure provides a measurement of absolute abundance (also referred to as "molecules per microliter" (MPM)).

[0189] Typically, the absolute abundance of a target nucleic acid (e.g., DNA or RNA) in a sample can be determined by normalizing the number of sequence reads of the target nucleic acid using an empirically determined diversity loss value.

[0190] In some cases, absolute abundance measurements may include nucleic acids of various lengths or single lengths and spiked samples at known concentrations. In some cases, the fraction of information from the sample that is actually observed in the sequencing data may be observed for each spike length (e.g., by comparing observed reads with reads associated with the spiked nucleic acids, or by separating observed reads by the spiked reads). The original number of non-host or pathogen molecules at each length may also be calculated in reverse (e.g., inferred in part from the number of spiked reads at each length). This load may be translated into a measurement of “molecules per microliter”.

[0191] In many cases, methods for detecting molecules per microliter (and others provided herein) may involve removing or isolating low-quality reads. Removing low-quality reads can improve the accuracy and reliability of the methods provided herein. In some cases, the method may include removing or isolating the following (in any combination): unmappable reads, reads generated by PCR repeat values, low-quality reads, adaptor dimer reads, sequencing adaptor reads, reads with non-unique mappings, and / or reads mapped to non-informative sequences.

[0192] In some cases, sequence reads may be mapped to a reference genome, and reads not mapped to such a reference genome may be mapped to one or more target or pathogen genomes. In some cases, reads may be mapped to a human reference genome (e.g., hg19), while the remaining reads are mapped to a select reference database of viruses, bacteria, fungi, and other eukaryotic pathogens (e.g., fungi, protozoa, parasites).

[0193] This disclosure provides methods that can be used to determine whether a measurement provided in this disclosure indicates that a subject is in a certain infection.Various controls and references for infection at stages or at localized sites.

[0194] Typically, the method includes processing a reference or control using the methods of this disclosure. In some cases, control or reference values ​​may be measured as concentrations or the number of sequencing reads. Levels may be qualitative or quantitative. Based on sequence reads from control or reference samples, baseline levels of target nucleic acids (e.g., pathogen species, genetic variants, contaminants introduced from the laboratory environment, or organ-derived contaminants) may be determined.

[0195] In some cases, controls or reference values ​​may be pathogen-dependent. For example, control values ​​for Helicobacter pylori may differ from those for Clostridium difficile. A database of levels or control values ​​may be generated based on samples obtained from one or more subjects, one or more pathogens, and / or one or more time points. This database may be curated or proprietary.

[0196] In some cases, control or reference values ​​are predetermined absolute values ​​that indicate the presence or absence of free pathogen nucleic acids or free organ-derived nucleic acids. Control or reference values ​​may be values ​​obtained by analyzing the free nucleic acid levels of uninfected subjects. In some cases, the control or reference value can be a positive control value and can be obtained by analyzing cell-free nucleic acids from subjects with a specific known infection or a specific organ with a specific known infection.

[0197] In some cases, the control may comprise a group of symbiotic or natural microbiota that cause or do not cause infection, identified using control samples from healthy individuals. Thresholds can be set based on a group of symbiotic microbes in the control sample. Specification 35 / 68 pages 38 CN 121575489 A

[0198] A Poisson model or other statistical models can be used to determine whether a defined baseline level in a clinical sample is significantly higher than that in a reference control. In cases where the sequence reads from a clinical sample are significantly higher than those in a reference control, this indicates that the reads are informative. In some cases, such informative reads can be selected to determine thresholds for two different clinical groups.

[0199] Depending on the level of the target nucleic acid and the background observed across samples, it may be desirable to use one or more references to subtract or filter out sequence reads. Filtering can be combined with selection and can be performed before or after selection. In some embodiments, at least one reference value is based on the level of pathogen nucleic acid detected in one or more samples selected from the group consisting of: water samples, blood samples, plasma samples, serum samples, urine samples, body fluid samples, reagent samples, samples from healthy subjects, or any combination thereof.

[0200] The control value may be the level of free pathogen or free organ-specific nucleic acid obtained from the subject at different time points.

[0201] In some cases, the control value may be at a later testing time point (e.g., after a treatment intervention or at some point after...).Samples are extracted at time points prior to the observation period (after which time has elapsed). In such cases, comparisons of levels at different time points can indicate the presence of infection, infection in a specific organ, improvement in infection, or worsening of infection. For example, an increase in pathogen- or organ-specific cell-free nucleic acids over time can indicate the presence of infection or worsening of infection, for example, an increase of at least 5%, 10%, 20%, 25%, 30%, 50%, 75%, 100%, 200%, 300%, or 400% compared to the original value can indicate the presence of infection or worsening of infection. In other instances, a decrease in pathogen- or organ-specific cell-free nucleic acids compared to the original value of at least 5%, 10%, 20%, 25%, 30%, 50%, 75%, 100%, 200%, 300%, or 400% can indicate the absence of infection or infection improvement (e.g., infection eradication).

[0202] Samples can be extracted over specific time periods, such as daily, every other day, weekly, every other week, monthly, or every other month. For example, an increase of at least 50% in pathogen- or organ-specific cell-free nucleic acids over a week can indicate the presence of infection.

[0203] The method may include determining a threshold or range of values. The threshold may be used to identify samples in a clinical group (colonizing stage vs. invasive disease stage or non-organ infection vs. infected organ). The threshold may be used to identify or select informative sequence reads from clinical samples. Typically, the desired threshold will be one that maximizes the number of true positives while minimizing the number of false positives. In some cases, the threshold may be selected using ROC curve analysis. In some cases, the threshold may be selected based on performance metrics.

[0204] Threshold Selection The threshold may be selected based on its performance using various statistical methods, such as receiver operating characteristic (ROC) curve analysis. ROC analysis may be used to evaluate the classifier’s performance over its entire operating range before selecting a cutoff threshold. To use ROC curves to determine which threshold cutoff values ​​should perform optimally, the threshold may be progressively moved across a range (e.g., 0 to 1.0) to find the cutoff values ​​that reduce the number of false positives and increase the number of true negatives.

[0205] ROC analysis may be performed by plotting data obtained from the methods of this disclosure as follows: TP (sensitivity) vs. FP (1-specificity). Using an ROC plot, a perfect or near-perfect classifier will typically move straight along the Y-axis and then along the X-axis, while a classifier unable to classify samples from different clinical groups will typically sit diagonally. Most classifiers will fall somewhere between these two extremes, and users can choose a threshold based on their most likely or desired performance.

[0206] A threshold can be selected using performance metrics such as accuracy, sensitivity, specificity, positive predictive value, or negative predictive value. In some cases, a single performance metric can be used to select a threshold. In other cases, multiple performance metrics can be used to select a threshold.

[0207] Any threshold applied to the dataset (where PP is the positive population and NP is the negative population) will produce true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN).

[0208] In some cases, accuracy performance metrics can be used to determine the probability of correct classification. Accuracy can be calculated by applying the following equation: (TP + TN) / (PP + NP). In some cases, accuracy is calculated using a trained algorithm.

[0209] In some cases, sensitivity performance metrics can be used to determine the ability of a test to detect disease in a population of infected individuals. Sensitivity percentage can be calculated by applying the following equation: TP / (TP + FN).

[0210] In some cases, specificity performance metrics can be used to determine the ability of a test to correctly exclude disease in a population of healthy individuals. Specificity can be calculated by applying the following equation: TN / (TN + FP).

[0211] When classifying samples for the diagnosis of infection, there are typically four possible outcomes from a binary classifier. If the predicted result is p, and the actual value is also p, it is called a true positive (TP); however, if the actual value is n, it is called a false positive (FP). Conversely, a true negative occurs when both the predicted result and the actual value are n, and a false negative occurs when the predicted result is n and the actual value is p. For tests that detect diseases or conditions such as infections, a false positive can occur in this case when a subject tests positive but is not actually infected. On the other hand, a false negative can occur when a subject actually has an infection but tests negative for that infection.

[0212] The positive predictive value (PPV), or accuracy rate or the post-test probability of a disease, is the proportion of patients with a positive test result who are correctly diagnosed. It can be calculated by applying the following equation: PPV = TP / (TP+FP) × 100. PPV can reflect the probability that a positive test reflects the underlying condition being tested. However, its value can indeed depend on the incidence of the disease, which can vary.

[0213] The negative predictive value (NPV) can be calculated by the following equation: TN / (TN+FN) × 100. The negative predictive value can be the proportion of patients with a negative test result who are correctly diagnosed. PPV and NPV measurements can be estimated using appropriate disease prevalence.

[0214] Thresholds can be set based on the user's desired performance in terms of specificity and sensitivity to differentiate between two clinical groups. In some cases, the specificity of the methods provided in this disclosure can be greater than 70%, 75%, 80%, or 85%.86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%, and the sensitivity can be greater than 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or more.

[0215] The methods provided in this disclosure can be used for a variety of purposes, such as diagnosing or detecting infection, determining the biological relationship between a microorganism and a host, the stage of infection, predicting whether an infection will progress to an invasive disease stage, monitoring the efficacy and response to treatment of an infection, modifying or optimizing a therapy for a better clinical response, and discontinuing treatment or therapy. Therefore, using the methods provided in this disclosure, individualized treatment can be provided to subjects based on data obtained by the methods.

[0216] The pathogens that are expected to cause an infection in a subject have several characteristics, such as, but not limited to, elevated absolute abundance levels compared to an asymptomatic reference or control, abnormal nucleic acid length distribution profiles, or may have two of these characteristics. Similarly, the pathogens in the organs of the subject to be infected have elevated absolute abundance levels, abnormal nucleic acid length distribution profiles, or may have both characteristics compared to asymptomatic references or controls. Pathogens causing infection in the subject may have several characteristics, such as, but not limited to, nucleic acid length distribution profiles comparable to symptomatic references or controls.

[0217] A: Stage of Infection The methods provided in this disclosure can be used to detect, diagnose, treat, monitor, predict, or predict the stage of infection in a subject. Pathogens causing infection may be bacteria, viruses, fungi, parasites, yeasts, or other microorganisms, especially infectious microorganisms. In some cases, the methods can be used to determine whether the subject is in a colonizing or invasive disease stage. Specification 37 / 68 pages 40 CN 121575489 A In some cases, the methods can be used to detect whether the subject is in a gestational stage, prodromal stage, disease stage, declining stage, healing stage, eradication stage, chronic stage, or invasive stage. In some cases, the methods can determine whether the infection is in an active or latent stage.

[0218] The methods of this disclosure can be used in conjunction with other medical tests. For example, the method can be used before or after fecal antigen testing, urea breath testing, serology, urease testing, histology, bacterial culture and susceptibility testing, biopsy, or endoscopy on the subject. In some cases, the method described herein is performed without performing fecal antigen testing, urea breath testing, serology, urease testing, histology, bacterial culture and susceptibility testing, biopsy, or endoscopy on the subject.

[0219] In some cases of the method described herein, the method reduces the risk of infection progressing to an invasive disease stage.The mortality rate was reduced by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In some cases of the methods described herein, the methods reduced mortality and / or complication-related mortality at the invasive disease stage by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%.

[0220] The methods described herein may further include RNA sequencing (RNA-Seq) of cell-free nucleic acids derived from an organ of a subject. Tissue damage caused by infection may result in the release of cell-free nucleic acids from an infected organ or tissue into the bloodstream. Figure 3 depicts an example of the release of cell-free DNA. An increase in organ-derived, for example, cell-free RNA in a sample may indicate that the subject's organ has been infected by a pathogen.

[0221] For example, a method may include analyzing circulating cell-free pathogen nucleic acids from pathogens associated with one or more clinical symptoms. The method may further include performing RNA-Seq to detect an increase in organ-derived cell-free RNA in the blood of a subject. The combination of these test results can indicate that a pathogen has infected the subject and determine which organ of the subject is infected.

[0222] RNA-Seq testing can be performed simultaneously with another clinical method for detecting infection, after the clinical method for detecting infection, or before the clinical method for detecting and infecting. In other cases, RNA-Seq can be used independently to study organ health or can provide increased confidence that an infection detected by another clinical method described herein is an infection of a specific organ.

[0223] In some cases, RNA-Seq testing may be able to determine whether an infection is in an invasive disease stage. In some cases, RNA sequencing tests can be repeated over time to determine whether an infection in a specific organ or tissue is worsening or improving, or whether it is spreading to different organs or tissues of the subject. Similarly, pathogen detection assays provided herein can be repeated over time in combination with organ infection assays.

[0224] RNA-Seq testing (or a series of RNA-Seq tests) can sometimes be performed after the methods described herein produce a positive test result (e.g., detection of pathogen infection). RNA-Seq testing is particularly useful for confirming infection or for identifying the location of infection. For example, the method can detect the presence of pathogens in a subject by analyzing circulating cell-free nucleic acids, but the site of infection may be unclear. In this case, the method may further include sequencing cell-free RNA from the subject to confirm infection within an organ.

[0225] Absolute Abundance of Organ-Specific RNA In some cases, the absolute abundance level of organ-specific RNA sequences can be used as an indicator of the subject's organ.Indicators of pathogen infection. Detection of organ infection may involve comparing the level of organ-specific nucleic acid with a control or reference value to determine the presence or absence of organ nucleic acid and / or the quantity of organ-specific nucleic acid. The level may be qualitative or quantitative. Specification 38 / 68 pages 41 CN 121575489 A

[0226] In some cases, the control or reference value is a predetermined absolute value that indicates the presence or absence of free organ-derived nucleic acid. For example, detecting a level of free pathogen nucleic acid higher than the control value may indicate the presence of infection in the organ, while a level lower than the control value may indicate the absence of infection in the organ.

[0227] The control value may be a value obtained by analyzing the level of free nucleic acid in uninfected subjects (e.g., healthy controls). In some cases, the control value may be a positive control value obtained by analyzing free nucleic acid from subjects with a specific infection or a specific organ with a specific infection.

[0228] The control or reference value may be measured as a concentration or the number of sequencing reads. The control or reference value may be pathogen-dependent, organ-dependent, or both pathogen-dependent and organ-dependent. A database of levels or control values ​​can be generated based on samples obtained from one or more subjects, one or more pathogens, and / or one or more time points. This database can be curated or proprietary.

[0229] In some embodiments, control or reference absolute abundance values ​​indicate the presence or absence of a localization site in the subject's body. For example, detecting an absolute abundance level of free pathogen nucleic acid higher than the control or reference value may indicate infection in an organ, while an absolute abundance value lower than the control or reference value may indicate infection not in an organ. In some cases, detecting an absolute abundance level of free pathogen nucleic acid higher than the control or reference value may indicate infection in an organ, while an absolute abundance value lower than the control or reference value may indicate infection not in an organ.

[0230] Distribution of fragment lengths of organ-specific RNA In some cases, the distribution of fragment lengths of organ-specific RNA sequences indicates that the subject's organ is infected with a pathogen.

[0231] For example, detecting an abnormal distribution of free organ-specific nucleic acid may indicate that the organ is infected, while a normal distribution of free organ-specific nucleic acid may indicate that the organ is not infected.

[0232] The control fragment length distribution can be predetermined by analyzing the free nucleic acid levels in uninfected subjects (e.g., healthy controls) in organs. The control fragment length distribution can be obtained in parallel by analyzing infection-independent free nucleic acid levels in subjects with organ infections.

[0233] In some embodiments, the control or reference distribution of fragment lengths indicates the presence or absence of a localization site. For example, detecting an abnormal distribution of free pathogen nucleic acids can indicate infection in an organ, while a normal distribution of free pathogen nucleic acids is present.This can indicate that the infection is not in the organ. In some cases, the detection of an abnormal distribution of free pathogen nucleic acid can indicate that the infection is in the organ, while the normal distribution of free pathogen nucleic acid can indicate that the infection is not in the organ.

[0234] Thresholds or ranges of values ​​for organ-specific RNA In some cases, threshold cutoff values ​​can be used as an indicator that a subject's organ is infected with the pathogens provided herein. Threshold cutoff values ​​can be determined as provided herein by using organ-specific RNA sequences from a subject infected with the pathogen and comparing these with controls or references.

[0235] In some cases, the accuracy of identifying a sample as an infected organ is greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or more. In some cases, the sensitivity of a sample to be identified as an infected organ is greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or more. In some cases, the specificity of a sample to be identified as an infected organ is greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or more.

[0236] In some cases, the positive predictive value of a sample to be identified as an infected organ is at least 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5% or more. In some cases, the sample is identified as an infected organ with a negative predictive value of at least 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5% or more, as per the specification on pages 39 / 68 of CN 121575489 A.

[0237] In some cases, the sample is identified as an infected organ with a sensitivity greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or more, and a specificity greater than 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or more.

[0238] B: Personalized Treatment and Monitoring This disclosure also provides for personalized treatment of infected subjects or those susceptible to or at risk of infection (e.g., ...).For example, methods for subjects with immunosuppressed, immunocompromised, living conditions, or genetic variations that increase susceptibility to infection. Personalized treatment may include predicting whether an infection will progress to an invasive disease stage, monitoring the efficacy of a subject's therapy, modifying the treatment regimen based on the subject's response to the therapy, and determining the pathogen's resistance to a specific therapeutic agent.

[0239] In some cases, the methods may be used to detect, diagnose, predict, or prognose pathogen resistance to a specific therapeutic agent. In some cases, the methods may further include sequencing the subject's DNA against genetic variations associated with resistance to a therapeutic agent or to a specific therapeutic agent.

[0240] In some cases, samples may be collected sequentially at various times before or during the course of infection to determine the pathogen and the subject's response to treatment, thereby providing a personalized regimen. In some cases, sequentially collected samples are compared to each other to determine whether the subject's infection is improving or worsening.

[0241] Treatment may involve administering drugs or other therapies to reduce or eliminate colonizing or invasive disease associated with the infection. In some cases, prophylactic treatment may be given to the subject to prevent the development of infection. Any medical procedure or treatment involving the administration of a drug can be used to improve or reduce symptoms of infection. Some non-limiting exemplary drugs that can be used are antibiotics (e.g., ampicillin, sulbactam, penicillin, vancomycin, gentamicin, aminoglycoside, clindamycin, cephalosporin, metronidazole, timentin, ticarcillin, clavulanic acid, cefoxitin), antiretroviral drugs (e.g., highly active antiretroviral therapy (HAART), reverse transcriptase inhibitors, nucleoside / nucleotide reverse transcriptase inhibitors (NRTI), non-nucleoside RT inhibitors and / or protease inhibitors), or immunoglobulins.

[0242] This disclosure also provides methods for modulating a treatment regimen. For example, the subject may have already been administered a drug for treating an infection. The methods provided herein can be used to track or monitor the efficacy of drug treatment. In some cases, the treatment regimen may be adjusted depending on the rise or fall of the infection. For example, if the methods provided herein indicate that the infection cannot be improved with drug treatment, the treatment regimen may be adjusted by changing the type of drug or treatment, discontinuing drug use, continuing drug use, increasing drug dosage, or adding a new drug or therapy to the subject's treatment regimen.

[0243] In some cases, the treatment regimen may involve specific procedures. For example, in some cases, the method may indicate the need for surgical procedures or invasive diagnostic procedures, such as tumor removal or biopsy to determine if an organ is infected. Similarly, if the method indicates that the infection is improving or has subsided through treatment intervention, adjusting the treatment regimen may involve reducing or discontinuing treatment. In other cases, no treatment regimen may be given, and instead, a “wait and see” or “observe and wait” approach may be used to observe whether the infection is cleared without any further medical intervention.

[0244] The methods of this disclosure may include detecting pathogens in a subject. In some cases, the method may include whole-genome sequencing of a sample. In some cases, the method may include targeted sequencing of a sample, wherein specific primers are used to detect a specific pathogen of interest. Typically, a pathogen may have a recommended treatment cycle. For example, a treatment cycle for Helicobacter pylori is shown in Figure 6. The methods provided in this disclosure may be used at any stage of the treatment cycle.

[0245] The methods of this disclosure may be applied to any pathogen having various stages of infection. The methods described herein are particularly applicable to pathogens having a colonization phase and an invasive disease phase. In some cases, the invasive disease phase may be caused by pathogen infection. In some cases, the invasive disease phase may be associated with pathogen infection.

[0246] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, monitoring, predicting, or preventing colonization of Helicobacter pylori (H. pylori). H. pylori colonization may be asymptomatic. In some cases, colonization may present as acute gastritis with abdominal pain (stomachache) or nausea. This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive Helicobacter pylori disease. Subjects with invasive Helicobacter pylori disease may develop complications such as chronic gastritis, peptic ulcer disease, gastric adenocarcinoma, gastric cancer, and / or lymphoma.

[0247] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of Clostridium difficile (CDI). CDI may be present in asymptomatic or symptomatic forms. The clinical spectrum of CDI infection can range from mild to moderate, severe, or complicated. Subjects with mild to moderate CDI may experience diarrhea, colitis, including fever, leukocytosis, and / or cramps. The severity of abdominal and systemic symptoms of CDI may increase with the severity of the infection. The methods described can be used to detect, monitor, diagnose, prognose, treat, or prevent invasive CDI disease. Subjects with complicated or invasive CDI disease may develop pseudomembranous colitis, toxic megacolon, colonic perforation, and / or sepsis.

[0248] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of Haemophilus influenzae. Typically, Haemophilus influenzae colonizes the upper respiratory tract of a subject. This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive Haemophilus influenzae disease. Subjects with invasive Haemophilus influenzae disease may develop complications such as sepsis and / or meningitis.

[0249] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization of Salmonella. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive Salmonella disease. Some non-limiting examples of Salmonella serotypes associated with invasive diseases include, but are not limited to, typhus, enteritis, Heidelberg, Dublin, paratyphoid A, swine cholera, and Schwarzkopf. Subjects with invasive Salmonella disease may develop bacteremia, meningitis, enteric fever, and / or invasive non-typhoidal Salmonella (iNTS) disease.

[0250] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of Streptococcus pneumoniae. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive Streptococcal pneumoniae disease. Subjects with invasive pneumococcal disease may develop bacteremia and / or meningitis.

[0251] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization of cytomegalovirus (CMV). Subjects infected with CMV may be asymptomatic because the virus can circulate to a dormant phase. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive CMV disease. Subjects with invasive CMV disease may develop complications in their eyes, lungs, and / or digestive system.

[0252] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of human papillomavirus (HPV). Subjects with HPV colonization may present with noninvasive cervical intraepithelial neoplasia and / or genital warts. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing invasive HPV diseases. Subjects with invasive HPV diseases may develop cervical cancer, anal squamous cell carcinoma, and / or anal carcinoma in situ.

[0253] This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of Epstein-Barr virus (EBV). Subjects colonized with EBV may be asymptomatic or present with fatigue, fever, sore throat, swollen cervical lymph nodes, splenomegaly, hepatomegaly, and / or rash. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive EBV diseases. Subjects with invasive EBV diseases may develop infectionIndividuals with mononucleosis (e.g., glandular fever), certain autoimmune diseases, or those who may develop cancers such as Hodgkin's lymphoma, Burkitt's lymphoma, gastric cancer, nasopharyngeal carcinoma, hairy leukoplakia, and / or central nervous system lymphoma may have a higher risk.

[0254] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of hepatitis B virus (HBV). HBV infection can be transient or chronic. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive diseases associated with HBV infection. Subjects with invasive HBV disease may develop cirrhosis, hepatocellular carcinoma, liver infection, and / or liver failure.

[0255] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of hepatitis C virus (HCV). HCV infection can be acute or chronic. Typically, HCV colonization can be asymptomatic. When signs and symptoms are present, they may include jaundice, along with fatigue, nausea, fever, and muscle pain. Some subjects may have spontaneous viral clearance, while others may progress to a chronic phase. However, in cases where HCV infection becomes chronic, it can lead to invasive HCV disease. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive HCV disease. Subjects with invasive HCV disease may develop cirrhosis, hepatocellular carcinoma, liver infection, and / or liver failure.

[0256] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing colonization of human T-cell lymphoma virus 1 (HTLV-1). HTLV-1 infects the T cells of a subject. Subjects infected with HTLV-1 may be asymptomatic for many years. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive HTLV-1 disease. Subjects with invasive HTLV-1 disease may develop T-cell (ATL) leukemia, HTLV-1-associated myelopathy / tropical spastic paraplegia (HAM / TSP) or other cancers.

[0257] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating or preventing colonization of gonorrhea. Subjects with colonization may be asymptomatic, while other subjects may exhibit symptoms such as burning urination, testicular or pelvic pain and / or discharge from the genitals. This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting or preventing invasive gonorrhea. Subjects with invasive gonorrhea may develop skin lesions, joint infections (e.g., joint pain and swelling), endocarditis and / or meningitis.

[0258] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating or preventing colonization of syphilis. Syphilis infectionInfection can be divided into first, second, latent, and third stages. Subjects in the first stage may experience aches and pains. Subjects in the second stage may experience skin rashes, swollen lymph nodes, and / or fever. During the latent or asymptomatic stages of syphilis, subjects are usually asymptomatic. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive syphilis. Subjects with third-stage or invasive syphilis may develop complications in other organ systems, including but not limited to the heart, blood vessels, brain, and / or nervous system.

[0259] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing trichomoniasis colonization. Subjects with colonization may be asymptomatic or may develop inflammation in their genital areas. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive trichomoniasis. Subjects with invasive trichomoniasis may develop cervical cancer and / or prostate cancer.

[0260] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization of human herpesvirus 8 (HHV-8), also known as Kaposi sarcoma-associated herpesvirus or KSHV. Healthy subjects with colonization are usually asymptomatic. However, subjects with weakened immune systems may develop invasive HHV-8 disease. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive HHV-8 disease. Subjects with invasive HHV-8 disease may develop Kaposi sarcoma and / or several lymphoproliferative disorders, such as primary exudative lymphoma, multicentric Castleman disease, or B-cell lymphoma.

[0261] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing colonization of Merkel cell polyomavirus. Subjects with colonization may be asymptomatic. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive Merkel cell polyomavirus disease. Subjects with invasive Merkel cell polyomavirus disease may develop Merkel cell carcinoma (MCC) tumors—a rare but aggressive form of skin cancer.

[0262] This disclosure provides methods for detecting, monitoring, diagnosing, prognosing, treating, or preventing chlamydia colonization. Subjects with colonization may be asymptomatic or may experience a burning sensation during urination or discharge from the genitals. This disclosure also provides methods for detecting, monitoring, diagnosing, prognosing, treating, predicting, or preventing invasive chlamydia disease.Treated chlamydia can progress to the invasive disease stage, spreading to the uterus and / or fallopian tubes of female subjects. Subjects with invasive chlamydial disease may develop pelvic inflammatory disease (PID), which can lead to chronic pelvic pain, infertility, and ectopic pregnancy.

[0263] In some cases, subjects are infected with or at risk of chlamydia infection at different stages of infection, such as the colonization stage and the invasive disease stage. Colonized subjects may not have clinical signs or symptoms. In other cases, colonized subjects may have clinical signs or symptoms. Subjects with invasive disease may exhibit clinical signs or symptoms. In other cases, subjects with invasive disease may exhibit no clinical signs or symptoms.

[0264] The subject may have another disease or condition or be at risk of having another disease or condition. For example, the subject may have cancer (e.g., breast cancer, lung cancer, stomach cancer, hematologic cancer), be at risk of having cancer, or be suspected of having cancer.

[0265] In some cases, the risk factors for the subject having an infectious disease or progressing to the invasive disease stage may be increased. In some cases, the risk factors are associated with living conditions. Some non-limiting examples of risk factors associated with living conditions include, but are not limited to, crowded living conditions, lack of reliable sources of clean water, living in or visiting a developing country, and / or cohabiting with an infected person.

[0266] In some cases, the risk factor for infection or progression to an invasive disease is a genetic variant of the subject's genomic DNA. Genetic variants that may be risk factors for infection include, but are not limited to, single nucleotide polymorphisms, deletions, insertions, etc. In some other cases, the subject may have a family history of diseases such as gastric cancer, lymphocytic gastritis, hyperplastic gastric polyps, or vomiting during pregnancy.

[0267] The subject may have another disease or be co-infected with more than one pathogen or be at risk of having another disease or being co-infected with more than one pathogen. In some cases, the subject is immunosuppressed (e.g., an organ transplant recipient). In some cases, the subject is immunocompromised (e.g., treated with chemotherapy, immunodeficiency caused by AIDS or a general disease such as diabetes or lymphoma).

[0268] In some cases, the subject may exhibit one or more clinical symptoms. Non-limiting examples of clinical symptoms may include abdominal pain or burning, abdominal pain that worsens upon emptying the tailbone, nausea, loss of appetite, frequent belching, bloating of the stomach area, weight loss, severe or persistent abdominal pain, dysphagia, bloody or black tarry stools and / or bloody or black vomit. Other clinical symptoms are known in the art.

[0269] In some cases, the subject may exhibit clinicopathological features such as atrophic gastritis, acute or chronic gastritis,Hyperacidity, antigenic stimulation, active peptic ulcer disease, history of PUD, low-grade gastric mucosa-associated lymphoid tissue lymphoma, history of endoscopic resection of early gastric cancer, dyspepsia, Barrett's esophagus, unexplained iron deficiency, or idiopathic thrombocytopenic purpura (ITP).

[0270] Subjects may be infected with any type of pathogen or microorganism, including bacteria, viruses, fungi, parasites, prokaryotes, eukaryotes, etc. In some cases, the pathogen is known, while in others it may be known to be symbiotic.

[0271] In some cases, subjects may have an active or latent infection. In some cases, subjects are infected, but the infection is below the diagnostic sensitivity level of other tests previously performed on the subject. In some cases, subjects are infected but asymptomatic, or the infection is at a subclinical level.

[0272] In some cases, the subject may have previously received treatment or may have been treated with drugs or medical procedures such as antimicrobial, antibacterial, antiviral, and / or antiparasitic drugs. In some cases, the subject may not have undergone biopsy, endoscopy, colonoscopy, blood culture, or other such procedures prior to using the methods described herein. In some cases, the subject may have undergone or may have undergone fecal antigen testing, urea breath testing, serology, urease testing, histology, bacterial culture and susceptibility testing, biopsy, or endoscopy prior to using the methods described herein.

[0273] This disclosure provides methods for determining the stage or site of infection in a subject using nucleic acids obtained from clinical samples (e.g., blood, serum, cells, or tissues). In some embodiments, the method includes preparing a spiked sample by adding synthetic nucleic acids provided in this disclosure; extracting nucleic acids from the spiked sample; generating a spiked sample library; enriching the spiked sample library of target nucleic acids of interest; performing sequencing assays to obtain sequence reads from the spiked sample library; and determining a measurement result based on the detected nucleic acids (e.g., DNA, RNA, cell-free DNA, or cell-free RNA) and comparing this measurement result with a control or reference to determine the stage of infection or localization site (e.g., organ or type of tissue) of the subject.

[0274] Embodiments of the method may include extracting nucleic acids or target nucleic acids from a sample or purifying nucleic acids or target nucleic acids from unwanted components in a reaction mixture (e.g., ligation, amplification, restriction enzymes, end-repair, etc.). Any apparatus known in the art for extracting nucleic acids may be used in the method of this application.

[0275] Extraction may include separating nucleic acids from other cellular components and contaminants that may be present in the sample.Liquid extraction (e.g., Trizol, DNAzol) techniques are used to extract nucleic acids from samples. In some cases, extraction is performed by means of phenol-chloroform extraction or precipitation with organic solvents (e.g., ethanol or isopropanol). In some cases, extraction is performed using nucleic acid binding columns.

[0276] In some cases, extraction is performed using commercially available kits, such as the Qiagen Qiamp Cyclic Nucleic Acid Kit, Qiagen Qubit dsDNA HS Assay Kit, Agilent™ DNA 1000 Kit, TruSeq™ Sequencing Library Preparation, QIAamp Cyclic Nucleic Acid Kit, Qiagen DNeasy Kit, QIAamp Kit, Qiagen Midi Kit, QIAprep Spin Kit) or nucleic acid binding spin columns (e.g., Qiagen DNA mini-prep Kit). In some cases, extraction of free nucleic acids may involve filtration or ultrafiltration.

[0277] Nucleic acids can be extracted or purified by using magnetic beads. For example, magnetic beads with an iron oxide core and a surface coated with molecules containing free carboxylic acids or synthetic polymers can be used. The salt concentration or polyalkylene glycol can be adjusted to control the strength of the bond between the functional group and the nucleic acid, thereby allowing controlled and reversible binding. Finally, the nucleic acid can be released from the magnetic beads using an elution buffer. In some cases, extraction or purification is performed using commercially available kits such as the Omega Biotek Mag-Bind® Magnetic Bead Kit, Agencourt®, RNAClean®, and / or XP Magnetic Beads.

[0278] The method may include purifying the target nucleic acid. Purification may be performed where the user desires to separate the target nucleic acid from unwanted components in the reaction mixture. Non-limiting exemplary purification methods include ethanol precipitation, isopropanol precipitation, phenol-chloroform purification and column purification (e.g., affinity-based column purification), dialysis, filtration, or ultrafiltration. Specification 44 / 68 pages 47 CN 121575489 A

[0279] Methods for generating nucleic acid libraries are known in the art.

[0280] Computer Control System This disclosure provides a computer control system programmed to implement the methods of this disclosure. Figure 7 illustrates a computer system 201 programmed or otherwise configured to implement the methods of this disclosure.

[0281] The computer system 201 includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") 205, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 201 also includes memory or memory location 210 (e.g., random access memory, read-only memory, flash memory), electronic storage unit 215 (e.g., hard disk), and communication interface 220 for communicating with one or more other systems (e.g., [example]).For example, a network adapter) and peripheral devices 225, such as cache, other memory, data storage areas, and / or electronic display adapters. Memory 210, storage unit 215, interface 220, and peripheral devices 225 communicate with CPU 205 via a communication bus (solid line), such as a motherboard. Storage unit 215 may be a data storage unit (or data repository) for storing data. Computer system 201 may be operatively coupled to computer network (“network”) 230 by means of communication interface 220. Network 230 may be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some cases, network 230 is a telecommunications network and / or a data network. Network 230 may include one or more computer servers that can implement distributed computing, such as cloud computing. In some cases, network 230 may implement a peer-to-peer network by means of computer system 201, which can enable devices to be coupled to computer system 201 to act as clients or servers.

[0282] CPU 205 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 210. The instructions may relate to CPU 205, which may subsequently be programmed or otherwise configured to implement the methods of this disclosure. Examples of operations performed by CPU 205 may include fetching, decoding, executing, and writing back.

[0283] CPU 205 may be a circuit, such as part of an integrated circuit. One or more other components of system 201 may be contained in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0284] Storage unit 215 may store files, such as drivers, libraries, and saved programs. Storage unit 215 may store user data, such as user preferences and user programs. In some cases, computer system 201 may include one or more additional data storage units located outside computer system 201, such as located on a remote server communicating with computer system 201 via an intranet or the Internet.

[0285] Computer system 201 can communicate with one or more remote computer systems via network 230. For example, computer system 201 can communicate with a user's (e.g., a healthcare provider's) remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. Users can access computer system 201 via network 230.

[0286] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of computer system 201, such as memory 210 or electronic storage unit 215. The machine-executable or machine-readable code may be provided in software form. During use, the code may be executed by processor 205. In some cases, the code may be retrieved from storage unit 215 and may be stored on memory 210 for immediate access by processor 205. In some cases, electronic storage unit 215 may be excluded, and machine-executable instructions may be stored on memory 210.

[0287] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or it may be compiled during operation. The code may be supplied in a programming language, which may be selected to enable execution of the code in a pre-compiled or as-compiled manner. Specification 45 / 68 pages 48 CN 121575489 A

[0288] Aspects of the systems and methods provided herein, such as server 201, may be embodied in programming. Various aspects of a technology can be considered as a “product” or “article of manufacture” typically in the form of machine (or processor) executable code and / or associated data, which is executed on or implemented in a type of machine-readable medium. Machine-executable code can be stored in electronic storage units such as memory (e.g., read-only memory, random access memory, flash memory) or hard disks. “Storage” type media can include any or all tangible memory or associated modules of a computer, processor, etc., such as various semiconductor memories, magnetic tape drives, hard disk drives, etc., that can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. Such communication, for example, enables the loading of software from one computer or processor to another, such as from a management server or host computer to a computer platform for an application server. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves used, such as those used via wired and optical terrestrial networks and physical interfaces between local devices over various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., can also be considered as media carrying software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine-readable medium refer to any medium involved in providing instructions to a processor for execution.

[0289] Therefore, machine-readable media (such as computer executable code) can take many forms, including but not limited to...Tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical discs or magnetic disks, any storage device in one or more storage devices in any computer, such as those used to implement the database shown in the accompanying drawings. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables, copper wires, and optical fibers, including conductors that include buses within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, floppy disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cards, paper tapes, any other physical storage media with a perforated pattern, RAM, ROM, PROM and EPROM, flash-EPROM, any other memory chip or cartridge, a carrier wave for transmitting data or instructions, a cable or link for transmitting such a carrier wave, or any other medium from which a computer can read program code and / or data. Many of these forms of computer-readable media may involve carrying one or more sequences of one or more instructions to a processor for execution.

[0290] Computer system 201 may include or be in communication with an electronic display 235, the electronic display including a user interface (UI) 240 for providing report output, the report output may include a diagnosis of a subject or a treatment intervention for a subject. Examples of UI include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. Analysis may be provided in report form. Reports may be provided to subjects, healthcare professionals, laboratory workers, or other individuals.

[0291] The methods and systems of this disclosure may be implemented by one or more algorithms. Algorithms may be implemented in software when executed by central processing unit 205. Algorithms may, for example, facilitate the enrichment, sequencing, and / or detection of pathogen nucleic acids.

[0292] Information about a may be input into the computer system, such as patient identifiers, such as information about the stage or risk of infection, patient background, patient medical history, previous infection, or ultrasound scans. Patient identifiers may be separated from clinical samples to obtain deidentified samples, for example, by the sample sender or sample recipient. Patient identifiers can be replaced with login numbers or other non-personally identifiable codes. Clinical samples can be sequenced using high-throughput sequencers. The deidentified sample sequence data generated by the sequencer can be uploaded to a server, such as a cloud server. Using the methods disclosed herein, pathogen nucleic acids within deidentified samples can be detected to obtain deidentified result data. Deidentified results can be downloaded from the server. (Instructions 46 / 68, pages 49, CN 121575489 A)Results data. Deidentified results data may be associated with patient identifiers, such as by the sample sender or sample recipient.

[0293] Electronic reports may be generated to indicate the stage of infection of the pathogen. Electronic reports may be generated to indicate prognosis. Electronic reports may be generated to indicate diagnosis. If the electronic report indicates the presence of a treatable infection, an electronic report may be generated to prescribe a treatment regimen or treatment plan. Computer systems may be used to analyze the results from the methods described herein, report the results to the patient or physician, or propose a treatment plan.

[0294] The kit also provides reagents and kits thereof for carrying out one or more of the methods described herein. The reagents and kits thereof of the present invention may vary considerably. The reagents of interest contain reagents specifically designed for the identification, detection, and / or quantification of one or more pathogen nucleic acids in samples obtained from subjects infected with or at risk of infection by the pathogen.

[0295] The kit may include reagents necessary for performing nucleic acid extraction and / or nucleic acid detection using the methods described herein, such as PCR and sequencing. The kit may further include software packages for data analysis, which may contain a reference spectrum for comparison with test spectra from clinical samples, and specifically may contain a reference database. The kit may include reagents such as buffers and water.

[0296] Such kits may also contain information indicating or establishing the activity and / or benefits of the composition and / or describing dosage, administration, side effects, drug interactions, such as scientific literature references, packaging inserts, clinical trial results and / or summaries of these, or other information useful to healthcare providers. Such kits may also include instructions for accessing databases. Such information may be based on the results of various studies, such as studies using laboratory animals involving in vivo models and studies based on human clinical trials. The kits described herein may be offered, sold and / or promoted to healthcare providers including physicians, nurses, pharmacists, prescribing officers and others. In some embodiments, the kits may also be sold directly to consumers.

[0297] It will be understood that references to the following examples are for illustrative purposes only and do not limit the scope of the claims.

[0298] The present invention provides embodiments including but not limited to the following: 1. A fragment length spectrum from a nucleic acid library, wherein the nucleic acid library is generated from an initial sample, and wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library, wherein the fragment length spectrum includes one or more characteristics selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, two or moreThe ratio of the maximum amplitude of each segment and the fragment length distribution within a subset of reads.

[0299] 2. A method for generating a fragment length spectrum of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample using a bias-corrected recovery method; (b) determining the number of reads of multiple fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of the maximum amplitude of two or more segments, and fragment length distribution within a subset of reads; and (d) generating a fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics.

[0300] 3. A method for generating a fragment length spectrum of a nucleic acid library, the method comprising the following steps: (a) preparing a nucleic acid library from an initial sample, the preparation of the nucleic acid library from the initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library; (b) determining the number of reads of multiple fragment lengths within the nucleic acid library; (c) Determine one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads; and (d) Generate a fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics.

[0301] 4. The method according to Embodiment 3, wherein generating the nucleic acid library from the initial sample comprises, consists of, or substantially comprises: (a) dephosphorylating the nucleic acid from the initial sample to generate a set of dephosphorylated nucleic acids; (b) denaturing the dephosphorylated nucleic acids to generate denatured nucleic acids; (c) ligating a 3' end adaptor to the denatured nucleic acids to generate adapted nucleic acids; (d) isolating the adapted nucleic acids; (e) attaching primers to the adapted nucleic acids and expanding the primers with a polymerase to generate complementary strands; (f) ligating a 5' end adaptor; (g) eluting the strands; and(h) Amplify the complementary strand.

[0302] 5. The method according to Embodiment 2, wherein the number of reads is a normalized number of reads.

[0303] 6. The method according to Embodiment 2, wherein the fragment length spectrum is used for at least a subset of reads, and the method further comprises: (a) identifying at least a subset of reads in the nucleic acid library; and (b) determining the fragment length spectrum within the at least a subset of reads.

[0304] 7. The method according to Embodiment 2, wherein the step of generating at least one fragment length spectrum further comprises using two or more fragment length characteristics.

[0305] 8. A method for identifying microorganisms present in a sample, the method comprising the steps of: (a) generating a fragment length spectrum of a nucleic acid library generated from the sample; (b) comparing the fragment length spectrum with a reference fragment length spectrum of one or more microorganisms; and (c) if the fragment length spectrum from the sample is similar to the reference fragment length spectrum of the microorganism, then identifying the microorganism as present in the sample.

[0306] 9. The method according to Embodiment 8, wherein generating the fragment length spectrum of the nucleic acid library comprises the following steps: (a) preparing a nucleic acid library from an initial sample, the preparation of the nucleic acid library from the initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample before the preparation of the nucleic acid library; (b) quantifying the number of reads of multiple fragment lengths within the nucleic acid library; (c) Determine one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads; and (d) Generate a fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics.

[0307] 10. The method according to embodiment 8, wherein the fragment length spectrum indicates that the microorganism exists in the form of a pathogen or a symbiotic microorganism.

[0308] 11. The method according to embodiment 8, wherein the fragment length spectrum includes at least one fragment length characteristic selected from the group consisting of: fragment count ratio of two or more peaks and fragment length distribution shape.

[0309] 12.A method for identifying a localization site of a subject, the method comprising the steps of: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample; (b) comparing the fragment length spectrum with a reference fragment length spectrum of one or more source sites; and (c) identifying the first site as a localization site if the fragment length spectrum from the sample is similar to the fragment length spectrum from a first source site; and identifying the second site as a localization site if the fragment length spectrum from the sample is similar to the fragment length spectrum from a second source site.

[0310] 13. The method according to embodiment 12, wherein generating the fragment length spectrum of the nucleic acid library comprises the steps of: (a) preparing a nucleic acid library from an initial sample, the preparation of the nucleic acid library from the initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library; and (b) quantifying the number of reads of multiple fragment lengths within the nucleic acid library; (c) Determine one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads; and (d) Generate a fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics.

[0311] 14. The method according to embodiment 12, wherein the location site is selected from the group consisting of: deep tissue, blood flow, skin, lungs, heart, brain, and blood.

[0312] 15. A method for monitoring graft status in a subject with a graft, the method comprising the steps of: (a) generating a baseline fragment length spectrum of a nucleic acid library generated from a sample obtained from the subject; (b) generating a second fragment length spectrum of a nucleic acid library generated from a second sample obtained from the subject; (c) comparing the second fragment length spectrum with the baseline fragment length spectrum; If the second fragment length spectrum differs from the baseline fragment length spectrum, an increased amount of anti-rejection therapy is administered internally, wherein the risk of rejection in the subject with the graft is reduced after administration of the anti-rejection therapy; if the second fragment length spectrum is similar to the baseline fragment length spectrum, the anti-rejection therapy is maintained or reduced, wherein the anti-rejection therapy...The side effect risk of the therapy in a patient is lower than the side effect risk of a patient receiving an increased dose of the anti-rejection therapy.

[0313] 16. A method for monitoring the toxicity of a compound administered to a subject, the method comprising the steps of: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample; and (b) comparing the fragment length spectrum with one or more reference fragment length spectra.

[0314] 17. The method of embodiment 16, wherein the one or more reference fragment length spectra are generated from a nucleic acid library obtained from a subject or cells exposed to the compound.

[0315] 18. The method of embodiment 16, wherein the subject has cancer, is at risk of developing cancer, or exhibits cancer-related symptoms.

[0316] 19. The method of embodiment 16, wherein the compound is a chemotherapeutic agent.

[0317] 20. The method according to embodiment 16, wherein generating the fragment length spectrum of the nucleic acid library comprises the following steps: (a) preparing a nucleic acid library from an initial sample using a bias-corrected recovery method; (b) determining the number of reads of multiple fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length distribution within a subset of reads; and (d) generating the fragment length spectrum of the nucleic acid library using the one or more fragment length characteristics.

[0318] 21. The method according to embodiment 16, wherein generating the fragment length spectrum of the nucleic acid library comprises the following steps: (a) preparing a nucleic acid library from an initial sample, the preparation of the nucleic acid library from the initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids for generating the nucleic acid library are not extracted from the initial sample prior to the preparation of the nucleic acid library; (b) quantifying the number of reads of multiple fragment lengths within the nucleic acid library; and (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group consisting of: shape of distribution, segment amplitude, peak shape, fragment count ratio of two or more segments, height of helical phasing peak, fragment count ratio at two different fragment lengths, fragment count ratio within two different fragment length ranges, fragment length range within a segment, ratio of maximum amplitude of two or more segments, and fragment length range within a subset of reads.Segment length distribution; and (d) using the one or more fragment length characteristics to generate a fragment length spectrum of the nucleic acid library.

[0319] 22. A method for determining the infection stage of a subject, the method comprising the steps of: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample obtained from the subject; (b) comparing the fragment length spectrum with a reference fragment length spectrum; and (c) if the fragment length spectrum from the sample is similar to the fragment length spectrum from a symptomatic subject, then determining that the infection stage indicates an increased risk of the subject exhibiting microbial-related symptoms; and if the fragment length spectrum from the sample is similar to the fragment length spectrum from an asymptomatic subject, then determining that the infection is in an asymptomatic stage.

[0320] 23. The method according to embodiment 22, wherein the fragment length spectrum is a non-microbial host or microbial subset of the fragment length spectrum of the nucleic acid library.

[0321] 24. The method of embodiment 22 further comprises the steps of: (a) determining the abundance of at least one significant microorganism in a sample from the subject; (b) comparing the abundance to a threshold and comparing the fragment length spectrum to a reference fragment length spectrum; and (c) if the fragment length spectrum from the sample is similar to the fragment length spectrum from a symptomatic subject, and the abundance is equal to or higher than the threshold, then determining that the infection stage indicates an increased risk of the subject exhibiting microbial-related symptoms; if the fragment length spectrum from the sample is similar to the fragment length spectrum from an asymptomatic subject, then determining that the infection is in an asymptomatic stage.

[0322] 25. The method of embodiment 22 further comprises administering an antimicrobial agent to a subject determined to have an increased risk of exhibiting microbial-related symptoms.

[0323] 26. A method for determining the stage of infection in a subject suspected of having a microbial infection, the method comprising: a) performing high-throughput sequencing on nucleic acids from a biological sample; b) performing bioinformatics analysis to identify free nucleic acid sequences present in the biological sample; and c) obtaining a measurement of the free nucleic acid and comparing the measurement with a control, thereby determining the stage of infection of the microorganism identified in the biological sample.

[0324] 27. The method according to embodiment 26, further comprising one or more steps selected from the group consisting of: (a) extracting free nucleic acids from a biological sample obtained from the subject; and (b) adding a synthetic nucleic acid spiking agent to the free portion.

[0325] 28. The method according to embodiment 26, wherein the nucleic acid comprises microbial nucleic acid, host nucleic acid, or both microbial nucleic acid and host nucleic acid.

[0326] 29.According to the method of embodiment 26, the measurement result is selected from the group consisting of: the absolute abundance of the free nucleic acid, the distribution of the fragment length of the free nucleic acid, and both the absolute abundance and the fragment length distribution of the target microorganism.

[0327] 30. According to the method of embodiment 26, the infection stage is selected from the asymptomatic stage, the symptomatic infection stage, the treatment stage, or the eradication stage.

[0328] 31. According to the method of embodiment 26, it further includes administering a treatment regimen to the subject, wherein the treatment regimen is suitable for the determined infection stage.

[0329] 32. According to the method of embodiment 26, it further includes repeating the method on samples obtained from the subject at multiple time points to monitor the infection or the efficacy of treatment for the infection.

[0330] 33. The method according to embodiment 26, wherein the microorganism is selected from the group comprising: Helicobacter pylori, Clostridium difficile, Haemophilus influenzae, Salmonella, Streptococcus pneumoniae, cytomegalovirus, hepatitis virus b, hepatitis virus c, human papillomavirus, Epstein-Barr virus, human T-cell lymphoma virus 1, Merkel cell polyomavirus, Kaposi's sarcoma virus, human herpesvirus 8, chlamydia, gonorrhea, syphilis, or trichomoniasis.

[0331] 34. The method according to embodiment 27, wherein adding the synthetic nucleic acid spike further comprises: (a) preparing a spiked sample by obtaining a sample comprising free nucleic acid from a subject and adding at least 1000 unique synthetic nucleic acids to the sample, wherein each of the 1000 unique synthetic nucleic acids comprises: (i) an identification tag; and (ii) a variable region comprising at least 5 degenerate bases; (b) extracting nucleic acids from the spiked sample; (c) generating a spiked sample library; (d) enriching the spiked sample library; and (e) performing high-throughput sequencing to obtain sequence reads from the spiked sample library;(f) Calculate the diversity loss value for 1,000 unique synthetic nucleic acids; and (g) calculate the measurement result of the free nucleic acid and compare the measurement result with a control, thereby determining the infection stage of the subject.

[0332] 35. A method for determining the Helicobacter pylori infection stage of a subject, the method comprising: (b) extracting free nucleic acid from a biological sample obtained from the subject; (c) adding a synthetic nucleic acid spike to the free portion; (d) performing high-throughput sequencing on the nucleic acid from the biological sample; (e) performing bioinformatics analysis to identify the free Helicobacter pylori nucleic acid sequence present in the biological sample; and (f) calculating the measurement result of the free Helicobacter pylori nucleic acid and comparing the measurement result with a control, thereby determining the Helicobacter pylori infection stage of the subject.

[0333] 36. A method for determining the stage of Helicobacter pylori infection in a subject, the method comprising: a) preparing a spiked sample by obtaining a sample comprising free nucleic acid from the subject and adding at least 1,000 unique synthetic nucleic acids to the sample, wherein each of the 1,000 unique synthetic nucleic acids comprises: (i) an identification tag; and (ii) a variable region comprising at least 5 degenerate bases; b) extracting nucleic acids from the spiked sample; c) generating a spiked sample library, wherein the generation comprises (i) ligating an adaptor to an end-repaired spiked sample; and (ii) amplification; d) enriching the spiked sample library; e) performing high-throughput sequencing assays to obtain sequence reads from the spiked sample library; f) calculating the diversity loss value of the 1,000 unique synthetic nucleic acids; and g) calculating the measurement result of the free nucleic acid and comparing the measurement result with a control, thereby determining the stage of Helicobacter pylori infection in the subject.

[0334] 37. A method for determining host-microbe biological interactions in a subject, the method comprising: (a) generating a fragment length spectrum of a nucleic acid library generated from a sample from the subject; (b) optionally, determining the abundance of a target nucleic acid and comparing the abundance with a threshold; (c) comparing the fragment length spectrum with a reference fragment length spectrum of one or more host-microbe biological interactions; and (d) identifying the host-microbe biological interaction if the fragment length spectrum is similar to the reference fragment length spectrum of the host-microbe biological interaction.

[0335] 38. The method according to embodiment 37, wherein the host-microbe biological interaction is identified if the fragment length spectrum and abundance of the target nucleic acid are similar to the reference fragment length spectrum and threshold of the host-microbe biological interaction.

[0336] 39.According to the method of embodiment 32, it further includes changing the treatment regimen.

[0337] 40. A method for identifying the presence of a viral infection in a subject suspected of having a microbial infection, the method comprising: a) generating a fragment length spectrum of a nucleic acid library generated from a sample from the subject; b) comparing the fragment length spectrum with a viral reference fragment length spectrum; c) optionally, quantifying the abundance of a target nucleic acid and comparing the abundance with a threshold; d) if the fragment length spectrum is similar to the reference spectrum, then identifying the presence of a viral infection in the subject.

[0338] Example Example 1: Distribution Shape and Microbial Status Processing biological samples using methods that lack bias or achieve correction for bias in the region of fragment length of interest allows for the measurement of endogenous fragment length distributions and generates the potential to inform diagnosis and treatment using endogenous fragment length distribution spectra. Thus, several different clinical samples processed show the diversity of fragment length distribution spectra. A direct-to-library method with no detectable length and GC bias within the studied fragment length range was applied to obtain the shape of the endogenous fragment length distribution.

[0339] Clinical plasma samples: Thirty-six diagnostically positive (i.e., presence of microorganisms confirmed by orthogonal testing, such as blood culture, targeted PCR, or Karius testing) samples were collected from 36 human subjects. For each sample, a single-centrifugation plasma extraction procedure was performed from whole blood within 24 hours of sample collection, as previously described (see the first centrifugation step in Fan HC et al., Proceedings of the National Academy of Sciences 2008; 105(42): 16266-16271, which is incorporated herein by reference in its entirety, including any figures), and stored at -80°C prior to use. The samples were then thawed and 2 µL of the spiked master mixture was spiked into 200 µL of each plasma sample (see below).

[0340] Positive assay control samples: For each group of 18 samples, two positive controls, referred to as assay control samples (AC), were processed. AC samples were prepared from asymptomatic human plasma, purchased in purified form from ATCC (American Type Culture Collection), and spiked with the genomes of human pathogens cleaved by enzymes. The selected human pathogens were *Aspergillus fumigatus*, *Escherichia coli*, *Pseudomonas aeruginosa*, and *Staphylococcus epidermidis*. 10 µL of the spiked master mixture was added to each 1 mL AC sample (see below).

[0341] Negative control sample: prepared from aqueous buffer (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) and 5...The µL spike master mixture (see below) is used to prepare four 500 µL negative control samples (EC) for every 18 samples and is used as a control for environmental contamination (e.g., microbial and pathogen nucleic acid contamination introduced during treatment via reagents, instruments, consumables, operators, and / or air). These synthetic nucleic acids are used to normalize the signal in the samples to address variations in sample treatment. Specification 53 / 68 pages 56 CN 121575489 A

[0342] Spike Master Mixture: A set of process control molecules are premixed together in a single spike master mixture, each spike master mixture containing a unique “ID spike” process control molecule, see, for example, U.S. Patent 9,976,181. The spike master mixture contains three classes of molecules: ID spike molecules, SPANK molecules, and SPARK molecules. The latter group of molecules consists of two classes of SPARK: GC dSPARK and long SPARK. The molar concentrations of ID-spiked, SPANK, and long SPARK molecules in the main spiking mixture were 500 pM per molecule, while GC dSPARK molecules were present at 50 pM per molecule.

[0343] "ID-spiked" molecules: Each sample received a unique ID-spiked single-stranded DNA molecule characterized by a unique 50-base-pair-long sequence that was not present in any reference genome available in public databases at the time of processing.

[0344] SPANK molecules: The SPANK molecules used were pools of single-stranded DNA molecules, each 50 base pairs long, with identical 3' and 5' end sequences that were not present in any reference genome available in public databases at the time of processing. In addition, there were two 8-base-pair extensions nested between the constant 3' and 5' end sequences and completely degenerated in the pool. The SPANK molecule pool contained 416 unique SPANK molecules. The two degenerate extensions were separated by four non-degenerate base extensions.

[0345] SPARK molecules: The GC spike group is a collection of molecules of 32, 42, 52, and 75 nt lengths, each containing seven distinct sequences with GC contents of 20%, 30%, 40%, 50%, 60%, 70%, and 80%, respectively. Like some of the other molecules provided above, the GC dSPARK sequences are not present in the available reference genome. The long SPARK sequence set is a group of four non-natural sequences, each with a GC content of 50% and lengths of 100 nt, 125 nt, 150 nt, and 175 nt. The complete collection of SPARK molecules contains 32 distinct sequences.

[0346] Library generation: Direct-to-library generation was achieved in U.S. Provisional Application 62 / 770, filed November 21, 2018.As described in 181, the entire contents of the aforementioned U.S. Provisional Application are incorporated herein by reference in their entirety. Here, a template-switching method using proteinase K is utilized. Briefly, 50.0 µL of each spiked sample was mixed with 20.0 µL of 10x terminal transferase reaction buffer (NEB, Ipswich, MA), 5.0 µL of proteinase K (Sigma), 2.0 µL of 10% Tween-20 (Thermo-Fisher Scientific, Waltham, MA), 2.0 µL of 10% Triton X100 (Thermo-Fisher Scientific, Waltham, MA), and 121.0 µL of nuclease-free water. The mixture was heated to 60°C for 20 minutes and then to 95°C for 10 minutes, and then placed on ice until cooled. To prepare the A-tail reaction, add 2.0 µL of 10 mM dATP, 2.0 µL of terminal transferase (20 u / µL, New England Biolabs, Ipswich, MD) and 6.0 µL of nuclease-free water, and incubate at 37 °C for 40 min. Add 300.0 µL of lysis / binding buffer (Thermo Fisher Scientific, Waltham, MA) to the reaction. Then add the entire volume to 50.0 µL of Dynabeads oligonucleotide (dT)25 (Thermo Fisher Scientific, Waltham, MA) and wash once with lysis / binding buffer (Thermo Fisher Scientific, Waltham, MA). Incubate the mixture at 25 °C and 600 RPM. The beads were then washed twice with 600.0 µL of wash buffer A (Thermo Fisher Scientific, Waltham, Massachusetts) and twice with 300.0 µL of wash buffer B (Thermo Fisher Scientific, Waltham, Massachusetts), followed by elution in 24.0 pF elution buffer (Thermo Fisher Scientific, Waltham, Massachusetts) at 80°C and 600 RPM for 3 minutes. The entire eluent was transferred to a new plate. 2.0 µL of 1 µM Poly dT primer (IDT) and 6 µL of SMARTScribe first-strand buffer (5x) (Takara, Kusatsu, Japan) were added to the eluent, and the resulting mixture was incubated at 95°C for 1 minute and then placed on ice. The solution was then eluted by adding 4.5 µL of SMARTScribe first-strand buffer (5x) (Takara, Kusatsu, Japan) and 0.5 µL of dNTP mixture (25 nucleotides per nucleotide).The extension and template switching mixture was prepared by combining 2.0 µL of 5 µM template-switching oligonucleotide (TS oligonucleotide) (IDT), 5.0 µL of DTT (20 M, Takara Corporation, Kusatsu City, Japan) and 4 µL of nuclease-free water. The resulting reaction mixture was incubated at 42 °C for 90 min and then subjected to thermal denaturation at 70 °C for 15 min. Next, 50.0 µL of NEBNext Ultra II Q5 (New England Biolabs, Ipswich City, MATLAB, USA) and 8.0 µL of a fractionating primer mixture (New England Biolabs, Ipswich City, MATLAB, USA) were added to the reaction from the previous step. Nucleic acid amplification was then performed using the following temperature cycling program: 98°C for 30 seconds, 8 cycles of 98°C for 10 seconds each, 65°C for 75 seconds, and a final extension at 65°C for 5 minutes. The final nucleic acid library was then pooled into groups of four ECs, two ACs, and eighteen clinical samples, and the pools were purified using RNAclean™ Ampure beads as described above. After purification, the concentration of nucleic acids in the library pools was measured using TapeStation as described above, and the libraries were loaded onto the sequencer according to the manufacturer's recommendations.

[0347] Sequencing: The samples were sequenced using an Illumina NextSeq™ 500 sequencer to obtain sequence reads. Sequencing was performed using 76 cycles as per the manufacturer's instructions.

[0348] Sequencing data analysis: The primary sequencing output was multiplexed using bcl2fastq v2.17.1.14 (with default parameters), and then template-switched oligonucleotides were removed using Cutadapt. Poly A tails were removed, and reads were quality trimmed and then filtered using Trimmomatic v 0.32 if shorter than 20 bases. Reads passing these filters were aligned to human and synthetic (including process control molecules and sequencing adaptors) references using Bowtie v2.2.4. Reads aligned to either sequence were set aside. Reads potentially representing human satellite DNA were also filtered using a k-mer-based approach. The remaining reads were aligned to a microbial reference database using BLAST v2.2.30. Reads exhibiting both high identity percentages and high query coverage were retained, except for those aligned to any mitochondrial or plasmid reference sequences. PCR repeats were removed based on the alignments. Relative abundance was assigned to samples based on sequencing reads and their alignments.For each taxonomic unit, a read sequence probability is defined for each combination of reads and taxonomic units, which resolves discrepancies between the microorganisms present in the sample and a reference assembly in the database. A mixture model is used to assign likelihoods to the complete set of sequencing reads, which incorporate the read sequence probability and the (unknown) abundance of each taxonomic unit in the sample. An expectation-maximization algorithm is applied to compute a maximum likelihood estimate of the abundance of each taxonomic unit. Based on these abundances, the number of reads generated by each taxonomic unit is summarized into a taxonomic tree. A set of libraries can be prepared from the corresponding negative control buffer and processed and sequenced within each batch. The estimated taxonomic unit abundances of negative control samples within a batch can be combined to parameterize the environmentally induced variants of the read abundance model driven by counting noise. Statistical significance values ​​for each estimated taxonomic unit abundance can be computed, and those at a high significance level within the CRR include candidate calls (i.e., significant calls). Final calls (i.e., reportable calls) are performed after applying additional filtering, which resolves read position consistency, read percentage identity, and cross-reactivity derived from higher abundance calls. The number of reads of multiple fragment lengths for each reportable microorganism within each processed nucleic acid library was determined, the fragment length distribution was evaluated, and the fragment length characteristics of the distribution shape were determined. Figure 8 shows examples of some fragment length distribution shapes among the different fragment length distribution shapes observed in the detected microorganisms in the tested clinical samples. Due to the minimum mapping length and the 68 bp set by combining the maximum read length in the described sequencing experiments and adapter trimming algorithms, the range of fragment lengths shown is limited to 22 bp at the shorter end. Therefore, fragments longer than 68 bp contribute to the counting in the 68 bp length bin. The three microorganisms detected in the three examples shown are Candida tropicalis, Aspergillus oryzae, and WU polyomavirus. The fragment length distribution shape varies greatly among these microorganisms and is not related to any particular species or superkingdom as shown by the remainder of the data (not shown).

[0349] Candida tropicalis was detected in three different clinical samples treated in this study. A subset of reads from each sample aligned to the *Candida tropicalis* reference genome (pages 55 / 68, CN 121575489 A) were identified, and their fragment length distributions were determined. Figure 9 shows the results of the *Candida tropicalis* fragment length distributions from each of the three samples in each individual group. Compared to the right panel, the left and middle panels show a distribution with higher short (<40 bp) and long (>65 bp) fractions relative to the 50 bp peak, while exhibiting a distinct peak at approximately 45–50 bp. The second left panel is from a patient with diffuse *Candida tropicalis* infection, unaffected by...Mechanistic limitations can explain the increase in the number of long and short fragments relative to the peak. Different fragment length distributions can indicate different states of disease or symptom. WU polyomavirus is another example of a microorganism that was treated in this study and detected in multiple clinical samples that exhibited different fragment length distributions in each sample (Figure 10). In one subject, WU polyomavirus showed only a “50 bp peak.” The second subject showed a considerable contribution of a short class index score as well as a higher score for reads longer than 68 bp. Although not mechanistically limited, WU polyomavirus may have been incorporated into the human genome in this sample or its genome may have been released into the body fluids, resulting in different fragmentation patterns. In a total of 36 clinical samples (see above), 60, 24, and 13 bacterial, fungal, and viral microorganisms were detected, respectively. The fragment length distributions of these microorganisms varied greatly, as demonstrated by the examples above. Next, the ratio of the read count detected in the “50 bp peak” to the short class index region of the distribution of all detected microorganisms or pathogens was determined. The obtained ratios were grouped by their superkingdoms, and histograms of the ratio characteristics for each superkingdom were generated. The results from one of this analysis are presented in Figure 11. The same analysis was performed on human DNA (i.e., host DNA) and human mitochondrial DNA (i.e., host mitochondrial DNA) as controls (Figure 11). Microbial behavior depends on the superkingdom, and these aspects must be addressed when using fragment length distribution shape and properties for diagnostic purposes.

[0350] Example 2: Analysis of plasma samples from pregnant subjects Many types of non-host nucleic acids can be found in samples obtained from the host. Fetal cell-free nucleic acids can be detected in maternal blood. In this sample, plasma samples were obtained from 15 consenting pregnant women and were de-identified. The samples were processed and sequenced according to the ligation-based direct-to-library method described in Example 1 of U.S. Provisional Application 62 / 770,181, filed November 21, 2018, which is incorporated herein by reference in its entirety. In this analysis, samples from subjects carrying male fetuses were considered only. Reads that aligned only to the Y chromosome were considered fetal. Bowtie2 was used to align reads to the human genome. Then, Bowtie2 was used to align reads mapped to chromosome Y to an index created from all chromosomes except Y. Any reads aligned to this index were discarded, thus preserving reads unique to chromosome Y only.

[0351] Figure 12 shows the fragment length distribution of cell-free nucleic acids from the mother (dashed lines) and fetus (solid lines) of an individual. In this example, the ratio of fetal to maternal reads in the “50 bp peak” region was higher than in nucleosome fragment regions (e.g., the 150–200 bp region). On average, in the “50 bp peak” region…The concentration of fetal fragments observed in the "bp peak" region was 4 times higher than that in the nucleosome length fragment region. The method used here can be used to enrich fetal fractions.

[0352] Example 3: Analysis of Microorganisms Using Fragment Length Profiling Nucleic Acid Libraries were prepared from over 4000 cell-free plasma samples and sequenced using a validated Karius test based on an extraction method that recovers double-stranded DNA fragments in an unbiased manner relative to their length and GC content within a fragment length range associated with cell-free nucleic acids. Fragment length profiling of the detected microorganisms was generated, and 33 taxa that were called 10 or more times within the studied sample groups were evaluated. More specifically, the ratio of short reads in low-probability calls to high-probability calls was evaluated. Results from one of these experiments are presented in Figure 13. In this experiment, the figure indicates that there are more short reads in low-probability calls compared to high-probability calls. Although not mechanistically limited, these results suggest that, given end-repairable cell-free double-stranded DNA, clinical infections have longer fragment length distributions compared to translocated colonies or non-pathogenic organisms in the bloodstream. Specification 56 / 68 pages 59 CN 121575489 A

[0353] Example 4: Analysis of colonization sites using fragment length profiling. Nineteen clinical samples were obtained from subjects identified as having an infection, such as those identified by positive urine (n = 19) and / or blood culture tests (n = 11). Nucleic acid libraries were prepared from these samples and sequenced using a validated Karius test, which is an extraction-based method that recovers double-stranded DNA fragments in an unbiased manner relative to their length and GC content within the fragment length range associated with free nucleic acids. Of the nineteen subjects, blood and urine cultures were identified as 19 and 11 microorganisms, respectively. The fragment length distribution profiles of the microorganisms detected by blood and urine cultures were evaluated. The results are shown in Table 14. Although not mechanistically limited, pathogen DNA from deep tissue infections (lungs, brain, etc.) can undergo different degradation mechanisms, thus affecting the fragment lengths observed as DNA from pathogens infects the blood.

[0354] Example 5: Host Nucleic Acid Length Distribution Profile and Infection Status. The fragment length distribution of host nucleic acids can help inform non-host nucleic acid signals within the host, such as microbial nucleic acid signals or the stage of infection in the host (e.g., asymptomatic vs. symptomatic). For example, the abundance of microbial nucleic acids in samples from human hosts can vary by several orders of magnitude (Blauwkamp et al., (2016)). Although samples obtained from asymptomatic individuals tend to show lower abundance of microbial nucleic acids compared to infected individuals, some asymptomatic samples show...The measured abundance may exceed the lowest abundance in infected individuals (Blauwkamp et al., (2016)). Additional properties of the nucleic acid pool obtained from the sample can help distinguish between different stages of infection or biological relationships (e.g., symbiosis and pathogen) between microorganisms and the host. Here, the utility of the length distribution of host nucleic acids in predicting the infection status of microorganisms in plasma from the host was tested. The method makes it possible to access endogenous fragment length profiles with fragment lengths that are not typically accessed in an unbiased manner by previous methods. The method makes it possible to access endogenous fragment length profiles with fragment lengths that are typically discarded, ignored, or considered unimportant by previous methods.

[0355] Clinical plasma samples: 100 asymptomatic (collection criteria: no infection-associated active healthy tissue and passed normal blood screening tests), 85 diagnostically positive (i.e., presence of microorganisms confirmed by orthogonal tests, e.g., blood culture, targeted PCR, Karius tests) and 45 diagnostically negative plasma samples were collected from human subjects. For each sample, a single-centrifugation plasma extraction procedure was performed from whole blood within 24 hours of sample collection, as previously described (see Fan HC et al., Proceedings of the National Academy of Sciences 2008; 105(42): 16266-16271, which is incorporated herein by reference in its entirety, including any figures), and stored at -80°C before use. The samples were then thawed and 5 µL of the spiked master mixture was spiked into 500 µL of each plasma (see above). If a smaller volume was obtained, a smaller volume of the spiked master mixture was added proportionally to maintain a constant concentration of process control molecules in all initial and control samples.

[0356] Positive and negative control samples were prepared as described above.

[0357] Nucleic acid library generation and sequencing directly from plasma: The generation of the direct-to-library is described in U.S. Provisional Application 62 / 770,181, filed November 21, 2018, which is incorporated herein by reference in its entirety. The library was prepared and sequenced as described in Example 1 above.

[0358] Results: The abundance of significant microorganisms present in each sample was determined as described above and given in units of molecular weight per microliter (MPM) of plasma sample concentration, which is a normalized number that gives an estimated number of unique nucleic acid fragments of an organism in 1 μL of plasma sample. This calculation is derived from the number of unique or deduplicated sequences of the presence of each organism normalized to a known number of unique synthetic spikes added to the plasma sample prior to the start of the process (see U.S. Patent No. 9,976,181). Figure 9A shows the distribution of MPM values ​​in asymptomatic (AP) and diagnostic positive (DP) sample types. The lower abundance values ​​in the DP sample type overlap with the range of MPMs observed in the AP samples, even if they contain microorganisms that are only orthogonally confirmed.The same applies to organisms. (DPNGS contains microorganisms confirmed by the Karius test, and DPmicro contains microorganisms confirmed by culture or by PCR-based methods as described on pages 57 / 68 of the manual, CN 121575489 A). Additionally, if the analysis is limited to the microbial species present in the collection of AP samples and also the collection of DP samples (in this dataset, the following species meet this description: Bacillus coagulans, Enterococcus spp., Enterococcus faecalis, Haemophilus influenzae, Haemophilus parainfluenzae, human mammalian adenovirus D, Neisseria myxophylla, Pediococcus lactis, Prevotella intermedia, Prevotella niger, Saccharomyces cerevisiae, Streptococcus agalactiae, Streptococcus salivarius, Streptococcus thermophilus), the abundance in the diagnostic positive group is still not always higher (Figure 15B). Therefore, the abundance is insufficient to distinguish the infection status of non-microbial hosts.

[0359] A combination of several measurable parameters can then be used to distinguish asymptomatic / healthy patients from those who have experienced infection. Therefore, as a potential classifier, a combination of MPM microbial abundance and nucleic acid fragment length distribution mapped to the host reference (i.e., the human reference in this sample cohort) was investigated.

[0360] Figure 15C shows an example of a typical distribution of nucleic acid fragments completed during the library generation process and as measured by a TapeStation instrument. Two main peaks in fragment length can be observed: (1) a “nucleosome” peak (ranging from 300 to 450 bp in the electrophoresis plot) and (2) a “subnucleosome” peak (ranging from 180 to 280 bp in the electrophoresis plot). This signal is determined by the nature of human (i.e., host) nucleic acid, as microbial (i.e., non-host) nucleic acid accounts for a small fraction of the total nucleic acid population in these samples, which contain DP sample types. The molar ratio and mass ratio of human fragments contributing to the two peaks vary between samples and are different between AP sample types and DP sample types (Figure 15D). The vast majority of AP samples (92%) showed a “nucleosome” peak molar fraction below 0.4, while the same value was evenly distributed across a wider range (< 0.7) in DP samples.

[0361] The properties of MPM microbial abundance and human fragment length distribution showed overlap between values ​​in AP and DP samples. The combination of two independent measurements can help distinguish between asymptomatic and infected calls in unknown samples where the stage of infection is unknown. Figure 15E shows the long human read fractions as measured from sequencing data (all reads mapped to a human reference longer than 65 bp after adaptor trimming) and the maximum MPM value measured in the same samples across all AP and DP samples. The area covered by coordinates [(0,3000),(0,0.4)] is exclusively filled by AP samples. Three of the 100 AP samples fall outside this space (arrows in 15E). The microorganisms detected in these three samples were Helicobacter pylori and human mammalian adenovirus.And Neisseria gonorrhoeae. All three microorganisms are known human pathogens, but it is unknown whether they are pathogenic in these individuals.

[0362] A comparison between the properties of microbial MPM and human fragment length distribution in AP and DN sample types (Fig. 15F) reveals that DN-free samples fall within the typical asymptomatic range, even if they are negative according to orthogonal testing.

[0363] The properties of fragment length distribution of non-microbial signals, such as non-microbial host nucleic acids, can be used to identify the asymptomatic or non-infectious status of subjects.

[0364] The data also indicate that asymptomatic individuals can be identified by combining abundance (e.g., maximum MPM) and fragment length distribution parameters, as presented herein, even if the MPM values ​​of microorganisms overlap with the ranges observable in diagnostically positive samples. This also suggests that early detection of infection is possible in the absence of standard symptoms. The region on this two-dimensional plane that can help distinguish between different infection states in an individual can be further optimized for, for example, the MPM or kingdom of a specific microbial species and the microbial fragment length to improve the performance of the test.

[0365] Finally, the normalized size distribution of fragments aligned to the human genome (nuclear genome predominant), human mitochondrial genome, all pathogens, significant pathogens, and bacteria, eukaryotes, viruses, and archaea for all samples was calculated. To distinguish AP from DP / DN samples, a classifier was trained on the fragment size distribution (features), in this case using logistic regression with L2 regularization. Logistic regression is a linear model for classification that multiplies features by a set of weights before transformation with a logistic function. The weights were determined using standard numerical optimization techniques with L2 regularization, providing additional constraints to minimize the sum of squares of the weights. This has the effect of reducing overfitting and multicollinearity in the features. The accuracy of this model was assessed by using the trained model to predict the probability that each sample was asymptomatic or symptomatic. A value > 0.5 indicates that the sample was predicted as asymptomatic, and a value < 0.5 indicates that the sample was predicted as symptomatic. Additionally, the trained model provides weights (coefficients). Positive coefficients indicate association with asymptomatic individuals, while negative coefficients indicate association with symptomatic individuals. Figure 16 shows the accuracy of training based on asymptomatic and symptomatic infection states using the normalized size distribution of fragments compared to the following: human genome (nuclear genome dominant), human mitochondrial genome, all pathogens, significant pathogens; and bacteria, eukaryotes, viruses, and archaea. Nucleic acid subgroups from the library used to train the model affect the model's accuracy. Additionally, nucleic acid subgroups from the library influence the regions with positive predictive values ​​for asymptomatic or symptomatic states in the fragment length distribution. For example, long human fragments (>60)The presence of 50 bp fragments predicted symptomatic status (Fig. 16A, right panel), as did short (< 30 bp) pathogen fragments (Fig. 16C, right panel). On the other hand, high concentrations of fragments around 50 bp predicted asymptomatic status (Fig. 16A, right panel), as did long (> 65 bp) pathogen fragments (Fig. 16C, right panel).

[0366] Example 6: Differentiation between asymptomatic patients colonized with Helicobacter pylori and patients with active Helicobacter pylori-associated inflammation Plasma processing and DNA extraction: Plasma was extracted from whole blood samples within 24 hours of sample collection, as previously described (Fan HC et al., Proceedings of the National Academy of Sciences 2008; 105(42): 16266-16271), and stored at -80 °C. When analysis was required, the plasma samples were thawed, and circulating DNA was immediately extracted from 0.5–1 ml of plasma.

[0367] Sequencing Library Preparation and Sequencing: Sequencing libraries were prepared from purified patient plasma DNA using a NEBNext DNA library master mix with a standard Illumina indexed adaptor (purchased from IDT) and end-repair purification (e.g., MagBind beads, NEBNext end-repair module) for Illumina setup, or using a microfluidic-based automated library preparation platform (Mondrian ST, Ovation SP ultra-low library system). The libraries were characterized using an Agilent 2100 Bioanalyzer (High Sensitivity DNA Kit) and quantified by qPCR.

[0368] qPCR Validation of Sequencing Results for Selected Bacterial Targets: Sequencing results for a subset of cell-free DNA samples were validated using a standard qPCR kit for quantification of selected bacterial targets (e.g., Helicobacter pylori). qPCR assays were run on cfDNA extracted from approximately 1 ml of plasma and eluted in 100 ml Tris buffer (50 mM [pH 8.1–8.2]). Plasma extraction and PCR experiments were performed at different facilities. Template-free controls were run to validate the PCR reagents included in each experiment.

[0369] After removing low-quality reads, the reads were mapped to a human reference genome. The remaining reads, presumably microbiome-derived, were mapped to a reference database of the target microbial genome. The relative abundance of each microorganism was calculated using a proprietary algorithm. The algorithm reported organisms present in statistically significant amounts compared to controls. Organisms with overrepresented sequences were reported as positive.

[0370] Quality control (QC) metrics included adding ID-spiked synthetic nucleic acids as a unique type of spike for each sample in the sequencing batch and other synthetic nucleic acid spikes (“SPANK”) spiked at constant concentrations across all libraries.The number of deduplicated SPANK molecules detected in a particular library is, in place of the minimum detectable concentration in that library. This can be used to set a threshold based on the minimum detectable concentration of SPANK molecules in the library. The threshold can be used to ensure sufficient sequencing depth for pathogen detection. The threshold can also be used to ensure that the pathogen signal is not due to cross-contamination from other samples. For example, the enrichment of pathogens relative to a threshold set by SPANK molecules can be compared between different samples. More generally, it is proportional to the efficiency with which the library converts DNA molecules in the original sample into reads in DNA sequencing data. The purpose of SPANK molecules is to help establish the relative abundance of pathogen molecules within the mixture represented in the sample, reported in “molecules per ml” (MPM). MPM data is used to construct heatmaps and correlation plots. The Sample Purity Ratio (SPR) is designed to capture how the number of reads associated with a taxonomic unit gives an estimate of the degree of cross-contamination in the sample. In the event of a failure of deduplicated SPANK and / or SPR, the sample is requeued and run again. If QC fails twice on the same sample, it is reported as “no result”. (Specification 59 / 68) Page 62 CN 121575489 A

[0371] Results. The method is capable of detecting Helicobacter pylori cell-free DNA in plasma obtained from patients with Helicobacter pylori-associated peptic ulcer disease. The method is capable of differentiating between patients with asymptomatic Helicobacter pylori and those with Helicobacter pylori disease. For the latter, samples were obtained from healthy (i.e., asymptomatic) and infected subjects, and said samples were analyzed using next-generation sequencing of cell-free plasma to detect pathogen DNA (Karius Test™, Karius, Redwood City, CA). In healthy volunteers, the test detected Helicobacter pylori in 8 / 106 samples tested. Some patients in the dataset were identified as having asymptomatic colonization of Helicobacter pylori (C) (n = 1) or symptomatic chronic infection of Helicobacter pylori (CI) (n = 7) (see Table 1 below). Helicobacter pylori-positive samples were associated with Afro-Weekly or Hispanic ethnicity, consistent with the epidemiology of Helicobacter pylori infection.

[0372] Table 1: Detection of Helicobacter pylori in plasma. Helicobacter pylori infection is a likely, probable, or unlikely cause of sepsis

[0373] Unrestricted by mechanism, free nucleic acids can be derived from dead and dying pathogens. Therefore, this method is uniquely suited for detecting organisms actively cleared by the immune system. In fact, the assay can differentiate between Helicobacter pylori in the context of active inflammation rather than asymptomatic colonization.

[0374] Example 7: Method for Detecting Helicobacter pylori Gastrointestinal Infection in High-Risk Patients The objective of this study was to evaluate the clinical utility of this method (i) to detect active Helicobacter pylori infection in symptomatic patients with peptic ulcer disease (Helicobacter pylori PUD) compared to conventional diagnostic tests; (ii) to confirm eradication of active Helicobacter pylori gastrointestinal infection after first-line therapy compared to conventional diagnostic tests; and (iii) to evaluate the optimal MPM threshold for distinguishing patients with active Helicobacter pylori PUD from those without (asymptomatic) infection. This non-invasive method allows physicians to make effective treatment decisions without resorting to traditional invasive diagnostic methods.

[0375] Study Design. As described below, the percentage of positive concordance (PPA) and percentage of negative concordance (PA) of this method compared to conventional non-serological Helicobacter pylori diagnostic tests were determined in two well-described adult study populations under specific testing conditions. At study entry, patients with symptomatic Helicobacter pylori pud met clinical criteria and had at least one positive protocol-approved, non-serological routine Helicobacter pylori diagnostic test prior to any administration of the first eradication instruction manual (pages 60 / 68, CN 121575489 A). Plasma testing was performed on all documented patients with symptomatic Helicobacter pylori pud. Subsequently, these PUD patients received a standard eradication protocol (according to standards of care) for 2–4 weeks, followed by a 1-month medication break. Within 30 days (+ / - 3 days) after completion of the first treatment, all PUD patients who participated in the study at the end underwent repeated plasma testing evaluation and at least one of the original non-serological, routine Helicobacter pylori diagnostic tests performed prior to treatment.

[0376] At study entry, negative control patients who underwent colonoscopy for any reason and showed no signs of active Helicobacter pylori gastrointestinal disease during screening, based on clinical criteria and at least one negative protocol-approved, non-serological routine Helicobacter pylori diagnostic test. Subsequently, negative control colonoscopy patients underwent plasma testing to complete all protocol requirements.

[0377] Data from these diagnostic test comparisons provide information on the utility of this method for detecting active Helicobacter pylori disease and confirming its eradication efficacy after first treatment compared to routine non-serological Helicobacter pylori diagnostic tests.

[0378] Methods and Materials: A quantitative assay method is used to detect microorganisms by analyzing non-human DNA in plasma. The analyte for this method is microbial cell-free nucleic acid, which is very short (average length less than 100 nucleotides) compared to human cfDNA.

[0379] Whole blood is centrifuged twice to present free (cf) plasma. To address potential environmental contamination, the non-volatile buffer can be heated to above 85°C and then cooled before use. After the first centrifugation, PCT-US2017-The method described in 024176 adds an internal control molecule to each sample. Plasma is extracted, and purified cell-free DNA (cfDNA) is used to prepare master mixtures of NEBNext DNA libraries using adaptors with standard Illumina indexes (purchased from IDT) for Illumina setup and post-end repair purification (e.g., MagBind beads, NEBNext end repair module) or sequencing libraries using microfluidic-based automated library preparation platforms (Mondrian ST, Ovation SP ultra-low library system). Adaptors are ligated, and purification is performed without heating using AMPure beads, followed by amplification by qPCR. Libraries are characterized using an Agilent HS TapeStation, and the total concentration of nucleic acids is measured to control loading volume for small selection steps by integrating the signal (e.g., between 50 bp and 1000 bp).

[0380] Sequencing cfDNA fragments are mapped to a reference database of microbial sequences to determine the identity of non-human, non-internal control material present in the sample at significant levels (above the assay background). First, the sequencing data is converted into reads representing DNA sequences, and then multiplexed based on the index sequence into a set of reads (read sets) derived from each library loaded into the sequencer. Reads aligned to human sequences are filtered, and the remaining reads aligned to internal control molecules are reserved for further analysis. Next, reads that do not match human or internal control references are aligned to known microbial genomes. Reads with one or more alignments to this database (pathogen reads) form the basis for subsequent analyses.

[0381] The relative abundance of each taxonomic unit associated with the reference sequence is inferred using the alignment of each pathogen read to the microbial genome database. These abundances are aggregated into a taxonomic tree to give abundance at all taxonomic unit levels. Finally, on the same sequencing run, the abundance in clinical samples is compared to the abundance in the negative control library to determine if it has risen above the expected background level due to environmental DNA contamination. Taxonomic units that meet this criterion are reported in molecules per milliliter (MPM) based on the ratio of the abundance of microbial reads to certain internal control reads obtained. Before results are obtained, the pipeline applies a set of filters to limit reportable organisms to, for example, 3-10% greater than the microorganism with the highest number of reads and, for example, 25-50% greater than any other taxonomically related organism. Filters are applied to all patient samples and assay controls.

[0382] Potential sources of performance deviation from sample-specific or microorganism-specific properties include: the category of microorganism-specific properties, including the category of microorganism (e.g., bacteria, viruses, eukaryotes, prokaryotes, fungi, etc.), GC content, and genetic specification.Pages 61 / 68, 64 CN 121575489 A Group size, abundance of endogenous microbiota, environmental contamination (EC) levels, and the quantity and quality of reference assemblies. To address these sources of bias, this method comprises using representative groups of 10–100 microorganisms that capture the full spectrum of potential performance biases along GC content, genome size, and strain. These representative organisms should be cross-kingdoms, ranging within GC content (e.g., 10%–80%), and have genomes ranging from kilobases to megabases. The representative population should contain a mixture of types, such as commensal and non-commensal, microorganisms typically present as environmental contaminants, and closely related strains. The method also incorporates standard quality control measures, such as reference intervals for microbial levels in healthy populations and EC negative controls.

[0383] A test is considered positive if it shows significant Helicobacter pylori relative to a negative background control. However, note that the percentage of negative identity (NPA) is unlikely to reflect the NPA of the test after resolving the quantitative MPM cutoff.

[0384] In addition to assessing PPA and NPA within each of the study cohorts, PUDs, and colonoscopies using positive and negative thresholds determined by the laboratory in MPM, other thresholds in the MPM should also be considered. First, the MPM will be summarized using the mean, standard deviation, median, and range for each study cohort. Second, receiver operating characteristic (ROC) curves will be used to identify the optimal cut point in the MPM to maximize PPA and NPA in the sample.

[0385] Finally, to assess the ability of this method to identify eradication at 30 days, successful eradication will be estimated using the proportion and 95% confidence interval within each study cohort.

[0386] Example 8: Fragment length distribution profiles and colonization sites The characteristics of fragment length distributions of microbial sequencing reads obtained from clinical samples from patients infected in the bloodstream and lungs were compared as an example of deep tissue infection. Fragment length distribution characteristics vary depending on the location of infection. Regardless of the mechanism, different host responses at different infection sites may contribute to altering fragment length distribution characteristics. Furthermore, regardless of the mechanism, different infection sites can exhibit different non-host nucleic acid fragmentation mechanisms.

[0387] Clinical plasma samples: Ten unlabeled clinical samples were collected from patients with confirmed bloodstream infections and ten unlabeled clinical samples were collected from patients with confirmed pulmonary infections. For each sample, a single-centrifugation plasma extraction procedure was performed from whole blood within 24 hours of sample collection, as previously described (see Fan HC et al., Proceedings of the National Academy of Sciences of the United States of America 2008; 105(42): 16266-16271, which is incorporated herein by reference in its entirety, including any figures).And stored at -80°C before use. The samples were then thawed and 1.5 µL of the spiked master mixture was spiked into 150 µL of each plasma (see below). If different volumes were obtained, smaller or larger volumes of the spiked master mixture were added proportionally to maintain a constant concentration of process control molecules in all initial and control samples.

[0388] Negative control samples: Four 500 pF negative control samples (ECs) were prepared from aqueous buffer (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) with 5 µL of the spiked master mixture (see below) and used as controls for environmental contamination (e.g., microbial and pathogen nucleic acid contamination introduced during treatment via reagents, instruments, consumables, operators and / or air). These synthetic nucleic acids were then used to normalize the signal in the samples to address variations in sample treatment.

[0389] The spiked master mixture was prepared with ID spiking molecules, SPANK molecules and SPARK molecules as described above.

[0390] A linker-based direct-to-library method, as described in Example 1 of U.S. Provisional Application 62 / 770,181, was used to prepare a sequencing library from 5 μl of spiked asymptomatic plasma. Sequencing and sequencing data analysis were performed as described in Example 8.

[0391] Results: Table 2 lists the infection sites and species of infectious microorganisms for each of the 20 clinical samples, along with the donated clinical samples, as part of this example. The fragment length distribution of infectious microorganisms in all tested samples is shown in Figure 17. Based on the following fragment length distribution spectral characteristics (e.g., short exponential decay fragments, peaks, long fragments) in the specification 62 / 68 pages 65 CN 121575489 A, the normalized fragment length distribution mapped to the reference reads of the infectious microorganisms was analyzed: (1) the short class exponential distribution score (“short” in Table 2), (2) the peak score (“peak” in Table 2), and (3) the score of reads longer than the experimental read length (75 bp; “long” in Table 2). Similarly, the typical length range is found in the microbial fragment length distribution. A comparison of fragment length distribution patterns reveals that bloodstream infections disproportionately exhibit fragment length distribution patterns characterized by: (1) a high fraction of fragments with a short pseudo-exponential distribution; (2) the absence of a peak between 20 bp and 75 bp read lengths; and (3) a fraction of long reads greater than 10% (>64 bp). Conversely, pulmonary infections disproportionately exhibit fragment length distribution patterns characterized by: (1) the presence of fragments with a short pseudo-exponential distribution; (2) a peak between 20 bp and 75 bp read lengths; and (3) a fraction of long reads less than 10%. This suggests that the characteristics of microbial fragment length distributions can be used to determine whether an infection is present in the blood or deep tissues.

[0392] Table 2: A list of properties of fragment length distribution of sequencing reads from clinical samples and infection sites, species of infecting microorganisms, and references mapped to the species of infecting microorganisms. For each property, its quantitative assessment (presence / absence) is indicated, and the score of total reads present in the said segment is given in parentheses. Here, short fragment segments contain reads from 22 bp to 29 bp and contain 29 bp; peak fragment length ranges contain reads from 30 bp to 59 bp and contain 59 bp; and long fragment ranges contain reads longer than 59 bp.

[0393]

[0394] Example 9: Fragment length distribution spectrum and colonization site 2 The characteristics of fragment length distribution of microbial sequencing reads obtained from clinical samples of patients infected with plasma located in the bloodstream (plasma drawn from venous blood) and from capillary blood that had come into contact with the skin on their fingertips before being collected in a capillary collection system were compared as an example of skin infection.

[0395] Clinical plasma samples: Blood from 20 healthy adult donors was collected into PPT tubes according to the manufacturer’s instructions, with K2EDTA as the anticoagulant (Becton Dickinson, Franklin Lakes, NJ). Immediately after venous blood collection, capillary blood was performed on the same group of 20 healthy donors using a Microvette CB300 blood sampling device with K2EDTA as the anticoagulant (Sarstedt Inc, Sparks, NV, Nevada). During capillary aspiration, the following steps were performed: (1) the donor’s finger was held in an upward position and a lancet of appropriate size was inserted into the palmar side of the finger; (2) pressure was avoided on the finger during puncture to prevent hemolysis of the collected blood; and (3) any blood droplets spreading above the fingertip were collected into a clean Microvette CB300 blood sampling device. According to the manufacturer's instructions, a single-centrifugation plasma extraction process was performed on each sample from whole blood within 12 hours after sample collection, and the plasma was stored at -80°C until use. The samples were then thawed, and a spiking master mixture equivalent to 1% of the plasma volume was added to each plasma.

[0396] Negative control samples: Four 500 µL negative control samples (ECs) were prepared from aqueous buffer (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) with 5 µL of spiking master mixture (see below), and were used as controls for environmental contamination (e.g., introduced during treatment via reagents, instruments, consumables, operators and / or air).Microbial and pathogen nucleic acid contamination). These synthetic nucleic acids were then used to normalize the signal in the sample to address variations in sample processing.

[0397] Negative Microvette samples: Four 300 µL aqueous buffers (10 mM Tris pH 8, 0.1 mM EDTA, 0.05 v / v% Tween-20) were added to four clean and unused Microvette CB300 blood sampling devices and incubated at room temperature for 6 hours. The contents were then quantitatively collected and 3 µL of the spiked master mixture was spiked (see below).

[0398] The spiked master mixture was prepared using ID spiking molecules, Spank molecules, and Spark molecules as described above.

[0399] Nucleic acid library generation directly from plasma: 25.0 µL of each spiked sample was mixed with 10.0 µL of 10x terminal transferase reaction buffer (New England Biolabs, Ipswich, MD), 2.5 µL of proteinase K (Sigma-Aldrich), 1.0 µL of 10% Tween-20 (Thermo Fisher Scientific, Waltham, MA), 1.0 µL of 10% Triton X100 (Thermo Fisher Scientific, Waltham, MA) and 60.5 µL of nuclease-free water. The mixture was heated to 60°C for 20 minutes and then to 95°C for 10 minutes, and then placed on ice until cooled. Add 1.0 µL of 10 mM dATP, 1.0 µL of terminal transferase (20 u / µL, New England Biolabs, Ipswich, MD) and 3.0 µL of nuclease-free water to prepare the A-tail reaction, and then incubate at 37 °C for 40 min. Add 150.0 µL of lysis / binding buffer (Thermo Fisher Scientific, Waltham, MA) to the reaction. Then add the entire volume to 25.0 µL of Dynabeads oligonucleotide (dT)25 (Thermo Fisher Scientific, Waltham, MA) and wash once with lysis / binding buffer (Thermo Fisher Scientific, Waltham, MA). Incubate the mixture at 25 °C and 600 RPM. The remainder of the procedure follows the steps of the protocol outlined in Example 1.

[0400] Sequencing: The sample was sequenced using an Illumina NextSeq™ 500 sequencer to obtain sequence reads. Sequencing was performed according to the manufacturer’s instructions. The sequencing analysis was performed as described in Example 1 above.

[0401] Results: Figure 18A shows the normalized fragment length distribution of microorganisms detected in venous aspirations from two donors in this study, and Figure 18B shows the detection in one of the capillary aspirations from replicas of the same two donors.The normalized fragment length distribution of the microorganisms was obtained. Two microorganisms (e.g., Haemophilus influenzae in donor 1 and Streptococcus thermophilus in donor 2) detected in venous sampling were also detected in biological samples obtained during the capillary collection process, as described on pages 64 / 68 of the manual (CN 121575489 A). These two microorganisms showed similar fragment length distributions in both collection types, i.e., peaked fragment length distributions (Figures 18A and 18B). Additional microorganisms detected in samples obtained using the methods applied during capillary collection comprised a more diverse group of microorganisms (Table 3). Most of these additional microorganisms were co-occurring in both replicas / each donor (Figure 18C). To confirm that these additional microorganisms were not caused by contaminants present in the Microvette CB300 blood sampling device used to collect samples obtained from the procedures applied during capillary collection or derived from process contamination, sequencing data obtained from negative Microvette samples were analyzed (see above). Figure 18D shows a comparison of the abundance (x-axis) of additional microorganisms in MPM in biological samples obtained from the procedure applied during capillary blood collection with the abundance of the same microorganisms in negative Microvette samples. The vast majority of the signals from the additional microorganisms in the data obtained from capillary blood collection were not caused by tube contamination spectra, and it can be concluded that the vast majority of the signals are derived from biological samples obtained by collecting blood drops from the fingertip. Since the signals of these microorganisms were not detected in venous aspiration, they must have originated from the skin surface on which blood spread after the fingertip skin was pricked. This suggests that skin-derived microbial nucleic acids exhibit different properties in their fragment length distribution, such as the absence of a peak between 20 bp and 75 bp, and a class-exponential decay in fragment frequency with fragment length. The same trend was observed in other sample donors (data not shown).

[0402] Table 3: List of microbial species detected in biological samples obtained from the process applied during capillary blood extraction from donor 1 and donor 2. Specification 65 / 68 pages 68 CN 121575489 A

[0403] Example 10: Post-transplant infection surveillance Ten transplant patients were monitored for possible infection after transplant surgery, and changes in the fragment length distribution of pathogens detected in the pre-symptomatic stage were monitored to correlate infection stage with observed fragment lengths. Specifically, as infection progressed to different stages, the presence of a peak between 20 bp and 75 bp and the fraction of fragments not associated with this peak were tracked. In addition to these 10 transplant patients, 10 delabeled consecutive sampling sets generated by Karius were selected to track the same behavior.

[0404] Example 11: Localization site assessmentA template-switching direct-to-library method was used to spike and process 1000 delabeled samples from Karius, along with assay and environmental controls, using proteinase K as described in U.S. Provisional 62 / 770,181. The 1000 delabeled samples comprised plasma samples from patients with pneumonia, immunocompromised conditions, endocarditis, sepsis, or invasive fungal infections. Umbilical cord abundance and microbial and host fragment length distributions were analyzed to correlate fragment length distribution characteristics (e.g., the presence or absence of peaks between 20 and 75 bp, the fraction of reads longer than 65 bp, and the fraction of reads shorter than 40 bp) with infection sites, particularly with the presence of or symbiotic peaks in deep tissue infections. (See specification 66 / 68, page 69, CN 121575489 A).

[0405] Example 12: Infection Stage Determination by Microbial Fragment Length Distribution To determine the diagnostic predictability value of the fragment length spectrum used to measure infection stage, a set of clinical plasma samples were collected from 16 different consenting subjects suspected of having an infection, following the manufacturer's instructions, by drawing blood into PPT tubes and extracting plasma through a single centrifugation step. The plasma samples were frozen or transported overnight at ambient temperature to the Karius lab in Redwood City, CA. For each subject, a first sample was obtained upon admission to the hospital, at which time orthogonal experiments (e.g., blood cultures) were also performed to identify the likely microbial species responsible for or partially responsible for the infection. Subsequently, additional samples were drawn from the subjects at various time points during treatment to monitor infection progression and treatment efficacy. In total, samples were collected at least two time points per subject (including the time point of admission). The maximum number of time points per subject was 7. The plasma samples and negative control samples were processed into nucleic acid libraries and sequenced as described above.

[0406] The subjects in this study included 3 patients orthogonally diagnosed with bloodstream infection, 8 patients orthogonally diagnosed with endocarditis, and 5 patients orthogonally diagnosed with febrile granulocytopenia. Figures 19A, 19B, and 19C show the changes in fragment length distribution in representative examples of bloodstream infection, endocarditis, and febrile granulocytopenia, respectively. The example fragment length distributions in Figure 19 indicate a high probability of short exponentially distributed fragments (range < 40 bp) and an increased probability of a peaked distribution of around 50 bp after treatment has begun. Therefore, the fractions of short exponentially distributed or near-exponentially distributed fragments in all treated samples were investigated. Figure 20A depicts the dynamics of this change in short read fractions. This suggests that invasive infections, in cases of bloodstream infection or bacteremia, can be diagnosed based on the presence of short and exponentially distributed read fractions.This is especially true in the following cases. High read fractions of >64 bp were observed in individual subjects, which may indicate saturation of the mechanisms for obtaining short, exponentially distributed fragments (data not shown). Simultaneous measurement of microbial abundance (Fig. 20B) allowed for the determination of the infection stage by combining abundance and fragment length spectrum measurements.

[0407] Sequencing data also indicated the presence of microorganisms not orthogonally confirmed by other performed microbial tests. In the case of these microorganisms, fragment length distributions can also be investigated. For example, *Haemophilus influenzae* and *Prevotella melanogaster* were detected, respectively, in admission samples from subjects RD-06 and RD-13 using the disclosed method (Fig. 21A). Although orthogonal detection of microorganisms was observed, the presumed cause of infection showed high short read scores in both cases, with other microorganisms showing variable trends; the fragment length distribution of Haemophilus influenzae was consistent with invasive or bacteremia-related infections, while Prevostonia niger only showed a peaked distribution, which is consistent with the elusive phase or symbiotic behavior of infection in asymptomatic patients (see, for example, Helicobacter pylori fragment length distribution in U.S. Provisional Application No. 62 / 770,181, filed November 21, 2018, entitled "Methods, Systems, and Compositions for Direct Access to Libraries"). Or the infection footprint is consistent with that managed. Additionally, new microorganisms may emerge during treatment, and fragment length analysis may also assist in diagnosing these infection states. For example, Figure 21B shows the fragment length distribution of reads compared to *Enterococcus quaternus*, which shows detectable fractions of reads with a short exponential distribution of a series of peaks. The decision to treat this infection can be based on the magnitude of the short read fractions. Examination of the clinical record confirms that the subject has indeed received treatment for this infection.

[0408] Finally, the changes in human fragment length distribution during the treatment period were analyzed as subjects moved from the symptomatic stage of infection at admission and diagnosis to the infection cycle and into the treatment stage. Figure 22 depicts the three main behavioral patterns of human fragment distribution in infected patients in this study: (1) the fraction of long (major nucleosome) human fragments decreased during treatment (left panel of Figure 22, 37.5% of all subjects in this study); (2) the fraction of long human reads fluctuated during treatment (middle panel of Figure 22, 37.5% of all subjects in this study); and (3) the fraction of long (major nucleosome) human reads increased during treatment (right panel of Figure 22, 37.5% of all subjects in this study). As shown above, the shape and nature of the human fragment length distribution can predict the stage of infection in subjects. Parameters derived from the human distribution can then be compared with those detected in the sample.Fragment lengths of the infecting microorganism or other microorganisms are used in combination to predict a subject's recovery trajectory, such as whether the subject is recovering, whether another microorganism will infect the subject during treatment of the initial infection, or whether an invisible infection or symbiotic relationship is recognized. Instruction manual page 68 / 68, page 71, CN 121575489 A, Figure 1; Instruction manual figure 1 / 47, page 72, CN 121575489 A, Figure 2; Instruction manual figure 2 / 47, page 73, CN 121575489 A, Figure 3; Instruction manual figure 3 / 47, page 74, CN 121575489 A, Figure 4; Instruction manual figure 4 / 47, page 75, CN 121575489 A, Figure 5; Instruction manual figure 5 / 47, page 76, CN 121575489 A, Figure 6; Instruction manual figure 6 / 47, page 77, CN 121575489 A, Figure 7; Instruction manual figure 7 / 47, page 78, CN 121575489 A, Figure 8; Instruction manual figure 8 / 47, page 79, CN 121575489 A, Figure 9; Instruction manual figure 9 / 47, page 80, CN 121575489 A, Figure 10. Figure 11, Figure 12, Figure 13, Figure 14, Figure 15A, Figure 15B, Figure 15C, Figure 15F, Figure 15E, Figure 15E, Figure 15F, Page 89, CN 121575489 A, Page 10 / 47, Page 81, CN 121575489 A ... Instruction manual figures 19 / 47, page 90, CN 121575489 A, Figure 16A; Instruction manual figures 20 / 47, page 91, CN 121575489 A, Figure 16B; Instruction manual figures 21 / 47, page 92, CN 121575489 A, Figure 16C; Instruction manual figures 22 / 47, page 93, CN 121575489 AFigure 16D Instructio...

Claims

1. A fragment length profile from a nucleic acid library, wherein the nucleic acid library is generated from an initial sample, and wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparation of the nucleic acid library, wherein the fragment length profile comprises one or more characteristics selected from the group comprising: shape of the distribution, amplitude of bins, peak shape, ratio of fragment counts of two or more bins, height of helical phasing peaks, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads.

2. A method of generating a fragment length profile of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample using a bias- corrected recovery method; (b) determining the number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of bins, peak shape, ratio of fragment counts of two or more bins, height of helical phasing peaks, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.

3. A method of generating a fragment length profile of a nucleic acid library, the method comprising the steps of: (a) preparing a nucleic acid library from an initial sample, the preparing a nucleic acid library from an initial sample comprising: (i) adding one or more process control molecules to the initial sample to provide a spiked initial sample; and (ii) generating a nucleic acid library from the spiked initial sample, wherein nucleic acids used to generate the nucleic acid library are not extracted from the initial sample prior to preparation of the nucleic acid library; (b) determining the number of reads of a plurality of fragment lengths within the nucleic acid library; (c) determining one or more fragment length characteristics of the nucleic acid library, wherein the one or more fragment length characteristics are selected from the group comprising: shape of the distribution, amplitude of bins, peak shape, ratio of fragment counts of two or more bins, height of helical phasing peaks, ratio of fragment counts at two different fragment lengths, ratio of fragment counts within two different fragment length ranges, fragment length range within a bin, ratio of maximum amplitudes of two or more bins, and fragment length distribution within a subset of reads; and (d) generating a fragment length profile of the nucleic acid library using the one or more fragment length characteristics.

4. A method of identifying a localization site in a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample; (b) comparing the fragment length profile to a reference fragment length profile of one or more source sites; and (c) if the fragment length profile from the sample is similar to a fragment length profile from a first source site, then the first site is identified as a localization site; if the fragment length profile from the sample is similar to a fragment length profile from a second source site, then the second site is identified as a localization site.

5. A method of monitoring toxicity of a compound administered to a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample; and (b) comparing the fragment length profile to one or more reference fragment length profiles.

6. A method of determining a stage of infection in a subject, the method comprising the steps of: (a) generating a fragment length profile of a nucleic acid library generated from a sample obtained from the subject; (b) comparing the fragment length profile to a reference fragment length profile; and (c) if the fragment length profile from the sample is similar to a fragment length profile from a symptomatic subject, then the stage of infection is determined to indicate an increased risk that the subject exhibits a symptom associated with the microbe; if the fragment length profile from the sample is similar to a fragment length profile from an asymptomatic subject, then the infection is determined to be in an asymptomatic stage.

7. A method for determining a stage of infection in a subject suspected of having a microbial infection, the method comprising: a) performing high-throughput sequencing of nucleic acids from a biological sample; b) performing bioinformatic analysis to identify free nucleic acid sequences present in the biological sample; and c) obtaining a measurement of the free nucleic acids and comparing the measurement to a control, whereby a stage of infection of the microbe identified in the biological sample is determined.

8. A method of determining a stage of infection of Helicobacter pylori in a subject, the method comprising: (b) extracting free nucleic acids from a biological sample obtained from the subject; (c) adding synthetic nucleic acid spike-in to the free fraction; (d) performing high-throughput sequencing of nucleic acids from the biological sample; (e) performing bioinformatic analysis to identify free Helicobacter pylori nucleic acid sequences present in the biological sample; and (f) calculating a measurement of the free Helicobacter pylori nucleic acids and comparing the measurement to a control, whereby a stage of infection of Helicobacter pylori in the subject is determined.

9. A method of determining a host-microbe biological interaction in a subject, the method comprising: (a) generating a fragment length profile of a nucleic acid library generated from a sample from the subject; (b) optionally, determining an abundance of a target nucleic acid and comparing the abundance to a threshold; (c) comparing the fragment length profile to one or more reference fragment length profiles of host-microbe biological interactions; and (d) if the fragment length profile is similar to a reference fragment length profile of a host-microbe biological interaction, then the host-microbe biological interaction is identified.

10. A method of identifying a presence of a viral infection in a subject suspected of having a microbial infection, the method comprising: a) generating a fragment length profile of a nucleic acid library generated from a sample from the subject; b) comparing the fragment length profile to a viral reference fragment length profile; c) optionally, quantifying the abundance of the target nucleic acid and comparing the abundance to a threshold value; d) identifying the presence of a viral infection in the subject if the fragment length profile is similar to the reference profile.