Methods for processing nucleic acid-containing samples
The method uses a template-free DNA polymerase to add nucleotides step-wise to DNA molecules, addressing contamination issues in nucleic acid sample processing and improving detection accuracy by distinguishing between native and contaminating nucleic acids.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NOSCENDO GMBH
- Filing Date
- 2024-10-14
- Publication Date
- 2026-04-23
AI Technical Summary
Current methods for processing nucleic acid samples, particularly in metagenomic next-generation sequencing, suffer from contamination issues due to laboratory and reagent contaminants, leading to false positive results and biases, and existing decontamination protocols are costly and time-consuming.
A method involving the use of a template-free DNA polymerase to enzymatically add nucleotides step-wise to the 3' end of DNA molecules in a sample, followed by sequence determination, allowing for accurate identification of nucleic acid sources without the need for costly decontamination procedures.
This approach effectively distinguishes between native and contaminating nucleic acids, reducing false positives and biases, and enhances the accuracy of nucleic acid detection and identification in samples.
Smart Images

Figure 00000036_0000 
Figure 00000036_0001 
Figure 00000036_0002
Abstract
Description
[0001] METHODS FOR PROCESSING NUCLEIC ACID-CONTAINING SAMPLES
[0002] FIELD
[0003] The present invention relates to a method for processing or analyzing samples containing nucleic acids, such as biological samples from a subject or samples obtained from the environment, using a DNA polymerase to add nucleotides in a step-wise fashion to the 3’ end of DNA molecules present in the sample, as well as to DNA molecules derived from other nucleic acids in the sample, such as those obtained by reverse-transcribing RNA molecules present in the sample. The methods of the invention allow for highly accurate levels of detection and identification of the source of the nucleic acids in the sample, even in the presence of contaminating nucleic acids without the need for expensive and time consuming controls or decontamination procedures.
[0004] BACKGROUND
[0005] Metagenomic next generation sequencing (mNGS) of cell free DNA (cfDNA) has emerged as a high sensitivity and specificity tool for the detection and identification of disease-causing microorganisms in patients with a suspected infection. Microbial nucleic acids are released to the bloodstream during infectious processes thus becoming part of the blood cfDNA fraction, which can be isolated and sequenced. Bioinformatics tools later allow for the taxonomic classification of those sequences being able to identify the microorganisms present within the sample and, in some instances, even to determine whether their abundance significantly differs from the one observed in cohorts of non-infectious control patients. Therefore, these approaches allow for the identification and potential clinical relevance of the microbial species identified. Nevertheless, sample processing steps are known to be critical for the appropriate diagnostic performance of these approaches and the abundance of microbial nucleic acids within laboratory environments and reagents, together with the high sensitivity of these technologies, are a common source of false positive results or of biases introduced during the attempt to correct potential false positive results. The low biomass of cfDNA samples had also been identified and extensively characterized as a critical factor to the increased effect of laboratory environment and reagent contaminants within cfDNA mNGS studies.
[0006] Most common strategies to address laboratory and reagent contaminants focus on utilizing costly microbial DNA free workflows or poorly scalable decontamination protocols, as well as on the extended use of batch negative controls, expected to reflect process-related contaminants that are then subtracted from the study sample results. The latter imposes additional costs per sample batch and, additionally, presents a risk to the overall workflow diagnostic performance due to the potential over- or undercorrection of tire results given the likely differences in the biomass of the study samples and the chosen negative control. Therefore, strategies capable of distinguishing nucleic acid molecules natively present within study samples would constitute the most appropriate approach. This has been proven by bisulfite treatment-based protocols, in which DNA molecules natively present within the study sample are subjected to DNA sequence modifications. Despite their robust performance, the rough treatment required to induce such sequence modifications, as well as the additional costs and time needed, still leave space for improvement in terms of the feasibility and attractiveness of its industrial application. Thus, there remains a need for improved methods to mitigate contamination of samples during handling, which need is fulfilled by the invention described herein.
[0007] SUMMARY
[0008] In an aspect, the present invention provides a method for processing nucleic acids in a sample, said method comprising: (a) contacting DNA molecules obtained from a sample with a template-free DNA polymerase in the presence of one or more types of nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of the DNA molecules present in the sample, in which each nucleotide is added to the 3’ end in a step-wise reaction; and (b) determining the sequence of the contacted DNA molecules.
[0009] In an embodiment, the sample is a biological sample obtained from a subject. In an embodiment, the processing method can be used for detecting the presence of non-subject-derived nucleic acid molecules in the biological sample. In an embodiment, the sample is one that has not been obtained from a subject, but has been obtained from the environment.
[0010] In an embodiment, the step of contacting can occur in the obtained sample. In an embodiment, the step of contacting can occur once the sample has been processed to provide for conditions allowing for the polymerase to add the nucleotides to the DNA molecules. In an embodiment, the step of contacting can occur once the nucleic acids have been isolated from the sample.
[0011] In an embodiment, the method can further comprise a step of reverse transcribing RNA molecules present in the obtained sample, such that the DNA molecules produced by reverse transcription can be contacted with the polymerase.
[0012] In an embodiment, the method can further comprise a step of amplifying the DNA molecules in or derived from the sample prior to the contacting step. In an embodiment, the method can further comprise a step of amplifying the contacted DNA molecules in the obtained sample prior to the step of determining the sequence. In an embodiment, the DNA molecules to be amplified are contacted DNA molecules comprising the one or more added nucleotides.
[0013] In an embodiment, the method can further comprise a step of enriching and / or isolating the contacted DNA molecules comprising the one or more added nucleotides prior to the step of determining the sequence.
[0014] In an embodiment, the method can further comprise a step of removing the one or more added nucleotides from the contacted DNA molecules comprising the one or more added nucleotides prior to the step of determining the sequence. In an embodiment, the step of determining the sequence comprises contacting the DNA molecules can comprise the one or more added nucleotides with at least one oligonucleotide that is complementary to the added nucleotides. In an embodiment, the at least one oligonucleotide can be conjugated to magnetic beads.
[0015] In an embodiment, at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides can be enzymatically added. In an embodiment, at least 4 nucleotides can be enzymatically added. In an embodiment, no more than 10, 9, 8, 7, 6, 5, 4, 3, 2 nucleotides are to be enzymatically added. In an embodiment, only one species of nucleotide can be added in step (a). In an embodiment, at least two different species of nucleotides can be added in step (a).
[0016] In an embodiment, the at least one nucleotide can be a deoxyribonucleotide or ribonucleotide or can be a modified deoxyribonucleotide or a modified ribonucleotide. In an embodiment, the deoxyribonucleotide can be selected from the group consisting of adenine, guanine, cytosine and thymidine. In an embodiment, the ribonucleotide can be selected from the group consisting of adenine, guanine, cytosine and uracil.
[0017] In an embodiment, the sample is a biological sample has been directly obtained from a subject. In an embodiment, the subject can be a mammal or a human.
[0018] In an embodiment, the sample is not a biological sample obtained from a subject, but a sample that has been obtained from locations / objects in the wider environment (outside / separate from the body of a subject) that are suspected of containing a pathogenic organism. For example, such samples can be obtained from the air conditioning system of a private homes or of a residential / commercial building such as a hotel, or from a hospital, or from a care home, or can be obtained from a municipal or private water source, or from open waters such as a pond, a lake, a stream, a river, or from the ocean, or can be obtained from a commercial building such as retail and wholesale stores, or a manufacturing plant such as a food manufacturing or food packaging plant, or from food storage facilities and food markets, and the like.
[0019] In an embodiment, the biological sample can be selected from the group consisting of blood, plasma, pleura, ascites, ocular fluid, urine, lymph, cerebrospinal fluid, and a cell scraping.
[0020] In an embodiment, the sample is not processed to specifically remove nucleic acids prior to step (a), but can be processed to be able to allow for the polymerase to add nucleotides in a stepwise manner to the DNA molecules contained in the sample or obtained by reverse transcription of the RNA molecules in the sample.
[0021] In an embodiment, the template-free DNA polymerase can be a terminal deoxynucleotidyl transferase or functional variant thereof retaining terminal deoxynucleotidyl transferase activity. The amino acid and encoding nucleotide sequences are readily available from databases such as GenBank and are also commercially available.
[0022] In an embodiment, the sequence of the contacted DNA molecules can be determined by high throughput sequencing methods, for example, by nanopore sequencing or the presence of a specific nucleotide sequence can be detected using any appropriate PCR analysis method.
[0023] In an embodiment, the method can further comprise in silico analysis of the obtained sequences. In an embodiment, the analysis can comprise determining the species of origin of the DNA molecules in the sample.
[0024] In an embodiment, the DNA molecules can be cell-free DNA molecules present in the sample.
[0025] In an embodiment, the non-subject-derived nucleic acid molecules can be microbial-, fungal, parasitic- and / or viral-derived nucleic acid molecules.
[0026] In an embodiment, the method can further comprise contacting one or more oligonucleotides to the contacted DNA molecules, wherein at least one oligonucleotide anneals to the nucleotides enzymatically added to the DNA molecules prior to the sequencing step.
[0027] Tn an embodiment, the method for identifying the species of origin of DNA molecules in a sample can comprise (a) contacting DNA molecules obtained from a sample with a template-free DNA polymerase in the presence of one or more types of nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of the DNA molecules present in the sample, in which each nucleotide is added to the 3’ end in a step-wise reaction; (b) determining the sequence of the contacted DNA molecules; and (c) determining the species of origin of the DNA molecules. The DNA molecules can be cell-free DNA molecules in the sample, and determining the species of origin can comprise comparing the obtained sequences to one or more sequence databases and identifying the species of origin of one or more DNA molecules in the sample. The species of origin of the DNA molecules can be a bacterial, viral, parasitic, or fungal species.
[0028] Tn an aspect, the present invention is directed to a method for identifying the pathogen causing a medical condition in a subject, said method comprising: (a) contacting DNA molecules present in a biological sample obtained from the subject with a template-free DNA polymerase in the presence of one or more nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of DNA molecules present in the biological sample, in which each nucleotide is added to the 3’ end in a step-wise reaction; (b) determining the sequence of the contacted DNA molecules to obtain the sequences of the DNA molecules; and (c) comparing the obtained sequences against one or more databases to determine the species identity of the DNA molecules in the sample, wherein when DNA molecules from a pathogenic species are present in the biological sample above a threshold value, the subject is diagnosed with a medical condition caused by a pathogen. In an embodiment, the pathogen can be a bacterium, a parasite, a virus, or a fungus.
[0029] In an aspect, the present invention is directed to a method for identifying a pathogen that could cause a medical condition in a subject, said method comprising: (a) contacting DNA molecules present in an sample obtained from the environment with a template-free DNA polymerase in the presence of one or more nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of DNA molecules present in the sample, in which each nucleotide is added to the 3’ end in a step- wise reaction; (b) determining the sequence of the contacted DNA molecules to obtain the sequences of the DNA molecules; and (c) comparing the obtained sequences against one or more databases to determine the species identity of the DNA molecules in the sample. In an embodiment, the pathogen can be a bacterium, a parasite, a virus, or a fungus.
[0030] In an embodiment of all aspects of the invention, the method can be carried out without any kind of batch correction or measures for laboratory reagent contamination control. In an embodiment, the contacting step can be carried out as soon as possible after obtaining the sample. In an embodiment, the contacting step is carried out as soon as possible after obtaining and processing the sample in order to create appropriate conditions for the polymerase to be able to add the nucleotides to the 3’ end of the DNA molecules.
[0031] DETAILED DESCRIPTION
[0032] Although the present invention is described in detail below, it is to be understood that this invention is not limited to the particular methodologies, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.
[0033] Preferably, the terms used herein are defined as described in “A multilingual glossary of biotechnological terms: (IUPAC Recommendations)”, Leuenberger et al., Eds., Helvetica Chimica Acta, CH-4010 Basel, Switzerland, (1995).
[0034] The practice of the present invention will employ, unless otherwise indicated, conventional methods of chemistry, biochemistry, cell biology, immunology, and recombinant DNA techniques which are explained in the literature in the field (cfy e.g., Molecular Cloning: A Laboratory Manual, 4thEdition, Green and Sambrook, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 2012).
[0035] In the following, the elements of the present invention will be described. These elements are listed with specific embodiments; however, it should be understood that they may be combined in any maimer and in any number to create additional embodiments. The variously described examples and preferred embodiments should not be construed to limit the present invention to only the explicitly described embodiments. This description should be understood to disclose and encompass embodiments which combine the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, any permutations and combinations of all described elements in this application should be considered disclosed by this description unless the context indicates otherwise. Unless otherwise specified or recognized to be inappropriate in achieving the use of the invention, each of the embodiments disclosed herein are applicable to any of the aspects disclosed herein.
[0036] The term “about” means approximately or nearly, and in the context of a numerical value or range set forth herein preferably means + / - 10 % of the numerical value or range recited or claimed.
[0037] The terms “a” and “an” and “the” and similar reference used in the context of describing the invention (especially in the context of the claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it was individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g, “such as”), provided herein is intended merely to better illustrate the invention and does not pose a limitation on the scope of the invention otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the invention.
[0038] Unless expressly specified otherwise, the term “comprising” is used in the context of the present document to indicate that further members may optionally be present in addition to the members of the list introduced by “comprising”. It is, however, contemplated as a specific embodiment of the present invention that the term “comprising” encompasses the possibility of no further members being present, i.e., for the purpose of this embodiment “comprising” is to be understood as having the meaning of “consisting of’.
[0039] Indications of relative amounts of a component characterized by a generic term are meant to refer to the total amount of all specific variants or members covered by said generic term. If a certain component defined by a generic term is specified to be present in a certain relative amount, and if this component is further characterized to be a specific variant or member covered by the generic term, it is meant that no other variants or members covered by the generic term are additionally present such that the total relative amount of components covered by the generic term exceeds the specified relative amount; more preferably no other variants or members covered by the generic term are present at all. Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, etc.), whether supra or infra, are hereby incorporated by reference in their entirety. Nothing herein is to be construed as an admission that the present invention was not entitled to antedate such disclosure.
[0040] Terms such as “reduce” or “inhibit” as used herein means the ability to cause an overall decrease, preferably of 5% or greater, 10% or greater, 20% or greater, more preferably of 50% or greater, and most preferably 75% or greater, in the level. The term “inhibit” or similar phrases includes a complete or essentially complete inhibition, i.e., a reduction to zero or essentially to zero.
[0041] Terms such as “increase” or “enhance” preferably relate to an increase or enhancement by about at least 10%, preferably at least 20%, preferably at least 30%, more preferably at least 40%, more preferably at least 50%, even more preferably at least 80%, and most preferably at least 100%.
[0042] The term “nucleic acid” according to the disclosure also comprises a chemical derivatization of a nucleic acid on a nucleotide base, on the sugar or on the phosphate, and nucleic acids containing non-natural nucleotides and nucleotide analogs. In some embodiments, the nucleic acid is a deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA). In general, a nucleic acid molecule or a nucleic acid sequence refers to a nucleic acid which is preferably deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). According to the disclosure, nucleic acids comprise genomic DNA, cDNA, mRNA, viral RNA, recombinantly prepared and chemically synthesized molecules. According to the disclosure, a nucleic acid may be in the form of a single-stranded or double-stranded and linear or covalently closed circular molecule.
[0043] According to the disclosure, “nucleic acid sequence” refers to the sequence of nucleotides in a nucleic acid, e.g., a ribonucleic acid (RNA) or a deoxyribonucleic acid (DNA). The term may refer to an entire nucleic acid molecule (such as to the single strand of an entire nucleic acid molecule) or to a part (e.g, a fragment) thereof.
[0044] According to the present disclosure, the term “RNA” or “RNA molecule” relates to a molecule which comprises ribonucleotide residues and which is preferably entirely or substantially composed of ribonucleotide residues. The term “ribonucleotide” relates to a nucleotide with a hydroxyl group at the 2’-position of a p-D-ribofiiranosyl group. The term “RNA” comprises double-stranded RNA single stranded RNA, isolated RNA such as partially or completely purified RNA, essentially pure RNA synthetic RNA, and recombinantly generated RNA such as modified RNA which differs from naturally occurring RNA by addition, deletion, substitution and / or alteration of one or more nucleotides.
[0045] “Fragment”, with reference to a nucleic acid sequence, relates to a part of a nucleic acid sequence, i.e., a sequence which represents the nucleic acid sequence shortened at the 5’- and / or 3’-end(s). Preferably, a fragment of a nucleic acid sequence comprises at least 80%, preferably at least 90%, 95%, 96%, 97%, 98%, or 99% of the nucleotide residues from said nucleic acid sequence.
[0046] The term “variant” with respect to, for example, nucleic acid and amino acid sequences, according to the disclosure includes any variants, in particular mutants, viral strain variants, splice variants, conformations, isoforms, allelic variants, species variants and species homologs, in particular those which are naturally present. An allelic variant relates to an alteration in the normal sequence of a gene, the significance of which is often unclear. Complete gene sequencing often identifies numerous allelic variants for a given gene. With respect to nucleic acid molecules, the term “variant” includes degenerate nucleic acid sequences, wherein a degenerate nucleic acid according to the disclosure is a nucleic acid that differs from a reference nucleic acid in codon sequence due to the degeneracy of the genetic code.
[0047] A nucleic acid is “capable of hybridizing” or “hybridizes” to another nucleic acid if the two sequences are complementary with one another. A nucleic acid is “complementary” to another nucleic acid if the two sequences are capable of forming a stable duplex with one another. According to the disclosure, hybridization is preferably carried out under conditions which allow specific hybridization between polynucleotides (stringent conditions). Stringent conditions are described, for example, in Current Protocols in Molecular Biology, F.M. Ausubel et al, Editors, John Wiley & Sons, Inc., New York and refer, for example, to hybridization at 65°C in hybridization buffer (3.5 x SSC, 0.02% Ficoll, 0.02% polyvinylpyrrolidone, 0.02% bovine serum albumin, 2.5 mM NaH2PO4 (pH 7), 0.5% SDS, 2 mM EDTA). SSC is 0.15 M sodium chloride / 0.15 M sodium citrate, pH 7. After hybridization, the membrane to which the DNA has been transferred is washed, for example, in 2 x SSC at room temperature and then in 0.1-0.5 x SSC / 0.1 x SDS at temperatures of up to 68°C.
[0048] A percent complementarity indicates the percentage of contiguous residues in a nucleic acid molecule that can form hydrogen bonds (e.g. , Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 being 50%, 60%, 70%, 80%, 90%, and 100% complementary). “Perfectly complementary” or “fully complementary” means that all the contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. Preferably, the degree of complementarity according to the disclosure is at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, even more preferably at least 90% or most preferably at least 95%, 96%, 97%, 98% or 99%. Most preferably, the degree of complementarity according to the disclosure is 100%.
[0049] The term “derivative” comprises any chemical derivatization of a nucleic acid on a nucleotide base, on the sugar or on the phosphate. The term “derivative” also comprises nucleic acids which contain nucleotides and nucleotide analogs not occurring naturally.
[0050] A “nucleic acid sequence which is derived from a nucleic acid sequence” refers to a nucleic acid which is a variant of the nucleic acid from which it is derived. Preferably, a sequence which is a variant with respect to a specific sequence, when it replaces the specific sequence in an RNA molecule retains RNA stability and / or translational efficiency.
[0051] “nt" is an abbreviation for nucleotide; or for nucleotides, preferably consecutive nucleotides in a nucleic acid molecule.
[0052] According to the disclosure, the term “codon” refers to a base triplet in a coding nucleic acid that specifies which amino acid will be added next during protein synthesis at the ribosome.
[0053] “3’ end of a nucleic acid" refers according to the disclosure to that end which has a free hydroxy group. In a diagrammatic representation of double-stranded nucleic acids, in particular DNA, the 3’ end is always on the right-hand side. “5’ end of a nucleic acid" refers according to the disclosure to that end which has a free phosphate group. In a diagrammatic representation of double-strand nucleic acids, in particular DNA, the 5’ end is always on the left-hand side.
[0054] “Upstream” describes the relative positioning of a first element of a nucleic acid molecule with respect to a second element of that nucleic acid molecule, wherein both elements are comprised in the same nucleic acid molecule, and wherein the first element is located nearer to the 58end of the nucleic acid molecule than the second element of that nucleic acid molecule. The second element is then said to be “downstream” of the first element of that nucleic acid molecule. An element that is located “upstream’8of a second element can be synonymously referred to as being located “5"’ of that second element. For a double-stranded nucleic acid molecule, indications like “upstream” and “downstream” are given with respect to the (+) strand.
[0055] A “polymerase” generally refers to a molecular entity capable of catalyzing the synthesis of a polymeric molecule from monomeric building blocks. An “RNA polymerase” is a molecular entity capable of catalyzing the synthesis of an RNA molecule from ribonucleotide building blocks. A “DNA polymerase” is a molecular entity capable of catalyzing the synthesis of a DNA molecule from deoxy ribonucleotide building blocks. For the case of DNA polymerases and RNA polymerases, the molecular entity is typically a protein or an assembly or complex of multiple proteins. Typically, a DNA polymerase synthesizes a DNA molecule based on a template nucleic acid, which is typically a DNA molecule. Typically, an RNA polymerase synthesizes an RNA molecule based on a template nucleic acid, which is either a DNA molecule (in that case the RNA polymerase is a DNA-dependent RNA polymerase, DdRP), or is an RNA molecule (in that case the RNA polymerase is an RNA-dependent RNA polymerase, RdRP). In the methods of the present invention, the DNA polymerase is able to, in a one-by-one (step-wise) manner, add nucleotides to the 3’ end of DNA molecules without the use of a template / o verhang sequence.
[0056] According to the disclosure, the term “gene” refers to a particular nucleic acid sequence which is responsible for producing one or more cellular products and / or for achieving one or more intercellular or intracellular functions. More specifically, said term relates to a nucleic acid section (typically DNA; but RNA in the case of RNA viruses) which comprises a nucleic acid coding for a specific protein or a functional or structural RNA molecule.
[0057] An “isolated molecule” as used herein, is intended to refer to a molecule which is substantially free of other molecules such as other cellular material. The term “isolated nucleic acid” means according to the disclosure that the nucleic acid has been (i) amplified in vitro, for example by polymerase chain reaction (PCR), (ii) recombinantly produced by cloning, (iii) purified, for example by cleavage and gel- electrophoretic fractionation, or (iv) synthesized, for example by chemical synthesis. An isolated nucleic acid is a nucleic acid available to manipulation by recombinant techniques.
[0058] The term “recombinant” in the context of the present disclosure means “made through genetic engineering”. Preferably, a “recombinant object” such as a recombinant cell in the context of the present disclosure is not occurring naturally.
[0059] The term “naturally occurring” as used herein refers to the fact that an object can be found in nature. For example, a peptide or nucleic acid that is present in an organism (including viruses) and can be isolated from a source in nature and which has not been intentionally modified by man in the laboratory is naturally occurring. The term “found in nature” means “present in nature” and includes known objects as well as objects that have not yet been discovered and / or isolated from nature, but that may be discovered and / or isolated in the future from a natural source.
[0060] According to the disclosure, a “base pair” is a structural motif of a secondary structure wherein two nucleotide bases associate with each other through hydrogen bonds between donor and acceptor sites on the bases. The complementary bases, A:U and G:C, form stable base pairs through hydrogen bonds between donor and acceptor sites on the bases; the A:U and G:C base pairs are called Watson-Crick base pairs. A weaker base pair (called Wobble base pair) is formed by the bases G and U (G:U). The base pairs A:U and G:C are called canonical base pairs. Other base pairs like Gill (which occurs fairly often in RNA) and other rare base-pairs (e.g. A:C; U:U) are called non-canonical base pairs.
[0061] According to the disclosure, “nucleotide pairing” refers to two nucleotides that associate with each other so that their bases form a base pair (canonical or non-canonical base pair, preferably canonical base pair, most preferably Watson-Crick base pair).
[0062] The following provides specific and / or preferred variants of the individual features of the disclosure. The present disclosure also contemplates as particularly preferred embodiments those embodiments, which are generated by combining two or more of the specific and / or preferred variants described for two or more of the features of the present disclosure.
[0063] The terms “genome” or “genomic DNA” or “genomic nucleic acid molecule” are meant to refer to any kind of DNA molecule that is propagated and equally distributed from mother to daughter cells. Genomic DNA refers to both chromosomal DNA and extra-chromosomal DNA such as episomes, preferably non-viral episomes.
[0064] The term “episome” is to be understood to refer to a DNA molecule that remains as a part of the eukaryotic genome without integration. Episomes manage this by replicating together with the rest of the genome and subsequently being distributed like chromosomes to each daughter cell equally.
[0065] Described herein is a method for processing nucleic acids in a sample, said method comprising: (a) contacting DNA molecules obtained from a sample with a template-free DNA polymerase in the presence of one or more types of nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of the DNA molecules present in the sample, in which each nucleotide is added to the 3’ end in a step-wise reaction; and (b) determining the sequence of the contacted DNA molecules. In an embodiment, the processing method can be used for detecting the presence of non- subject-derived nucleic acid molecules in a biological sample obtained from a subject.
[0066] The template-free DNA polymerase can be a terminal deoxynucleotidyl transferase (TdT), also known as DNA nucleotidylexotransferase (DNTT) or terminal transferase, is a specialized DNA polymerase found to be expressed in immature, pre-B, pre-T lymphoid cells, and acute lymphoblastic leukemia / lymphoma cells. TdT adds N-nucleotides to the V, D, and J exons of the TCR and BCR genes during antibody gene recombination, enabling the phenomenon of junctional diversity. Generally, TdT catalyzes the addition of nucleotides to the 3' terminus of a DNA molecule. Unlike most DNA polymerases, it does not require a template. The preferred substrate of this enzyme is a 3'overhang, but it can also add nucleotides to blunt or recessed 3* ends. Further, TdT is the only polymerase that is known to catalyze the synthesis of2-15nt DNA polymers from free nucleotides in solution in vivo. In vitro, this behavior catalyzes the general formation of DNA polymers without specific length. Like many polymerases, TdT requires a divalent cation cofactor, however, TdT is unique in its ability to use a broader range of cations such as Mg2+, Mn2+, Zn2+ and Co2+. The rate of enzymatic activity depends on the available divalent cations and the nucleotide being added.
[0067] Several isoforms of TdT have been observed in mice, bovines, and humans. To date, two variants have been identified in mice while three have been identified in humans. The amino acid and encoding nucleotide sequences are known in the art, for example, GenBank Accession Nos. NP_001017520.1, NP_004079.3, P04053.3, AAA36726.1, AAF26284.1, NP_033371.2, AAK07884.1 and is available commercially.
[0068] The nucleotides to be added to the 3’ end of the DNA molecules can be any nucleotide so long as the added nucleotides can assist in the ability to identify the DNA molecules in the sample. For example, any one or more of adenosine, cytosine, guanosine, thymidine can be used. When only a single kind of nucleotide is used, the result is the addition of an homopolymer of the nucleotide. Modified deoxynucleotides and / or modified ribonucleotides can also be used, including those that have a modification of the sugar and / or backbone and / or base modifications. In certain embodiments, the modified nucleotides added to the 3 ’ end of DNA molecules are able to base-pair with other nucleotides, preferably via Watson-Crick pairing, allowing for other nucleic acid molecules to anneal to the added nucleosides / nucleotides.
[0069] Sugar Modifications: The modified nucleosides and nucleotides, which may be added to the 3’ end of DNA molecules by a TdT can be modified in the sugar moiety. For example, the 2’ hydroxyl group (OH) can be modified or replaced with a number of different “oxy” or “deoxy” substituents. Examples of “oxy” -2’ hydroxyl group modifications include, but are not limited to, alkoxy or aryloxy (-OR, e.g., R = H, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or sugar); polyethyleneglycols (PEG), - O(CH2CH2O)nCH2CH2OR; "locked" nucleic acids (LNA) in which the 2’ hydroxyl is connected, e.g., by a methylene bridge, to the 4' carbon of the same ribose sugar; and amino groups (-O-amino, wherein the amino group, e.g., NRR, can be alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroaryl amino, ethylene diamine, polyamino) or aminoalkoxy. “Deoxy” modifications include hydrogen, amino (e.g. NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diaryl amino, heteroaryl amino, diheteroaryl amino, or amino acid); or the amino group can be attached to the sugar through a linker, wherein the linker comprises one or more of the atoms C, N, and O. The sugar group can also contain one or more carbons that possess the opposite stereochemical configuration than that of the corresponding carbon in ribose.
[0070] Backbone Modifications: The phosphate backbone may further be modified in the modified nucleosides and nucleotides, which may be added to the 3’ end of a DNA molecule by a TdT. The phosphate groups of the backbone can be modified by replacing one or more of the oxygen atoms with a different substituent. Further, the modified nucleosides and nucleotides can include the full replacement of an unmodified phosphate moiety with a modified phosphate as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioate, phosphoroselenates, borano phosphates, borano phosphate esters, hydrogen phosphonates, phosphoroamidates, alkyl or aryl phosphonates and phosphotriesters. Phosphorodithioates have both non-linking oxygens replaced by sulfur. The phosphate linker can also be modified by the replacement of a linking oxygen with nitrogen (bridged phosphoroamidates), sulfur (bridged phosphorothioates) and carbon (bridged methylene - phosphonates).
[0071] Base Modifications: The modified nucleosides and nucleotides, which may be added to the 3' end of a DNA molecule by a TdT can further be modified in the nucleobase moiety. Examples of nucleobases found in RNA include, but are not limited to, adenine, guanine, cytosine and thymidine. For example, the nucleosides and nucleotides described herein can be chemically modified on the major groove face. In some embodiments, the major groove chemical modifications can include an amino group, a thiol group, an alkyl group, or a halo group. In particular embodiments of the present disclosure, the nucleotide analogues / modifications are selected from base modifications, which are preferably selected from 2-amino-6-chloropurineriboside-5’- triphosphate, 2-aminopurine-riboside-5 ’-triphosphate; 2-aminoadenosine-5’-triphosphate, 2’-amino-2’- deoxy- cytidine-triphosphate, 2-thiocytidine-5’ -triphosphate, 2 ’-fluorothymidine-5’ -triphosphate, 5- aminoallylcytidine-5'-triphosphate, 5-bromocytidine-5'-triphosphate, 5-bromo-2'-deoxycytidine-5'- triphosphate, 5 -iodocytidine-5 '-triphosphate, 5-iodo-2'-deoxycytidine-5'-triphosphate, 5- methylcytidine-5 '-triphosphate, 5 -propynyl-2'-deoxycytidine-5 '-tri-phosphate, 6-azacytidine-5'- triphosphate, 6-chloropurineriboside-5'-triphosphate, 7-deaza-adenosine-5'-triphosphate, deazaguanosine-5'-triphosphate, 8-azaadenosine-5'-triphosphate, 8-azidoadenosine-5'-triphosphate, benzimidazole-riboside-5'-triphosphate, N1 -methyladenosine-5 '-triphosphate, N1 -methylguanosine-5'- triphosphate, ' N6-methyladenosine-5'-triphosphate, 06-methylguanosine-5'-triphosphate, N6- methylguanosine-5 '-triphosphate, or puromycin-5 '-triphosphate, xanthosine-5 '-triphosphate. Particular preference may be given to nucleotides for base modifications selected from the group of base-modified nucleotides consisting of S-methylcytidine-S'-triphosphate, 7-deazaguanosine-5'-triphosphate, and 5- bromocytidine-5 -triphosphate.
[0072] In some embodiments, modified nucleosides include 5-aza-cytidine, pseudoisocytidine, 3-methyl- cytidine, N4-acetylcytidine, 5-formylcytidine, N4- methylcytidine, 5 -hydroxymethylcytidine, 1-methyl- pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5 -methyl- cytidine, 4-thio-pseudoisocytidine, 4-thio- 1 -methyl-pseudoisocytidine, 4-thio- 1 -methyl- 1 -deaza- pseudoisocytidine, 1-methyl-l-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl- zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5 -methyl- cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-l-methyl-pseudoisocytidine.
[0073] In other embodiments, modified nucleosides include 2-aminopurine, 2,6-diaminopurine, 7-deaza- adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6- diaminopurine, 7-deaza-8-aza-2,6-diamino-purine, 1 -methyladenosine, N6-methyladenosine, N6- isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6 -(cis- hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2- methyl-thio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2- methylthio-adenine, and 2-methoxy-adenine. In other embodiments, modified nucleosides include inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio- guanosine, 6-thio-7 -deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7- methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1 -methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, l-methyl-6-thio-guanosine, N2-methyl-6- thio-guanosine, and N2,N2-dhnethyl-6-thio-guanosine.
[0074] In specific embodiments, a modified nucleoside is 5'-O-(l-thiophosphate)-adenosine, 5'-O-(l- thiophosphate)-cytidine, or 5'-O-(l-thiophosphate)-guanosine. In some embodiments, modified nucleosides include 5-aza-uridine, 2-thio-5 -aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5 -hydroxyuridine, 3-methyluridine, 5 -carboxymethyl- uridine, 1 -carboxymethyl-pseudouridine, 5-propynyl-uridine, 1 -propynyl-pseudouridine, 5- taurinomethyluridine, 1 -taurinomethyl-pseudouridine, 5-taurinomethyl-2 -thiouridine, 1-taurinomethyl- 4-thio-uridine, 5 -methyl -uridine, 1 -methyl -pseudouridine, 4-thio-l-methyl-pseudouridine, 2-thio-l- methyl-pseudouridine, 1 -methyl- 1 -deaza-pseudouridine, 2-thio- 1 -methyl- 1 -deaza-pseudouridine, dihydrouridine, dihydro-pseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2- methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio- pseudouridine.
[0075] As appropriate, any of the foregoing listed modified nucleotide / nucleosides can be either the deoxy- or ribo-version of the modified nucleotide.
[0076] The biological sample includes any biological sample obtained from a subject, e.g., from the body of the subject. Examples of such biological samples include whole blood, blood fractions such as plasma, serum, smears or swabs of a tissue, sputum, bronchial aspirate, urine, semen, stool, bile, gastrointestinal secretions, reproductive system secretions, lymph fluid, liquor, bone marrow, organ aspirates and tissue biopsies, including punch biopsies. Optionally, the biological sample can be obtained from a mucous membrane of the patient. Biological samples can also include processed biological samples such as fractions or isolates, e.g., nucleic acids or isolated cells. Preferably, the biological sample contains nucleic acids, e.g., genomic DNA or mRNA, such that the sequence of the nucleic acids can be determined. In an embodiment, the biological sample can be one that is obtained from a tissue showing signs of a disease state, e.g., showing signs of infection. In a preferred embodiment, the biological sample is blood or blood plasma obtained from the subject. The sample is analyzed according to the methods of the invention and during the method or thereafter is not normally returned to the body. In most embodiments, the presence of the subject’s body is not necessary in order to carry out the methods of the invention.
[0077] In one embodiment, the biological sample is blood plasma, preferably obtained directly from the subject. The blood plasma is preferably cell-free, preferably mainly / mostly cell-free, e.g., fewer than 10,000, 1,000, 100, or 10 cells per mL. The biological sample, e.g., blood plasma, may contain free circulating nucleic acids, comprising nucleic acids of the subject and nucleic acids not of the subject, e.g., those of a microorganism. In one embodiment the biological sample can be diluted or concentrated. In another embodiment the sample is processed prior to the contacting step, preferably the sample is purified to remove cellular components, such as lipids and proteins, prior to contacting with the DNA polymerase.
[0078] Tissues of the patient from which the biological sample can be obtained include, but are not limited to, throat, mouth, nasal, stomach, intestinal, skin, liver, pancreatic, lung, neuronal cervical, vaginal, urethral, rectal, penial, and muscle. Any suitable method for obtaining the biological sample from the patient and / or from an appropriate tissue can be used in connection with the present invention. The non-biological sample, i.e., a sample not obtained from a subject, can be a sample that has been obtained from objects outside / separate from the body of a subject that are suspected of containing a pathogenic organism. For example, such samples can be obtained from the air conditioning system from private homes or residential / commercial buildings such as hotels, or from hospitals, or from care homes, or can be obtained from municipal or private water sources, or from open waters such as ponds, lakes, streams, rivers, or from the ocean, or can be obtained from commercial buildings such as retail and wholesale stores, or manufacturing plants such as food manufacturing or food packaging plants, or from food storage facilities and food markets, and the like.
[0079] The sample that has been obtained can be processed in order to perform the steps of the method of the invention, for example, providing for conditions for the DNA polymerase to be enzymatically active and / or isolate the nucleic acids in the sample. Methods for doing so are known in the art, see, for example, Niu et al. 2018, Recent advances in sample preparation methods coupled with chromatography, spectrometry and electrochemistry analysis techniques, TrAC Trends in Analytical Chemistry 102 : 123- 146.
[0080] Methods for determining the sequence of the contacted DNA molecules are known in the art. In context of the present invention, the term “sequencing” means to determine the sequence of at least one nucleic acid, and it includes any method that is used to determine the order of the bases in a strand of at least one nucleic acid. A preferred method of sequencing is high-throughput sequencing, such as next- generation sequencing or third generation sequencing.
[0081] For clarification purposes: the terms “Next Generation Sequencing” or “NGS” in the context of the present invention mean all high throughput sequencing technologies which, in contrast to the “conventional” sequencing methodology known as Sanger chemistry, read nucleic acid templates randomly in parallel along the entire genome by breaking the entire genome into small pieces. Such NGS technologies (also known as massively parallel sequencing technologies) are able to deliver nucleic acid sequence information of a whole genome, exome, transcriptome (all transcribed sequences of a genome) or methylome (all methylated sequences of a genome) in very short time periods, e.g., within 1-2 weeks, preferably within 1-7 days or most preferably within less than 24 hours and allow, in principle, single cell sequencing approaches. Multiple NGS platforms which are commercially available or which are mentioned in the literature can be used in the context of the present invention, e.g. , those described in detail in Zhang et al., 2011, The impact of next-generation sequencing on genomics. J. Genet Genomics 38:95-109; or in Voelkerding et al., 2009, Next generation sequencing: From basic research to diagnostics, Clinical chemistry 55:641-658. Non-limiting examples of such NGS technologies / platfonns are
[0082] 1) The sequencing-by-synthesis technology known as pyrosequencing implemented, e.g., in the GS-FLX 454 Genome Sequencer™ of Roche-associated company 454 Life Sciences (Branford, Connecticut), first described in Ronaghi et al., 1998, A sequencing method based on real-time pyrophosphate, Science 281:363-365. This technology uses an emulsion PCR in which single- stranded DNA binding beads are encapsulated by vigorous vortexing into aqueous micelles containing PCR reactants surrounded by oil for emulsion PCR amplification. During the pyrosequencing process, light emitted from phosphate molecules during nucleotide incorporation is recorded as the polymerase synthesizes the DNA strand.
[0083] 2) The sequencing-by-synthesis approaches developed by Solexa (now part of Illumina Inc., San Diego, California) which is based on reversible dye-terminators and implemented, e.g. , in the Illumina / Solexa Genome Analyzer™ and in the Illumina HiSeq 2000 Genome Analyzer™. In this technology, all four nucleotides are added simultaneously into oligo-primed cluster fragments in flow-cell channels along with DNA polymerase. Bridge amplification extends cluster strands with all four fluorescently labeled nucleotides for sequencing.
[0084] 3) Sequencing-by-ligation approaches, e.g., implemented in the SOLid™ platform of Applied Biosystems (now Life Technologies Corporation, Carlsbad, California). In this technology, a pool of all possible oligonucleotides of a fixed length are labeled according to the sequenced position. Oligonucleotides are annealed and ligated; the preferential ligation by DNA ligase for matching sequences results in a signal informative of the nucleotide at that position. Before sequencing, the DNA is amplified by emulsion PCR. The resulting bead, each containing only copies of the same DNA molecule, are deposited on a glass slide. As a second example, the Polonator™ G.007 platform of Dover Systems (Salem, New Hampshire) also employs a sequencing-by-ligation approach by using a randomly arrayed, bead-based, emulsion PCR to amplify DNA fragments for parallel sequencing.
[0085] 4) Single-molecule sequencing technologies such as, e.g., implemented in the PacBio RS system of Pacific Biosciences (Menlo Park, California) or in the HeliScope™ platform of Helicos Biosciences (Cambridge, Massachusetts). The distinct characteristic of this technology is its ability to sequence single DNA or RNA molecules without amplification, defined as Single- Molecule Real Time (SMRT) DNA sequencing. For example, HeliScope uses a highly sensitive fluorescence detection system to directly detect each nucleotide as it is synthesized. A similar approach based on fluorescence resonance energy transfer (FRET) has been developed from Visigen Biotechnology (Houston, Texas). Other fluorescence-based single-molecule techniques are from U.S. Genomics (GeneEngine™) and Genovoxx (AnyGene™).
[0086] 5) Nano-technologies for single-molecule sequencing in which various nanostructures are used which are, e.g., arranged on a chip to monitor the movement of a polymerase molecule on a single strand during replication. Non-limiting examples for approaches based on nano- technologies are the GridON™ platform of Oxford Nanopore Technologies (Oxford, UK), the hybridization-assisted nano-pore sequencing (HANS™) platforms developed by Nabsys (Providence, Rhode Island), and the proprietary ligase-based DNA sequencing platform with DNA nanoball (DNB) technology called combinatorial probe-anchor ligation (cP AL™).
[0087] 6) Electron microscopy based technologies for single-molecule sequencing, e.g., those developed by LightSpeed Genomics (Sunnyvale, California) and Halcyon Molecular (Redwood City, California)
[0088] 7) Ion semiconductor sequencing which is based on the detection of hydrogen ions that are released during the polymerization of DNA. For example, Ion Torrent Systems (San Francisco, California) uses a high-density array of micro-machined wells to perform this biochemical process in a massively parallel way. Each well holds a different DNA template. Beneath the wells is an ion-sensitive layer and beneath that a proprietary Ion sensor.
[0089] Other sequencing methods useful in the context of the invention include tunneling currents sequencing (Xu et al., 2007, The electronic properties of DNA bases, Small 3:1539-1543, Di Ventra, 2013, Fast DNA sequencing by electrical means inches closer, Nanotechnology 24:342501). Particularly preferable next-generation sequencing (NGS) methodologies include Illumina, lONTorrent and NanoPore sequencing.
[0090] Once the nucleic acids have been sequenced, the resulting sequences (sequenced reads) can be compared to one or more databases comprising the genetic information preferably from multiple species, such that the sequenced reads can be determined to be from a particular species, such as the subject and / or from a particular microorganism. Methods for mapping sequenced reads to provide information on their species of origin are well known in the art, and any such suitable method can be used in connection with the present invention. For example, the Kraken ultrafast metagenomics sequence classification methodology described in Wood and Salzberg, 2014, Genome Biol 15:R46 can be used. Another exemplary method is NextGenMap which is described in Sedlazeck et al., 2013, Bioinfonnatics 29:2790-2791. Yet another exemplary method is a cloud-compatible bioinformatics pipeline for ultra- rapid pathogen identification from next-generation sequencing of clinical samples as described in Naccache et al., 2014, Genome Res 24: 1180-1192. Addition methods known in the art and useful in the present invention include, but are not limited to those described in Huson et al., 2007, Genome Res 17:377-386; Freitas et al, 2015, Nucl Acids Res 43:e69; and Kim et al., 2016, Genome Res 26:1721- 1729.
[0091] In certain embodiments of the invention, in order to reduce the number of false positive findings in detecting and comparing sequences, it is preferred to determine / compare the sequences in replicates. Thus, it is preferred that nucleic acid sequences in a sample be determined twice, three times or more. Technical repeats of a sample should generate identical results and any detected mutation in this “same vs. same comparison” is a false positive. Furthermore, various quality related metrics (e.g., coverage or SNP quality) may be combined into a single quality score using a machine learning approach. In a preferred embodiment, the method is carried out without performing any controls for subsequent contamination of the sample, such a batch control.
[0092] Database relates to an organized collection of data, preferably as an electronic filing system. In an embodiment, a sequence database is a type of database that is composed of a collection of computerized (“digital”) nucleic acid sequences, protein sequences, or other polymer sequences stored on a computer. Preferably, the database is a collection of nucleic acid sequences, i.e., the genetic information from a number of species. The genetic information can be derived from the genome and / or the exome and / or the transcriptome of a species. Exemplary nucleic acid databases useful in the present invention include, but are not limited to, International Nucleotide Sequence Database (INSD), DNA Data Bank of Japan (National Institute of Genetics), EMBL (European Bioinformatics Institute), GenBank (National Center for Biotechnology Information), Bioinformatic Harvester, Gene Disease Database, SNPedia, CAMERA Resource for microbial genomics and metagenomics, EcoCyc (a database that describes the genome and the biochemical machinery of the model organism E. coli K-12), Ensembl (provides automatic annotation databases for human, mouse, other vertebrate and eukaryote genomes) Ensembl Genomes (provides genome-scale data for bacteria, protists, fungi, plants and invertebrate metazoa, through a unified set of interactive and programmatic interfaces (using the Ensembl software platform)), Exome Aggregation Consortium (ExAC) (exome sequencing data from a wide variety of large-scale sequencing projects (Broad Institute)), PATRIC (PathoSystems Resource Integration Center), MGI Mouse Genome (Jackson Laboratory), JGI Genomes of the DOE-Joint Genome Institute (provides databases of many eukaryote and microbial genomes), National Microbial Pathogen Data Resource (a manually curated database of annotated genome data for the pathogens Campylobacter, Chlamydia, Chlamydophila, Haemophilus, Listeria, Mycoplasma, Neisseria, Staphylococcus, Streptococcus, Treponema, Ureaplasma and Vibrio\ RegulonDB (a model of the complex regulation of transcription initiation or regulatory network of the cell E. coli K-12), Saccharomyces Genome Database (genome of the yeast model organism), Viral Bioinformatics Resource Center (curated database containing annotated genome data for eleven virus families), The SEED platform (includes all complete microbial genomes, and most partial genomes, the platform is used to annotate microbial genomes using subsystems), WormBase ParaSite (parasitic species), UCSC Malaria Genome Browser (genome of malaria causing species (Plasmodium falciparum and others)), Rat Genome Database (genomic and phenotype data for Rattus norvegicus)', INTEGRALL (database dedicated to integrons, bacterial genetic elements involved in the antibiotic resistance), VectorBase (NIAED Bioinformatics Resource Center for Invertebrate Vectors of Human Pathogens), EzGenome, comprehensive information about manually curated genome projects of prokaryotes (archaea and bacteria), GeneDB (Apicomplexan Protozoa, Kinetoplastid Protozoa, Parasitic Helminths, Parasite Vectors as well as several bacteria and viruses), EuPathDB (eukaryotic pathogen database resources includes amoeba, fungi, plasmodium, trypanosomatids ete.); The 1000 Genomes Project (providing the genomes of more than a thousand anonymous participants from a number of different ethnic groups), Personal Genome Project (providing human genomes). Other databases can include personalized databases, such as databases comprising the genetic information of healthy and diseased tissues of the same subject.
[0093] In context of the present invention, the terms “sequence read” or “read” are used interchangeably and refer to a specific nucleic acid of any size for which the nucleotide sequence has been determined by sequencing, and which is preferably assigned to a species, preferably mapped to the genome of the respective species. In a preferred embodiment, the reads are classified to a specific species, such as the subject and / or microorganisms, preferably classified to specific microorganisms. In an embodiment, reads can be normalized by their abundance.
[0094] The step of determining the sequence of the contacted DNA molecules, in addition to the above- sequencing techniques, can also be carried out by detecting the presence of a specific nucleotide sequence using enzymatic amplification processes, such as PCR or any of its variations, or hybridization-based procedures. In both cases, oligonucleotides complementary to a sequence of interest, or a fragment of it, are utilized to assess the presence or absence of the specific nucleotide sequence of interest. In an embodiment of the present invention, an oligonucleotide that hybridizes / anneals to a DNA sequence suspected of being present among the contacted DNA molecules, e.g., a sequence derived from a particular pathogenic organism, can be used to detect the presence of such suspected sequence using standard PCR or hybridization methods.
[0095] The present invention in a further embodiment relates to a method for diagnosis of a disease state or a disease, e.g, infectious disease, in a subject, wherein a method for determining a disease state or disease in said subject according to the present invention is carried out.
[0096] In an embodiment, the invention provides a method for the identification of a subject suffering from a disease, preferably to a screening for a disease, preferably to a preventive medical analysis. In a preferred embodiment such methods identify correlation of the occurrence of a microorganism and the development of a disease in a subject. The present invention preferably relates to a method, wherein the pathogenic condition is characterized by abnormal, especially pathogenic quantities of nucleic acids of at least one microorganism, e.g., at least one viral, bacterial, fungal or parasitic organism.
[0097] Any microorganism, preferably one whose nucleic acid sequence is known, can be determined to be present in a subject, as well as be determined as the causative agent of a disease in the subject. Exemplary microorganisms, the presence of which that can be determined in a subject, include viruses, bacteria, fungi and parasites. Exemplary bacteria include, but are not limited to, Neisseria meningitis Streptococcus pneumoniae, Streptococcus pyogenes, Moraxella catarrhalis, Bordetella pertussis, Staphylococcus aureus, Clostridium tetani, Corynebacterium diphtheria, Haemophilus influenza, Pseudomonas aeruginosa, Streptococcus agalactiae, Chlamydia trachomatis, Chlamydia pneumoniae, Helicobacter pylori, Escherichia coli, Bacillus anthracis, Yersinia pestis, Staphylococcus epidermis, Clostridium perfringens, Clostridium botulinum, Legionella pneumophila, Coxiella burnetii, Brucella spp. such as B. abortus, B. canis, B. melitensis, B. neotomae, B. ovis, B. suis, B. pinnipediae, Francisella spp. such as F. novicida, F. philomiragia, F. tularensis, Neisseria gonorrhoeae, Treponema pallidum, Haemophilus ducreyi, Enterococcus faecalis, Enterococcus faecium, Staphylococcus saprophyticus, Yersinia enterocolitica, Mycobacterium tuberculosis, Rickettsia spp., Listeria monocytogenes, Vibrio cholera, Salmonella typhi, Borrelia burgdorferi, Porphyromonas gingivalis, Klebsiella spp., Klebsiella pneumoniae.
[0098] Exemplary viruses include, but are not limited to, Orthomyxoviridae, such as influenza A, B or C virus; Paramyxoviridae viruses, such as Pneumoviruses (e.g. , respiratory syncytial virus, RSV), Rubulaviruses (e.g., mumps virus), Paramyxoviruses (e.g., parainfluenza virus), Metapneumoviruses and Morbilliviruses (e.g., measles); Poxviridae, such as Orthopoxvirus (e.g, Variola vera, including Variola major and Variola minor); Picomaviridae, such as Enteroviruses (e.g., poliovirus e.g. a type 1, type 2 and / or type 3 poliovirus, EV71 enterovirus, coxsackie A or B virus), Rhinoviruses, Hepamavirus, Cardioviruses and Aphthoviruses; Bunyaviruses, such as Orthobunyavirus (e.g., California encephalitis virus), Phlebovirus (e.g, Rift Valley Fever virus), or Neurovirus (e.g., Crimean-Congo hemorrhagic fever virus); Hepamaviruses (e.g., hepatitis A virus (HAV), B and C); Filoviridae (e.g., Ebola virus (including a Zaire, Ivory Coast, Reston or Sudan ebolavirus) or Marburg virus); Togaviruses (e.g., Rubivirus, Alphavirus, and Arterivirus, including rubella virus); Flaviviruses (e.g., Tick-borne encephalitis (TBE) virus, Dengue (types 1, 2, 3 or 4) virus, Yellow Fever virus, Japanese encephalitis virus, Kyasanur Forest Virus, West Nile encephalitis virus, St. Louis encephalitis virus, Russian spring- summer encephalitis virus, and Powassan encephalitis virus); Pestiviruses (e.g., Bovine viral diarrhea (BVDV), Classical swine fever (CSFV) and Border disease (BDV)); Hepadnavirus (e.g., Hepatitis B virus, hepatitis C virus, delta hepatitis virus, hepatitis E virus, or hepatitis G virus); Rhabdoviruses (e.g., Lyssavirus, Rabies virus and Vesiculovirus (VSV)); Caliciviridae (e.g., Norwalk virus (Norovirus), and Norwalk-like Viruses, such as Hawaii Virus and Snow Mountain Virus); Coronavirus (e.g., SARS coronavirus, avian infectious bronchitis (IBV), Mouse hepatitis virus (MHV), and Porcine transmissible gastroenteritis virus (TGEV)); Retroviruses (e.g., Oncovirus, Lentivirus (e.g. HIV-1 or HIV-2) or a Spuma virus); Reoviruses (e.g., Orthoreovirus, Rotavirus, Orbivirus, and Coltivirus); Parvoviruses (e.g., Parvovirus Bl 9); Herpesviruses (e.g., human herpesvirus, such as Herpes Simplex Viruses (HSV), e.g., HSV types 1 and 2, Varicella-zoster virus (VZV), Epstein-Barr virus (EBV), Cytomegalovirus (CMV), Human Herpesvirus 6 (HHV6), Human Herpesvirus 7 (HHV7), and Human Herpesvirus 8 (HHV8)); Papovaviridae (e.g., Papillomaviruses and Polyomaviruses, e.g., serotypes 1, 2, 4, 5, 6, 8, 11, 13, 16, 18, 31, 33, 35, 39, 41, 42, 47, 51, 57, 58, 63 or 65, preferably from one or more of serotypes 6, 11, 16 and / or 18); Adenoviruses, such as adenovirus serotype 36 (Ad-36).
[0099] Exemplary fungi include, but are not limited to, Dermatophytres, including Epidermophyton floccusum, Microsporum audouini, Microsporum canis, Microsporum distortum, Microsporum equinum, Microsporum gypsum, Microsporum nanum, Trichophyton concentricum, Trichophyton equinum, Trichophyton gallinae, Trichophyton gypseum, Trichophyton naegnini, Trichophyton mentagrophytes, Trichophyton quinckeanum, Trichophyton rubrum, Trichophyton schoenleini, Trichophyton tonsurans, Trichophyton verrucosum, T. verrucosumvar. album, var. discoides, var. ochraceum, Trichophyton violaceum, and / or Trichophyton favifonne; Aspergillus fumigatus, Aspergillus flavus, Aspergillus niger, Aspergillus nidulans, Aspergillus terreus, Aspergillus sydowi, Aspergillus flavatus, Aspergillus glaucus, Blastoschizomyces capitatus, Candida albicans, Candida enolase, Candida tropicalis, Candida glabrata, Candida krusei, Candida parapsilosis, Candida stellatoidea, Candida kusei, Candida parakwsei, Candida lusitaniae, Candida pseudotropicalis, Candida guilliermondi, Cladosporium carrionii, Coccidioides immitis, Blastomyces dermatidis, Cryptococcus neofonnans, Geotrichum clavatum, Histoplasma capsulatum, Microsporidia, Encephalitozoon spp., Septata intestinalis and Enterocytozoon bieneusi; Brachiola spp., Microsporidium spp., Nosema spp., Pleistophora spp., Trachipleistophora spp., Vittaforma spp., Paracoccidioides brasiliensis, Pneumocystis carinii, Pythiumn insidiosum, Pityrosporum ovale, Sacharomyces cerevisae, Saccharomyces boulardii, Saccharomyces potnbe, Scedosporium apiosperum, Sporothrix schenckii, Trichosporon beigelii, Toxoplasma gondii, Penicillium mameffei, Malassezia spp., Fonsecaea spp., Wangiella spp., Sporothrix spp., Basidiobolus spp., Conidiobolus spp., Rhizopus spp., Mucor spp., Absidia spp., Mortierella spp., Cunninghamella spp., Saksenaea spp., Altemaria spp., Curvularia spp., Helminthosporium spp., Fusarium spp., Aspergillus spp., Penicillium spp., Monolinia spp., Rhizoctonia spp., Paecilomyces spp., Pithomyces spp., and Cladosporium spp.
[0100] Exemplary parasites include, but are not limited to, Plasmodium, such as P. falciparum, P. vivax, P. malariae and P. ovale, as well as those parasites from the Caligidae family, particularly those from the Lepeophtheirus and Caligusgenera, e.g., sea lice such as Lepeophtheirus salmonis and Caligus rogercresseyi.
[0101] Accordingly, the present invention provides a complete diagnostic workflow for the determination of the presence of microorganisms in a sample based on sequence analysis of nucleic acids, for example, free circulating DNA, after contacting the DNA with TdT since it allows to distinguish sequences derived from nucleic acids originally present in the sample at the time of collection from nucleic acids later contaminating the sample coming from reagents and laboratory equipment without the need for contamination and / or batch controls. The method advantageously provides a data-driven diagnosis without knowing the suspected microorganism, does not require specific primer design, and provides the opportunity to detect multiple viral, bacterial, fungal and parasitic microorganism in a single assay without the need for extensive controls to exclude contaminating nucleic acids.
[0102] The method of the present invention is preferably not restricted to the determination of a specific microorganism. In one embodiment, the present method determines the presence of all microorganisms, preferably all microorganisms relevant for a disease state in the subject, such as an infection. Thus, the present invention provides a usefill method for identification of the cause of an infection in a subject within short time, such that an appropriate therapy for the identified infection state can be selected within a short time.
[0103] Citation of documents and studies referenced herein is not intended as an admission that any of the foregoing is pertinent prior art. All statements as to the contents of these documents are based on the information available to the applicants and do not constitute any admission as to the correctness of the contents of these documents.
[0104] The description (including the following examples) is presented to enable a person of ordinary skill in the art to make and use the various embodiments. Descriptions of specific devices, techniques, and applications are provided only as examples. Various modifications to the examples described herein will be readily apparent to those of ordinary skill in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the various embodiments. Thus, the various embodiments are not intended to be limited to the examples described herein and shown, but are to be accorded the scope consistent with the claims.
[0105] DESCRIPTION OF THE FIGURES
[0106] Please provide brief descriptions of the figures
[0107] Figures 1 A-1C. Figure 1 A is a read-out demonstrating that DNA isolated from plasma treated with TdT exhibited a progressive increase of the DNA fragments sizes that positively correlated with enzyme incubation times with the TdT. Figure IB is a read-out of sequence data corroborating that the presence and the length of the homopolymeric tails positively correlated with TdT incubation times. Figure 1C is a chart setting out that the frequency of sequencing reads bearing the homopolymeric tails allows for DNA species introduced prior to or after the incubation with TdT to be distinguished from each other.
[0108] Figures 2A-2C. Figure 2A is a read-out demonstrating that plasma samples treated with TdT exhibited a progressive increase of the DNA fragments sizes that positively correlated with enzyme incubation times with the TdT. Figure 2B is a read-out of sequence data corroborating that the length of the homopolymeric tails positively correlated with TdT incubation times. Figure 2C is a chart setting out that DNA species introduced into the plasma prior to the TdT treatment can be positively identified by the frequency of sequencing reads bearing the homopolymeric tails, thus confirming that the TdT treatment in plasma was successful.
[0109] Figure 3 is a chart setting out that performing on-molecule DNA label synthesis within human plasma samples allows for the correct identification of microbial species known to be present within the samples with a 96.1% sensitivity across different sample biomass values. Figure 4 is a chart setting out that performing on-molecule DNA label synthesis within human plasma samples allows for the correct identification of contaminant microbial species known to be absent in the samples with a 99.3% specificity across different sample biomass values.
[0110] Figure 5 is a chart setting out that the on-molecule DNA label synthesis based approach also allows for high sensitivity and specificity values at reduced library preparation inputs in human plasma.
[0111] Figure 6 is a chart setting out that performing the on-molecule DNA label synthesis approach within a blood cfDNA mNGS workflow can identify disease-causing pathogens in plasma samples from patients with a suspected infection.
[0112] EXAMPLES
[0113] The present examples demonstrate that the use of a terminal deoxynucleotidyl transferase (TdT) allows for the identification and quantification of non-native DNA molecules present in a biological sample in a more highly efficient and cost-effective manner without the need for any kind of batch control correction or measures for laboratory reagent control contamination. Such results indicate that the claimed methods can be readily applied to any sample, whether obtained from a subject or not, which is suspected of containing non-subject-derived or potentially disease-causing pathogen-derived nucleic acids.
[0114] Results
[0115] On-molecule DNA label synthesis to distinguish nucleic acid molecules natively present in a biological specimen
[0116] Terminal deoxynucleotidyl transferase (TdT) is an enzyme specialized in the polymerization of nucleotides to the 3* end of DNA molecules without requiring a template. Given a mixture of nucleic acids, this enzyme would drive the on-molecule stepwise synthesis of a nucleotide sequence at their 3 ’end and, this newly synthesized sequence can be used as label to allow for the distinction of the initial nucleic acid molecule population from any other molecule added after the labelling step.
[0117] On-molecule DNA label synthesis was initially assessed on previously isolated human plasma cell free DNA samples analyzing DNA size distributions by capillary electrophoresis. When compared to baseline control, samples treated with TdT exhibited a progressive increase of the DNA fragments sizes that positively correlated with enzyme incubation times (Figure 1 A). From an initial average size of 114 base pair (bp) at baseline, the peak size shifted to 136 bp after 3 minutes of incubation with the enzyme. After library preparation, sequencing and bioinformatic analyses, the presence of an homopolymeric tail of adenosines (read as thymidines in the sequencing dataset due to strand complementarity) at the 3 ’end of the DNA fragments was confirmed. Sequence data corroborated that the length of the homopolymeric tails positively correlated with TdT incubation times with peak lengths shifting from 10 nucleotides, upon 30 seconds incubation, to 32 nucleotides at the 3 minutes reaction time (Figure IB). Therefore, TdT was able to perform on-molecule label polymerization by the stepwise addition of nucleotides to the 3 ’end of the isolated cell-free DNA (cfDNA) molecules and reaction times were identified as a factor determining the length of the label synthetized.
[0118] We next investigated whether on-molecule DNA label synthesis could be used to identify nucleic acid molecules natively present in a sample from fragments added throughout sample processing steps. To mimic this scenario, two synthetic DNA spikes were used, one fragment corresponding to a region of the Legionella sainthelensi genome and the other one to Candida orthopsilosis. The Candida orthopsilosis spike was introduced in previously isolated human cfDNA prior to TdT treatment, thus resembling a sample native DNA species, while the fragment corresponding to Legionella sainthelensi was spiked after TdT inactivation to mimic an environmental or laboratory contaminant. These samples yielded an average number of 37,061,327 reads per sample and 78.5% (SD=0.17) showed the newly synthetized label (Figure 1C). For the synthetic spikes, a 75.1% of the sequencing reads (SD=1.72) assigned to the Candida orthopsilosis fragment exhibited the homopolymeric tail while for Legionella sainthelensi sequences, the spike added post-polymerization, only 5.1% (SD=0.14) showed a sequence on the 3 ’end compatible with the label. This demonstrates that the approach allows distinguishing nucleic acid fragments natively present in a sample from those added at steps following the label polymerization event by assessing the frequency of sequencing reads carrying the label for a given DNA species.
[0119] Considering that fragments exogenous to a sample would only be recognized when introduced after the label synthesis, performing this step as early as possible is preferred. Since the TdT enzyme activity has a physiological role within primary lymphoid organs driving the junctional diversity required for adequate T and B cell receptor generation in the lymphoid lineage, we hypothesized that this enzyme could be capable of driving template-free DNA polymerization within human biological matrices, i.e., biological samples. To assess this, human plasma was utilized, since it constitutes a complex biological matrix, and plasma samples were spiked with three different synthetic DNA fragments (Candida orthopsilosis, Legionella sainthelensi and human herpesvirus-3) mimicking a scenario with three known sequence species natively present in plasma. Plasma samples were denatured and subjected to TdT treatment, followed by cfDNA isolation and library preparation prior to sequencing. At baseline, isolated size distribution peaked at 137 bp and this shifted to 147 bp upon 5 minutes of in-plasma TdT incubation (Figure 2A), This size shift on the capillary electrophoresis profiles also correlated with the length of the newly synthesized labels observed within the sequencing dataset (Figure 2B). Upon in-plasma TdT treatment, 49.37% of the sequencing reads (total of 46,762,203 sample reads) showed the homopolymeric tail upon within-plasma TdT treatment as opposed to the 1 .46% (total of 50,745,668 sample reads) observed at baseline. Previous experiments on isolated cfDNA confirmed that this approach allows distinguishing DNA molecules added at post-polymerization steps and here we determined whether the labelling efficiency achieved in plasma is sufficient to identify as true positives DNA species that were natively present in the sample prior to the labelling reaction. For this, true positive and false positive separation is based on the percentage of labelled reads obtained for a given DNA species. Upon TdT treatment, the three synthetic species presented percentages of labelled reads ranging from 70.8% to 75.3% while 0% was obtained for the unlabeled matched control (Figure 2C).
[0120] In summary, these results show that TdT is able to perform on-molecule DNA label synthesis within a biological sample, here plasma matrix, and that the frequency of labeled reads for a given DNA species can be utilized to distinguish nucleic acids, natively present in a sample prior to the labelling step, from molecules introduced at later steps. Thus, allowing for the ability to exclude nucleic acids that can contaminate the sample after TdT treatment.
[0121] Evaluation of on-molecule DNA label synthesis for contaminant filtering in low biomass samples within a cfDNA mNGS workflow
[0122] The on-molecule DNA label synthesis approach was implemented within a blood cfDNA mNGS workflow in order to determine its feasibility and performance for contaminant correction in a pathogen identification scenario. Since sample biomass is recognized as a critical factor for the increased incidence of laboratory and reagent contaminants, we tested serial dilutions of three different plasma samples simulating plasma inputs ranging from 220 ul to 6.88 ul. The results from previous cfDNA mNGS analyses of those samples were used as a baseline in order to define true positive species present within the samples. Focusing on the identification of true positive species, the on-molecule DNA label synthesis approach correctly identified as natively present in the sample 96.2% of the species (125 / 130) thus showing a robust sensitivity (Figure 3). Three out of the 5 missed species exhibited low read counts (1-4 reads for those specific species) while the other two presented higher numbers (17 and 23 reads). However, the broad range of read counts achieved for the species correctly identified (reads counts ranging from 1 to 491) shows that the missed hits are not likely due to species read count, thus indicating the suitability of the approach for the correction of metagenomics datasets with both high and low abundance species. Similarly, we evaluated the specificity of the approach in terms of its ability to identify as contaminants species not present within the sample (likely laboratory and reagent contaminants). The approach of utilizing plasma serial dilutions to resemble the increased incidence of laboratory and reagent contaminants was determined to be successful when considering the number and diversity of contaminant species detected in high versus low plasma inputs (Figure 4). The strategy correctly identified as contaminants the 99.3% of the contaminant species (141 / 142), again independently of the read count from the species (range 5-291) and just missing a single event present with 32 reads (Figure 4). In addition to the assessment of serially diluted plasma volumes, we also investigated whether reduced library preparation inputs could impose hurdles to the approach performance. For this purpose, we utilized three different inputs for library preparation (from 2 to 0.5 nanograms) from a previously in- plasma TdT labelled and isolated cfDNA sample for which a ground truth was established. Potentially due to the limited testing dataset, almost no differences could be observed in terms of contaminant incidence between the three inputs tested (Figure 5). Under these conditions, the approach exhibited an excellent performance in terms of sensitivity and specificity being able to correctly assign 100% of the hits identified (18 / 18 and 13 / 13 for sensitivity and specificity, respectively).
[0123] Therefore, the collected evidence shows the high performance of on-molecule DNA label synthesis, performed within the human plasma matrix, for the appropriate differentiation between DNA species natively present in the sample and those introduced during sample processing in the context of cfDNA mNGS for the identification of potentially disease causing microorganisms.
[0124] Implementation of on-molecule DNA label synthesis for contaminant filtering within a cfDNA mNGS workflow identifying disease-causing pathogens in human plasma samples from patients with suspected infection
[0125] Within metagenomics workflows, the usage of batch controls for contaminant correction constitutes one of the most extended strategies. However, besides the risk of over- or under-correction of contaminants, batch controls can also impose hurdles to cost efficiency, particularly for laboratories handling batches with low sample numbers. On the other hand, recently proposed approaches focusing on sample intrinsic DNA molecule sequence modification, require additional time for sample processing (>1.5 hours) thus posing additional hurdles for time-sensitive applications such as bloodstream infections. The herein disclosed strategy based on on-molecule DNA synthesis, allows for the substitution of control-based approaches by providing contaminant correction within each individual sample and with a performance independent form the sample biomass similarity to a batch control. Furthermore, the labelling process requires less than 50 minutes thus adding reduced times upon implementation within mNGS workflows.
[0126] We implemented the on-molecule DNA label synthesis approach within a blood cfDNA mNGS workflow identifying disease-causing pathogens in human plasma samples from patients with suspected infection. A total of 97 previously analyzed human plasma samples were processed within 36 sample pools using the on-molecule DNA label synthesis approach for contaminant filtering. Without any kind of batch control correction or measures for laboratory reagent contamination control, the species identified by the mNGS software, were then evaluated based on their individual labelled read frequencies. A threshold of 0.125 was used to classify positive and potential contaminant species. The baseline was established considering the results previously obtained for the patient samples using batch control and relevance assessment of the samples, the results of repeated analysis and clinical feedback available. Within the 97 samples, a total of 158 true positive species and 157 true negative species were identified and correctly classified by the labelling approach, while 21 false positives and 2 false negatives were wrongly categorized (Figure 6). Therefore, this approach allows for cost and / or time saving within mNGS workflows while ensuring a robust performance for contaminant species filtering with a sensitivity of 98.8% and a specificity of 88.2%.
[0127] MATERIALS AND METHODS
[0128] On-molecule DNA label synthesis in isolated cell free DNA
[0129] Human cell free DNA (cfDNA) was previously isolated from 1 mL of human plasma samples with the QIAsymphony Circulating DNA DSP kit (Qiagen, product #937556) on the QIAsymphony liquid handling platform (Qiagen) according to manufacturer instructions. Isolated material was stored at - 80°C until utilization. Where applicable, synthetic double-stranded DNA fragments with known sequences (IDT, custom gBlocks) were introduced prior and / or after the labelling step.
[0130] Thawed cell free DNA was quantified using the Qubit™ dsDNA HS Assay Kit (Invitrogen, product #Q32851). Isolated human blood cfDNA was denatured at 98°C for 4 minutes followed by cooling. Template free polymerization was carried out on 40 pL of denatured cfDNA utilizing Terminal Transferase (New England Biolabs, product #M0315) and a dATP solution (New England Biolabs, product #N0440). Reaction was prepared by adding 6 pL of lOx reaction buffer, 8 pL of 2.5 mM CoCb, 20 units of terminal deoxynucleotidyl transferase (TdT) enzyme, 0.5 pL of 100 mM dATP. The reaction was filled up with nuclease free water to a final volume of 20 pL and incubated at 37°C for 5 minutes or the indicated incubation times. The labelling reaction was stopped by addition of 5 pL of 0.5 M EDTA pH 8.0 (VWR, product #E177). The resulting labelled cell free DNA was purified using AMPure XP beads at a ratio of 1.6x following the manufacturer instructions. The purified labelled cfDNA is then quantified using the Qubit™ ssDNA Assay-Kit (Invitrogen, product #Q10212) and fragment size distribution is assessed using the Agilent HS NGS Fragment Kit (1-6000 bp) (Agilent, product #DNF- 473-0500) on the Fragment Analyzer 5200 platform (Agilent). Sample sequencing libraries were prepared -using the xGen™ ssDNA & Low-Input DNA Library Preparation Kit (IDT, product #10009817) and xGen™ UDI Primers (IDT, product #10005922) following the manufacturer’s instructions, but replacing recommended SPRISelect beads with AMPure XP beads (Beckman Coulter, product #A63882) and adjusting bead-to-sample ratios to 1.5x, 0.9x and 0.7x, respectively. The resulting sample libraries were then pooled equimolarly and sequenced on an Illumina sequencer using paired- end sequencing mode.
[0131] On-molecule DNA label synthesis within human plasma samples
[0132] Peripheral blood samples were collected in Cell-Free DNA BCT tubes (Streck, product #230471).
[0133] Plasma separation was performed, by centrifugation at 1600 x g for 10 minutes at 4°C. After transferring the plasma fraction to a fresh tube, it is again centrifuged at 1600 x g for 10 minutes at 4°C. The supernatant was stored at -80°C until utilization. Human plasma samples were diluted 1 :2 in lx PBS buffer pH7.4 (Invitrogen, product SAM9625) for subsequent denaturation at 98°C for 4 minutes or the indicated incubation followed by cooling. Denatured samples were centrifuged at 1600 x g for 10 minutes to pellet debris and the supernatant was transferred into a fresh tube. Where applicable, synthetic double-stranded DNA fragments with known sequences (IDT, custom gBlocks) were introduced in the plasma prior to the labelling step. Template free polymerization was carried out on 400 pL of denatured diluted plasma supernatant utilizing Terminal Transferase (New England Biolabs, product #M0315) and a dATP solution (New England Biolabs, product #N0440). Reaction was prepared by adding 60 pL of lOx reaction buffer, 80 pL of 2.5 mM CoC12, 200 units of terminal deoxynucleotidyl transferase (TdT) enzyme, 5 pL of 100 mM dATP. The reaction was filled up with nuclease free water to a final volume of 600 pL and incubated at 37°C for 20 minutes. The labelling reaction was stopped by addition of 50 pL of 0.5 M EDTA pH 8.0 (VWR, product #E177). After cooling, samples were diluted in lx PBS buffer pH 7.4 (Invitrogen, product #AM9625) to a final volume of 1 .1 mL for subsequent cell free DNA isolation with the QIAsymphony Circulating DNA DSP kit (Qiagen, product #937556) on the QIAsymphony liquid handling platform (Qiagen) according to manufacturer instructions. The resulting isolated cell free DNA was denatured for 2 min at 98 °C and quantified using the Qubit™ ssDNA Assay- Kit (Invitrogen, product #Q10212). The fragment size distribution is assessed using the Agilent HS NGS Fragment Kit (1-6000 bp) (Agilent, product #DNF-473-0500) on the Fragment Analyzer 5200 platform (Agilent). Sample sequencing libraries were prepared using the xGen™ ssDNA & Low-Input DNA Library Preparation Kit (IDT, product #10009817) and xGen™ UDI Primers (IDT, product #10005922) following the manufacturer instructions, but replacing recommended SPRISelect beads with AMPure XP beads (Beckman Coulter, product #A63882) and adjusting bead-to-sample ratios to 1.5x, 0.9x and 0.7x, respectively. The resulting sample libraries were then pooled equimolarly and sequenced on an Illumina sequencer using paired-end sequencing mode.
[0134] Processing of sequencing files
[0135] Paired end sequencing files were obtained from the Illumina sequencer. Forward FASTQ file (Rl) was quality processed using BBDuk (BBDuk version 38.90), to trim adapter, low entropy and low quality sequence regions. A remaining minimal read of length of 50 base pairs was required after trimming for a sequence to proceed. Further processing of the forward sequences which passes the quality filter, included the removal of human sequences with kraken2 (kraken2 version 2.1.1) against the Human genome. For this, a classification database was built using the data deposited in the Genome Reference consortium in NCBI (RefSeq GCF_000001405.40).
[0136] Identification of the labelled sequences and microbial identification
[0137] During sequencing library preparation a „CT" motif of variable length is introduced at the 3 ’-end of the single-stranded DNA fragments which preferably carry a poly-A tail As the identification of the label is being done in the second reverse FASTQ file (R2), this poly A is read to its complementary nucleotide T, as well as the „CT“ motif is read to G and A. Therefore the identification of the label was done considering a "GA" motif and a T stretch, of minimum 4 consecutive T's. Thus, a regular expression (r,A([GA]+T {4,})') was defined to divide the reverse file (R2), into label and not-labelled sequences. To achieve this, a script was implemented in python (version 3.9.0) using BioPython library (Bio, Bio.Seq 6 Bio.SeqRecord). Once these two files were generated, the tool “filterbyname” from the BBDuk tools suite was employed to obtain the labelled and not-labelled reverse (R2) and forward (Rl) FASTQ files. Microbial identification was carried out through a clinical metagenomics pipeline performing taxonomic classification and clinical relevance assessment. Labelling ratios were calculated on species level by dividing the labelled read count by the total count of reads being classified to a particular microbial species and a threshold of 0.125 was used for categorization.
Claims
We claim:
1. A method for processing nucleic acids in a sample, said method comprising:- (a) contacting DNA molecules obtained from a sample with a template-free DNA polymerase in the presence of one or more types of nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of the DNA molecules present in the sample, in which each nucleotide is added to the 3 ’ end in a step-wise reaction; and- (b) determining the sequence of the contacted DNA molecules.
2. The method according to claim 1 , wherein the sample is a biological sample obtained from a subject.
3. The method according to claim 1 or 2, wherein the step of contacting occurs in the obtained sample.
4. The method according to claim 1 or 2, further comprising a step of reverse transcribing RNA molecules present in the obtained sample.
5. The method according to any one of claims 1 to 4, further comprising a step of amplifying the DNA molecules in the sample prior to the contacting step.
6. The method according to any one of claims 1 to 5, further comprising a step of amplifying the contacted DNA molecules in the obtained sample prior to the sequencing step.
7. The method according to claim 6, wherein the DNA molecules to be amplified are contacted DNA molecules comprising the one or more added nucleotides.
8. The method according to any one of claims 1 to 7, further comprising a step of enriching and / or isolating the contacted DNA molecules comprising the one or more added nucleotides prior to the sequencing step.
9. The method according to any one of claims 1 to 8, further comprising removing the one or more added nucleotides from the contacted DNA molecules comprising the one or more added nucleotides prior to the sequencing step.
10. The method according to any one of claims 1 to 9, wherein the step of sequencing comprises contacting the DNA molecules comprising one or more added nucleotides with at least one oligonucleotide that is complementary to the added nucleotides.
11. The method according to claim 10, wherein the at least one oligonucleotide is conjugated to magnetic beads.
12. The method according to any one of claims 1 to 11, wherein at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides are enzymatically added.
13. The method according to claim 12, wherein at least 4 nucleotides are enzymatically added.
14. The method according to any one of claims 1 to 11, wherein no more than 10, 9, 8, 7, 6, 5, 4, 3, 2 nucleotides are enzymatically added.
15. The method according to any one of claims 1 to 14, wherein only one species of nucleotide is added in step (a).
16. The method according to any one of claims 1 to 15, wherein at least two different species of nucleotides are added in step (a).
17. The method according to any one of claims 1 to 16, wherein the at least one nucleotide is a deoxyribonucleotide or ribonucleotide.
18. The method according to claim 17, wherein the deoxyribonucleotide or ribonucleotide is a modified deoxyribonucleotide or ribonucleotide.
19. The method according to claim 18, wherein the deoxyribonucleotide is selected from the group consisting of adenine, guanine, cytosine and thymidine.
20. The method according to claim 19, wherein the sample is a biological sample that has been directly obtained from a subject.
21. The method according to claim 20, wherein the subject is a mammal.
22. The method according to claim 20 or 21, wherein the subject is a human.
23. The method according to any one of claims 1 to 22, wherein the biological sample is selected from the group consisting of blood, plasma, pleura, ascites, ocular fluid, urine, lymph, cerebrospinal fluid, and a cell scraping.
24. The method according to any one of claims 1 to 23, wherein the sample is not processed to specifically remove nucleic acids prior to step (a).
25. The method according to any one of claims 1 to 24, wherein the template-free DNA polymerase is a terminal deoxynucleotidyl transferase or functional variant thereof retaining terminal deoxynucleotidyl transferase activity.
26. The method according to any one of claims 1 to 25, wherein the sequence is determined by high throughput sequencing methods, for example, by nanopore sequencing.
27. The method according to any one of claims 1 to 25, wherein the sequence is determined by PCR analysis.
28. The method according to any one of claims 1 to 27, further comprising in silica analysis of the obtained sequences.
29. The method according to claim 28, wherein the analysis comprises determining the species of origin of the DNA molecules in the sample.
30. The method according to any one of claims 1 to 29, wherein the DNA molecules are cell-free DNA molecules present in the sample.
31. The method according to any one of claims 1 to 30, wherein the processing method is for detecting the presence of non-subject-derived nucleic acid molecules in a biological sample obtained from a subject.
32. The method according to claim 31 , wherein the non-subject-derived nucleic acid molecules are microbial-, fungal, parasitic-, or viral-derived nucleic acid molecules.
33. The method according to any one of claims 1 to 32, further comprising contacting one or more oligonucleotides to the contacted DNA molecules, wherein at least one oligonucleotide anneals to the nucleotides enzymatically added to the DNA molecules prior to the sequencing step.
34. A method for identifying the species of origin of DNA molecules in a biological sample, said method comprising:- (a) contacting DNA molecules obtained from a biological sample with a template-free DNA polymerase in the presence of one or more types of nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3* end of the DNA molecules present in the biological sample, in which each nucleotide is added to the 3’ end in a step-wise reaction;- (b) determining the sequence of the contacted DNA molecules; and- (c) determining the species of origin of the contacted DNA molecules.
35. The method according to claim 34, wherein the DNA molecules are cell-free DNA molecules in the sample.
36. The method according to claim 34 or 35, wherein determining the species of origin comprises comparing the obtained sequences to one or more nucleotide sequence databases and identifying the species of origin of one or more DNA molecules in the sample.
37. The method according to any one of claims 34 to 36, wherein the species of origin of the DNA molecules is determined to be bacterial, viral, parasitic, or fungal.
38. A method for diagnosing a medical condition in a subject caused by a pathogen, said method comprising:- (a) contacting DNA molecules present in a biological sample obtained from the subject with a template-free DNA polymerase in the presence of one or more nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of DNA molecules present in the biological sample, in which each nucleotide is added to the 3’ end in a step-wise reaction;- (b) determining the sequence of the contacted DNA molecules to obtain the sequences of the DNA molecules; and- (c) comparing the obtained sequences against one or more databases to determine the species identity of the DNA molecules in the sample, wherein when DNA molecules from a pathogenic species are present in the biological sample above a threshold value, the subject is diagnosed with a medical condition caused by a pathogen.
39. The method according to claim 38, wherein the pathogen is a bacterium, a parasite, a virus, or a fungus.
40. A method for detecting pathogen-derived nucleic acids in a sample, said method comprising:- (a) contacting DNA molecules present in a sample not obtained from a subject with a template- free DNA polymerase in the presence of one or more nucleotides and under conditions sufficient to enzymatically add at least one nucleotide to the 3’ end of DNA molecules present in the sample, in which each nucleotide is added to the 3’ end in a step-wise reaction;- (b) determining the sequence of the contacted DNA molecules to obtain the sequences of the DNA molecules; and- (c) comparing the obtained sequences against one or more databases to determine the species identity of the DNA molecules in the sample.
41. The method according to claim 40, wherein the pathogen is a bacterium, a parasite, a virus, or a fungus.
42. The method according to any one of claims 1 to 41, wherein the method is carried out without any kind of batch correction or measures for laboratory reagent contamination control.
Citation Information
Patent Citations
Polynucleotide amplification
WO2007062445A1
Direct oligonucleotide synthesis on cells and biomolecules
WO2020120442A2
Method of amplifying mrnas and for preparing full length mRNA libraries
WO2021130151A1
Method for performing multiple analyses on same nucleic acid sample
WO2021141852A1
Methods and kits for analyzing nucleosomes and plasma proteins
WO2023131939A1