How to repair a 3' overhang

The method preserves 3' overhang sequence information in DNA fragments by using primers with random target hybridization sequences and DNA polymerase lacking exonuclease activity, allowing for effective ligation and subsequent processing without data loss.

JP7850658B2Active Publication Date: 2026-04-23GUARDANT HEALTH INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GUARDANT HEALTH INC
Filing Date
2020-10-23
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for repairing 3' overhangs in DNA molecules, such as those found in sheared or cell-free DNA, result in information loss as they rely on exonucleases that remove nucleotides, making it impossible to determine the sequence between the 5' and 3' ends or the positioning of nucleosomes.

Method used

A method involving the use of primers with random target hybridization sequences, extended by DNA polymerase lacking exonuclease activity, to ligate the 3' end of extended primers to the 5' end of the DNA fragment, preserving the 3' overhang sequence.

Benefits of technology

Preserves the sequence information of the 3' overhang, enabling subsequent steps like ligation of tags, barcodes, or adapters without losing sequence data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850658000001
    Figure 0007850658000001
  • Figure 0007850658000002
    Figure 0007850658000002
  • Figure 0007850658000003
    Figure 0007850658000003
Patent Text Reader

Abstract

Provided is a method for repairing a partially double-stranded DNA fragment. In some embodiments, the method comprises the steps of: (a) contacting the partially double-stranded DNA fragment with one or more primers from a primer population, wherein the partially double-stranded DNA fragment comprises a 3' overhang, and the primer population comprises a random target hybridization sequence; (b) using a DNA polymerase to extend one or more primers from the primer population along the DNA fragment, thereby producing one or more extended primers annealed to the DNA fragment; and (c) ligating the 3' end of one or more extended primers to the 5' end of the extended primer or a strand of the partially double-stranded DNA fragment, thereby obtaining a repaired DNA fragment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Citation of Related Applications This application claims the benefit of priority based on U.S. Provisional Patent Application No. 62 / 926,093, filed Oct. 25, 2019, which is hereby incorporated by reference in its entirety for all purposes.

[0002] Sequence Listing This application includes a sequence listing, which is submitted electronically in ASCII format and is hereby incorporated by reference in its entirety. The ASCII copy was created on Oct. 21, 2020, has the file name 2020 - 10 - 23_GH0054WO_Sequence_Listing_ST25.txt, and is 970 bytes in size.

Background Art

[0003] Introduction and Summary Repair of overhangs that occur with certain DNA molecular types is an important step in the preparation of molecules for further analysis such as sequencing and / or amplification. For example, a 3’ overhang where several nucleotides near the 3’ end of the molecule are single-stranded can occur with sheared DNA and cell-free DNA (obtained from a blood sample). It may be desirable to convert the 3’ overhang to a blunt end or a one-base overhang to be compatible with subsequent steps such as ligation of tags, barcodes, or adapters.

[0004] Existing overhang repair methods use 3'→5' exonucleases to excise the 3' overhang. This approach is information-lossy in that nucleotides are removed, making it impossible to determine the position of the original molecule's ends. Therefore, repairing a 3' overhang by exonucleolisis does not provide information, for example, about the sequence of bases between the 5' and 3' ends, as well as the positioning of nucleosomes. Thus, an improved method for repairing 3' overhangs is needed. This disclosure aims to meet this need, provide other benefits, or at least publicly offer a useful alternative. [Overview of the project] [Means for solving the problem]

[0005] Therefore, the following embodiments are provided. Embodiment 1 is a method for repairing a partially double-stranded DNA fragment, (a) A step of contacting the partially double-stranded DNA fragment with one or more primers of a primer population, wherein the partially double-stranded DNA fragment includes a 3' overhang and the primer population includes a random target hybridization sequence; (b) a step of extending one or more primers of the primer population along the DNA fragment using a DNA polymerase, thereby producing one or more extended primers annealed to the DNA fragment; and (c) A step of ligating the 3' end of one or more elongated primers to the 5' end of the strand of the elongated primer or a partially double-stranded DNA fragment, thereby obtaining a repaired DNA fragment. This method includes [something].

[0006] Embodiment 2 is a method according to any one of the prior embodiments, wherein the DNA polymerase lacks 3'→5' exonuclease activity.

[0007] Embodiment 3 is a method according to any one of the prior embodiments, wherein the DNA polymerase lacks 5'→3' exonuclease activity.

[0008] Embodiment 4 is a method according to any one of the prior embodiments, wherein the DNA polymerase lacks strand displacement activity.

[0009] Embodiment 5 is a method according to any one of the prior embodiments, wherein the DNA polymerase is a Klenow fragment.

[0010] Embodiment 6 is the method described in the preceding embodiment, wherein the DNA polymerase is an exo-Klenow fragment.

[0011] Embodiment 7 is a method according to any one of the prior embodiments, wherein the partially double-stranded DNA fragment has a 3' overhang at each end.

[0012] Embodiment 8 is a method according to any one of the prior embodiments, wherein the partially double-stranded DNA fragment has a 3' overhang and (i) a blunt end or (ii) a 5' overhang.

[0013] Embodiment 9 is the method described in the preceding embodiment, wherein the 5' overhang is repaired by extending the 3' end along the 5' overhang.

[0014] Embodiment 10 is a method according to any one of the prior embodiments, wherein the partially double-stranded DNA fragment is derived from a bodily fluid sample.

[0015] Embodiment 11 is the method described in the immediately preceding embodiment, wherein the body fluid is whole blood, serum, plasma, or urine.

[0016] Embodiment 12 is a method according to any one of the prior embodiments, wherein the partially double-stranded DNA fragment is a cfDNA fragment.

[0017] Embodiment 13 is the method according to any one of the preceding embodiments, wherein the partially double-stranded DNA fragment is of mammalian origin.

[0018] Embodiment 14 is the method according to any one of the preceding embodiments, wherein the partially double-stranded DNA fragment is of human origin.

[0019] Embodiment 15 is the method according to any one of the preceding embodiments, wherein the partially double-stranded DNA fragment is part of a population of DNA fragments in a composition.

[0020] Embodiment 16 is the method according to the immediately preceding embodiment, wherein the population of DNA fragments comprises sheared DNA.

[0021] Embodiment 17 is the method according to Embodiment 15 or 16, wherein the population of DNA fragments comprises epigenetically modified DNA.

[0022] Embodiment 18 is the method according to any one of Embodiments 15 to 17, wherein the population of DNA fragments comprises fragments from multiple genomic loci.

[0023] Embodiment 19 is the method according to any one of Embodiments 15 to 18, wherein the population of DNA fragments is not enriched.

[0024] Embodiment 20 is the method according to any one of Embodiments 15 to 19, wherein the population of DNA fragments is not amplified.

[0025] Embodiment 21 is the method according to any one of the preceding embodiments, wherein the length of the random target hybrid-forming sequence is at least 4, 5, 6, 7, 8, 9, or 10 nucleotides.

[0026] Embodiment 22 is the method according to any one of the preceding embodiments, wherein the length of the random target hybrid-forming sequence is about 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides.

[0027] Embodiment 23 is the method described in any one of the preceding embodiments, and the primers in the primer population are single-stranded.

[0028] Embodiment 24 is the method described in any one of Embodiments 1 to 20, the primers in the primer population are double-stranded with 3'-overhangs, and the random target hybrid-forming sequences are present in the 3'-overhangs.

[0029] Embodiment 25 is the method described in any one of Embodiments 1 to 20, the primers in the primer population are hairpins with 3'-overhangs, and the random target hybrid-forming sequences are present in the 3'-overhangs.

[0030] Embodiment 26 is the method described in any one of Embodiments 21 to 22, and the hairpin or double-stranded region of the primer contains an adapter.

[0031] Embodiment 27 is the method described in any one of Embodiments 21 to 23, and the hairpin or double-stranded region of the primer contains a tag.

[0032] Embodiment 28 is the method described in Embodiment 24, and the tag contains a barcode.

[0033] Embodiment 29 is the method described in Embodiment 25, and the primer population contains a plurality of different barcodes.

[0034] Embodiment 30 is the method described in any one of the preceding embodiments, and at least steps (a) to (c) are performed in a single tube.

[0035] Embodiment 31 is the method described in any one of the preceding embodiments, the double-stranded DNA fragment is present in a composition, and for at least steps (a) to (c), the components are not removed from the composition.

[0036] Embodiment 32 is a method according to any one of the prior embodiments, wherein the repaired DNA fragment comprises one or two blunt ends, or the method further comprises the step of blunt-ending one or two ends of the repaired DNA fragment.

[0037] Embodiment 33 is the method of the preceding embodiment, further comprising the step of end-tailing the repaired DNA fragment using a polymerase that adds nucleotides to the 3' end of the blunt-ended nucleic acid in a non-template-directed manner, wherein A is preferred over G and G is preferred over C or T, if necessary.

[0038] Embodiment 34 is a method according to any one of the prior embodiments, further comprising the step of ligating the repaired DNA fragment with a tag (on both ends, if applicable), wherein the tag optionally includes a barcode and / or an adapter.

[0039] Embodiment 35 is a method according to any one of the prior embodiments, further comprising the step of purifying the repaired DNA fragment.

[0040] Embodiment 36 is a method according to any one of the prior embodiments, further comprising the step of denaturing one or more enzymes used in step (b) and / or (c) after step (b) and / or (c).

[0041] Embodiment 37 is a method according to any one of the prior embodiments, further comprising the step of amplifying the repaired DNA fragment.

[0042] Embodiment 38 is the method of the immediately preceding embodiment, wherein the repaired DNA fragment comprises one or more adapters (e.g., two adapters), and one or more (e.g., two) amplification oligomers are used to amplify the repaired DNA fragment by annealing to the one or more adapters.

[0043] Embodiment 39 is a method according to any one of the prior embodiments, further comprising the step of enriching a target fragment from the repaired DNA fragments, thereby obtaining an enriched DNA fragment, wherein the enrichment step is performed after the amplification step, if necessary.

[0044] Embodiment 40 is the method described in the preceding embodiment, wherein the fragment of interest comprises a locus that varies in a manner associated with a disease or disorder, and optionally the disease or disorder is cancer.

[0045] Embodiment 41 is the method described in the preceding embodiment, wherein the variation is one or more of a single-nucleotide variation, a copy number variation, a gene fusion, or an indel.

[0046] Embodiment 42 is the method described in the preceding embodiment, wherein the fragment of interest comprises one, two, three, or four fragments that exhibit a single-nucleotide variation associated with a disease or disorder; a copy number variation associated with a disease or disorder; a gene fusion associated with a disease or disorder; or an indel associated with a disease or disorder, wherein the disease or disorder is cancer, if applicable.

[0047] Embodiment 43 is a method according to any one of Embodiments 36 to 39, further comprising the step of amplifying the enriched DNA fragment.

[0048] Embodiment 44 is a method according to any one of the prior embodiments, further comprising the step of sequencing the repaired DNA fragment.

[0049] Embodiment 45 is the method described in the preceding embodiment, wherein the sequencing step involves sequencing nucleotides that have formed a 3' overhang in a partially double-stranded DNA fragment.

[0050] Embodiment 46 is the method described in Embodiment 41 or 42, wherein the sequence determination step is high-throughput sequence determination.

[0051] Embodiment 47 is a method according to any one of Embodiments 44 to 46, wherein a plurality of repaired DNA fragments are generated, at least a portion of the repaired DNA fragments include a tag, and the sequencing generates a plurality of sequence reads from the plurality of repaired DNA fragments.

[0052] Embodiment 48 is the method described in Embodiment 47, wherein the sequence read includes a sequence of nucleotides and a tag sequence that form a 3' overhang in the partially double-stranded DNA fragment.

[0053] In each aspect of the present invention and in some embodiments of any aspect thereof, the results of the systems and / or methods disclosed herein are used as material for preparing a report. The report may be in paper or electronic format. For example, such a report may include information regarding the presence or absence of cancer as determined by the methods or systems disclosed herein. Alternatively, the report may further include information relating to or derived from the identification of nucleic acid bases in the sample. The methods or systems disclosed herein may further include the step of communicating the report to a third party, such as the subject who provided the sample or a medical professional.

[0054] The various steps of the methods disclosed herein, or the steps performed by the systems disclosed herein, may be performed simultaneously or at different times, and / or in the same geographical location or in different geographical locations (e.g., countries). The various steps of the methods disclosed herein may be performed by the same person or by different persons. [Brief explanation of the drawing]

[0055] [Figure 1] Figure 1 shows 5' overhang, 3' overhang, and existing terminal repair methods.

[0056] [Figure 2]Figure 2 shows an embodiment of 3' overhang repair according to this disclosure, in which primers containing a 3' random sequence and a 5' double-stranded sequence (e.g., including a barcode, adapter, or tag) are annealed to a partially double-stranded DNA fragment. Extension (e.g., using a Klenowexo) and ligation yield a repaired molecule in which the 3' overhang sequence is preserved.

[0057] [Figure 3] Figure 3 shows an embodiment of 3' overhang repair according to this disclosure, in which a primer having a random sequence is annealed to a partially double-stranded DNA fragment. Extension (e.g., using Klenowexo) and ligation yield a repaired molecule in which the 3' overhang sequence is preserved. [Modes for carrying out the invention]

[0058] Detailed description of a particular embodiment The section headings used herein are for organizational purposes only and should not be construed as limiting the scope of the inventions described herein.

[0059] All references herein, including patent applications, patent publications, and Genbank reference numbers, are incorporated herein by reference in the same manner as if each entire reference were specifically and individually indicated.

[0060] definition Unless otherwise defined, scientific and technical terms used in connection with this disclosure shall have the meanings generally understood by those skilled in the art. Furthermore, unless otherwise explicitly stated and there is no need for context, singular terms shall include plural forms and plural terms shall include singular forms. In the event of any inconsistency between definitions from various sources or documents, the definitions provided herein shall prevail.

[0061] Before describing this instruction in detail, it should be understood that this disclosure is not limited to a specific composition or process steps, and may vary as such. It should be noted that, where used herein and in the appended claims, the singular forms “a,” “an,” and “the” include the plural form unless the context explicitly states otherwise. Therefore, for example, a reference to “oligomer” includes multiple oligomers, etc. In this application, unless explicitly stated or understood by those skilled in the art, the use of “or” means “and / or.” When “or” is used in the context of multiple dependent claims, it means referring back to one or more preceding independent or dependent claims.

[0062] The “approximately” preceding temperature, concentration, time, etc., as discussed in this disclosure should be understood to mean that slight deviations are within the scope of the teachings herein. In general, the term “approximately” refers to slight variations in the amounts of components of a composition that do not have any significant effect on the activity or stability of the composition. Furthermore, the use of “comprise,” “comprises,” “comprising,” “contain,” “contains,” “containing,” “include,” “includes,” and “including” is not intended to be limiting. The general and detailed descriptions herein are for illustrative and explanatory purposes only and should be understood not to limit the teachings herein. In the event that any material invoked by reference contradicts the content of the representations herein, the representations shall prevail.

[0063] Unless otherwise specified, embodiments in this specification that "include" various components are also intended to "consist of" or "essentially consist of" the enumerated components; embodiments in this specification that "consist of" various components are also intended to "include" or "essentially consist of" the enumerated components; embodiments in this specification that "essentially consist of" various components are also intended to "consist of" or "include" the enumerated components (this interchangeability does not apply to the use of these terms in the claims).

[0064] The subject refers to an animal such as a mammalian species (preferably human) or a bird (e.g., bird), or another organism (such as a plant). More specifically, the subject may be a vertebrate, such as a mammal like a mouse, primate, monkey, or human. Animals include livestock, sports animals, and pets. The subject may be a healthy individual, an individual with or suspected to have a disease or a predisposition to a disease, or an individual that requires or is suspected to require treatment.

[0065] A genetic variant refers to a change, variant, or polymorphism in a subject's nucleic acid sample or genome. Such a change, variant, or polymorphism may be relative to a reference genome, which may be the subject's or another individual's reference genome. Variations include one or more single-nucleotide variations (SNVs), insertions, deletions, repeats, small insertions, small deletions, small repeats, structural variant junctions, variable-length tandem repeats, and / or adjacent sequences, and copy number variants (CNVs), transversions, and other rearrangements are also forms of genetic variation. Variations can be changes in base, insertions, deletions, repeats, copy number variations, transpositions, or combinations thereof.

[0066] Cancer markers are genetic variants associated with the presence or risk of developing cancer. Cancer markers can indicate that a subject has cancer or has a higher risk of developing cancer than similar subjects of the same age and sex. Cancer markers may or may not be the cause of cancer.

[0067] Nucleic acid tags are typically artificial sequences, short nucleic acids (e.g., less than 100, 50, or 10 nucleotides long), usually DNA, and are used to label repaired DNA fragments to identify nucleic acids that (i) originate from different samples (e.g., indicate a sample), (ii) are of a different type, or (iii) have undergone different processing. Tags can be single-stranded or double-stranded. Nucleic acid tags can be decoded to reveal information such as the originating sample, the morphology or processing of the nucleic acid. Tags can be used to pool multiple nucleic acids with different tags and process them in parallel, after which the nucleic acids can be analyzed by reading the tags. Tags can also be called molecular identifiers or barcodes.

[0068] An adapter is a short nucleic acid (e.g., less than 500, 100, or 50 nucleotides long, and typically DNA) for ligating to one or both ends of a repaired DNA fragment molecule. Adapters may, but not necessarily, be supplied in a double-stranded form for ligation. Alternatively, adapters may be supplied as 5' elements in primers (e.g., which may be single-stranded, hairpin, or double-stranded) (e.g., members of a primer population having a randomized target hybridization sequence). Adapters may include primer binding sites for amplifying a repaired DNA fragment molecule with both ends adjacent to the adapter, and / or sequencing primer binding sites (including primer binding sites for next-generation sequencing techniques). Adapters may also include binding sites for capture probes (e.g., oligonucleotides attached to a flow cell support). Adapters may also include tags as described herein. Tags may be positioned relative to primer and sequencing primer sites so that they are included in the amplicon and sequencing reads of the repaired DNA fragment. Identical and different adapters can be ligated to each end of the same molecule. Occasionally, identical adapters are attached to each end, except for different tags. An example of an adapter type is a Y-type adapter in which one end is blunt-ended or tailed as described herein for ligation to nucleic acids (e.g., repaired DNA fragments) that are also blunt-ended or tailed with complementary nucleotides. Another example of an adapter type is a ball-type adapter with a blunt-ended or tailed end for ligation to nucleic acids to be analyzed.

[0069] A "partially double-stranded DNA fragment" refers to linear DNA that is partially double-stranded and partially single-stranded.

[0070] A "3' overhang" refers to one or more consecutive nucleotides at the 3' end of a partially double-stranded DNA fragment that do not anneal to a complementary nucleotide.

[0071] The interchangeable terms “oligomer,” “oligo,” and “oligonucleotide” generally refer to nucleic acids with fewer than 1,000 nucleotide (nt) residues (including polymers ranging from approximately 5 nt to 500–900 nt residues). In some embodiments, the size of oligonucleotides ranges from approximately 12–15 nt to 50–600 nt, while in other embodiments, it ranges from approximately 15–20 nt to 22–100 nt. Oligonucleotides can be purified from naturally occurring sources or synthesized using any of the various well-known enzymatic and chemical methods. The term oligonucleotide does not imply any specific function for any reagent; rather, it is used generally to refer to all such reagents described herein. Oligonucleotides can perform a variety of different functions. For example, they can function as primers if they can hybridize with a complementary chain and can be further extended in the presence of nucleic acid polymerase; they can function to detect a target nucleic acid if they can hybridize with a target nucleic acid or its amplicon and can further provide a detectable portion (e.g., a fluorophore).

[0072] A "primer" is an oligonucleotide that contains a 3' end that can be extended by polymerase.

[0073] A primer population "containing a random target hybridizing sequence" is a group of primers in which the target hybridizing sequence may vary more than the constant sequence. For example, the random target hybridizing sequence may include at least four positions that vary between populations, as will be discussed in detail elsewhere herein.

[0074] DNA polymerase is an enzyme that can extend a primer annealed to a template by adding a nucleotide to the 3' end of a primer complementary to the template (as is well known in the field, polymerases are generally understood to have some degree of error).

[0075] "3'→5' exonuclease activity" refers to the enzymatic activity that removes a nucleotide from the 3' end of a nucleic acid.

[0076] "5'→3' exonuclease activity" refers to the enzymatic activity that removes a nucleotide from the 5' end of a nucleic acid.

[0077] If enzyme activity cannot be detected in a standard assay for enzyme activity, the enzyme or polypeptide is considered to "lack" enzymatic activity. For example, exonuclease activity can be assayed by preparing a suitable nucleic acid substrate, which may have nucleotides labeled at the 3' or 5' ends, and determining whether the enzyme detectably removes the label. The designation "exo-" is used herein to omit polymerases that lack exonuclease activity.

[0078] A "tag" refers to any sequence added to a nucleic acid molecule, for example, by the incorporation of the 5' element of a primer or by ligation. Tags can have various functions, including serving as a primer binding site in subsequent reactions, providing information about the processing of a sample or molecule, or serving as a barcode or index to identify a molecule (either independently or in combination with its endogenous sequence) and its replicated or amplified products.

[0079] "Cell-free DNA" ("cfDNA") refers to DNA that is not contained within cells when isolated from a subject.

[0080] "Purification" refers to the separation of the analyte of interest (such as a repaired DNA molecule) from at least one other component of the composition (e.g., primers, enzymes, salts, and nucleotides). "Purification" includes any procedure that, by separation, yields a composition containing the analyte of interest at a higher concentration ratio compared to the other components (one or more) in the starting composition.

[0081] "Epigenetically modified" DNA involves one or more modifications to its nucleotides that originate in vivo. 5-methylation and 5-hydroxymethylation of cytosine are examples of epigenetically modified DNA.

[0082] Detailed explanation 1. Overview Nucleic acid samples often contain partially double-stranded nucleic acid fragments with single-stranded overhangs that require processing to prepare nucleic acid samples for sequencing, such as high-throughput sequencing or next-generation sequencing. While 5' overhangs can be repaired with a simple extension reaction (see Figure 1), which does not result in loss of sequence information, conventional 3' overhang repair relies on exonucleolisis, which removes sequence from the fragment.

[0083] The present invention provides an improved method for repairing a 3' overhang that can preserve the sequence of all or substantial portion of the 3' overhang, for example. In some embodiments, the method comprises the step of contacting a partially double-stranded DNA fragment containing a 3' overhang with one or more primers of a primer population, where the primer population includes a random target hybrid-forming sequence. One or more members of the primer population can be annealed to the partially double-stranded DNA fragment to extend, thereby producing one or more extended primers annealed to the DNA fragment. Another extended primer, or an extended primer annealed to the DNA fragment together with the 5' end of a strand of the partially double-stranded DNA fragment, can form a substrate for ligation. Ligation is then performed to obtain the repaired DNA fragment. Those skilled in the art will be familiar with the appropriate conditions for each of the individual operations of such a method (e.g., annealing of primers having a random target hybrid-forming sequence to a DNA molecule, extension of primers to the 5' end of another segment of DNA, and ligation of the primers extended to the 5' end).

[0084] In some embodiments, the primer is extended to the 5' end of the strand of a partially double-stranded DNA fragment and then ligated to the 5' end of the strand of the partially double-stranded DNA fragment. In some embodiments, the first primer is extended to the 5' end of the strand of a partially double-stranded DNA fragment; the second primer is extended to the 5' end of the first primer; then the first primer is ligated to the 5' end of the strand of the partially double-stranded DNA fragment, and the second primer is ligated to the 5' end of the first primer.

[0085] In some embodiments, at least steps (a) to (c) of the method described herein are carried out in a single tube.

[0086] In some embodiments, double-stranded DNA fragments are present in the composition, and for at least steps (a) to (c), the components are not removed from the composition.

[0087] 2. Partially double-stranded DNA fragments In some embodiments, a partially double-stranded DNA fragment has a 3' overhang at each end. In some embodiments, a partially double-stranded DNA fragment has a 3' overhang and (i) a blunt end or (ii) a 5' overhang. In some embodiments, the 5' overhang is repaired by extending the 3' end along the 5' overhang.

[0088] In some embodiments, the partially double-stranded DNA fragment is a cfDNA fragment. In some embodiments, the partially double-stranded DNA fragment is of mammalian origin. In some embodiments, the partially double-stranded DNA fragment is of human origin. In some embodiments, the partially double-stranded DNA fragment is derived from a bodily fluid sample. In further embodiments, the bodily fluid is whole blood, serum, or plasma.

[0089] In some embodiments, partially double-stranded DNA fragments are part of a population of DNA fragments in the composition. In further embodiments, the population of DNA fragments includes sheared DNA. In further embodiments, the population of DNA fragments includes epigenetically modified DNA. In some embodiments, the population of DNA fragments includes fragments derived from multiple genomic loci (e.g., at least 10, 100, 1000, or 10000 genomic loci). In some embodiments, the population of DNA fragments is unenriched. A population is unenriched if it has not been subjected to a procedure that increases the abundance of one fragment compared to another (such as amplification using sequence-specific primers or capture of a target using a sequence-specific capture probe). In some embodiments, an unamplified population of DNA fragments means that the population has not undergone any amplification procedure. Unamplified DNA can be used to preserve epigenetic genetic information, such as DNA methylation.

[0090] In some embodiments, partially double-stranded DNA fragments are present in or obtained from the sample. The sample may be any biological sample isolated from a subject. The sample may include body tissues, e.g., solid tumors of known or suspected presence, whole blood, platelets, serum, plasma, stool, red blood cells, leukocytes or leukocytes, endothelial cells, tissue biopsy, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, intercellular fluids (including gingival crevicular exudate), bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions, as well as urine. The sample may be in the form initially isolated from the subject, or it may be subjected to further processing to remove or add components (such as cells), or to enrich one component compared to others. In some embodiments, partially double-stranded DNA fragments are derived from a body fluid sample. In further embodiments, the body fluid is whole blood, serum, or plasma.

[0091] In some embodiments, partially double-stranded DNA fragments are derived from a plasma sample. The volume of plasma may depend on the desired read depth of the region to be sequenced. Exemplary volumes are 0.4–40 mL, 5–20 mL, and 10–20 mL. For example, the volume may be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of sampled plasma may be, for example, 5–20 mL.

[0092] Samples can contain varying amounts of DNA, including genome equivalents. For example, a DNA sample of about 30 ng can contain about 10,000 haploid human genome equivalents, or about 200 billion individual nucleic acid molecules in the case of cell-free DNA (cfDNA). Similarly, a DNA sample of about 100 ng can contain about 30,000 haploid human genome equivalents, or about 600 billion individual molecules in the case of cell-free DNA (cfDNA). Some samples contain 1-500, 2-100, or 5-150 ng of cell-free DNA (e.g., 5-30 ng or 10-150 ng of cell-free DNA).

[0093] The sample may contain DNA from different sources. For example, the sample may contain germline DNA or somatic DNA. The sample may contain mutant DNA. For example, the sample may contain DNA with germline mutations and / or somatic mutations. The sample may also contain DNA with cancer-related mutations (e.g., cancer-related somatic mutations).

[0094] Examples of cell-free DNA (cfDNA) amounts in the sample before amplification range from approximately 1 fg to approximately 1 ug, for example, 1 pg to 200 ng, 1 ng to 100 ng, and 10 ng to 1000 ng. For example, the amount may be up to approximately 600 ng, up to approximately 500 ng, up to approximately 400 ng, up to approximately 300 ng, up to approximately 200 ng, up to approximately 100 ng, up to approximately 50 ng, or up to approximately 20 ng of cell-free nucleic acid molecules. The amount may be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity can be 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or up to 200 ng of cell-free DNA molecules. The method may include a step to obtain 1 femtogram (fg) to 200 ng.

[0095] In some embodiments, the bodily fluid sample is 5-10 ml of whole blood, plasma, or serum, and the sample contains about 30 ng of DNA or about 10,000 haploid human genome equivalents.

[0096] In some embodiments, partially double-stranded DNA fragments are part of a population of DNA fragments in the composition. In further embodiments, the population of DNA fragments includes sheared DNA. In some embodiments, the partially double-stranded DNA fragments are cfDNA fragments.

[0097] Cell-free DNA is DNA that is not contained within a cell or otherwise not bound to a cell; in other words, nucleic acids remain in the sample after the removal of intact cells. Cell-free DNA can be double-stranded, single-stranded, or a hybrid thereof. In some embodiments, cfDNA fragments contain double-stranded DNA molecules, at least some of which have single-stranded overhangs. Cell-free DNA can be released into the body fluid by secretion or by cell death processes (e.g., cell necrosis and apoptosis). Some cell-free DNA is released into the body fluid from cancer cells (e.g., circulating tumor DNA (ctDNA)). Some is released from healthy cells.

[0098] Cell-free DNA can have one or more epigenetic modifications; for example, cell-free nucleic acids may be acetylated, methylated, ubiquitinated, phosphorylated, sumoated, ribosylated, and / or citrullinated. In some embodiments, partially double-stranded DNA fragments are part of a population of DNA fragments in the composition, and this population of DNA fragments includes epigenetically modified DNA. Cell-free DNA has a size distribution of about 100–500 nucleotides, particularly 110–about 230 nucleotides (mode is about 168 nucleotides), and a second small peak in the range of 240–440 nucleotides.

[0099] Cell-free DNA can be isolated from body fluids by a splitting step, and if found in solution, cell-free DNA is separated from intact cells and other insoluble components of the body fluid. Splitting may involve techniques such as centrifugation or filtration. Alternatively, cells in the body fluid can be lysed, and cell-free nucleic acids and cellular nucleic acids can be processed together. Generally, after buffer addition and washing steps, nucleic acids can be precipitated with alcohol. Further cleaning steps, such as silica-based columns, may be used to remove impurities or salts. Nonspecific bulk carrier nucleic acids may be added, for example, during the reaction to optimize certain aspects of the procedure, such as yield.

[0100] After such processing, the sample may contain various DNA forms (including double-stranded and single-stranded DNA). If necessary, single-stranded DNA can be converted to double-stranded DNA, and therefore, the double-stranded DNA form will be included in subsequent processing and analysis steps.

[0101] 3. DNA polymerase In some embodiments, the DNA polymerase lacks 3'→5' exonuclease activity. In some embodiments, the DNA polymerase lacks 5'→3' exonuclease activity. In some embodiments, the DNA polymerase lacks strand displacement activity. Any DNA polymerase known in the art that can exhibit 5'-3' polymerase activity and lacks 3'→5' exonuclease activity, 5'→3' exonuclease activity, and strand displacement activity may be used as the DNA polymerase for step (b) of the method described herein. In some embodiments, the DNA polymerase is a Klenow fragment. In some embodiments, the DNA polymerase is an exo-Klenow fragment.

[0102] 4. Primer In some embodiments, the primers in the primer population are single-stranded.

[0103] In some embodiments, the primers of the primer population include a hairpin region or a double-stranded region. In some embodiments, the primers of the primer population are double-stranded with a 3' overhang, and the random target hybridizing sequence is located within the 3' overhang. In some embodiments, the primers of the primer population are hairpins with a 3' overhang, and the random target hybridizing sequence is located within the 3' overhang.

[0104] In some embodiments, the primers of the primer group include adapters. In some embodiments, the hairpin region or double-stranded region of the primer includes adapters.

[0105] In some embodiments, the adapter includes a tag as described herein. In some embodiments, the adapter includes a tag. In further embodiments, the tag includes a barcode. In further embodiments, the primer group includes a plurality of different barcodes.

[0106] Identical or different adapters can be attached to each end of the repaired DNA fragment. Occasionally, identical adapters, except for different tags, are attached to each end. In some embodiments, the adapters are Y-shaped adapters. In some embodiments, the adapters are bell-shaped adapters. The adapters used herein are further described below (Section 6).

[0107] In some embodiments, the length of the random target hybrid sequence is at least 4, 5, 6, 7, 8, 9, or 10 nucleotides. In some embodiments, the length of the random target hybrid sequence is at least 4 nucleotides. In some embodiments, the length of the random target hybrid sequence is at least 5 nucleotides. In some embodiments, the length of the random target hybrid sequence is at least 6 nucleotides. In some embodiments, the length of the random target hybrid sequence is at least 7 nucleotides. In some embodiments, the length of the random target hybrid sequence is at least 8 nucleotides. In some embodiments, the length of the random target hybrid sequence is at least 9 nucleotides. In some embodiments, the length of the random target hybrid sequence is at least 10 nucleotides.

[0108] In some embodiments, the length of the random target hybrid sequence is approximately 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 4 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 5 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 6 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 7 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 8 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 9 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 10 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 11 nucleotides. In some embodiments, the length of the random target hybrid sequence is approximately 12 nucleotides.

[0109] In any embodiment described herein, the primer population may include members in which each of four different bases (e.g., A, C, T, and G) appears at each position of the random target hybridization sequence. In other words, in different members of the primer population, each of the four bases can appear at each position of the random target hybridization sequence. Those skilled in the art will recognize that U can be used in place of T, and / or modified bases (e.g., methylated cytosine, pseudouridine, etc.) having the same base pairing priority as the unmodified bases can be used; as such, T encompasses U, and each of A, C, T, and G encompasses modified forms that retain the same base pairing priority as the unmodified bases.

[0110] In some embodiments, a partially double-stranded DNA fragment includes two 3' overhangs, each of which is repaired as disclosed herein.

[0111] In some embodiments, a partially double-stranded DNA fragment includes a 3' overhang and a 5' overhang. The 5' overhang can be repaired, for example, by extending the 3' end incorporated along the 5' overhang. In some embodiments, this extension is performed by the same polymerase that extends one or more primers along the 3' overhang.

[0112] In some embodiments, the repaired DNA fragment contains one or two blunt ends. In some embodiments, after the ligation reaction described herein, any remaining overhangs are further repaired, for example, using a suitable exonuclease (e.g., a 3'→5' exonuclease). This results in the loss of a small amount of sequence, but still provides a molecule with a blunt end that can be subjected to further manipulation, while retaining a substantial amount of the sequence that was the original portion of the 3' overhang.

[0113] In some embodiments, the repaired DNA fragment is subjected to end tailing using a polymerase that adds nucleotides to the 3' end of a blunt-ended nucleic acid in a non-template-directed manner, for example, with A preferred over G, and G preferred over C or T, as needed. This polymerase may lack proofreading function and / or may be thermally stable, retaining activity at high temperatures, for example. Taq, Bst large fragment, and Tth polymerases are examples of such polymerases. The reaction mixture typically contains equimolar amounts of each of the four standard nucleotide types derived from the previous step, but the four nucleotide types are not added to the 3' end in equal proportions. Rather, A is added most frequently, followed by G, and then C and T.

[0114] Where applicable, blunt-ending and tailing of repaired DNA fragments can be performed in a single tube. Blunt-ended nucleic acids do not need to be separated from the enzyme(s) performing the blunt-ending before the tailing reaction occurs. If necessary, all enzymes, nucleotides, and other reagents are supplied together before the blunt-ending reaction occurs. Supplying together means introducing everything into the sample at a sufficiently close time so that everything is present when the sample is incubated for blunt-ending. If necessary, nothing is removed from the sample after supplying the enzymes, nucleotides, and other reagents until incubation for at least both blunt-ending and end-tailing is complete. Often, the end-tailing reaction is performed at a higher temperature than the blunt-ending reaction. For example, the blunt-end reaction can be carried out at ambient temperature where the 5'-3' polymerase and 3'-5' exonuclease are active and the thermostable polymerase is inactive or has minimal activity, while the end-tailing reaction can be carried out at high temperatures (above 60°C) where the 5'-3' polymerase and 3'-5' exonuclease are inactive and the thermostable polymerase is active.

[0115] In some embodiments, after repair and / or after any further processing steps, the enzyme(s) (e.g., polymerase, ligase, and / or exonuclease) are denatured, for example, by thermal denaturation. For example, denaturation can be achieved by raising the temperature to 75°C to 80°C.

[0116] 5. Linking the repaired DNA fragment to the adapter. In some embodiments, after repair, the tailed sample molecule is brought into contact with an adapter, either after purification or without purification. For example, after tailing a repaired DNA fragment, the tailed sample molecule may be brought into contact with an adapter whose one end is tailed with complementary T and C nucleotides. In another example, a repaired DNA fragment with a blunt end may be brought into contact with an adapter with a blunt end.

[0117] Adapters can be formed by the individual synthesis and annealing of each of their strands. Therefore, additional T-tails and C-tails, if used, can be added as additional nucleotides in the synthesis of one of the strands. Typically, adapters tailed with G and A are not included because, while these adapters may anneal with sample molecules tailed with C and T respectively, they will also anneal with other adapters. Adapter molecules and sample molecules possessing complementary nucleotides (i.e., TA and CG) at their 3' ends can anneal and link together. The percentage of C-tailed adapters compared to T-tailed adapters may range from approximately 5–40% on a molar basis (e.g., 10–35%, 15–25%, 20–35%, 25–35%, or approximately 30%). Since the non-template-directed addition of a single nucleotide to the 3' end of the sample molecule does not proceed to completion, the sample may also contain some untailed, blunt-ended sample molecules. These molecules can be recovered by supplying the sample with one, preferably sole, blunt-end adapters. T-tailed and C-tailed adapters can be supplied to the blunt-end adapters in adapter molar ratios of 0.2–20%, 0.5–15%, or 1–10%. The blunt-end adapters can be supplied simultaneously, before, or after the T-tailed and C-tailed adapters. When the blunt-end adapters are re-ligated with the blunt-end sample molecules, sample molecules are obtained with adapters on both sides. These molecules lack an AT nucleotide pair or a CG nucleotide pair between the samples, and the adapter appears when the tailed sample molecule is ligated to the tailed adapter.

[0118] The adapters used in these reactions preferably have a single end that is T or C-tailed or a single blunt end so that the adapter can ligate to the sample molecule in only one direction. The adapter may be, for example, a Y-shaped adapter with one end being tailed or blunt and the other end having two single strands. An exemplary Y-shaped adapter has a sequence having a tag (6 bases) as shown below. The oligonucleotide above contains a single base T tail.

[0119] Universal adapter: 5'AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCT(Sequence ID 1).

[0120] Adapter, index 1-12: 5' GATCGGAAGAGCACACGTCTGAACTCCAGTCAC (6 bases)ATCTCGTATGCCGTCTTCTGCTTG (Sequence ID 2)

[0121] Another Y-shaped adapter with a C-tail has the following arrangement:

[0122] 5'AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCC (Sequence ID 3) and adapter, indices 1-12: 5' GATCGGAAGAGCACACGTCTGAACTCCAGTCAC (6 bases)ATCTCGTATGCCGTCTTCTGCTTG (Sequence ID 2)

[0123] Customized combinations of such oligonucleotides (including oligonucleotides having both T-tails and C-tails) can be synthesized for use in this method.

[0124] Shortened versions of these adapter sequences are described in Rohland et al., Genome Res. 2012 May;22(5):939-946.

[0125] The adapter may also be bell-shaped with a single end that is either tailed or blunt-ended. The adapter may include a primer binding site for amplification, a binding site for sequencing primers, and / or a nucleic acid tag for identification. The same or different adapters can be used in a single reaction.

[0126] If the adapter contains an identification tag and nucleic acids in the sample are attached to the adapter at each end, the number of potential identifier combinations increases exponentially with the number of unique tags supplied (i.e., n n The number of unique tag combinations is n, where n is the number of unique identification tags. In some methods, the number of unique tag combinations is statistically sufficient to ensure that all or substantially all (e.g., at least 90%) of the different double-stranded DNA molecules in the sample receive different tag combinations. In some methods, the number of unique identification tag combinations is less than the number of unique double-stranded DNA molecules in the sample (e.g., 5 to 10,000 different tag combinations).

[0127] A kit providing enzymes suitable for carrying out the above method is the NEBNext® Ultra® II DNA Library Prep Kit for Illumina®. This kit provides the following reagents: NEBNext Ultra II-terminated enzyme mix, NEBNext Ultra II-terminated reaction buffer, NEBNext ligation enhancer, NEBNext Ultra II ligation master mix-20, and NEBNext® Ultra II Q5® master mix.

[0128] A population of adapter-treated nucleic acids can be obtained by attaching the described T-tailed and C-tailed adapters, the population comprising multiple nucleic acid molecules, each nucleic acid molecule comprising a nucleic acid fragment, the nucleic acid fragment flanked on both sides by adapters containing barcodes, and forming A / T or G / C base pairs between the nucleic acid fragment and the adapter. The number of nucleic acid molecules may be at least 10,000, 100,000, or 1,000,000 molecules. Most of the nucleic acids in the population can be flanked by adapters having different barcodes (e.g., at least 99%). If blunt-ended adapters are also included, the population comprises nucleic acid molecules in which adapters are directly attached to either end or both ends of a nucleic acid fragment (i.e., without intervening A / T or G / C pairs).

[0129] 6. Amplification The repaired DNA fragment adjacent to the adapter can be amplified by PCR or another amplification method, which primes it with a primer that binds to a primer binding site in the adapter adjacent to the nucleic acid to be amplified. The amplification method may include a cycle of extension, denaturation, and annealing by thermocycling, or it may be isothermal amplification such as transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and autonomous sequence-based replication.

[0130] In some embodiments, the method described herein further includes a step of amplifying the repaired DNA fragment after steps (a) to (c). In further embodiments, the repaired DNA fragment comprises one or more adapters (e.g., two adapters), and the amplification of the repaired DNA fragment is performed using one or more (e.g., two) amplification oligomers that anneal to the one or more adapters.

[0131] If an enrichment step is performed (as discussed elsewhere in this specification), amplification may precede and / or follow the enrichment step. In some embodiments, the method includes the steps of amplifying the repaired fragment, performing an enrichment step to obtain an added fragment, performing an enrichment step, and further amplifying the subsequently added repaired fragment.

[0132] 7. Tags In some embodiments, nucleic acid molecules (derived from polynucleotide samples) may be tagged with a sample index and / or molecular barcode (commonly referred to as a “tag”). The tag may be incorporated into an adapter, or otherwise ligated, in particular by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap-extension polymerase chain reaction (PCR). Such an adapter may ultimately be ligated to a target nucleic acid molecule. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are generally applied to introduce the sample index to the nucleic acid molecule using a conventional nucleic acid amplification method. Amplification may occur in one or more reaction mixtures (e.g., multiple microwells in an array). The molecular barcode and / or sample index may be introduced simultaneously or in any order. In some embodiments, the molecular barcode and / or sample index are introduced before and / or after the sequence capture step. In some embodiments, only the molecular barcode is introduced before probe capture, and the sample index is introduced after the sequence capture step. In some embodiments, both the molecular barcode and sample index are introduced before the probe-based capture step. In some embodiments, the sample index is introduced after the sequence capture step has been performed. In some embodiments, the molecular barcode is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample via an adapter by ligation (e.g., blunt-end ligation or adherent-end ligation). In some embodiments, the sample index is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample by overlap extension polymerase chain reaction (PCR). Typically, the sequence capture protocol involves the introduction of a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence (e.g., the coding sequence of a genomic region), where mutations in such a region are associated with a particular type of cancer.

[0133] In some embodiments, the tag may be positioned at one or both ends of the sample nucleic acid molecule. In some embodiments, the tag is an oligonucleotide with a predetermined, random, or semi-random sequence. In some embodiments, the tag may be about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or less than 1 nucleotide in length. The tag may be ligated to the sample nucleic acid randomly or non-randomly.

[0134] In some embodiments, each sample is uniquely tagged using a sample index or a combination of sample indices. In some embodiments, each nucleic acid molecule in a sample or subsample is uniquely tagged using a molecular barcode or a combination of molecular barcodes. In other embodiments, multiple molecular barcodes may be used, such that the molecular barcodes are not necessarily unique to one another (e.g., non-unique molecular barcodes). In these embodiments, the molecular barcodes are generally attached to individual molecules (e.g., by ligation), and as a result, a unique sequence is created that can be individually tracked when the molecular barcode and the sequence to which it can be attached are combined. When the detection of non-uniquely tagged molecular barcodes is combined with endogenous sequence information (e.g., start and / or stop positions corresponding to the sequence of the original nucleic acid molecule in the sample, subsequences of the sequence reads at one or both ends, the length of the sequence reads, and / or the length of the original nucleic acid molecule in the sample), it is typically possible to assign a unique identity to a particular molecule. The length or number of base pairs of individual sequence reads may also be used to assign a unique identity to a given molecule, if necessary. As described herein, a single-stranded nucleic acid fragment to which a unique identity has been assigned may subsequently be able to identify the parent strand fragment and / or complementary strand.

[0135] In some embodiments, molecular barcodes are introduced in an expected ratio of identifier sets (e.g., combinations of unique or non-unique molecular barcodes) for molecules in the sample. One example of the format is to use about 2 to about 1,000,000 different molecular barcodes, or about 5 to about 150 different molecular barcodes, or about 20 to about 50 different molecular barcodes. Alternatively, about 25 to about 1,000,000 different molecular barcodes may be used. Molecular barcodes can be ligated to both ends of the target molecule. For example, 20-50 × 20-50 molecular barcodes can be used. In some embodiments, 20-50 different molecular barcodes can be used. In some embodiments, 5-100 different molecular barcodes can be used, and in some embodiments, 5-150 different molecular barcodes can be used. In some embodiments, 5-200 different molecular barcodes can be used. The members of such identifiers are typically sufficient to increase the probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) that different molecules having the same start and end points receive different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of molecules have the same combination of molecular barcodes.

[0136] In some embodiments, the assignment of unique or non-unique molecular barcodes during the reaction is carried out using the methods and systems described, for example, in U.S. Patent Applications Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patents Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992 (each of which is incorporated herein by reference in whole). Alternatively, in some embodiments, different nucleic acid molecules in a sample may be identified using only endogenous sequence information (e.g., start and / or end positions, subsequences at one or both ends of the sequence, and / or length).

[0137] Accordingly, this disclosure also provides compositions of repaired and tagged DNA fragments produced by the methods described herein. The polynucleotides may include fragmented DNA (e.g., cfDNA). A set of polynucleotides in a composition can be non-uniquely tagged, that is, the number of different identifiers may be at least two and less than the number of polynucleotides mapped to the mappable base positions. A composition between about 10 ng and about 10 μg (e.g., any of about 10 ng to 1 μg, about 10 ng to 100 ng, about 100 ng to 10 μg, about 100 ng to 1 μg, or about 1 μg to 10 μg) may have any number of different identifiers ranging from 2, 5, 10, 50, or 100 to any of 100, 1,000, 10,000, or 100,000. For example, polynucleotides in such a composition can be tagged using different identifiers between 5 and 100, or between 100 and 4000.

[0138] The event in which different molecules are mapped to the same coordinates (in this case, having the same start / end positions) and possess the same tag rather than different tags is called a “molecular collision.” In certain cases, the actual number of molecular collisions may be greater than the theoretical number of collisions calculated as described above. This may correlate with the non-uniform distribution of molecules across the coordinates, differences in ligation efficiency between barcodes, and other factors. In this case, an empirical method can be used to determine the number of barcodes required to approach the theoretical number of collisions. In one embodiment, a method is provided herein for determining the number of barcodes required to reduce barcode collisions of a given haploid genome equivalent based on the distribution of lengths and sequence uniformity of the sequenced molecules. The method includes the steps of creating multiple nucleic acid molecule pools; tagging each pool with an increasing number of barcodes; and determining the optimal number of barcodes to reduce the number of barcode collisions to a theoretical level (e.g., due to differences in effective barcode concentration resulting from differences in pool efficiency and ligation efficiency).

[0139] In one embodiment, the number of identifiers required to substantially uniquely tag polynucleotides mapped to a given region can be determined empirically. For example, a selected number of different identifiers can be attached to molecules in the sample, and the number of different identifiers needed to map molecules to the region can be counted. If the number of identifiers used is insufficient, several polynucleotides mapped to the region will have the same identifier. In that case, the number of identifiers counted will be less than the number of original molecules in the sample. The number of different identifiers used can be iteratively increased for each sample type until no further identifiers indicating novel original molecules are detected. For example, in a first iteration, five different identifiers may be counted, indicating at least five different original molecules. In a second iteration, using more barcodes, seven different identifiers are counted, indicating at least seven different original molecules. In a third iteration, using more barcodes, ten different identifiers are counted, indicating at least ten different original molecules. In a fourth iteration, using more barcodes, ten different identifiers are counted again. At this point, adding more barcodes is unlikely to increase the number of detected molecules from the original molecule.

[0140] 8. Enrichment In some embodiments, a sample containing DNA fragments is enriched with respect to the fragment of interest. For example, enrichment may occur after ligation of a primer extended to the 5' end of a partially double-stranded DNA fragment, or after an amplification step following such ligation and adapter attachment (either as part of the ligation or in a subsequent step). Enrichment refers to any procedure that increases the relative abundance of the fragment of interest to other fragments, and includes procedures in which the fragment of interest is preferentially retained in the sample while other fragments are removed. Enrichment may be, for example, a capture step using a capture probe set having a target hybrid-forming sequence specific to the target of interest. The target of interest may include one or more or all of a single nucleotide variant, a copy number variable region, a fusion, and an indel. In some embodiments, one or more or all of a single nucleotide variant, a copy number variable region, a fusion, and an indel are associated with a disease or disorder (e.g., cancer, such as any cancer discussed elsewhere herein).

[0141] As discussed above, nucleic acids in a sample can be subjected to a capture step to capture molecules having a target sequence for subsequent analysis. Target capture may include the use of a bait set containing oligonucleotide baits labeled with a capture moiety (such as biotin or other examples below). The probe may have a sequence selected to explore a group of regions, such as genes. In some embodiments, as discussed elsewhere herein, the bait set can increase or decrease the capture yield of target region sets (such as the capture yields of sequence-variable target region sets and epigenetic target region sets, respectively). Such a bait set is combined with the sample under conditions that allow the target molecules to hybridize with the bait. The captured molecules are then isolated using the capture moiety. For example, the biotin capture moiety is a bead-based streptavidin. Such methods are further described, for example, in U.S. Patent No. 9,850,523 issued December 26, 2017 (as incorporated herein by reference).

[0142] The capture portion includes, but is not limited to, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetic particles. The extraction portion may be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, the capture portion attached to the analyte is captured by its binding pair attached to an isolateable portion (such as magnetic particles or large particles that can be precipitated by centrifugation). The capture portion may be any type of molecule capable of affinity separation of nucleic acids possessing the capture portion from nucleic acids lacking the capture portion. Exemplary capture portions are biotin capable of affinity separation by binding to streptavidin linked to or linkable to a solid phase, or oligonucleotides capable of affinity separation by binding to complementary oligonucleotides linked to or linkable to a solid phase. 9. Exemplary workflow for preparing samples for sequencing.

[0143] In some embodiments, the methods described herein include a step of producing a DNA fragment repaired according to one of the embodiments described above, wherein the adapter is incorporated by a ligation step or a subsequent step. The adapter includes a primer binding site and optionally a barcode. The repaired DNA fragment containing the adapter is subjected to an amplification reaction. An enrichment step may follow the amplification reaction as described herein. A further amplification step may follow the enrichment step, if desired. In some embodiments, an additional tag (e.g., a sample index) is added during the amplification reaction or a further amplification step. These workflows allow for the preparation of a sample for sequencing by including either or both a barcode and a sample index in the repaired fragment, and by enriching the repaired fragment with respect to the fragment of interest. 10. Sequence determination

[0144] In some embodiments, the methods described herein further include the step of sequencing the repaired DNA fragment. In further embodiments, sequencing involves sequencing nucleotides that have formed 3' overhangs in a partially double-stranded DNA fragment.

[0145] Repaired DNA fragments adjacent to pre-amplified or unamplified adapters can be subjected to sequencing. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, synthesis sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, synthetic single-molecule sequencing (SMSS) (Helicos), massively parallel sequencing, clonal single-molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions can be performed in various sample processing units, which may include other means for processing multiple lanes, multiple channels, multiple wells, or multiple sample sets substantially simultaneously. The sample processing unit may also include multiple sample chambers to allow multiple processing operations to run simultaneously. In some embodiments, multiple repaired DNA fragments are generated, at least a portion of which include a tag, and sequencing generates multiple sequence reads from the multiple repaired DNA fragments. In some embodiments, the sequence reads include sequences of nucleotides and tags that form a 3' overhang in a partially double-stranded DNA fragment.

[0146] The sequencing reaction can be performed on one or more fragment types known to contain markers for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. The sequencing reaction may provide a genome sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100%. In other cases, the genome sequence coverage may be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100%.

[0147] Simultaneous sequencing reactions can be performed using multiple sequencing. In some cases, cell-free nucleic acids can be sequenced using at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions. In other cases, cell-free polynucleotides can be sequenced using fewer than 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions. Sequencing reactions can be performed sequentially or simultaneously. Subsequently, data analysis may be performed on all or part of the sequencing reactions. In some cases, data analysis may be performed on sequencing reactions with at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequences. In other cases, data analysis may be performed on sequencing reactions with fewer than 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequences.

[0148] The sequencing method may be massively parallel sequencing (i.e., sequencing at least 100, 1,000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion nucleic acid molecules simultaneously (or sequentially)). 11.Analysis

[0149] The method of the present invention can be used to diagnose the presence or absence of a condition (particularly cancer) in a subject, characterize the condition (e.g., determine the staging of cancer or the heterogeneity of cancer), monitor the response of the condition to treatment, or determine the risk of recurrence of the condition in the subject.

[0150] Various types of cancer can be detected using the method of the present invention. As with most cells, cancer cells can be characterized by their rate of metabolic turnover (the death of old cells and replacement with new ones). Generally, dead cells in contact with vascular structures in a given subject may release DNA or DNA fragments into the bloodstream. This is also true for cancer cells at various stages of the disease. Furthermore, cancer cells may be characterized by various genetic abnormalities, such as copy number variations and rare mutations, depending on the stage of the disease. This phenomenon can be used to detect the presence or absence of cancer in an individual using the method and system described herein.

[0151] The types and number of cancers that may be detected may include blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, pharyngeal cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, gastric cancers, solid tumors, mixed tumors, and homogeneous tumors.

[0152] Cancer can be detected from genetic variations (including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, changes in chromosome structure, gene fusions, chromosome fusions, gene shortening, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation due to infection and cancer).

[0153] Furthermore, genetic data can be used to characterize specific cancer morphologies. Cancers are often heterogeneous in both composition and staging. Genetic profile data may enable the characterization of specific cancer subtypes, which can be important in the diagnosis or treatment of specific cancer subtypes. This information can also provide subjects or physicians with clues to predict the prognosis of specific cancer types, allowing them to adopt treatment options as the diagnosis progresses. Some cancers progress, become more invasive, and become genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of this disclosure may be useful in determining disease progression.

[0154] Furthermore, this analysis is useful in determining the effectiveness of specific treatment choices. Since a successful treatment can lead to the death of more cancer cells and the release of DNA, a well-chosen treatment can increase the detection rate of copy number variations or rare mutations in the subject's blood. In other cases, this may not occur. In other cases, a particular treatment choice may correlate with the genetic profile of the cancer over time. This correlation can be useful in selecting a treatment. Moreover, if cancer remission is observed after treatment, the method of the present invention can be used to monitor residual disease or disease recurrence.

[0155] Furthermore, the method of the present invention can be used to detect genetic variations in conditions other than cancer. Immune cells, such as B cells, can rapidly proliferate clones in the presence of certain diseases. Clonal proliferation can be monitored using the detection of copy number variation, and a particular immune state can be monitored. In this example, the analysis of copy number variation over time can be used to obtain a profile of the progression of a particular disease. Using the detection of copy number variation or even rarer mutations, it is possible to determine how the pathogen population is changing during infection. This may be particularly important in chronic infections (such as HIV / AIDS or hepatitis) in which viruses can change the state of their life cycle during infection and / or mutate into more virulent forms. Using the method of the present invention, rejection activity in the host body can be determined or profiled, the state of transplanted tissue can be monitored if immune cells are attempting to destroy the transplanted tissue, and the treatment or preventive measures for rejection can be modified.

[0156] Furthermore, the methods of this disclosure can be used to characterize heterogeneity of abnormal conditions in a subject, the methods comprising the step of creating a gene profile of extracellular polynucleotides in the subject, the gene profile including multiple data from the analysis of copy number variation and rare mutations. In some cases, including but not limited to cancer, disease can be heterogeneous. Disease cells may not be identical. In the case of cancer, it is known that some tumors contain different types of tumor cells, and some cells have different stages of cancer. In other cases, heterogeneity may include multiple lesions. Furthermore, in the case of cancer, there may be multiple tumor lesions, and perhaps one or more lesions are the result of metastasis extending from the primary site.

[0157] Using the method of the present invention, a fingerprint or dataset can be created or profiled that aggregates genetic information from different cell origins in heterogeneous diseases. This dataset may include, alone or in combination, analyses of copy number variation and rare mutations.

[0158] This method can be used to diagnose, predict the prognosis of, monitor, or observe cancer or other diseases of fetal origin. In other words, these methodologies can be used in pregnant subjects to diagnose, predict the prognosis of, monitor, or observe cancer or other diseases in prenatal subjects where DNA and other nucleic acids can circulate simultaneously with maternal molecules.

[0159] Unless otherwise specifically indicated, any feature, process, element, embodiment, or aspect of the present invention may be used in combination with any other. While the present invention has been described in some detail as illustrations and examples for the purposes of clarity and understanding, it will be apparent that certain changes and modifications may be made within the scope of the appended claims. In certain embodiments, for example, the following items are provided: (Item 1) A method for partially repairing double-stranded DNA fragments, (a) A step of contacting the partially double-stranded DNA fragment with one or more primers of a primer population, wherein the partially double-stranded DNA fragment includes a 3' overhang and the primer population includes a random target hybridization sequence; (b) a step of extending one or more primers of the primer population along the DNA fragment using a DNA polymerase, thereby producing one or more extended primers annealed to the DNA fragment; and (c) A step of ligating the 3' end of one or more elongated primers to the 5' end of the strand of the elongated primer or a partially double-stranded DNA fragment, thereby obtaining a repaired DNA fragment. Methods that include... (Item 2) The method according to any one of the preceding items, wherein the DNA polymerase lacks 3'→5' exonuclease activity. (Item 3) The method according to any one of the preceding items, wherein the DNA polymerase lacks 5'→3' exonuclease activity. (Item 4) The method according to any one of the preceding items, wherein the DNA polymerase lacks strand displacement activity. (Item 5) The method according to any one of the preceding items, wherein the DNA polymerase is a Klenow fragment. (Item 6) The method according to the preceding item, wherein the DNA polymerase is an exo-Klenow fragment. (Item 7) The method according to any one of the preceding items, wherein the partially double-stranded DNA fragment has a 3' overhang at each end. (Item 8) The method according to any one of the preceding items, wherein the partially double-stranded DNA fragment has a 3' overhang and (i) a blunt end or (ii) a 5' overhang. (Item 9) The method according to the preceding item, wherein the 5' overhang is repaired by extending the 3' end along the 5' overhang. (Item 10) The method according to any one of the preceding items, wherein the partially double-stranded DNA fragment is derived from a bodily fluid sample. (Item 11) The method according to the preceding item, wherein the body fluid is whole blood, serum, plasma, or urine. (Item 12) The method according to any one of the preceding items, wherein the partially double-stranded DNA fragment is a cfDNA fragment. (Item 13) The method according to any one of the preceding items, wherein the partially double-stranded DNA fragment is of mammalian origin. (Item 14) The method according to any one of the preceding items, wherein the partially double-stranded DNA fragment is of human origin. (Item 15) The method according to any one of the preceding items, wherein the partially double-stranded DNA fragment is part of a population of DNA fragments in the composition. (Item 16) The method described in the preceding item, wherein the DNA fragment population includes sheared DNA. (Item 17) The method according to item 15 or 16, wherein the DNA fragment population includes epigenetically modified DNA. (Item 18) The method according to any one of items 15 to 17, wherein the DNA fragment population includes fragments derived from multiple genomic loci. (Item 19) The method according to any one of items 15 to 18, wherein the aforementioned DNA fragment population is not enriched. (Item 20) The method according to any one of items 15 to 19, wherein the aforementioned DNA fragment population is not amplified. (Item 21) The method according to any one of the preceding items, wherein the length of the random target hybrid-forming sequence is at least 4, 5, 6, 7, 8, 9, or 10 nucleotides. (Item 22) The method according to any one of the preceding items, wherein the length of the random target hybrid-forming sequence is approximately 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides. (Item 23) The method according to any one of the preceding items, wherein the primers of the aforementioned primer population are single-stranded. (Item 24) The method according to any one of items 1 to 20, wherein the primers of the primer population are double-stranded having a 3' overhang, and the random target hybridization sequence is located within the 3' overhang. (Item 25) The method according to any one of items 1 to 20, wherein the primers of the primer population are hairpins having a 3' overhang, and the random target hybridizing sequence is located within the 3' overhang. (Item 26) The method according to any one of items 21 to 22, wherein the hairpin or double-stranded region of the primer includes an adapter. (Item 27) The method according to any one of items 21 to 23, wherein the hairpin or double-stranded region of the primer includes a tag. (Item 28) The method described in item 24, wherein the aforementioned tag includes a barcode. (Item 29) The method according to item 25, wherein the primer group includes multiple different barcodes. (Item 30) The method described in any one of the preceding items, wherein at least steps (a) to (c) are performed in a single pipe. (Item 31) The method according to any one of the preceding items, wherein the double-stranded DNA fragment is present in the composition and the component is not removed from the composition for at least steps (a) to (c). (Item 32) The method according to any one of the preceding items, wherein the repaired DNA fragment comprises one or two blunt ends, or the method further comprises the step of blunt-ending one or two ends of the repaired DNA fragment. (Item 33) The method according to the preceding item, further comprising the step of end-tailing the repaired DNA fragment using a polymerase that adds nucleotides to the 3' end of the blunt-ended nucleic acid in a non-template-directed manner, wherein A is added in preference to G and G is added in preference to C or T, if necessary. (Item 34) The method according to any one of the preceding items, further comprising the step of ligating a tag to the repaired DNA fragment (on both ends, if necessary), wherein the tag optionally includes a barcode and / or an adapter. (Item 35) The method according to any one of the preceding items, further comprising the step of purifying the repaired DNA fragment. (Item 36) The method according to any one of the preceding items, further comprising the step of denaturing one or more enzymes used in step (b) and / or (c) after step (b) and / or (c). (Item 37) The method according to any one of the preceding items, further comprising the step of amplifying the repaired DNA fragment. (Item 38) The method according to the preceding item, wherein the repaired DNA fragment comprises one or more adapters (e.g., two adapters), and the step of amplifying the repaired DNA fragment uses one or more (e.g., two) amplified oligomers that anneal to the one or more adapters. (Item 39) The method according to any one of the preceding items, further comprising the step of enriching the repaired DNA fragment with respect to a target fragment, thereby obtaining an enriched DNA fragment, wherein the enrichment step is performed after an amplification step, if necessary. (Item 40) The method of the preceding item, wherein the fragment for the purpose comprises a gene locus that varies in a manner related to a disease or disorder, and optionally the disease or disorder is cancer. (Item 41) The method described in the preceding item, wherein the variation is one or more of a single-nucleotide variation, a copy number variation, a gene fusion, or an indel. (Item 42) The method according to the preceding item, wherein the aforementioned fragment of interest comprises one, two, three, or four fragments that represent a single-nucleotide variation associated with a disease or disorder; a copy number variation associated with a disease or disorder; a gene fusion associated with a disease or disorder; or an indel associated with a disease or disorder, and optionally the disease or disorder is cancer. (Item 43) The method according to any one of items 36 to 39, further comprising the step of amplifying the enriched DNA fragment. (Item 44) The method according to any one of the preceding items, further comprising the step of sequencing the repaired DNA fragment. (Item 45) The method described in the preceding item, wherein the sequencing step involves sequencing nucleotides that have formed a 3' overhang in a partially double-stranded DNA fragment. (Item 46) The method according to item 41 or 42, wherein the sequencing step is a high-throughput sequencing step. (Item 47) The method according to any one of items 44 to 46, wherein a plurality of repaired DNA fragments are generated, at least a portion of the repaired DNA fragments include a tag, and the sequencing step generates a plurality of sequence reads from the plurality of repaired DNA fragments. (Item 48) The method according to item 47, wherein the sequence read comprises a sequence of nucleotides and a tag sequence forming a 3' overhang in the partially double-stranded DNA fragment.

Claims

1. A method for partially repairing double-stranded DNA fragments, (a) A step of contacting the partially double-stranded DNA fragment with one or more primers of a primer population, wherein the partially double-stranded DNA fragment includes a 3' overhang, the primer population includes a random target hybridization sequence, and one or more primers of the primer population anneal to the 3' overhang; (b) a step of extending one or more of the primers of the primer population along the DNA fragment using DNA polymerase, thereby producing one or more extended primers annealed to the DNA fragment; and (c) A step of ligating the 3' end of one or more extended primers to the 5' end of the strand of the extended primer or the partially double-stranded DNA fragment, thereby obtaining a repaired DNA fragment, wherein the 3' end of one of the one or more extended primers is ligated to the 5' end of the strand of the partially double-stranded DNA fragment. Includes, The primers of the aforementioned primer group are a) It is single-stranded; b) A double-stranded structure having a 3' overhang, wherein the random target hybridizing sequence is located within the 3' overhang; or c) A hairpin having a 3' overhang, wherein the random target hybrid-forming sequence is located within the 3' overhang; The aforementioned DNA polymerase a) Lacking 3'→5' exonuclease activity; b) Lacking 5'→3' exonuclease activity; and / or c) A method lacking chain substitution activity.

2. The DNA polymerase is a Klenow fragment. The method according to claim 1.

3. The aforementioned partially double-stranded DNA fragment, a) Each end has a 3' overhang; b) Having a 3' overhang and (i) a blunt end or (ii) a 5' overhang; c) Derived from bodily fluid samples; d) It is a cf DNA fragment; e) of mammalian origin; f) of human origin; and / or g) A portion of the DNA fragment population in the composition, The method according to claim 1 or claim 2.

4. The method according to any one of claims 1 to 3, wherein the random target hybrid-forming sequence has a length of 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides.

5. The hairpin or double-stranded region of the primer is a) Includes an adapter; and / or b) Including tags, The method according to claim 4.

6. The method according to any one of claims 1 to 5, wherein at least steps (a) to (c) are performed in a single pipe.

7. The method according to any one of claims 1 to 6, wherein the double-stranded DNA fragment is present in the composition, and the component is not removed from the composition for at least steps (a) to (c).

8. The method according to any one of claims 1 to 7, wherein the repaired DNA fragment comprises one or two blunt ends, or the method further comprises the step of blunt-ending one or two ends of the repaired DNA fragment.

9. The method according to claim 8, further comprising the step of end-tailing the repaired DNA fragment using a polymerase that adds nucleotides to the 3' end of the blunt-ended nucleic acid in a non-template-directed manner.

10. The method according to claim 9, wherein A is given priority over G, and G is given priority over C or T when added.

11. a) further comprising the step of ligating a tag onto the repaired DNA fragment; b) further comprising the step of purifying the repaired DNA fragment; c) further comprising a step of denaturing one or more enzymes used in step (b) and / or (c) after step (b) and / or (c); and / or d) The method according to any one of claims 1 to 10, further comprising the step of amplifying the repaired DNA fragment.

12. The method according to claim 11, further comprising the step of ligating a tag to the repaired DNA fragment, wherein the tag includes a barcode and / or an adapter.

13. The method according to any one of claims 1 to 12, further comprising the step of enriching the repaired DNA fragment with respect to a target fragment, thereby obtaining an enriched DNA fragment.

14. The method according to claim 13, wherein the enrichment step is performed after the amplification step.

15. a) The fragment of interest includes a locus that varies in a manner related to a disease or disorder; and / or b) The fragment of interest comprises one, two, three, or four fragments that exhibit a single-nucleotide variation related to a disease or disorder; a copy number variation related to a disease or disorder; a gene fusion related to a disease or disorder; or an indel related to a disease or disorder. The method according to claim 13 or claim 14.

16. The method according to claim 15, wherein the fragment of interest comprises a locus that is altered in a manner associated with a disease or disorder, and the alteration is one or more of a single-nucleotide alteration, a copy number alteration, a gene fusion, or an indel.

17. The method according to any one of claims 13 to 16, further comprising the step of amplifying the enriched DNA fragment.

18. The process further includes sequencing the repaired DNA fragment. The method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Target sequence enrichment

    JP2017537657A

  • Nucleic acid sample preparation methods and compositions

    US20120172258A1

  • Methods of attaching adapters to sample nucleic acids

    WO2018191702A2