Methods for targeted single-molecule genetic and epigenetic sequencing
Patent Information
- Application Number
- PCT/US2024/061178
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-19
- Publication Date
- 2025-08-21
AI Technical Summary
Current single-base resolution methylation (SBRM) sequencing methods are inefficient, laborious, costly, and not compatible with cell-free DNA (cfDNA) due to the need for base conversion steps and are limited to high molecular weight genomic DNA.
The development of new library prep methodologies for PCR-free target capture and direct nanopore sequencing, which involve denaturing double-stranded nucleic acids, hybridizing probes to single-stranded DNA, enriching using binding pairs, ligating directional adapters, and extending with DNA polymerase for sequencing on a nanopore sequencer.
This approach provides improved information on DNA variations and modifications, including copy number and sequence variations in cfDNA, enabling more accurate disease diagnosis and treatment.
Smart Images

Figure US2024061178_21082025_PF_FP_ABST
Abstract
Description
METHODS FOR TARGETED SINGLE-MOLECULE GENETIC AND EPIGENETIC SEQUENCINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority of US Provisional Patent Application No. 63 / 613,917, filed December 22, 2023, which is incorporated by reference herein in its entirety for all purposes.SUMMARY
[0002] The present disclosure aims to meet the need for improved DNA analysis, such as analysis of cfDNA from tumor cells. In some embodiments, the methods herein may include steps that can provide information about DNA variations and modifications, including copy number, and sequence variations in cfDNA. Such methods comprising DNA analysis may provide improved information about the likelihood of a particular disease state of a subject. In some embodiments, improved detection of cancer markers in blood allows for more accurate detection of disorders (diagnosis) and therefore improved treatments.
[0003] Presently, single-base resolution methylation (SBRM) sequencing methods on clonal single-based substitution (SBS) systems require a base conversion step that is lossy, inefficient in modification calling, laborious / time-consuming, costly, and / or not easily compatible with simultaneous high sensitivity / specificity genomic base calling.
[0004] Other SBRM sequencing methods that do not require a base conversion step and are compatible at sensing modified bases in PCR-free libraries are limited to high molecular weight genomic DNA (gDNA) and are typically performed on a low target plexity. While adaptive sampling is possible with SBRM based methods, this procedure requires libraries of at least 400 bp, and thus are not compatible with cell-free DNA (cfDNA).
[0005] Hybrid capture-based enrichment methods are flexible in ability to target a broad number of regions, from single region to larger than an exome, and are compatible with many input types, including cfDNA. However, such methods are not directly compatible with existing PCR-free library preps.
[0006] The present disclosure aims to describe new library prep methodologies for PCR-free target capture and direct nanopore sequencing to provide valuable information of samples.
[0007] One of the provided embodiments includes a method for sequencing a nucleic acid composition in which double-stranded nucleic acids are denatured to single strands. Probes are then hybridized to the single stranded DNA, where in the probe is a first member of a binding pair so as to produce a probe annealed DNA. The probe annealed DNA is then enriched using a second member of the binding pair. The prob annealed DNA is then ligated with a directional adapter to produce an adapter ligated DNA. The adapter ligated DNA is subsequently contacted with a DNA polymerase and deoxynucleotide triphosphates so as to produce a double stranded primer extended sample. The primer extended sample is subsequently sequenced. The sequencing may be on a nanopore sequencer.
[0008] One of the provided embodiments includes a method for sequencing a nucleic acid composition in which double-stranded nucleic acids are denatured to single strands. Probes are then hybridized to the single stranded DNA, where in the probe is a first member of a binding pair so as to produce a probe annealed DNA, wherein the probe comprises a splint region that that is not complementary to the single stranded DNA, wherein the splint region comprises a hybridized oligomer. The probe annealed DNA is then enriched using a second member of the binding pair. The probe annealed DNA is then ligated with a directional adapter to produce an adapter ligated DNA. The adapter ligated DNA is subsequently contacted with a DNA polymerase and deoxynucleotide triphosphates so as to produce a double stranded primer extended sample. The primer extended sample is subsequently sequenced. The sequencing may be on a nanopore sequencer.
[0009] One of the provided embodiments includes a method for sequencing a nucleic acid composition in which double-stranded nucleic acids are denatured to single strands. Probes are then hybridized to the single stranded DNA, where in the probe is a first member of a binding pair so as to produce a probe annealed DNA. The probe annealed DNA is then enriched using a second member of the binding pair. The probe annealed DNA is then ligated with a directional adapter to produce an adapter ligatedDNA. The adapter ligated DNA is subsequently contacted with a DNA polymerase and deoxynucleotide triphosphates so as to produce an adapter extended sample. The adapter-extended sample is subsequently sequenced. The sequencing may be on a nanopore sequencer.
[0010] One of the provided embodiments includes a method for sequencing a nucleic acid composition in which double-stranded nucleic acids are denatured to single strands. A primer is hybridized to the single stranded DNA to produce primer anneal DNA. A directional adapter is then ligated to the primer anneal DNA so as to produce an adapter- ligated DNA. The adapter ligated DNA is then sequenced. The sequencing may be on a nanopore sequencer.
[0011] One of the provided embodiments includes a method for sequencing a nucleic acid composition in which double-stranded nucleic acids are denatured to single strands. A directional adapter is ligated to the single stranded DNA so it's produce adapter- like A to DNA. The adapter-ligated DNA is subsequently sequenced. The sequencing may be on a nanopore sequencer.
[0012] One of the provided embodiments includes a method for sequencing a nucleic acid composition in which double-stranded nucleic acids are denatured to single strands. A directional adapter is ligated to the single stranded DNA so as to produce adapter ligated to the DNA. The adapter-ligated DNA is subsequently sequenced. The sequencing may be on a nanopore sequencer.
[0013] One of the provided embodiments includes a method for sequencing a single stranded DNA composition. A poly A tail is added to the single stranded DNA. A poly T oligonucleotide is annealed to the poly A tail of single stranded DNA to produce a 3’ extended DNA. The 3’ extended DNA is contacted with an exonuclease. A directional adapter is like added to the 3’ extended DNA to produce an adapter- ligated DNA. The adapter- ligated DNA is subsequently sequenced. The sequencing may be on a nanopore sequencer.
[0014] One of the provided embodiments includes a method for sequencing a single stranded DNA composition. The single stranded DNA is extended with a polymerase to produce a 3’ extended DNA. A tail complementary oligonucleotide is hybridized to the 3’ extended DNA. A directional adapter is then annealed to the 3’extended DNA so as to produce an adapter- ligated DNA, where in the directional adapter comprises a region complementary to the 3’ extended DNA. Adapter- ligated DNA is then sequenced. The sequencing may be on a nanopore sequencer.
[0015] One of the provided embodiments includes a method for sequencing a nucleic acid composition in which a 3’ nanopore adapter is ligated to the double stranded DNA to produce a 3’ nanopore adapter ligated DNA. The 3’ nanopore adapter ligated DNA is denatured to produce single stranded nanopore adapter ligated DNA. A probe is then hybridized to the single stranded nanopore adapter ligated DNA, wherein the probe is one part of a binding pair, so as to produce a probe-annealed DNA. The probe-annealed DNA is then enriched using the second part of the binding pair. The probe-annealed DNA is subsequently contacted with a DNA polymerase and deoxyribonucleoside triphosphate to produce a double-stranded primer extended sample. The nanomotor protein is then loaded onto the primer- extended sample, then the primer- extended sample is then sequenced.
[0016] One of the provided embodiments includes a method for sequencing a nucleic acid composition. Double stranded-DNA is ligated to an LP adapter (library preparation adapter) to produce LP adapter ligated DNA. The LP adapter ligated DNA is denatured to produce a single stranded LP adapter ligated DNA. A probe is hybridized to the single stranded LP adapter ligated DNA, wherein the probe is the first member of a binding pair, to produce a probe-annealed DNA. The probe-annealed DNA is enriched using a second member of the binding pair. A universal primer is then hybridized to the LP adapter to produce primer adapted DNA. The primer adapted DNA is then contacted with a DNA polymerase and deoxyribonucleoside triphosphates to produce a double stranded primer-extended sample. The primer extended sample members are then ligated with a directional adapter to produce an adapter-ligated DNA. The adapter-ligated DNA is then sequenced. The sequencing may be on a nanopore sequencer.
[0017] One of the provided embodiments includes a method for sequencing a nucleic acid composition, wherein a double-stranded DNA is ligated with an LP adapter (library preparation adapter) to produce LP adapter ligated DNA. The LP adapter ligated DNA is denatured to produce a single stranded LP adapter ligated DNA. Aprobe is hybridized to the single stranded LP adapter ligated DNA, wherein the probe is a first member of a binding pair, to produce a probe-annealed DNA. The probe- annealed DNA is enriched using a second member of the binding pair. A universal primer is then hybdized to the LP adapter to produce primer adapted DNA. The primer adapted DNA is then ligated with a directional adapter to produce an adapter-ligated DNA. The adapter-ligated DNA is then sequenced. The sequencing may be on a nanopore sequencer.
[0018] One of the provided embodiments includes a method for sequencing a nucleic acid composition, wherein double-stranded DNA is ligated with an LP adapter, wherein the LP adapter comprises a 3’ overhang region, to produce LP adapter ligated DNA. The LP adapter ligated DNA is denatured to produce a single stranded LP adapter ligated DNA. A probe is hybridized to the single stranded LP adapter ligated DNA, wherein the probe is the first member of a binding pair, to produce a probe- annealed DNA. The probe-annealed DNA is then enriched using a second member of the binding pair. A universal primer is hybridized to the LP adapter to produce primer adapted DNA. The primer adapted DNA is ligated with a directional adapter to produce an adapter-ligated DNA. The adapter-ligated DNA is sequenced. The sequencing may be on a nanopore sequencer.DETAILED DESCRIPTION
[0019] The following are definitions of terms used in the present application.
[0020] As used herein, the singular terms “a,” “an,” and “the” include the plural reference unless the context clearly indicates otherwise.
[0021] The phrase “and / or,” as used herein, means “either or both” of the elements so conjoined, i.e. , elements that are conjunctively present in some cases and disjunctively present in other cases. Thus, as a non-limiting example, “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in some embodiments, to A only (optionally including elements other than B); in other embodiments, to B only (optionally including elements other than A); in yet other embodiments, to both A and B (optionally including other elements); etc.
[0022] As used herein, “at least one” means one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0023] As used herein, “about” when used in connection with doses, amounts, or ratios, include the value of a specified dose, amount, or ratio or a range of the dose, amount, or ratio that is recognized by one of ordinary skill in the art to provide a desired effect equivalent to that obtained from the specified dose, amount, or ratio. The term “about” may refer to an acceptable error for a particular value as determined by one of skill in the art, which depends in part on how the values is measured or determined.
[0024] “Cell-free DNA,” “cfDNA molecules,” or simply “cfDNA” include DNA molecules that naturally occur in a subject in extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). While the cfDNA previously existed in a cell or cells in a large complex biological organism, e.g., a mammal, it has undergone release from the cell(s) into a fluid found in the organism, and may be obtained from a sample of the fluid without the need to perform an in vitro cell lysis step. cfDNA molecules may occur as DNA fragments.
[0025] As used herein, “primer-annealed DNA,” when referring to primers that anneal to at least one target region, means DNA molecules annealed to primers, wherein the DNA molecules comprise at least one target region. In some embodiments, the target region of primer-annealed DNA is tiled with primers such that there are nogaps between the primers annealed to the target region or such that the gaps are small enough to prevent significant primer extension by a DNA polymerase, wherein the target region consists of a wild type sequence. In some embodiments, the target region of primer-annealed DNA has a gap between primers large enough to permit primer extension by a DNA polymerase, such as a 5’ to 3’ exonuclease negative, strand displacement negative DNA polymerase, wherein the target region comprises a rearrangement (e.g., rearrangement breakpoint).
[0026] In some embodiments, the DNA polymerase is a 5’ to 3’ exonuclease positive polymerase and does not have significant 3’ to 5’ exonuclease activity. A DNA polymerase that is “strand displacement positive” is able displace a strand, such as a primer or probe, annealed to DNA. In some embodiments, a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase is used for extension of a nucleic acid sequence in order to remove a primer or probe annealed to a region. In some embodiments, the use of a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase elutes a nucleic acid strand bound to a member of a binding pair. In some embodiments, the member of the binding pair is a streptavidin bead.
[0027] As used herein, a “primer-extended sample,” when referring to primers that anneal to at least one target region, means a nucleic acid strand formed by extension of a primer annealed to a DNA target region. In some embodiments, a primer- extended sample is a significant primer- extended sample or is formed by significant primer extension, meaning that the resulting nucleic acid strand has sufficient additional length (e.g., at least 10, 15, 20, 30, 40, 50, 60, 75, or 100 nucleotides in addition to the length of the original primer) to be detected and / or identified using methods described herein. In some embodiments, significant primer extension results in primer- extended samples comprising a capture moiety present at a low percentage in the deoxynucleoside triphosphate mixture. In some embodiments, primer extension that results in no primer-extended samples or only short primer-extended samples that do not comprise the capture moiety occurs on target regions comprising a completely tiled primer-annealed target region.
[0028] As used herein, a “blocking oligonucleotide” is an oligonucleotide having a sequence that is complementary to a portion of a DNA molecule that is outside of atarget region of the DNA molecule, or to a portion of an intron within or adjacent to an exonic target region, or to a portion of an exon adjacent to an intronic target region. Blocking oligonucleotides can be used in addition to primers, which are generally complementary to sequence within a target region. In some embodiments, blocking oligonucleotides help prevent significant primer extension of primers annealed to a wild type sequence within a DNA molecule. For example, blocking may occur after a short primer-extended product of up to 50 nucleotides (e.g., 1 , 2, 5, 10, 15, 20, 30, 40, or 50 nucleotides) in addition to the length of the primer has formed. In some embodiments, a DNA end, such as a 3’ DNA end or a 5’ DNA end, is said to be blocked by a blocking oligonucleotide. In some embodiments, blocking oligonucleotides help prevent ligation, e.g. ligation of an adapter. In some embodiments, blocking oligonucleotides help prevent degradation or exonucleolysis of DNA molecules originally isolated from a sample. In some embodiments, blocking oligonucleotides anneal to adapters ligated to DNA molecules. In some embodiments, blocking oligonucleotides anneal to a portion of a wild type exon adjacent to an intron in an intronic target region. In some embodiments, blocking oligonucleotides anneal to a portion of a wild type intron adjacent to an exon in an exonic target region.
[0029] As used herein, “adjacent” nucleosides or oligonucleotides are nucleosides or oligonucleotides that are next to each other, with no intervening nucleosides. For example, “adjacent” nucleosides may be covalently linked together within a nucleic acid or oligonucleotide, or they may be unlinked but are next to each other because they are annealed to or hybridized to adjacent linked nucleosides of a nucleic acid. “Adjacent” oligonucleotides may likewise be linked together or unlinked to each other but annealed to or hybridized to adjacent, linked portions of a nucleic acid.
[0030] A “target region” in the context of a nucleic acid refers to a genetic locus or genetic region comprising multiple loci pursued for capture, identification, and / or detection, for example, by using probes (e.g., through sequence complementarity). A “target region set” or “set of target regions” refers to a plurality of genomic loci targeted for identification and / or capture, for example, by using a set of probes (e.g., through sequence complementarity). In some embodiments, a target region is an intronic region. In some embodiments, a target region is an exonic region. In some embodiments, atarget region is an exon-exon junction region. In some embodiments, a target region comprises a rearrangement, such as a rearrangement associated with cancer.
[0031] An “intronic region,” of DNA herein encodes an intron or a portion thereof. Intronic regions include regions that encode introns removed during post- transcriptional splicing of pre- mRNA to generate mRNA and also “J intronic regions,” which are sequences that intervene between germ line J gene segments (e.g., in immunoglobulin or T cell receptor loci) and are removed during somatic V(D)J recombination. An “exonic region” of DNA herein encodes at least one exon or a portion thereof in pre-mRNA or mRNA. In some embodiments, an exonic region is a VDJ exonic region, meaning that it encodes one or more V, D, or J exons in pre- mRNA or mRNA. An “exon-exon junction region” of DNA herein encodes at least one exon- exon junction in pre-mRNA (e.g., immunoglobulin or T cell receptor pre-mRNA) before any introns are removed by splicing. Exon-exon junction regions can be formed by V(D)J recombination, e.g., when a J segment is joined to a D or V segment or a V segment is joined to a D or J segment.
[0032] As used herein, a DNA “structural variation” is a mutation comprising a DNA sequence not present in the wild-type genome other than a point mutation (e.g., in which at least 5, 10, 20, or 50 contiguous nucleotides are different relative to the wild type sequence at the corresponding locus). Examples of DNA structural variations include rearrangements, such as translocations, insertions, deletions, duplications, copy -number variants, and inversions. As used herein, a DNA “rearrangement” is a structural variation, wherein the DNA sequence comprises two adjacent sequence portions that are not adjacent to each other in the germline genomic DNA. In some embodiments, a rearrangement is a translocation, gene fusion, insertion, deletion, or inversion. Exemplary rearrangements include products of a translocation, gene fusion, and VDJ recombination. In some embodiments, the rearrangement is the product of a translocation comprising fusion of two intronic regions. In some embodiments, the rearrangement is a product of VDJ recombination comprising adjacent J exonic regions. A molecule comprising a structural variation may be referred to as a structural variant. In some embodiments, an insertion is an insertion of at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides. In some embodiments, a deletion affects sequencespanning the end of the target region, so as to result in a primer being unblocked and undergoing extension in a method described herein.
[0033] “Sequence-variable target regions” refer to target regions that may exhibit changes in sequence such as nucleotide substitutions (i.e. , single nucleotide variations), insertions, deletions, or gene fusions or transpositions in neoplastic cells (e.g., tumor cells and cancer cells) relative to normal cells. A sequence-variable target region set is a set of sequence-variable target regions. In some embodiments, the sequence-variable target regions are target regions that may exhibit changes that affect less than or equal to 50 contiguous nucleotides, e.g., less than or equal to 40, 30, 20, 10, 5, 4, 3, or 2 nucleotides, or that affect 1 nucleotide.
[0034] “Epigenetic target regions” refers to target regions that may show sequence-independent differences in different cell or tissue types (e.g., a target region having a different extent of methylation in a solid tissue type than in hematopoietic cells) or differences in neoplastic cells, such as tumor cells or cancer cells, relative to normal cells. In some embodiments, epigenetic target regions show sequence-independent differences in cfDNA originating from tissue types that ordinarily do not substantially contribute to cfDNA, such as lung, colon, etc., relative to background cfDNA, such as cfDNA that originated from hematopoietic cells. In some embodiments, epigenetic target regions show sequence-independent differences in cfDNA from subjects having cancer relative to cfDNA from healthy subjects. Examples of sequence- independent changes include, but are not limited to, changes in methylation (increases or decreases), nucleosome distribution, cfDNA fragmentation patterns, CCCTC-binding factor (“CTCF”) binding, transcription start sites, and regulatory protein binding regions. An “epigenetic target region set” is a set of epigenetic target regions. Epigenetic target region sets thus include, but are not limited to, hypermethylation variable target region sets, hypomethylation variable target region sets, and fragmentation variable target region sets, such as CTCF binding sites and transcription start sites. For present purposes, loci susceptible to neoplasia-, tumor-, or cancer-associated focal amplifications and / or gene fusions may be analyzed in the same manner as an epigenetic target region set because detection of a change in copy number by sequencing or a fused sequence that maps to more than one locus in a reference genome tends to be more similar todetection of exemplary epigenetic changes discussed above than detection of nucleotide substitutions, insertions, or deletions, e.g., in that the focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing because their detection does not depend on the accuracy of base calls at one or a few individual positions.
[0035] As used herein, an “epigenetic feature” refers to any feature of DNA or chromatin other than primary sequence (i.e., the sequence of A, C, G, and T bases). Epigenetic features include covalent modifications of bases, such as methylation, and modifications and positioning of histones and other stably DNA-associated proteins.
[0036] As used herein, a “differentially methylated region” refers to a region of DNA having a detectably different degree of methylation in at least one type of tissue relative to the degree of methylation in another type of tissue; or in a sample from a healthy subject relative to the degree of methylation in a subject having pre-cancer, cancer, or a neoplasm. In some embodiments, a differentially methylated region has a detectably higher degree of methylation in at least one type of tissue relative to the degree of methylation in cell-free DNA from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation in at least one type of tissue relative to the degree of methylation in cell-free DNA from a healthy subject. In some embodiments, differentially methylated regions are hypomethylated in the erythrocyte lineage or in an immature red blood cell (e g., reticulocyte) and hypermethylated in at least one non-erythrocyte cell or tissue type (e.g., a leukocyte or a solid tissue cell type, such as epithelial cells, muscle cells, etc.).
[0037] As used herein, “type-specific” in the context of an epigenetic variation means an epigenetic variation that is present at a detectably different degree in one cell or tissue type, or in a plurality of related cell or tissue types, relative to other cell or tissue types. Similarly, a “type- specific epigenetic target region” is an epigenetic target region that has a detectably different epigenetic characteristic in one cell or tissue type, or in a plurality of related cell or tissue types, relative to other cell or tissue types. Exemplary epigenetic characteristics are discussed in the definition of epigenetic target regions set forth above. For example, a “type-specific differentially methylated region” is a region of DNA that has a detectably different degree of methylation in one cell ortissue type, or in a plurality of related cell or tissue types, relative to other cell or tissue types. Examples of a type-specific differentially methylated region include tissue-specific differentially methylated regions, including those associated with copy-number gain in early cancer. In some embodiments, capturing, identification, and / or detection of typespecific differentially methylated regions facilitates identification of the cell or tissue type from which the DNA originated. The cell or tissue from which a type-specific differentially methylated region originated may be a wild type cell or tissue or a neoplastic cell or tissue. In another example, a “type-specific fragment” of DNA is a DNA fragment arising from a type- specific fragmentation pattern that is present at a detectably different degree in one cell or tissue type, or in a plurality of related cell or tissue types, relative to other cell or tissue types. In some embodiments, a type-specific fragment is only present in the specific cell or tissue type(s). In some embodiments, a type-specific fragment is present to a detectably greater extent in the specific cell or tissue type(s).
[0038] As used herein, a “blood sample” refers to a sample comprising whole blood or a component thereof (e.g., plasma, serum, huffy coat, plasma pellet).
[0039] As used herein, “partitioning” of nucleic acids, such as DNA molecules, means separating, fractionating, sorting, or enriching a sample or population of nucleic acids into a plurality of subsamples or subpopulations of nucleic acids based on one or more modifications or features that is in different proportions in each of the plurality of subsamples or subpopulations. Partitioning may include physically partitioning nucleic acid molecules based on the presence or absence of one or more methylated nucleobases. A sample or population may be partitioned into one or more partitioned subsamples or subpopulations based on a characteristic that is indicative of a genetic or epigenetic change or a disease state.
[0040] As used herein, “base pairing specificity” refers to the standard DNA base (A, C, G, or T) for which a given base most preferentially pairs. For example, unmodified cytosine and 5- methylcytosine have the same base pairing specificity (i.e. , specificity for G) whereas uracil and cytosine have different base pairing specificity because uracil has base pairing specificity for A while cytosine has base pairingspecificity for G. The ability of uracil to form a wobble pair with G is irrelevant because uracil nonetheless most preferentially pairs with A among the four standard DNA bases.
[0041] “Capturing” one or more target molecules, such as one or more nucleic acids comprising at least one target region refers to preferentially isolating or separating the one or more target molecules from non-target molecules.
[0042] As used herein, a “label” is a capture moiety, fluorophore, oligonucleotide, or other moiety that facilitates detection, separation, or isolation of that to which it is attached.
[0043] As used herein, a “capture moiety” is a molecule that allows affinity separation of molecules linked to the capture moiety from molecules lacking the capture moiety. Exemplary capture moieties include biotin, which allows affinity separation by binding to streptavidin linked or linkable to a solid phase or an oligonucleotide, which allows affinity separation through binding to a complementary oligonucleotide linked or linkable to a solid phase.
[0044] As used herein, a “target-specific probe” means a probe that specifically binds to a target region, such as an epigenetic target region or a sequence-variable target region. In some embodiments, target-specific probes comprise a capture moiety to facilitate capture of the target region to which it specifically binds.
[0045] As used herein, a “tag” is a molecule, such as a nucleic acid, label, fluorophore, or peptide, containing information that indicates a feature of the molecule to which the tag is associated. For example, molecules can bear a sample tag (which distinguishes molecules in one sample from those in a different sample), a molecular tag / molecular barcode / barcode (which distinguishes different molecules from one another (in both unique and non-unique tagging scenarios), a purification tag, and / or a detectable tag or label.
[0046] As used herein, a “target molecule” is a molecule, such as a protein, carbohydrate, nucleic acid, or lipid, that is targeted for capture, identification, and / or detection. In some embodiments, a target molecule is a nucleic acid comprising an epigenetic target region and / or a sequence-variable target region.
[0047] “Specifically binds” in the context of a primer, probe, or other oligonucleotide, a protein, or other binding molecule and a target sequence means thatunder appropriate hybridization conditions, the primer, oligonucleotide, or probe hybridizes to its target sequence, or replicates thereof, to form a stable hybrid, while at the same time formation of stable non-target hybrids is minimized. Thus, a primer or probe hybridizes to a target sequence or replicate thereof to a sufficiently greater extent than to a non-target sequence, to ultimately enable capture or detection of the target sequence. Appropriate hybridization conditions are well-known in the art, may be predicted based on sequence composition, or can be determined by using routine testing methods (see, e.g., Sambrook et ah, Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) at §§I .90-1.91 , 7.37-7.57, 9.47-9.51 and 11.47-11.57, particularly §§ 9.50-9.51 , 11.12-11.13,I I .45-11 .47 and 11 .55-11 .57, incorporated by reference herein).
[0048] A molecule is “produced by a tumor” if it originated from a tumor cell. cfDNA that originated from a tumor cell is “circulating tumor DNA” (“ctDNA”). Tumor cells are neoplastic cells that originated from a tumor, regardless of whether they remain in the tumor or become separated from the tumor (as in the cases, e.g., of metastatic cancer cells and circulating tumor cells).
[0049] The “capture yield” of a collection of probes for a given target region set refers to the amount (e.g., amount relative to another target region set or an absolute amount) of nucleic acid corresponding to the target region set that the collection of probes captures under typical conditions. Exemplary typical capture conditions are an incubation of the sample nucleic acid and probes at 65 °C for 10-18 hours in a small reaction volume (about 20 pL) containing stringent hybridization buffer. The capture yield may be expressed in absolute terms or, for a plurality of collections of probes, relative terms. When capture yields for a plurality of sets of target regions are compared, they are normalized for the footprint size of the target region set (e.g., on a per-kilobase basis). Thus, for example, if the footprint sizes of first and second target regions are 50 kb and 500 kb, respectively (giving a normalization factor of 0.1 ), then the DNA corresponding to the first target region set is captured with a higher yield than DNA corresponding to the second target region set when the mass per volume concentration of the captured DNA corresponding to the first target region set is more than 0.1 times the mass per volume concentration of the captured DNA correspondingto the second target region set. As a further example, using the same footprint sizes, if the captured DNA corresponding to the first target region set has a mass per volume concentration of 0.2 times the mass per volume concentration of the captured DNA corresponding to the second target region set, then the DNA corresponding to the first target region set was captured with a two-fold greater capture yield than the DNA corresponding to the second target region set.
[0050] The term “methylation” or “DNA methylation” refers to addition of a methyl group to a nucleobase in a nucleic acid molecule. In some embodiments, methylation refers to addition of a methyl group to a cytosine at a CpG site (cytosine- phosphate-guanine site (i.e., a cytosine followed by a guanine in a 5’ - 3’ direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to addition of a methyl group to adenine, such as in N6- methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation refers to addition of a methyl group to the 5C position of the cytosine to create 5-methylcytosine (5mC). In some embodiments, methylation comprises a derivative of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5- caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation comprises addition of a methyl group to the 3C position of the cytosine to generate 3-methylcytosine (3mC). Methylation can also occur at non CpG sites, for example, methylation can occur at a CpA, CpT, or CpC site. DNA methylation can change the activity of methylated DNA region. For example, when DNA in a promoter region is methylated, transcription of the gene may be repressed. DNA methylation is critical for normal development and abnormality in methylation may disrupt epigenetic regulation. The disruption, e.g., repression, in epigenetic regulation may cause diseases, such as cancer. Promoter methylation in DNA may be indicative of cancer.
[0051] The term “hypermethylation” refers to an increased level or degree of methylation of DNA relative to the other DNA molecules within a population (e.g., sample) of DNA molecules. In some embodiments, hypermethylated DNA can include DNA molecules comprising at least 1 methylated residue, at least 2 methylatedresidues, at least 3 methylated residues, at least 5 methylated residues, or at least 10 methylated residues. As used herein, “type-specific hypermethylation” means an increased level or degree of methylation of DNA in at one cell or tissue type, or in a plurality of related cell or tissue types, relative to other cell or tissue types. In some embodiments, capturing, identification, and / or detection of type-specific hypermethylated regions facilitates identification of the cell or tissue type from which the DNA originated. The cell or tissue from which a type-specific hypermethylated region originated may be a wild type cell or tissue or a neoplastic cell or tissue.
[0052] The term “hypomethylation” refers to a decreased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules comprising 0 methylated residues, at most 1 methylated residue, at most 2 methylated residues, at most 3 methylated residues, at most 4 methylated residues, or at most 5 methylated residues. As used herein, “type-specific hypomethylation” means a decreased level or degree of methylation of DNA in at one cell or tissue type, or in a plurality of related cell or tissue types, relative to other cell or tissue types. In some embodiments, capturing, identification, and / or detection of type-specific hypomethylated regions facilitates identification of the cell or tissue type from which the DNA originated. The cell or tissue from which a type-specific hypomethylated region originated may be a wild type cell or tissue or a neoplastic cell or tissue.
[0053] The terms “agent that recognizes a modified nucleobase in DNA,” such as an “agent that recognizes a modified cytosine in DNA” refers to a molecule or reagent that binds to or detects one or more modified nucleobases in DNA, such as methyl cytosine. A “modified nucleobase” is a nucleobase that comprises a difference in chemical structure from an unmodified nucleobase.
[0054] In the case of DNA, an unmodified nucleobase is adenine, cytosine, guanine, or thymine. In some embodiments, a modified nucleobase is a modified cytosine. In some embodiments, a modified nucleobase is a methylated nucleobase. In some embodiments, a modified cytosine is a methyl cytosine, e.g., a 5-methyl cytosine.In such embodiments, the cytosine modification is a methyl. Agents that recognize a methyl cytosine in DNA include but are not limited to “methyl binding reagents,” which refer herein to reagents that bind to a methyl cytosine. Methyl binding reagents include but are not limited to methyl binding domains (MBDs) and methyl binding proteins (MBPs) and antibodies specific for methyl cytosine. In some embodiments, such antibodies bind to 5-methyl cytosine in DNA. In some such embodiments, the DNA may be single-stranded or double-stranded.General Exemplary Methods
[0055] In some embodiments, the present disclosure provides method of preparing DNA samples for Oxford Nanopore (ONT) sequencing. ONT technology can provide both nucleotide sequence information and epigenetic information, such as methylation status.
[0056] The ONT sequencing method includes using a double stranded DNA adapter ligation kit. In some embodiments, the present methods include hybrid capture enrichment prior to sequencing. Hybrid capture enrichment results in single-stranded DNA. In some embodiments the present methods include steps before enrichment that will be used post enrichment to regenerate a double-stranded DNA substrate for ONT adapter ligation. Because the ONT adapter includes a motor protein that is used in the ONT sequencing process, Adapter ligation should occur typically prior to enrichment. If ligation of the adapter preceded enrichment, the enrichment process could denature the bound motor protein. Such adapters may be referred to, for the sake of convenience, LP adapters (Library preparation adapters).
[0057] Thus, in some embodiments, the present methods involve preparing a PCR- free, targeted library for ONT sequencing. In some embodiments, the methods described herein make the library preparation process more efficient, motor protein.
[0058] Thus, in some embodiments, the present methods involve preparing a PCR- free, targeted library for ONT sequencing. In some embodiments, the methods described herein make the library preparation process more efficient.Probe-Capture Method 1
[0059] Fig. 1 depicts methods according to some embodiments. As shown in Fig. 1 , an optional first step is to perform end-repair prior to denaturation to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0060] As shown in Fig. 1 , the next step after the optional end-repair steps includes denaturing the DNA to provide single strands. A probe that is attached to biotin is then hybridized to the target region. The probe includes a 5’ ligation block. An exemplary 5’ ligation block is a 5’ OH group.
[0061] As shown in Fig. 1 , the next step involves 3’ to 5’ exonuclease treatment to remove the portion of the target strand that is 3’ of the probe. That step is followed by A-tailing treatment. The A-tailing step is optional if blunt end ligation of the adapter is performed subsequently. With probes of ~20-40nt and Tm 40- ~55, two step I two tube exonuclease and A-tailing steps can be performed. 3’->5’ exonuclease treatment can be performed with mesophilic enzymes possessing 3’->5’ exonuclease and no 5’->3” exonuclease activity. Exonuclease III is an example. The reaction can be performed near 37C. DNA Pol I can also be used, and the 3’ end of the probe can be blocked to prevent extension if undesired (see Probe-Capture Method 4 in which it is desired). The reaction is then cleaned up to remove the exonuclease and a mesophilic DNA polymerase possessing A-tailing activity is added with ATP, and incubated at similarly moderate temperature (~ 37C).
[0062] The target region is then enriched with streptavidin-magnetic beads that bind to the biotin-probe / target region complex. In some embodiments, this streptavidin enrichment step can precede the exonuclease and A-tailing steps.
[0063] As shown in Fig. 1 , after enrichment, the ONT adapter ligation step is performed. The adapter includes a motor protein and T-tailed region that hybridizes to the poly-A tail on the target region.
[0064] After ligation, an optional cleanup step can be performed to remove enzymes, buffers, and adapter dimers that may have been formed. The adapter dimers may be removed using streptavidin magnetic beads (e.g., PHi29).
[0065] As shown in Fig. 1 , after the ligation step, a DNA extension step is performed with a strand displacing 5’ exonuclease polymerase to remove the biotinylated probe from the target region. The removed biotinylated probes can then be removed using streptavidin beads. The double-stranded target region can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.
[0066] Adapter dimers should be well-tolerated by workflow by adding a magnetic streptavidin bead based cleanup step after ligation that will remove adapter dimers.Probe-Capture Method 2
[0067] Fig. 2 depicts methods according to some embodiments. As shown in Fig. 2, an optional first step is to perform end-repair prior to denaturation to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0068] As shown in Fig. 2, the next step after the optional end-repair steps includes denaturing the DNA to provide single strands. A probe that is attached to biotin is then hybridized to the target region. As shown in Fig. 2, the probe includes a splint region that is not complementary the target DNA. In some embodiments, the noncomplementarity is accomplished using a low complexity sequence that is not likely to be in the target DNA. In some embodiments, the non-complementarity is accomplished by probe placement design. (One can design a non-complementary splint region if the sequence of the target DNA is known.) The splint region is complimentary to an overhang on the adapter as described below. The probe also includes a 5’ ligation block. An exemplary 5’ ligation block is a 5’ OH group.
[0069] As shown in Fig. 2, the next step involves 3’ to 5’ exonuclease treatment to remove the portion of the target strand that is 3’ of the probe.
[0070] The target region is then enriched with streptavidin-magnetic beads that bind to the biotin-probe / target region complex. In some embodiments, this streptavidin enrichment step can precede the exonuclease step.
[0071] As shown in Fig. 2, after enrichment, the ONT adapter ligation step is performed. The adapter includes a motor protein and an overhang that is complementary to the splint region of the probe.
[0072] After ligation, an optional cleanup step can be performed to remove enzymes, buffers, and adapter dimers that may have been formed. The adapter dimers may be removed using streptavidin magnetic beads (e.g., PHi29).
[0073] As shown in Fig. 2, after the ligation step, a DNA extension step is performed with a strand displacing 5’ exonuclease polymerase to remove the biotinylated probe from the target region. The removed biotinylated probes can then be removed using streptavidin beads. The double-stranded target region can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.
[0074] Adapter dimers should be well-tolerated by workflow by adding a magnetic streptavidin bead based cleanup step after ligation that will remove adapter dimers.
[0075] Probe-capture Method 2 can include the following variations: extend the bound primer on the target DNA strand and elevate the temperature before the enrichment to increase enrichment stringency (OTR). do not include a separate enrichment step, just use primer binding as selectively for On- targets -> only these will be prepped with nanopore adaptersthis can be done this with variant 1 ) extend the adapted ligated DNA to create a dsDNA molecule, which may be preferred for ONT seq. This can be done by extending from the adapter 3’ end (adjacent biotin primer) with strand displacement polymerase.this can be the elution step from streptavidin beads too.Probe-Capture Method 3a
[0076] Fig. 3 depicts methods according to some embodiments. As shown in Fig. 3, an optional first step is to perform end-repair prior to denaturation to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0077] As shown in Fig. 3, the next step after the optional end-repair steps includes denaturing the DNA to provide single strands. A probe that is attached to biotin is then hybridized to the target region. As shown in Fig. 3, the probe includes a doublestranded splint region that has a short strand hybridized to the 5’ end region of the strand attached to the biotin. The hybridized short region includes bases that can be digested with an enzyme. For example, the short region can include some or all RNA bases that can be digested by RnaseH2 or some uracils only for digestion by USER enzyme mix. (for example, is not complementary the target DNA. In some optional embodiments, the short strand is hybridized to the probe after the probe is hybridized to the target DNA, and before exonuclease treatment. In some instances, this allows the splint probe to have a lower Tm than when probe non-complementarity is accomplished using a low complexity sequence that is not likely to be in the target DNA. The splint region is complimentary to an overhang on the adapter as described below. The probe also includes a 5’ ligation block. An exemplary 5’ ligation block is a 5’ OH group.
[0078] As shown in Fig. 3, the next step involves 3’ to 5’ exonuclease treatment to remove the portion of the target strand that is 3’ of the probe.
[0079] As shown in Fig. 3, the next step involves digesting the short region hybridized to the probe. In some embodiments, this is accomplished using RNA bases in the short region that can be digested with an appropriate enzyme, for example, RNase2.
[0080] The target region is then enriched with streptavidin-magnetic beads that bind to the biotin-probe / target region complex. In some embodiments, this streptavidin enrichment step can precede the short region digestion step.
[0081] As shown in Fig. 3, after enrichment, the ONT adapter ligation step is performed. The adapter includes a motor protein and an overhang that is complementary to the splint region of the probe.
[0082] After ligation, an optional cleanup step can be performed to remove enzymes, buffers, and adapter dimers that may have been formed. The adapter dimers may be removed using streptavidin magnetic beads (e.g., PHi29).
[0083] As shown in Fig. 3, after the ligation step, a DNA extension step is performed with a strand displacing 5’ exonuclease polymerase to remove the biotinylated probe from the target region. The removed biotinylated probes can then be removed using streptavidin beads. The double-stranded target region can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.
[0084] Adapter dimers should be well-tolerated by workflow by adding a magnetic streptavidin bead based cleanup step after ligation that will remove adapter dimers.
[0085] Probe-capture Method 2 can include the following variations:1 ) extend the bound primer on the target DNA strand and elevate the temperature before the enrichment to increase enrichment stringency (OTR).2) do not include a separate enrichment step, just use primer binding as selectively for On-targetsonly these will be prepped with nanopore adapters - -> this can be done this with variant 1 )3) extend the adapted ligated DNA to create a dsDNA molecule, which may be preferred for ONT seq. This can be done by extending from the adapter 3’ end (adjacent biotin primer) with strand displacement polymerase.this can be the elution step from streptavidin beads too.Probe-Capture Method 3b
[0086] Fig. 4 depicts methods according to some embodiments. As shown in Fig. 4, an optional first step is to perform end-repair prior to denaturation to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencingyield. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0087] As shown in Fig. 4, the next step after the optional end-repair steps includes denaturing the DNA to provide single strands. A probe that is attached to biotin is then hybridized to the target region. As shown in Fig. 4, the probe includes a doublestranded splint region that has a short strand hybridized to the 5’ end region of the strand attached to the biotin. The splint region is complimentary to an overhang on the adapter as described below.
[0088] As shown in Fig. 4, the next step involves 3’ to 5’ exonuclease treatment to remove the portion of the target strand that is 3’ of the probe.
[0089] As shown in Fig. 4, the next step involves ligating the biotinylated splint probe to the target DNA.
[0090] The target region is then enriched with streptavidin-magnetic beads that bind to the biotin-probe / target region complex. As shown in Fig. 4, after binding to the streptavidin beads, an enrichment wash is performed and the target DNA strand ligated to the biotinylated splint probe is eluted from the streptavidin beads. In some alternative embodiments, this streptavidin enrichment step can precede the probe ligation step. In some such embodiments, the enrichment step is performed prior to the exonuclease and probe ligation steps to reduce generation of off-target molecules with the biotinylated splint probe ligated to them. In some embodiments, the ligation is performed after the enrichment washes, but before elution from the streptavidin beads. The biotinylated splint probe may be reannealed after enrichment.
[0091] As shown in Fig. 4, the biotinylated probe is then eluted from the target DNA strand.
[0092] As shown in Fig. 4, the ONT adapter ligation step is then performed. The adapter includes a motor protein and an overhang that is complementary to the short region of the splint probe that has been ligated to the target DNA strand.
[0093] After ligation, an optional cleanup step can be performed to remove enzymes, buffers, and adapter dimers that may have been formed. The adapter dimers may be removed using streptavidin magnetic beads (e.g., PHi29).
[0094] As shown in Fig. 4, after the ligation step, a DNA extension step is performed to create a double-stranded molecule for ONT sequencing. The doublestranded target region can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing. In some alternate embodiments, the extension step is not used, because ONT sequencing can be performed without a full second strand.
[0095] Adapter dimers should be well-tolerated by workflow by adding a magnetic streptavidin bead based cleanup step after ligation that will remove adapter dimers.Probe-Capture Method 4
[0096] Fig. 5 depicts methods according to some embodiments. Probe-capture Method 4 can be applied to any of Probe-capture Methods 1 , 2, and 3a. The probe is extended to provide double-stranded DNA prior to enrichment. This example shows the method applied to Probe-capture Method 1 .
[0097] As shown in Fig. 5, an optional first step is to perform end-repair prior to denaturation to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0098] As shown in Fig. 5, the next step after the optional end-repair steps includes denaturing the DNA to provide single strands. A probe that is attached to biotin is then hybridized to the target region. The probe includes a 5’ ligation block. An exemplary 5’ ligation block is a 5’ OH group.
[0099] As shown in Fig. 5, the next step involves 3’ to 5’ exonuclease treatment to remove the portion of the target strand that is 3’ of the probe.
[0100] As shown in Fig. 5, after the exonuclease step, the probe is extended to provide double-stranded DNA including the probe in the new strand.
[0101] As shown in Fig. 5, the target strand is subjected to A-tailing treatment. The A-tailing step is optional if blunt end ligation of the adapter is performed subsequently.
[0102] In some embodiments, the three enzymatic reactions (exonuclease treatment, probe extension, and A-tailing) can be performed in a single tube using different combinations of enzymes at two reaction temperatures as illustrated by the following two exemplary procedures:
[0103] (1 ) exonuclease III at 37C to degrade 3’end of target DNA bound to the probe, and (2) Taq DNA polymerase at ~ 70+C to extend the probe and A-tail.
[0104] DNA Pol I, T4 DNA pol or other mesophilic enzymes with 3’>5’ exo activity (and no 5’->3” exo) at 37C to degrade 3’end of target DNA bound to probe and extend the probe, and (2) Taq DNA polymerase or other thermophilic polymerase at 70+C to extend the primer and A-tail.
[0105] If A-tailing is not performed (e.g. blunt ligation), then only a mesophilic enzymes with 3’>5’ exo activity (and no 5’->3” exo) at 37C can be used (e.g. DNA Pol I, T4 DNA Pol).
[0106] The double-stranded DNA is then enriched with streptavidin-magnetic beads that bind to the biotin-DNA complex. In some embodiments, the temperature is increased for this step to improve enrichment stringency (for example, a 65-75C hybridization temperature may be used).
[0107] As shown in Fig. 5, after enrichment, the ONT adapter ligation step is performed. The adapter includes a motor protein and T-tailed region that hybridizes to the poly-A tail on the target region.
[0108] After ligation, an optional cleanup step can be performed to remove enzymes, buffers, and adapter dimers that may have been formed. The adapter dimers may be removed using streptavidin magnetic beads (e.g., PHi29).
[0109] As shown in Fig. 5, after the ligation step, a DNA extension step is performed with a strand displacing 5’ exonuclease polymerase to remove the biotinylated probe from the target region. The removed biotinylated probes can then be removed using streptavidin beads. The double-stranded target region can then be subjected to ONT sequencing. This sequencing step can include an optional stepknown to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.
[0110] Adapter dimers should be well-tolerated by workflow by adding a magnetic streptavidin bead based cleanup step after ligation that will remove adapter dimers.
[0111] In some embodiments, the A-tailing step or an additional A-tailing step may be desired after hybridization capture if there is significant exonuclease activity in enrichment (contamination).
[0012] Probe-capture Method 4 can include the following variations:1 ) extend the bound primer on the target DNA strand and elevate the temperature before the enrichment to increase enrichment stringency (OTR).2) do not include a separate enrichment step, just use primer binding as selectively for On-targetsonly these will be prepped with nanopore adapters - this can be done this with variant 1 )3) extend the adapted ligated DNA to create a dsDNA molecule, which may be preferred for ONT seq. This can be done by extending from the adapter 3’ end (adjacent biotin primer) with strand displacement polymerase.this can be the elution step from streptavidin beads too.Probe-Capture Method 5
[0113] Fig. 6 depicts methods according to some embodiments. Probe-capture Method 5 can be applied to any of Probe-capture Methods 1 , 2, 3a, and 4. The probe does not include biotin. This example shows the method applied to Probe-capture Method 1 .
[0114] As shown in Fig. 6, an optional first step is to perform end-repair prior to denaturation to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0115] As shown in Fig. 6, the next step after the optional end-repair steps includes denaturing the DNA to provide single strands. A probe is then hybridized to thetarget region. In some embodiments, tile probes are used (overlapping and nonoverlapping).
[0116] As shown in Fig. 6, the next step involves 3’ to 5’ exonuclease treatment to remove the portion of the target strand that is 3’ of the probe.
[0117] As shown in Fig. 6, the target strand is subjected to A-tailing treatment. The A-tailing step is optional if blunt end ligation of the adapter is performed subsequently.
[0118] As shown in Fig. 6, the target region ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.
[0119] In some embodiments, an optional DNA Polymerase-mediated second strand synthesis may be performed in 3’->5’ exonuclease treatment step or separate from the exonuclease treatment step for higher probability of completion. In some embodiments, all three reactions can be performed in a single-tube using commercially available end-repair / dA tail reactions (such as NEBNext’s lllra II module ER / dA-tail module). These typically involve a 20 C incubation in which a mesothillic DNA polymerase with 3’->5’ exonuclease (and 5’->3’ DNA polymerase) capability is active, then incubation at 65C+ that inactivates the first mesophilic enzyme(s) and activated the A-tailing of a Taq DNA polymerase.Post-Enrichment LP Method 1.1 - 3’ Adapter Random Overhang T4DNA Ligation
[0120] Fig. 7 depicts methods according to some embodiments. As shown in Fig. 7, oligonucleotide-based hybrid capture is performed, which results in elution of single-stranded DNA molecules from regions of interest.
[0121] As shown in Fig. 7, an optional first step is to perform end-repair prior to hybrid capture to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. Performing this optional first step, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0122] As shown in Fig. 7, the denatured single-stranded DNA is stabilized by adding single-stranded binding protein to limit secondary structure formation. In some embodiments, one or both strands of the target are hybridized to tile probes (overlapping and non-overlapping).
[0123] As shown in Fig. 7, the ONT adapter ligation step is then performed. The adapter includes a motor protein and a random 3’ overhang sequence. In some embodiments, an optional phosphatase treatment is performed to dephosphorylate 3’ ends of the target DNA molecules. In some embodiments, this step is performed if the end repair step was not performed.
[0124] In some embodiments, to limit adapter dimer species formation, the adapter ends or at least the 3’ ends are modified to block undesired ligation at these junctions (for example, dideoxy base or C3 spacer). If phosphatase treatment is not performed at the ligation step, 3’ P (and 5’ OH) modifications are suitable for blocking undesired ligation. In some embodiments, the random overhang length is greater than or equal to 9 bases (Tm = 30), and between 9 to 15 bases for the adapter to anneal to the DNA target ends at room temperature during ligation.
[0125] As shown in Fig. 7, The target region ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.Post-Enrichment ssDNA LP Method 1.2 - 3’ Adapter Random Overhang T4DNA Ligation
[0126] Fig. 8 depicts methods according to some embodiments. As shown in Fig. 8, oligonucleotide-based hybrid capture is performed, which results in elution of single-stranded DNA molecules from regions of interest.
[0127] As shown in Fig. 8, an optional first step is to perform end-repair prior to hybrid capture to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. Performing this optional first step, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0128] As shown in Fig. 8, the denatured single-stranded DNA is stabilized by adding single-stranded binding protein to limit secondary structure formation. In some embodiments, one or both strands of the target are hybridized to tile probes (overlapping and non-overlapping).
[0129] As shown in Fig. 8, the ONT adapter ligation step is then performed. The adapter includes a motor protein and a random 3’ overhang sequence. In some embodiments, an optional phosphatase treatment is performed to dephosphorylate 3’ ends of the target DNA molecules. In some embodiments, this step is performed if the end repair step was not performed.
[0130] As shown in Fig. 8, after the adapter ligation, the 3’ ends are de-blocked (for example, with PNK treatment) and extension by DNA polymerase is performed (for example, with 3’ exo minus polymerase).
[0131] In some embodiments, to limit adapter dimer species formation, the adapter ends or at least the 3’ ends are modified to block undesired ligation at these junctions (for example, dideoxy base or C3 spacer). If phosphatase treatment is not performed at the ligation step, 3’ P (and 5’ OH) modifications are suitable for blocking undesired ligation. In some embodiments, the random overhang length is greater than or equal to 9 bases (Tm = 30), and between 9 to 15 bases for the adapter to anneal to the DNA target ends at room temperature during ligation.
[0132] As shown in Fig. 8, after extension, the double-stranded DNA can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.Post-Enrichment ssDNA LP Method 2.1 : TdT-Mediated Ligation and Library Preparation — Sequential Anneal and Adapter Ligation Steps
[0133] Fig. 9 depicts methods according to some embodiments. As shown in Fig. 9, oligonucleotide-based hybrid capture is performed, which results in elution of single-stranded DNA molecules from regions of interest.
[0134] As shown in Fig. 9, an optional first step is to perform end-repair prior to hybrid capture to fix any internal nicks in the DNA. In some instances, internal nicks candecrease sequencing yield. Performing this optional first step, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0135] As shown in Fig. 9, the next step after the optional end-repair steps includes performing terminal transferase non-templated 3’ extension of homopolymers. In some embodiments, tail length distribution can be controlled by altering the ratio of 3’ DNA ends and dNTPs. In the embodiments illustrated in Fig. 9, the extension is a poly- A sequence, denaturing the DNA to provide single strands. A probe is then hybridized to the target region.
[0136] As shown in Fig. 9, the next step includes anealing an oligonucleotide probe that is complementary to the extended section. The probe length should be shorter than the expected length of the extended tail generated in the last step. According to some embodiments, as depicted in Fig. 9, the probe has a tail sequence that preferentially binds and extends from tail bases adjacent to the 3’ end of the target DNA strand. Fig. 9 shows an exemplary probe that is a polyT VN primer (probe).
[0137] As shown in Fig. 9, the single-stranded section of the sequence generated the terminal transferase step. An exemplary exonuclease is Exonuclease III.
[0138] As shown in Fig. 9, the target strand is subjected to A-tailing treatment. The A-tailing step is optional if blunt end ligation of the adapter is performed subsequently.
[0139] As shown in Fig. 9, the ONT adapter ligation step is then performed. The adapter includes a motor protein and T-tailed region that hybridizes to the poly-A tail on the target region.
[0140] After ligation, an optional cleanup step can be performed to remove enzymes, buffers, and adapter dimers that may have been formed. The adapter dimers may be removed using streptavidin magnetic beads (e.g., SPRI beads).
[0141] As shown in Fig. 9, the target region ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.Post-Enrichment ssDNA LP Method 2.2: TdT-Mediated Ligation and Library Preparation — Single Adapter Ligation
[0142] Fig. 10 depicts methods according to some embodiments. As shown in Fig. 10, oligonucleotide-based hybrid capture is performed, which results in elution of single-stranded DNA molecules from regions of interest.
[0143] As shown in Fig. 10, an optional first step is to perform end-repair prior to hybrid capture to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. Performing this optional first step, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy. The DNA is denatured (for example, by heating to 95C).
[0144] As shown in Fig. 10, after the hybrid capture step, or the optional endrepair step, a single-tube process is performed. The single-tube includes Tdt polymerase, dATP, and reaction buffer. The single-tube further includes a nanoporecompatible adapter with the following features: 1 ) an attached motor protein; 2) a polyT 3’ end overhang tail that is complementary to Tdt-extended target DNA 3’ end; 3) 3’ end(s) phosphate group or other modification to block extension by TdT; and 4) optionally, an attached tether protein.
[0145] As shown in Fig. 10, the single-tube process results in extending the target region with the TdT polymerase. In some embodiments, a different templateindependent nucleic acid polymerase is used. The extension step uses a precisely defined dNTP pool (for example, dATP only creating a polyA tail) to generate a low complexity 3’ extension on the target DNA molecule (for example, a polyA tail).
[0146] As shown in Fig. 10, the single-tube process also results in DNA hybridization of the nanopore-compatible adapter through the 3’ tail (for example, polyT), which is complementary to the target region extended 3’ tail (for example, polyA). This results in attenuation / arrest of the TdT 3’ extension of the target region. The single-tube process also results in ligation of the extended 3’ end tail of the target region and the 5’ top strand of the nanopore-compatible adapter.
[0147] As shown in Fig. 10, the target region ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves an optional ONT “tether” attached to the adapter prior to sequencing.Post-Hybrid Capture Motor Protein Loading Method
[0148] Fig. 11 depicts methods according to some embodiments. As shown in Fig. 11 , end-repair is performed to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. Performing this optional first step, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0149] As shown in Fig. 11 , both strands of the target DNA are subjected to A- tailing treatment.
[0150] As shown in Fig. 11 , a nanopore adapter is ligated to 3’ end of both strands of the target DNA. The nanopore adapter includes a 3’ dideoxy polyT overhang, but no motor protein attached. The 5’ end of the target DNA strands are blocked by either (1 ) the base modification on the adapter (for example, the 3’ dideoxy polyT overhang) or (2) using phosphatase in the prior end-repair step to remove the 5’ phosphates on the target strands.
[0151] As shown in Fig. 11 , the target DNA is denatured and enriched using biotinylated probes that hybridize to the target DNA strands and streptavidin beads that bind to the bound biotinylated probes. In some embodiments, the probes are greater to or equal to 80 bases for high specificity enrichment.
[0152] As shown in Fig. 11 , the target DNA strands are eluted from the streptavidin beads. As shown in Fig. 11 , a primer that is complementary to a section of the nanopore adapter sequence is used with DNA pol-mediated extension to generate double-stranded DNA. These steps can be performed simultaneously or sequentially by either (1 ) annealing the primers to the target DNA strands ligated to the adapters while bound to the streptavidin beads and using strand-displacing 5’ to 3’ exonuclease polymerase to generate double-stranded DNA free in the solution, or (2) eluting thesingle-stranded DNA from the streptavidin beads using denaturing conditions, and then annealing and extending the primer using 5’ to 3’ exonuclease polymerase, such as DNA Pol.
[0153] As shown in Fig. 11 , the motor protein is then loaded onto the target DNA molecules ligated to the nanopore adapter.
[0154] As shown in Fig. 11 , the target DNA ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves an optional ONT “tether” attached to the adapter prior to sequencing.Enzymatic Removal of Residual Adapters
[0155] Fig. 12 depicts methods according to some embodiments. This method can be applied to any of prior methods. This example shows the method applied to Post-Enrichment ssDNA LP method 1.1 . As shown in Fig. 12, oligonucleotide-based hybrid capture is performed, which results in elution of single-stranded DNA molecules from regions of interest.
[0156] As shown in Fig. 12, an optional first step is to perform end-repair prior to hybrid capture to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. Performing this optional first step, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the optional end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0157] As shown in Fig. 12, if end-repair is not performed prior to denaturation, after denaturation the target DNA is subjected to phosphotase treatment to remove phosphate groups from the 5’ ends of the target DNA (and from the 3’ end if any phosphates are on 3’ ends).
[0158] As shown in Fig. 12, the denatured single-stranded DNA is stabilized by adding single-stranded binding protein to limit secondary structure formation. In some embodiments, one or both strands of the target are hybridized to tile probes (overlapping and non-overlapping).
[0159] As shown in Fig. 12, the ONT adapter ligation step is then performed. The adapter includes a motor protein and a random 3’ overhang sequence. In some embodiments, to limit adapter dimer species formation, the adapter ends or at least the 3’ ends are modified to block undesired ligation at these junctions (for example, dideoxy base or C3 spacer). In some embodiments, the random overhang length is greater than or equal to 9 bases (Tm = 30), and between 9 to 15 bases for the adapter to anneal to the DNA target ends at room temperature during ligation. As shown in Fig. 12, the adapter includes a 5’ OH group on the Y end to prevent degradation by lambda exonuclease.
[0160] As shown in Fig. 12, after the ligation step, lambda exonuclease treatment is performed. The 5’ to 3’ exonuclease activity will specifically degrade residual adapters with free 5’ P groups.
[0161] As shown in Fig. 12, The target region ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves attaching ONT’s “tether” by annealing to the adapter oligo prior to sequencing.Two-Ligation Step Method V1 - Extend Second Strand Prior to Adapter Ligation
[0162] Fig. 13 depicts methods according to some embodiments. As shown in Fig. 13, end-repair is performed to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. Performing this optional first step with phosphotase, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0163] As shown in Fig. 13, both strands of the target DNA are subjected to A- tailing treatment.
[0164] As shown in Fig. 13, an LP adapter is ligated to both strands of the target DNA using T4DNA ligase. The LP adapter includes a sample index, flanking constant regions for universal priming, and an internal base modification that terminates DNA polymerase synthesis (for example, a uracil base). In some embodiments, the adapteralso includes a 3’ end that will form and overhang for the subsequent adapter ligation step (for example, Fig. 13 shows an LP adapter with a 3’ T tail that hybridizes to the polyA tail on the target DNA strands.
[0165] As shown in Fig. 13, the target DNA is denatured and enriched using biotinylated probes that hybridize to the target DNA strands and streptavidin beads that bind to the bound biotinylated probes. In some embodiments, the probes are greater to or equal to 80 bases for high specificity enrichment.
[0166] As shown in Fig. 13, the target DNA strands are eluted from the streptavidin beads. As shown in Fig. 13, DNA Pol-mediated extension is then performed from the universal priming region of the ligated LP adapters to generate doublestranded DNA. In some embodiments, as depicted in Fig. 13, the priming occurs from the Poly A tail. These steps can be performed simultaneously or sequentially by either (1 ) using strand-displacing 5’ to 3’ exonuclease polymerase while the probes are hybridized to the target strands and are bound to the streptavidin beads to generate double-stranded DNA free in the solution, or (2) eluting the single-stranded DNA from the streptavidin beads using denaturing conditions, and then annealing and extending the primer using 5’ to 3’ exonuclease polymerase, such as DNA Pol. As shown in Fig. 13, the DNA Pol is intolerant to synthesize across the modified base in the 5’ LP- adapter sequence (for example, uracil intolerance), exonuclease 3’ to 5’ deficient, and / or generate non-templated single 3’ base overhang (A). Certain exemplary polymerases intolerant to uracil modification include Vent exo- and Deep Vent exo-. In some embodiments, if A-tailing polymerase is used, a 3’ overhang on the LP adapter is not needed, because it will be added during the extension step. In some embodiments, the LP adapter 3’ overhang can have high efficiency for generating library molecules with 3’ overhang.
[0167] As shown in Fig. 13, the nanopore adapter is then ligated to both strands of the target DNA using T4 DNA ligase. As shown in Fig. 13, the nanopore adapter has a T-tailed y strand and a bound motor protein.
[0168] As shown in Fig. 13, there can then be an optional cleanup step to remove enzymes, buffers, and adapter dimers using streptavidin beads (for example, SPRI beads).
[0169] As shown in Fig. 13, the target DNA ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves an optional ONT “tether” attached to the adapter prior to sequencing.Two-Ligation Step Method V2 - Adapter Ligation to Partially Double- Stranded DNA Library
[0170] Fig. 14 depicts methods according to some embodiments. As shown in Fig. 13, end-repair is performed to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. Performing this optional first step, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0171] As shown in Fig. 1 , both strands of the target DNA are subjected to A- tailing treatment.
[0172] As shown in Fig. 14, an LP adapter is ligated to both strands of the target DNA using T4DNA ligase. The LP adapter includes a sample index and flanking constant regions for universal hybridization. In some embodiments, the ligated strand of the adapter contains a sequence of bases at the 3’ end that will form an overhang for the subsequent nanopore adapter ligation step (for example, A).
[0173] As shown in Fig. 14, the target DNA is denatured and enriched using biotinylated probes that hybridize to the target DNA strands and streptavidin beads that bind to the bound biotinylated probes. In some embodiments, the probes are greater to or equal to 80 bases for high specificity enrichment.
[0174] As shown in Fig. 14, the target DNA strands are eluted from the streptavidin beads and hybridized to a universal hybridizing sequence creating a double-stranded DNA end with a 3’ A tail. In some embodiments, instead of using an LP adapter with an A tail, the universal sequences on the LP adapters can be designed to generate a 3’ sticky end (>1 base) rather than a single A tail and the nanopore adapter (discussed below) has a complementary sticky end to hybridize to LP adapter tail. An example of such an embodiment uses the sticky end sequence and nanopore adaptercommercially available from ONT in their native barcoding kits (https: / / store.nanoporetech.com / native-barcoding-kit-24.html, https: / / store.nanoporetech.com / native-barcoding-expansion-96.html) and utilized in recent publication(https: / / www.biorxiv.Org / content / 10.1101 / 2022.06.22.497080v1 full.pdf)
[0175] As shown in Fig. 14, the nanopore adapter is then ligated to both strands of the target DNA using T4 DNA ligase. As shown in Fig. 14, the nanopore adapter has a T-tailed y strand and a bound motor protein.
[0176] As shown in Fig. 14, there can then be an optional cleanup step to remove enzymes, buffers, and adapter dimers using streptavidin beads (for example, SPRI beads).
[0177] As shown in Fig. 14, the target DNA ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves an optional ONT “tether” attached to the adapter prior to sequencing.
[0178] As shown in Fig. 14, in an alternate method, after ligation of the nanopore adapter, a universal primer is used to extend from the universal hybridizing sequence added with the LP adapter to create fully double stranded DNA molecules. For A-tail overhang, extending after ligation ensure that library will remain directional (adapter only attached to 3’ end of original DNA strand).Two-Ligation Step Method V3 - Overhang / Sticky End for High Efficiency Adapter Ligation
[0179] Fig. 15 depicts methods according to some embodiments. As shown in Fig. 15, end-repair is performed to fix any internal nicks in the DNA. In some instances, internal nicks can decrease sequencing yield. In some instances, internal nicks can decrease sequencing yield. Performing this first step with phosphotase, also generates a 3’ OH on the strands for downstream ligation. In some embodiments, at least one modified dNTP is used during the end-repair step to mark the added nucleotide to improve methylation / somatic calling accuracy.
[0180] As shown in Fig. 15, both strands of the target DNA are subjected to A- tailing treatment.
[0181] As shown in Fig. 15, a phosphatase step is then performed to remove the 5’ P from DNA molecules to prevent 5’ end ligation.
[0182] As shown in Fig. 15, an LP adapter is ligated to both strands of the target DNA using T4DNA ligase. The LP adapter includes a sample index, flanking constant regions for universal priming, and a 3’ overhang tail on a single stranded end (for example, a T tail that hybridizes to the A tail added to the target DNA).
[0183] As shown in Fig. 15, the target DNA is denatured and enriched using biotinylated probes that hybridize to the target DNA strands and streptavidin beads that bind to the bound biotinylated probes. In some embodiments, the probes are greater to or equal to 80 bases for high specificity enrichment.
[0184] As shown in Fig. 15, the target DNA strands are eluted from the streptavidin beads and extended with a universal primer that hybridizes to the constant region added to the target DNA with the LP adapter to generate double-stranded DNA. In some embodiments, as depicted in Fig. 15, DNA pol (3’ to 5’ exo) is used to create double-stranded DNA strands with 3’ sticky ends. In some embodiments, the extension step can be carried out while the probes are hybridized to the target strands and are bound to the streptavidin beads using a strand displacing polymerase to elute the target DNA from the streptavidin beads. In some embodiments, the extension step is not used. In some such embodiments, the partially double-stranded DNA / single-stranded DNA targets can be loaded on a sequencer.
[0185] As shown in Fig. 15, the nanopore adapter is then ligated to one strand of the target DNA using T4 DNA ligase. As shown in Fig. 13, the nanopore adapter has a bound motor protein and a sticky end y adapter that hybridizes to the sticky end attached to the target DNA strand.
[0186] As shown in Fig. 15, there can then be an optional cleanup step to remove enzymes, buffers, and adapter dimers using streptavidin beads (for example, SPRI beads).
[0187] As shown in Fig. 15, the target DNA ligated to the adapter can then be subjected to ONT sequencing. This sequencing step can include an optional step known to be used with ONT sequencing that involves an optional ONT “tether” attached to the adapter prior to sequencing. In some embodiments, commercially available ONTadapters from a ‘native barcoding’ kit are used: https: / / store.nanoporetech.com / native- barcoding-expansion-96.html) and utilized in recent publication(https: / / www.biorxiv.Org / content / 10.1101 / 2022.06.22.497080v1 .full.pdf).GENERAL DESCRIPTION OF EXEMPLARY MATERIALS AND METHODSExemplary Sources of DNA
[0188] In some embodiments, the DNA (e.g., cfDNA) or nucleic acids from a subject are obtained and / or derived from a sample obtained from a subject having a cancer or a precancer. In some embodiments, the sample is obtained from a subject suspected of having a cancer or a precancer. In some embodiments, the sample is obtained from a subject having a tumor. In some embodiments, the sample is obtained from a subject suspected of having a tumor. In some embodiments, the sample is obtained from a subject having neoplasia. In some embodiments, the sample is obtained from a subject suspected of having neoplasia. In some embodiments, the sample is obtained from a subject in remission from a tumor, cancer, or neoplasia (e.g., following chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the pre-cancer, cancer, tumor, or neoplasia or suspected pre-cancer, cancer, tumor, or neoplasia may be of the bladder, head and neck, lung, colon, rectum, kidney, breast, prostate, skin, or liver. In some embodiments, the pre- cancer, cancer, tumor, or neoplasia or suspected pre-cancer, cancer, tumor, or neoplasia is of the lung. In some embodiments, the pre-cancer, cancer, tumor, or neoplasia or suspected pre-cancer, cancer, tumor, or neoplasia is of the colon or rectum. In some embodiments, the pre-cancer, cancer, tumor, or neoplasia or suspected pre cancer, cancer, tumor, or neoplasia is of the breast. In some embodiments, the pre-cancer, cancer, tumor, or neoplasia or suspected pre-cancer, cancer, tumor, or neoplasia is of the prostate. In any of the foregoing embodiments, the subject may be a human subject. In some embodiments, the sample is obtained from a subject having a stage I cancer, stage II cancer, stage III cancer or stage IV cancer.Methods of analyzing DNA
[0189] Current methods for analyzing DNA typically require amplification steps that filter out epigenetic information or require high concentrations of DNA, thus limiting the type of genomic samples and data. In some embodiments, the methods described herein provide a new method for analyzing DNA with improved sensitivity and may be applied to samples with low concentrations of DNA, e.g. cell free DNA (cfDNA).Target-specific probes
[0190] In some embodiments, methods of analyzing DNA herein comprise contacting the DNA with a plurality of target-specific probes specific for members of an epigenetic target region set comprising target regions that have a type-specific epigenetic variation and a copy number variation. In some embodiments, the epigenetic target region set comprises target regions that have a type-specific epigenetic variation and a copy number variation. In some such embodiments, the plurality of target specific probes comprises probes specific for members of the epigenetic target region set. In some embodiments, the target regions comprise type-specific differentially methylated regions and copy number variants. In some embodiments, the target regions comprise type-specific fragments arising from type-specific fragmentation patterns and copy number variants. In some embodiments, the type-specific epigenetic variation is present in greater proportion in wild type genomes of one or a plurality of cell or tissue types relative to wild type genomes of other cell or tissue types. In some embodiments, the DNA is from a blood sample obtained from a subject. In some embodiments, the methods herein detect abnormal levels of type-specific DNA, such as cfDNA, in a sample. For example, detection of higher than normal levels of DNA, such as cfDNA, originating from a solid tissue in a blood sample may be indicative of the presence of disease related to the solid tissue. In some embodiments, the DNA analysis is used to determine the likelihood that the subject has cancer or pre-cancer.
[0191] In some embodiments, the copy number of the target regions is amplified or aberrantly high, thereby facilitating increased sensitivity of detection of the target regions. In some embodiments, the target regions are type-specific hypermethylated regions. In some embodiments, the target regions are type-specific hypomethylatedregions. In some embodiments, type-specific differentially methylated regions are differentially methylated in one or a plurality of related cell types. In some such embodiments, the target regions are differentially methylated in immune cells relative to non-immune cells. In some embodiments, type-specific differentially methylated regions are differentially methylated in one or a plurality of related tissue types. In some such embodiments, the target regions are differentially methylated in one or more solid tissue types relative to cell types normally found in a sample, such as a blood sample. In some such embodiments, the target regions are differentially methylated in one or more solid tissue types other than bladder tissue relative to cell types normally found in a sample, such as a blood sample. In some embodiments, the target regions exclude regions differentially methylated in bladder. In some embodiments, the target regions are typespecific fragments. In some embodiments, type-specific fragments arise from fragmentation patterns specific to one or a plurality of related cell types. In some such embodiments, the fragmentation patterns are specific to immune cells relative to non- immune cells. In some embodiments, type-specific fragmentation patterns are specific to one or a plurality of related tissue types. In some such embodiments, the fragmentation patterns are specific to one or more solid tissue types relative to cell types normally found in a sample, such as a blood sample.
[0192] In some embodiments, the copy number variation of one or more target regions is a focal amplification. In some such embodiments, the focal amplification is associated with cancer. In some embodiments, the plurality of target regions comprises regions of one or more of AR, BRAF, CCND1 , CCND2, CCNE1 , CDK4, CDK6, EGFR, ERBB2, FGFR1 , FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAFI. For example, in some embodiments, the plurality of target regions comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, or 18 of the foregoing genes. Accordingly, in some embodiments, the plurality of target-specific probes comprises probes that each specifically bind to one of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, or 18 of the foregoing genes.
[0193] In some embodiments, target regions comprise a type-specific epigenetic variation specific to DNA, such as cfDNA, originating from immune cells relative to DNA, such as cfDNA, originating from non-immune cells. In some such embodiments, theplurality of target regions comprises target regions that are differentially methylated in immune cells relative to non-immune cells. In some such embodiments, the plurality of target regions comprises target regions that are hypermethylated in at least some types of immune cells relative to non-immune cells. In some embodiments, the plurality of target regions comprises target regions that are hypomethylated in at least some types of immune cells relative to non-immune cells. In some embodiments, the plurality of target regions comprises fragmentation patterns present in greater proportion in immune cells relative to non-immune cells. In some embodiments, the target regions comprise a type-specific epigenetic variation specific to cfDNA originating from a plurality of immune cell types relative to other immune cell types and non-immune cells present in the sample. In some such embodiments, the plurality of immune cell types comprises naive and activated lymphocytes; monocytes and macrophages; or myelocytes, neutrophils, and eosinophils. In some embodiments, the plurality of immune cell types comprises naive T cells, naive B cells, effector CD4 T cells, effector CD8 T cells, Treg cells, plasma cells, and memory cells. In some embodiments, the plurality of immune cell types comprises metamyelocytes. In some embodiments, the plurality of immune cell types comprises natural killer (NK) cells.
[0194] In some embodiments, target regions comprise a type-specific epigenetic variation specific to DNA, such as cfDNA, originating from a solid tissue relative to DNA, such as cfDNA, originating from other cell or tissue types, such as other cell types found in the sample. In some such embodiments, the plurality of target regions comprises target regions that are differentially methylated in a solid tissue type relative to other tissue types or cell types in the sample. In some embodiments, the plurality of target regions comprises fragmentation patterns present in greater proportion in a solid tissue type relative to other tissue types or cell types in the sample. In some embodiments, the solid tissue type is colon, lung, breast, liver, kidney, prostate, skin, bladder, or pancreas.
[0195] In some embodiments, the plurality of target regions comprises typespecific hypermethylated regions. In some embodiments, the hypermethylated target regions are methylated to an extent that is at least 10%, at least 20%, at least 30% or at least 40% greater than the average methylation of the target regions in the sample. In some embodiments, the hypermethylated target regions are methylated to an extentthat is 5-10%, 10-20%, 10-30%, 20- 30%, 30-40%, 40-50%, or 10-50% greater than the average methylation of the target regions in the sample. In some embodiments, the hypermethylated target regions are methylated to an extent that is at least 10%, at least 20%, at least 30% or at least 40% greater than the average methylation of the DNA in the sample. In some embodiments, the hypermethylated target regions are methylated to an extent that is 5-10%, 10-20%, 10-30%, 20-30%, 30-40%, 40-50%, or 10-50% greater than the average methylation of the DNA in the sample. In some embodiments, the hypermethylated target regions are methylated to an extent that is at least 10%, at least 20%, at least 30% or at least 40% greater than the average methylation of the corresponding target regions in DNA originating from cell or tissue types than the one or more related cell or tissue types of the type-specific hypermethylated target regions. In some embodiments, the hypermethylated target regions are methylated to an extent that is 5-10%, 10- 20%, 10-30%, 20-30%, 30-40%, 40-50%, or 10-50% greater than the average methylation of the corresponding target regions in DNA originating from cell or tissue types than the one or more related cell or tissue types of the type-specific hypermethylated target regions. In some embodiments, type-specific hypermethylated target regions are hypermethylated in healthy cells (e.g., healthy cells of one or more solid tissue types) or healthy subjects. In such embodiments, the methylation status per se of such target regions may not be directly indicative of the presence of disease. The methylated status of such target regions is indicative of the cell or tissue type from which the DNA originated, and if the cell or tissue types from which the DNA originated is not expected to be present at significant levels in a given sample, e.g., DNA from colon tissue in a blood sample, it may be indicative of the presence of disease in the subject from which the sample was obtained. In some embodiments, type-specific hypermethylated target regions are hypermethylated in healthy cells and subjects and in diseases cells and in a subject having a disease. In some such embodiments, the extent of methylation is further increased in diseased cells compared to healthy cells, thereby further increasing the sensitivity of detection of type- specific DNA that may be indicative of diseases in the subject from which it was obtained.
[0196] In some embodiments, the plurality of target regions comprises typespecific fragmentation patterns. In some embodiments, fragments produced by suchtype-specific fragmentation patterns are present at levels at least 10%, at least 20%, at least 30% or at least 40% greater than the average levels of the fragments in samples obtained from healthy subjects. In some embodiments, fragments produced by such type-specific fragmentation patterns are present at levels 5-10%, 10-20%, 10-30%, 20- 30%, 30-40%, 40-50%, or 10-50% greater than the average levels of the fragments in samples obtained from healthy subjects. In some embodiments, type-specific fragmentation patterns are present in healthy cells or healthy subjects. In such embodiments, the presence of the corresponding fragments may not be directly indicative of the presence of disease. The presence of such fragments is indicative of the cell or tissue type from which the DNA originated, and the presence of DNA that originated from cell or tissue types not expected to be present in a given sample, e.g., DNA from colon tissue in a blood sample, may be indicative of the presence of disease in the subject from which the sample was obtained. In some such embodiments, the levels of fragments corresponding to a type- specific fragmentation pattern are further increased in diseased cells compared to healthy cells, thereby further increasing the sensitivity of detection of type-specific DNA that may be indicative of diseases in the subject from which it was obtained. Exemplary approaches for analysis of DNA fragmentation patterns are provided in, e.g., W02022040163A1 , US11352670B2, US20200056245A1 , US10297342B2, US10741270B2, US10453556B2, US9892230B2, EP3617324A1 , and EP2860266B1 , which are each incorporated herein by reference in their entireties for all purposes.
[0197] In some embodiments, the plurality of target regions comprises copy number variants having an aberrantly high copy number, e.g., a focal amplification or duplication. In some such embodiments, the increased copy number of the target regions further increases the sensitivity of the methods described herein. In some embodiments, the copy number variants are type-specific copy number variants. In some embodiments, the copy number variants are aberrantly high and associated with pre-cancer or cancer. In some embodiments, the copy number variants are copy number amplifications known to occur in early cancer or pre-cancer. In some such embodiments, the copy number variants are present in subjects having a disease. In some embodiments, the copy number variants are aberrantly high in diseased cellscompared to healthy cells, thereby increasing the sensitivity of detection of type-specific DNA that may be indicative of diseases in the subject from which it was obtained.Enriching / Capturing
[0198] In some embodiments, methods disclosed herein can comprise enriching or capturing DNA, such as type- specific cfDNA target regions that are also copy number variants. In some embodiments, the capturing comprises contacting the DNA with probes specific for the target regions to form probe-annealed DNA. Enrichment or capture may be performed on any sample or subsample described herein using any suitable approach known in the art.
[0199] In some embodiments, the probes specific for the target regions (i.e. , target-specific probes) comprise a capture moiety that facilitates the enrichment or capture of the DNA hybridized to the probes (probe-annealed DNA). In some embodiments, a probe specific for a target region may be a first member of a binding pair. In some embodiments, the second member of the binding pair is a capture moiety, wherein the capture moiety has high affinity to the first member of the binding pair. In some embodiments, the capture moiety is biotin. In some such embodiments, streptavidin attached to a solid support, such as magnetic beads, is used to bind to the biotin. Nonspecifically bound DNA that does not comprise a target region is washed away from the captured DNA. In some embodiments, DNA is then dissociated from the probes and eluted from the solid support using salt washes or buffers comprising another DNA denaturing agent. In some embodiments, the probes are also eluted from the solid support by, e.g., disrupting the biotin-streptavidin interaction.
[0200] In some embodiments, the probe annealed to the DNA comprises a splint region, which is complimentary to the splint overhang of an adapter. In some embodiments, the splint region further comprises a hybridized oligomer. In some embodiments, the hybridized oligomer comprises digestible bases.
[0201] In some embodiments, the methods herein comprise enriching for or capturing DNA comprising epigenetic and / or sequence-variable target regions. Such regions may be captured from an aliquot, portion, or subsample of a sample (e.g., a sample that has undergone attachment of adapters and amplification), while the step ofpartitioning the DNA with an agent that recognizes methyl cytosine is performed on a separate aliquot, portion, or subsample of the sample. Enriching for or capturing DNA comprising epigenetic and / or sequence-variable target regions may comprise contacting the DNA with a first or second set of target-specific probes. Such target-specific probes may have any of the features described herein for sets of target- specific probes, including but not limited to in the embodiments set forth above and the sections relating to probes below. Capturing may be performed on one or more subsamples prepared during methods disclosed herein. In some embodiments, DNA is captured from a first subsample or a second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture. Exemplary methods for capturing DNA comprising epigenetic and / or sequencevariable target regions can be found in, e.g., WO 2020 / 160414, which is hereby incorporated by reference.
[0202] In some embodiments, the capturing step or steps may be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on features of the probes such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given general knowledge in the art regarding nucleic acid hybridization.
[0203] In some embodiments, methods described herein comprise capturing a plurality of sets of target regions of cfDNA obtained from a subject. In some embodiments, the target regions may comprise differences depending on whether they originated from a tumor or from healthy cells or from a certain cell type. The capturing step produces a captured set of cfDNA molecules. In some embodiments, cfDNA molecules corresponding to a sequence-variable target region set are captured at a greater capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to an epigenetic target region set. In some embodiments, a method described herein comprises contacting cfDNA obtained from a subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set. For additionaldiscussion of capturing steps, capture yields, and related aspects, see W02020 / 160414, which is incorporated herein by reference for all purposes.
[0204] In some embodiments, it can be beneficial to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set because a greater depth of sequencing may be necessary to analyze the sequence-variable target regions with sufficient confidence or accuracy than may be necessary to analyze the epigenetic target regions. The volume of data needed to determine fragmentation patterns (e.g., to test for perturbation of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated partitions) is generally less than the volume of data needed to determine the presence or absence of cancer-related sequence mutations. Capturing the target region sets at different yields can facilitate sequencing the target regions to different depths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell). Although copy number variations such as focal amplifications are somatic mutations, they can be detected by sequencing based on read frequency in a manner analogous to approaches for detecting certain epigenetic changes such as changes in methylation. Thus, they can be considered epigenetic target regions for functional reasons. Additionally, regions showing copy number variation that are also hypermethylation-variable or fragmentation-variable target regions are considered epigenetic target regions because they may show epigenetic variation.
[0205] In some embodiments, a capturing step is performed with probes for a sequence-variable target region set and probes for an epigenetic target region set in the same vessel at the same time, e.g., the probes for the sequence-variable and epigenetic target region sets and capture probes are in the same composition. In some embodiments, this approach provides a relatively streamlined workflow.
[0206] In some embodiments, adapters are included in the DNA as described herein. In some embodiments, tags, which may be or include barcodes, are included in the DNA. In some embodiments, such tags are included in adapters. In some embodiments, tags can facilitate identification of the origin of a nucleic acid. For example, barcodes can be used to allow the origin (e.g., subject) whence the DNAcame to be identified following pooling of a plurality of samples for parallel sequencing. This may be done concurrently with an amplification procedure, e.g., by providing the barcodes in a 5’ portion of a primer, e.g., as described herein. In some embodiments, adapters and tags / barcodes are provided by the same primer or primer set. For example, the barcode may be located 3’ of the adapter and 5’ of the target-hybridizing portion of the primer. Alternatively, barcodes can be added by other approaches, such as ligation, optionally together with adapters in the same ligation substrate.
[0207] Additional details regarding amplification, tags, and barcodes are discussed herein, which can be combined to the extent practicable with any of these embodiments.Capture Moieties and Binding Pairs
[0208] As discussed above, nucleic acids in a sample can be subject to a capture step, in which molecules having certain characteristics are captured and analyzed. Target capture can involve use of a bait set comprising oligonucleotide baits labeled with a capture moiety, such as biotin or the other examples noted below. The probes can have sequences selected to tile across a panel of regions, such as genes. In some embodiments, a bait set can have higher and lower capture yields for sets of target regions such as those of the sequence-variable target region set and the epigenetic target region set, respectively, as discussed elsewhere herein. Such bait sets are combined with a sample under conditions that allow hybridization of the target molecules with the baits. Then, captured molecules are isolated using the capture moiety. DNA capture can involve use of oligonucleotides labeled with a capture moiety, such as target-specific probes labeled with biotin, and a second moiety or binding partner that binds to the capture moiety, such as streptavidin. In some embodiments, a capture moiety and binding partner can have higher and lower capture yields for different sets of probes, such as those used to capture a sequence- variable target region set and an epigenetic target region set, respectively, as discussed elsewhere herein. Methods comprising capture moieties are further described in, for example, U.S. patent 9,850,523, issuing December 26, 2017, which is incorporated herein by reference.
[0209] Capture moieties include, without limitation, biotin, avidin, streptavidin, a nucleic acid comprising a particular nucleotide sequence, a hapten recognized by an antibody, and magnetically attractable particles. The extraction moiety can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, a capture moiety that is attached to an analyte is captured by its binding pair which is attached to an isolatable moiety, such as a magnetically attractable particle or a large particle that can be sedimented through centrifugation. In some embodiments, the capture moiety can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties are biotin which allows affinity separation by binding to streptavidin linked or linkable to a solid phase or an oligonucleotide, which allows affinity separation through binding to a complementary oligonucleotide linked or linkable to a solid phase.Primer Extension
[0210] In some embodiments, the methods disclosed herein comprise steps of contacting the DNA with a primer that anneals in a parallel orientation to a target region, thereby producing primer-annealed DNA.
[0211] In some embodiments, the primer-annealed DNA is treated with an exonuclease. In some embodiments, the exonuclease is a 5’ to 3’ exonuclease. In some embodiments, the exonuclease is a 3’ to 5’ exonuclease.
[0212] In some embodiments, the primer-annealed DNA is ligated to an adapter to produce adapter-ligated DNA. In some embodiments, the adapter is added by blunt- end ligation. In some embodiments, the adapter is added by sticky-end ligation. In some embodiments, the adapter comprises an overhang region, e.g. a splint overhang, which mediates sticky-end ligation. In some embodiments, sticky-end ligation is preceded by an A-tailing treatment to promote sticky-end ligation.
[0213] In some embodiments, the adapter-ligated DNA is contacted by a DNA polymerase. In some embodiments, the adapter-ligated DNA is contacted by a DNA polymerase and deoxynucleoside triphosphates, thereby producing a primer-extended sample. In some embodiments, the DNA polymerase is a 5’ to 3’ exonuclease positivepolymerase and does not have significant 3’ to 5’ exonuclease activity. A DNA polymerase that is “strand displacement positive” is able displace a strand, such as a primer or probe, annealed to DNA. In some embodiments, a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase is used for extension of a nucleic acid sequence in order to remove a primer or probe sequence annealed to a region. In some embodiments, the use of a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase elutes a nucleic acid strand bound to a member of a binding pair. In some embodiments, the member of the binding pair is a streptavidin bead.
[0214] In some embodiments, one or more target regions are amplified, e.g., as part of sequencing library preparation (e.g., LP-PCR) and / or adapters are ligated to the DNA prior to annealing the DNA to primers used in primer extension. In some embodiments, the primers used in primer extension are shorter than 100 nucleosides in length and at least 20 nucleotides in length. In some embodiments, the primers used in primer extension are 20-60, 25-60, or 30-40 nucleotides in length. In some embodiments, a majority (e.g., at least 60%, 70%, 80%, 90%, or all) of the primers used in primer extension anneal to their corresponding target regions in a parallel orientation. In such embodiments, the 3’ end of each primer points toward the 5’ end of another primer and / or the 5’ end of each primer points toward the 3’ end of another primer when they are annealed to the same DNA strand. Thus, in such embodiments, only one strand of DNA in target region anneals to any of the plurality of primers. In some embodiments, a majority (e.g., at least 60%, 70%, 80%, 90%, or all) of the primers anneal to a site separated from the site to which another primer anneals by less than or equal to about 50 nucleotides (e.g., less than or equal to 40, 30, 20, 10, 5, 4, 3, 2, or 1 nucleotides).
[0051] In some embodiments, a majority (e.g., at least 60%, 70%, 80%, 90%, or 95%) of the primers annealed to a wild type target region are blocked by a downstream primer. This means of the distinct species of primer in the plurality, the specified portion can anneal to wild-type sequence upstream of another primer so as to be blocked from extension (a small amount of extension may be permitted depending on the amount of separation between primers, as discussed above). For any given target region, there will be one last primer that is not upstream of another primer andthus is not blocked by a downstream primer; this primer may be extended or may be blocked by a non-extendable blocking oligonucleotide, as discussed elsewhere herein.
[0215] In some embodiments, the plurality of primers have approximately uniform melting temperatures with respect to annealing to their complementary sequences, e.g., so that substantially all of the primers are capable of annealing under the same condition. For example, the plurality of primers may each anneal to their respective complementary sequence with a melting temperature within a range of 10°C, 9°C, 8°C, 7°C, 6°C, 5°C, 4°C, 3°C, 2°C, or 1 °C. Variation of parameters such as primer length, GC content, and nucleotide modifications (e.g., methylation, base analogs, or LNA modifications) are known approaches for adjusting primer melting temperatures to desired values.
[0216] Extension being blocked by a downstream primer when the primers are annealed to a wild-type sequence means that the 5’ to 3’ exonuclease negative, strand displacement negative DNA polymerase is unable to continue extending a primer when it encounters the 5’ end of another primer. This may occur after a short amount of extension, such as up to 50 nucleotides, e.g., up to 40, 30, 20, 10, 5, 4, 3, 2, or 1 nucleotides. The amount of permitted extension before becoming blocked will depend on the distance between the primers when annealed to the template sequence. In some embodiments, at least 60%, 70%, 80%, 90%, or 95% of the primers are blocked in this way when annealed to wild-type sequence. In some embodiments, the sample comprises a mixture of DNA having a wild-type sequence and a structural variation mutant sequence, and extension of at least one primer is blocked on the wild-type sequence but not on the structural variation mutant sequence. Thus, such methods can preferentially generate products in a primer-extended sample complementary to the structural variation mutant sequence. This can facilitate detection of such sequences because they are more common in the primer-extended sample than in the initial sample.
[0027] In some embodiments, the methods comprise contacting the DNA with a plurality of primers. In some embodiments, each of the plurality of primers anneal to different portions of the same target region. In some embodiments, two or more of the plurality of primers anneal to a first target region and two or more other primers of theplurality of primers anneal to a second target region. In some such embodiments, the plurality of primers comprises primers that each anneal to one of two, three, four, five, 10, 20, 30, 40, 50, 100, or more different target regions. For example, the DNA may comprise a plurality of target regions, and the plurality of primers may comprise at least two, three, four, five, 10, 20, 30, 40, 50, 100, or more primers configured to anneal to each of the plurality of target regions, resulting in each of the plurality of target regions annealing to at least two primers comprising two different sequences.
[0218] In some embodiments, the methods comprise contacting the DNA with blocking oligonucleotides that anneal to an exon adjacent to an intron of the target region in wild type sequences. In some embodiments, the methods comprise contacting the DNA with blocking oligonucleotides that anneal to an intron adjacent to an exon of the target region in wild type regions. In some embodiments, such as embodiments in which possible breakpoints in the target region are unknown, the primers used in the primer extension step are configured so that the entire wild type sequence of the target region is tiled with primers but that the sequence will not be completely tiled with primers if it comprises a rearrangement. In some embodiments, such as embodiments in which possible breakpoints in the target region are known, the primers used in the primer extension step are configured to be adjacent to a possible breakpoint when annealed to the DNA. In some such embodiments, the methods comprise contacting the DNA with blocking oligonucleotides that anneal to the wild type sequence adjacent to the breakpoint such that the blocking oligonucleotide and primer are adjacent when annealed to a wild type sequence but the blocking oligonucleotide does not anneal to DNA comprising the rearrangement.
[0219] In some embodiments, the deoxynucleoside triphosphates comprise a non-zero percentage of a modified deoxynucleoside triphosphate (e.g., a modified deoxycytidine triphosphate, deoxyadenosine triphosphate, or deoxyuridine triphosphate). The modification can facilitate capture of molecules into which the modified deoxynucleoside triphosphate has been incorporated. In some embodiments, the modified deoxynucleoside triphosphate comprises a capture moiety, which may be linked to the nucleobase of the dNTP. In some embodiments, the capture moiety is biotin. In any of the foregoing embodiments, the modified deoxynucleoside triphosphatemay be a modified deoxycytidine triphosphate. In some embodiments, the percentage of deoxynucleoside triphosphate comprising a capture moiety minimizes the number of short primer-extended products produced on a wild type DNA molecule template that comprise the capture moiety and / or maximizes the number of primer-extended products produced on a DNA molecule comprising a rearrangement that comprise the capture moiety. In some embodiments, the ratio of the modified:unmodified versions of the modified deoxynucleoside triphosphate is from 1 :1 to 1 :100, e.g., from 1 :9 to 1 :19.
[0220] In some embodiments, the 5’ to 3’ exonuclease negative, strand displacement negative DNA polymerase is also a 3’ to 5’ exonuclease positive DNA polymerase that degrades any unannealed 3’ primer ends. In some such embodiments, the methods comprise contacting the DNA with blocking oligonucleotides that anneal to the 3’ end adapter of the DNA to prevent degradation of the 3’ end of the DNA molecule isolated from the sample, and / or contacting the DNA with blocking oligonucleotides that anneal to the 5’ end adapter of the DNA to block extension of primers that anneal at or near the junction between a wild-type intron (or portion thereof) and the adapter, n some embodiments, the blocking oligonucleotides comprise a modified nucleoside at the 3’ end of the blocking oligonucleotide. In some such embodiments, the modified nucleoside comprises an inverted nucleobase, such as inverted thymine. In some embodiments, the modified nucleoside is an abasic nucleoside.
[0221] In some embodiments, the primer-extended products are enriched and / or captured using a solid support linked to a binding partner of the capture moiety, thereby also enriching and / or capturing the DNA molecules isolated from the sample that are hybridized to the primer- extended products. In some embodiments, the DNA molecules are denatured from their corresponding primer-extended products and sequenced in order to determine if they comprise a rearrangement and, if so, the breakpoint(s) of the rearrangement and further analysis as described herein. In some embodiments, the DNA molecules isolated from the sample that are hybridized to the primer-extended products comprise adapters and / or barcodes. In some embodiments, the adapters are used to amplify such molecules after denaturation. The barcodes can be used to identify sequence reads originating from the same molecule. When blocking oligonucleotides that anneal to adapters are present during production of the primer-extended sample, these can prevent incorporation of adapter and / or barcode sequences into the primer-extended products, e.g., so that only or substantially only the original DNA molecules and not the primer-extended products are amplified and sequenced.
[0222] In some embodiments, the primers that anneal to the target region of the DNA for use in a primer extension procedure do not comprise a tail that does not bind to a target region. In some embodiments, the primers that anneal to the target region of the DNA for use in a primer extension procedure do comprise a tail that does not bind to a target region but is short enough to allow the primers to anneal to the target region. In some such embodiments, the primer tail is at the 5’ end of the primer. In such embodiments, the primer tail binds to the 5’ end of the adapter ligated to the 3’ end of the DNA molecule. In some embodiments, the primer tail comprises a capture moiety and the deoxynucleoside triphosphates used in primer extension may not comprise a capture moiety. In other embodiments, the primer tail does not comprise a capture moiety and the deoxynucleoside triphosphates used in primer extension comprise a capture moiety. The primer-extended products are captured using a solid support linked to a binding partner of the capture moiety and amplified on the solid support using PCR primers that anneal to a sequence within the primer tail and to the adapter ligated to the 5’ end of the DNA. Alternatively, the primer tail comprises a modification at the 5’ end that protects the primer from exonuclease activity, such as phosphorothioate intemucleoside linkages. In some such embodiments, primer-extended products are enriched by contacting the primer-extended sample with a 5’ to 3’ exonuclease to degrade wild type sequences. The remaining sequences are amplified by PCR. In some embodiments, TdT ddATP tailing is performed prior to amplification in order to prevent truncated sequences from acting as primers in the PCR.Ligating Adapters for Sequencing
[0223] In some embodiments, adapters are added to the nucleic acids. In some embodiments, this may be done concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer (where PCR is used, this can be referred to as library prep-PCR or LP-PCR). In some embodiments, adapters are addedby other approaches, such as ligation. In some such methods, prior to primer annealing, first adapters are added to the nucleic acids by ligation to the 3’ ends thereof, which may include ligation to single-stranded DNA. The adapter can be used as a priming site for second-strand synthesis, e.g., using a universal primer and a DNA polymerase. A second adapter can then be ligated to at least the 3’ end of the second strand of the now double- stranded molecule. In some embodiments, the first adapter comprises an affinity tag, such as biotin, and nucleic acid ligated to the first adapter is bound to a solid support (e.g., bead), which may comprise a binding partner for the affinity tag such as streptavidin. For further discussion of a related procedure, see Gansauge et al., Nature Protocols 8:737-748 (2013). Commercial kits for sequencing library preparation compatible with single-stranded nucleic acids are available, e.g., the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adapter ligation, nucleic acids are amplified.
[0224] Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites, but preferably adapters include the same primer binding site.
[0225] In some embodiments, the adapter is a directional adapter and comprises a motor protein.
[0226] In some embodiments, following attachment of adapters, the nucleic acids are subject to amplification. The amplification can use, e.g., universal primers that recognize primer binding sites in the adapters.
[0227] In some embodiments, following attachment of adapters, the nucleic acids are contacted with an agent that preferentially binds to nucleic acids bearing an epigenetic modification. The nucleic acids are partitioned into at least two subsamples differing in the extent to which the nucleic acids bear the modification from binding to the agents. For example, if the agent has affinity for nucleic acids bearing the modification, nucleic acids overrepresented in the modification (compared with median representation in the population) preferentially bind to the agent, whereas nucleic acidsunderrepresented for the modification do not bind or are more easily eluted from the agent. The nucleic acids are then amplified from primers binding to the primer binding sites within the adapters. Partitioning may be performed instead before adapter attachment, in which case the adapters may comprise differential tags that include a component that identifies which partition a molecule occurred in.Tagging
[0228] “Tagging” DNA molecules is a procedure in which a tag is attached to or associated with the DNA molecules. Tags can be molecules, such as nucleic acids, containing information that indicates a feature of the molecule with which the tag is associated. In some embodiments, tags can allow one to differentiate molecules from which sequence reads originated. For example, molecules can bear a sample tag (which distinguishes molecules in one sample from those in a different sample) or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from one another (in both unique and non-unique tagging scenarios). For methods that involve an optional partitioning step, in some embodiments, a partition tag (which distinguishes molecules in one partition from those in a different partition) may be included. In some embodiments, adapters added to DNA molecules comprise tags. In certain embodiments, a tag can comprise one or a combination of barcodes. As used herein, the term “barcode” refers to a nucleic acid molecule having a particular nucleotide sequence, or to the nucleotide sequence, itself, depending on context. In some embodiments, a barcode can have, for example, between 10 and 100 nucleotides. In some embodiments, a collection of barcodes can have degenerate sequences or can have sequences having a certain hamming distance, as desired for the specific purpose. So, for example, in some embodiments, a molecular barcode can be comprised of one barcode or a combination of two barcodes, each attached to different ends of a molecule. Additionally or alternatively, for different partitions and / or samples, in some embodiments, different sets of molecular barcodes, or molecular tags can be used such that the barcodes serve as a molecular tag through their individual sequences and also serve to identify the partition and / or sample to which they correspond based the set of which they are a member. In some embodiments, tagscomprising barcodes can be incorporated into or otherwise joined to adapters. In some embodiments, tags can be incorporated by ligation, overlap extension PCR among other methods.
[0229] Tagging strategies include unique tagging and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample bear a different tag, so that reads can be assigned to original molecules based on tag information alone. Tags used in such methods are sometimes referred to as “unique tags”. In non -unique tagging, different molecules in the same sample can bear the same tag, so that other information in addition to tag information is used to assign a sequence read to an original molecule. In some embodiments, such information may include start and stop coordinate, coordinate to which the molecule maps, start or stop coordinate alone, etc. Tags used in such methods are sometimes referred to as “nonunique tags”. Accordingly, In some embodiments, it is not necessary to uniquely tag every molecule in a sample. It suffices to uniquely tag molecules falling within an identifiable class within a sample. Thus, molecules in different identifiable families can bear the same tag without loss of information about the identity of the tagged molecule.
[0230] In certain embodiments of non-unique tagging, the number of different tags used can be sufficient that there is a very high likelihood (e.g., at least 99%, at least 99.9%, at least 99.99% or at least 99.999% that all molecules of a particular group bear a different tag. It is to be noted that when barcodes are used as tags, and when barcodes are attached, e.g., randomly, to both ends of a molecule, the combination of barcodes, together, can constitute a tag. This number, in term, is a function of the number of molecules falling into the calls. For example, the class may be all molecules mapping to the same start-stop position on a reference genome. The class may be all molecules mapping across a particular genetic locus, e.g., a particular base or a particular region (e.g., up to 100 bases or a gene or an exon of a gene). In certain embodiments, the number of different tags used to uniquely identify a number of molecules, z, in a class can be between any of 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, 9*z, 10*z, 11 *z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z or 100*z (e.g., lower limit) and any of 100,000*z, 10,000*z, 1000*z or 100*z (e.g., upper limit).
[0231] For example, in a sample of about 5 ng to 30 ng of cell free DNA, one expects around 3000 molecules to map to a particular nucleotide coordinate, and between about 3 and 10 molecules having any start coordinate to share the same stop coordinate. Accordingly, about 50 to about 50,000 different tags (e.g., between about 6 and 220 barcode combinations) can suffice to uniquely tag all such molecules. To uniquely tag all 3000 molecules mapping across a nucleotide coordinate, about 1 million to about 20 million different tags would be used.
[0232] In some embodiments, assignment of unique or non-unique tags barcodes in reactions follows methods and systems described by US patent applications 20010053519, 20030152490, 20110160078, and U.S. Pat. No. 6,582,908 and U.S. Pat. No. 7,537,898 and US Pat. No. 9,598,731 , incorporated by reference herein. Tags can be linked to sample nucleic acids randomly or non-randomly.
[0233] In some embodiments, the tagged nucleic acids are sequenced after loading into a microwell plate. In some embodiments, the microwell plate can have 96, 384, or 1536 microwells. In some cases, they are introduced at an expected ratio of unique tags to microwells. For example, the unique tags may be loaded so that more than about 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1 ,000,000, 10,000,000, 50,000,000 or 1 ,000,000,000 unique tags are loaded per genome sample. In some cases, the unique tags may be loaded so that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1 ,000,000, 10,000,000, 50,000,000 or 1 ,000,000,000 unique tags are loaded per genome sample. In some cases, the average number of unique tags loaded per sample genome is less than, or greater than, about 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1 ,000,000, 10,000,000, 50,000,000 or 1 ,000,000,000 unique tags per genome sample.
[0234] In some embodiments, a format uses 20-50 different tags (e.g., barcodes) ligated to both ends of target nucleic acids. For example 35 different tags (e.g., barcodes) ligated to both ends of target molecules creating 35 x 35 permutations, which equals 1225 for 35 tags. Such numbers of tags are sufficient so that different molecules having the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags. Otherbarcode combinations include any number between 10 and 500, e g., about 15x15, about 35x35, about 75x75, about 100x100, about 250x250, about 500x500.
[0235] In some cases, unique tags may be predetermined or random or semirandom sequence oligonucleotides. In some cases, a plurality of barcodes may be used such that barcodes are not necessarily unique to one another in the plurality. For example, in some embodiments, barcodes may be ligated to individual molecules such that the combination of the barcode and the sequence it may be ligated to creates a unique sequence that may be individually tracked. As described herein, detection of non-unique barcodes in combination with sequence data of beginning (start) and end (stop) portions of sequence reads may allow assignment of a unique identity to a particular molecule. The length or number of base pairs, of an individual sequence read may also be used to assign a unique identity to such a molecule. As described herein, fragments from a single strand of nucleic acid having been assigned a unique identity, may thereby permit subsequent identification of fragments from the parent strand.
[0236] In some embodiments, two or more partitions, e.g., each partition, is / are differentially tagged. Tags can be used to label the individual polynucleotide population partitions so as to correlate the tag (or tags) with a specific partition. Alternatively, tags can be used in embodiments of the invention that do not employ a partitioning step. In some embodiments, a single tag can be used to label a specific partition. In some embodiments, multiple different tags can be used to label a specific partition. In embodiments employing multiple different tags to label a specific partition, the set of tags used to label one partition can be readily differentiated for the set of tags used to label other partitions. In some embodiments, the tags may have additional functions, for example the tags can be used to index sample sources or used as unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations, for example as in Kinde et al., Proc NatT Acad Sci USA 108: 9530-9535 (2011 ), Kou et al., PLoS ONE, 11 : e0146638 (2016)) or used as non unique molecule identifiers, for example as described in US Pat. No. 9,598,731. Similarly, in some embodiments, the tags may have additional functions, for example the tags can be used to index sample sources or used as non-uniquemolecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations).
[0237] In some embodiments, partition tagging comprises tagging molecules in each partition with a partition tag. After re-combining partitions (e.g., to reduce the number of sequencing runs needed and avoid unnecessary cost) and sequencing molecules, the partition tags identify the source partition. In some embodiments, different partitions are tagged with different sets of molecular tags, e.g., comprised of a pair of barcodes. In this way, each molecular barcode indicates the source partition as well as being useful to distinguish molecules within a partition. For example, a first set of 35 barcodes can be used to tag molecules in a first partition, while a second set of 35 barcodes can be used tag molecules in a second partition.
[0238] In some embodiments, after partitioning and tagging with partition tags, the molecules may be pooled for sequencing in a single run. In some embodiments, a sample tag is added to the molecules, e.g., in a step subsequent to addition of partition tags and pooling. Sample tags can facilitate pooling material generated from multiple samples for sequencing in a single sequencing run.
[0239] In some embodiments, partition tags may be correlated to the sample as well as the partition. As a simple example, a first tag can indicate a first partition of a first sample; a second tag can indicate a second partition of the first sample; a third tag can indicate a first partition of a second sample; and a fourth tag can indicate a second partition of the second sample.
[0240] While tags may be attached to molecules already partitioned based on one or more characteristics, the final tagged molecules in the library may no longer possess that characteristic. For example, while single stranded DNA molecules may be partitioned and tagged, the final tagged molecules in the library are likely to be double stranded. Similarly, while DNA may be subject to partition based on different levels of methylation, in the final library, tagged molecules derived from these molecules are likely to be unmethylated. Accordingly, the tag attached to molecule in the library typically indicates the characteristic of the “parent molecule” from which the ultimate tagged molecule is derived, not necessarily to characteristic of the tagged molecule, itself. As an example, barcodes 1 , 2, 3, 4, etc. are used to tag and label molecules inthe first partition; barcodes A, B, C, D, etc. are used to tag and label molecules in the second partition; and barcodes a, b, c, d, etc. are used to tag and label molecules in the third partition.
[0241] In some embodiments, differentially tagged partitions can be pooled prior to sequencing. Differentially tagged partitions can be separately sequenced or sequenced together concurrently, e.g., in the same flow cell of an Illumina sequencer.
[0242] In some embodiments, after sequencing, analysis of reads to detect genetic variants can be performed on a partition-by-partition level, as well as a whole nucleic acid population level. Tags are used to sort reads from different partitions. Analysis can include in silico analysis to determine genetic and epigenetic variation (one or more of methylation, chromatin structure, etc.) using sequence information, genomic coordinates length, coverage, and / or copy number. In some embodiments, higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or a nucleosome depleted region (NDR).
[0243] In some instances, partitioning procedures may result in imperfect sorting of DNA molecules among the subsamples. For example, a minority of the molecules in the second subsample may be highly modified (e.g., hypermethylated), and / or a minority of the molecules in the first subsample may be unmodified or mostly unmodified (e.g., unmethylated or mostly unmethylated). Highly modified molecules in the second subsample and unmodified or mostly unmodified molecules in the first subsample are considered nonspecifically partitioned. In some embodiments, the methods described herein comprise steps that can reduce technical noise from nonspecifically partitioned DNA, e.g., by converting certain bases such that nonspecifically partitioned DNA can be identified following sequencing and / or by degrading it.Captured DNA; target regions
[0244] In some embodiments, nucleic acids captured or enriched using a method described here comprise “captured DNA” or “probe-annealed DNA”. As used herein, the terms “captured DNA” and “probe-annealed DNA” are used interchangeably.In some embodiments, captured DNA comprises a region comprising both a typespecific epigenetic variation and a copy number variation. In some embodiments, the variations are present in healthy cells but not normally present in the sample type, such as a blood sample. In some embodiments, the variations are present in aberrant cells (e.g., hyperplastic, metaplastic, or neoplastic cells).
[0245] In some embodiments, a first captured epigenetic target region set captured from a sample or first subsample comprises hypermethylation variable target regions. In some embodiments, the hypermethylation variable target regions show typespecific hypermethylation in healthy cfDNA from one or more related cell or tissue types. Without wishing to be bound by any particular theory, the presence of cancer cells may increase the shedding of DNA into the bloodstream (e.g., from the cancer and / or the surrounding tissue). As such, the distribution of tissue of origin of cfDNA may change upon carcinogenesis. Thus, in some instances, an increase in the level of hypermethylation variable target regions in the first subsample can be an indicator of the presence (or recurrence, depending on the history of the subject) of cancer.
[0246] In some embodiments, the methods herein comprise capturing a second captured epigenetic target region set from a sample or second subsample. In some embodiments, the second epigenetic target region set comprises hypomethylation variable target regions. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. As such, the distribution of tissue of origin of cfDNA may change upon carcinogenesis. Thus, in some instances an increase in the level of hypomethylation variable target regions in the second subsample can be an indicator of the presence (or recurrence, depending on the history of the subject) of cancer.
[0247] Additionally, in some instances captured target region sets may comprise DNA corresponding to a sequence-variable target region set. The captured sets may be combined to provide a combined captured set.
[0248] In some embodiments in which a captured set comprising DNA corresponding to a sequence-variable target region set and the epigenetic target region set includes a combined captured set as discussed above, the DNA corresponding to the sequence-variable target region set may be present at a greater concentration thanthe DNA corresponding to the epigenetic target region set, e.g., a 1 .1 to 1.2-fold greater concentration, a 1 .2- to 1.4-fold greater concentration, a 1 .4- to 1 .6-fold greater concentration, a 1 .6- to 1.8-fold greater concentration, a 1 .8- to 2.0-fold greater concentration, a 2.0- to 2.2-fold greater concentration, a 2.2- to 2.4-fold greater concentration a 2.4- to 2.6-fold greater concentration, a 2.6- to 2.8-fold greater concentration, a 2.8- to 3.0-fold greater concentration, a 3.0- to 3.5-fold greater concentration, a 3.5- to 4.0, a 4.0- to 4.5-fold greater concentration, a 4.5- to 5.0-fold greater concentration, a 5.0- to 5.5-fold greater concentration, a 5.5- to 6.0-fold greater concentration, a 6.0- to 6.5-fold greater concentration, a 6.5- to 7.0-fold greater, a 7.0- to 7.5-fold greater concentration, a 7.5- to 8.0-fold greater concentration, an 8.0- to 8.5- fold greater concentration, an 8.5- to 9.0-fold greater concentration, a 9.0- to 9.5-fold greater concentration, 9.5- to 10.0-fold greater concentration, a 10- to 11 -fold greater concentration, an 11- to 12-fold greater concentration a 12- to 13-fold greater concentration, a 13- to 14-fold greater concentration, a 14- to 15-fold greater concentration, a 15- to 16-fold greater concentration, a 16- to 17-fold greater concentration, a 17- to 18-fold greater concentration, an 18- to 19-fold greater concentration, a 19- to 20-fold greater concentration, a 20- to 30-fold greater concentration, a 30- to 40-fold greater concentration, a 40- to 50-fold greater concentration, a 50- to 60-fold greater concentration, a 60- to 70-fold greater concentration, a 70- to 80-fold greater concentration, a 80- to 90-fold greater concentration, or a 90- to 100-fold greater concentration. The degree of difference in concentrations accounts for normalization for the footprint sizes of the target regions, as discussed in the definition section.
[0249] In some embodiments, the captured DNA comprises target regions having a type-specific epigenetic variation and a copy number variation. In some embodiments, an epigenetic target region set consists of target regions having a typespecific epigenetic variation and a copy number variation. In some embodiments, the type-specific epigenetic variations, e.g., differential methylation or a type-specific fragmentation pattern, are likely to differentiate DNA from one or more related cell or tissue types cells from DNA from other cell or tissue types present in a sample or in a subject.
[0250] In some embodiments, a captured epigenetic target region set captured from a sample or first subsample comprises hypermethylation variable target regions. In some embodiments, the hypermethylation variable target regions are differentially or exclusively hypermethylated in one or more related cell or tissue types. Such hypermethylation variable target regions may be hypermethylated in other cell or tissue types but not to the extent observed in the one or more related cell or tissue types. In some embodiments, the hypermethylation variable target regions show even higher methylation in cfDNA from a diseased cell of the one or more related cell or tissue types. In some embodiments, target regions comprise hypermethylated regions with aberrantly high copy number. In some such embodiments, the target regions are hypermethylated in healthy and diseased colon tissue and have aberrantly high copy number in pre-cancerous or cancerous colon tissue.
[0251] In some embodiments, a captured epigenetic target region set captured from a sample or subsample comprises hypomethylation variable target regions. In some embodiments, the hypomethylation variable target regions are exclusively hypomethylated in one or more related cell or tissue types. Such hypomethylation variable target regions may be hypomethylated in other cell or tissue types but not to the extent observed in the one or more related cell or tissue types.
[0252] Without wishing to be bound by any particular theory, in an individual with cancer, proliferating or dying cancer cells may shed more DNA into the bloodstream than cells in a healthy individual and / or healthy cells of the same tissue type, respectively. As such, in some instances, the distribution of cell type and / or tissue of origin of cfDNA may change upon carcinogenesis. Thus, in some instances, the presence and / or levels of cfDNA originating from certain cell or tissue types can be an indicator of disease.
[0253] Further exemplary hypermethylation variable target regions and hypomethylation variable target regions useful for distinguishing between various cell types have been identified by analyzing DNA obtained from various cell types via whole genome bisulfite sequencing, as described, e.g., in Scott, C.A., Duryea, J.D., MacKay, H. etal. , “Identification of cell type- specific methylation signals in bulk whole genome bisulfite sequencing data,” Genome Biol 21 , 156 (2020) (doi.org / 10.1186 / sl3059-020-02065-5), which is incorporated by reference herein in its entirety. Whole-genome bisulfite sequencing data is available from the Blueprint consortium, available on the internet at dec. blueprint- epigenome. eu.
[0254] In some embodiments, first and second captured target region sets comprise, respectively, DNA corresponding to a sequence-variable target region set and DNA corresponding to the epigenetic target region set, for example, as described in WO 2020 / 160414. The first and second captured sets may be combined to provide a combined captured set. In some embodiments, the sequence-variable target region set and epigenetic target region set may have any of the features described for such sets in WO 2020 / 160414, which is incorporated by reference herein in its entirety. In some embodiments, the epigenetic target region set comprises a hypermethylation variable target region set. In some embodiments, the epigenetic target region set comprises a hypomethylation variable target region set. In some embodiments, the epigenetic target region set comprises CTCF binding regions. In some embodiments, the epigenetic target region set comprises fragmentation variable target regions. In some embodiments, the epigenetic target region set comprises transcriptional start sites.
[0255] In some embodiments, the sequence-variable target region set comprises a plurality of regions known to undergo somatic mutations in cancer. In some aspects, the sequence-variable target region set targets a plurality of different genes or genomic regions (“panel”) selected such that a determined proportion of subjects having a cancer exhibits a genetic variant or tumor marker in one or more different genes or genomic regions in the panel. The panel may be selected to limit a region for sequencing to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA, e.g., by adjusting the affinity and / or amount of the probes as described elsewhere herein. The panel may be further selected to achieve a desired sequence read depth. The panel may be selected to achieve a desired sequence read depth or sequence read coverage for an amount of sequenced base pairs. The panel may be selected to achieve a theoretical sensitivity, a theoretical specificity, and / or a theoretical accuracy for detecting one or more genetic variants in a sample.
[0256] In some embodiments, probes for detecting the panel of regions can include those for detecting genomic regions of interest (hotspot regions). In someembodiments, information about chromatin structure can be taken into account in designing probes, and / or probes can be designed to maximize the likelihood that particular sites (e.g., KRAS codons 12 and 13) can be captured, and may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation impacted by nucleosome binding patterns and GC sequence composition. Regions used herein can also include non-hotspot regions optimized based on nucleosome positions and GC models.
[0257] In some embodiments, probes for detecting the panel of regions can include those for detecting genomic regions of interest (hotspot regions). In some embodiments, information about chromatin structure can be taken into account in designing probes, and / or probes can be designed to maximize the likelihood that particular sites (e.g., KRAS codons 12 and 13) can be captured, and may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation impacted by nucleosome binding patterns and GC sequence composition. In some embodiments, regions used herein can also include non-hotspot regions optimized based on nucleosome positions and GC models.
[0258] Examples of listings of genomic locations of interest may be found in WO 2020 / 160414, e.g., at Table 4, which is incorporated by reference herein in its entirety.In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes of Table 3 of WO 2020 / 160414. In some embodiments, a sequence- variable target region set used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes of Table 4 of WO 2020 / 160414. Additionally or alternatively, suitable target region sets are available from the literature. For example, Gale et ak, PLoS One 13: e0194630 (2018), which is incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of a sequence-variable target region set. These 35 targets are AKT1 , ALK, BRAF, CCND1 , CDK2A, CTNNB1 , EGFR, ERBB2, ESR1 ,FGFR1 , FGFR2, FGFR3, F0XL2, GAT A3, GNA11 , GNAQ, GNAS, HRAS, IDH1 , IDH2, KIT, KRAS, MED 12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11 , TP53, and U2AF1.
[0259] In some embodiments, the sequence-variable target region set comprises target regions from at least 10, 20, 30, or 35 cancer-related genes, such as the cancer-related genes listed above and in WO 2020 / 160414.Sequencing
[0260] In general, sample nucleic acids, including nucleic acids flanked by adapters can be subject to sequencing. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, Digital Gene Expression (Helicos), Next generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), enzymatic methyl sequencing (EM-Seq), Tet-assisted pyridine borane sequencing (TAPS), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore (ONT), Roche Genia, Maxim-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms.
[0261] In some embodiments, sequencing comprises detecting and / or distinguishing unmodified and modified nucleobases. For example, single-molecule real-time (SMRT) sequencing and nanopore sequencing can facilitate direct detection of, e.g., 5-methylcytosine and 5- hydroxymethylcytosine as well as unmodified cytosine. See, e.g., Schatz., Nature Methods. 14(4): 347-348 (2017); US 9,150,918; and Simpson et al., Nature Methods. 14:407-410, doi:10.1038 / nmeth.4184 (2017). Sequencing reactions can be performed in a variety of sample processing units, which may multiple lanes, multiple channels, multiple wells, or other mean of processing multiple sample sets substantially simultaneously. Sample processing unit can also include multiple sample chambers to enable processing of multiple runs simultaneously.
[0262] The sequencing reactions can be performed on one or more forms of nucleic acids, such as those known to contain markers of cancer or of other disease.The sequencing reactions can also be performed on any nucleic acid fragments present in the sample. In some embodiments, sequence coverage of the genome may be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100%. In some embodiments, the sequence reactions may provide for sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the genome. Sequence coverage can be performed on at least 5, 10, 20, 70, 100, 200 or 500 different genes, or at most 5000, 2500, 1000, 500 or 100 different genes.
[0263] Simultaneous sequencing reactions may be performed using multiplex sequencing. In some cases, cell-free nucleic acids may be sequenced with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases cell-free nucleic acids may be sequenced with less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. Sequencing reactions may be performed sequentially or simultaneously. Subsequent data analysis may be performed on all or part of the sequencing reactions. In some cases, data analysis may be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, data analysis may be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. An exemplary read depth is 1000- 50000 reads per locus (base).Analyzing nucleic acids using target-specific probes
[0264] Methods of analyzing DNA or nucleic acids from a sample herein comprise contacting the DNA with a plurality of target-specific probes specific for member of an epigenetic target region set comprising or consisting of target regions that have both a type-specific epigenetic variation and a copy number variation. In some embodiments, the type-specific epigenetic variation is differential methylation. In some embodiments, the type-specific variation is a type - specific fragmentation pattern. In some embodiments, the copy number variation is aberrantly high copy number.
[0265] In some embodiments, the methods disclosed herein further comprise contacting DNA from the sample with probes specific for sequence variable targetregions and / or a second epigenetic target region set and capturing DNA that anneals to the probes specific for the sequence variable and / or second epigenetic set of target regions. Such embodiments further comprise sequencing the captured DNA using methods such as those disclosed herein.
[0266] The present methods can be used to diagnose presence of conditions, particularly cancer or precancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may decrease the amount of copy number variation or rare mutations detected in subject’s blood if the treatment is successful as there will be fewer cancer cells to shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy.
[0267] Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.
[0268] The types and number of cancers that may be detected may include blood cancers, brain cancers, lung cancers, skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, bowel cancers, rectal cancers, colon cancers, prostate cancers, thyroid cancers, bladder cancers, head and neck cancers, kidney cancers, mouth cancers, stomach cancers, solid state tumors, heterogeneous tumors, homogenous tumors and the like. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, recombination, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0269] In some embodiments, a method described herein comprises identifying the presence of nucleic acids, such as DNA, produced by a tumor (or neoplastic cells, or cancer cells) or by precancer cells.
[0270] Genetic data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.
[0271] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a profile of DNA derived from the subject, wherein the profile comprises a plurality of data resulting from epigenetic and mutation analyses. In some embodiments, an abnormal condition is cancer or precancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.
[0272] The present methods can be used to generate a profile, fingerprint or set of data that is a summation of information derived from different cells in a heterogeneous disease. This set of data may comprise structural variation identities and levels, copy number variation, epigenetic variation, or other mutation analyses alone or in combination.
[0273] The present methods can be used to diagnose, prognose, monitor or observe pre-cancers, cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such arenot directed to non-invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.Samples
[0274] In some embodiments, a sample can be any biological sample isolated from a subject. A sample can be a bodily sample. Samples can include body tissues or fluids, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, gingival crevicular fluid, bone marrow, pleural effusions, pleura fluid, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, and urine. Samples are preferably body fluids, particularly blood and fractions thereof, cerebrospinal fluid, pleura fluid, saliva, sputum, or urine. A sample can be in the form originally isolated from a subject or can have been subjected to further processing to remove or add components, such as cells, or enrich for one component relative to another. Thus, a preferred body fluid for analysis is plasma or serum containing cell-free nucleic acids.
[0275] In some embodiments, a population of nucleic acids is obtained from a serum, plasma or blood sample from a subject suspected of having neoplasia, a tumor, precancer, or cancer or previously diagnosed with neoplasia, a tumor, precancer, or cancer. The population includes nucleic acids having varying levels of sequence variation, epigenetic variation, and / or post replication or transcriptional modifications. Post-replication modifications include modifications of cytosine, particularly at the 5- position of the nucleobase, e.g., 5-methylcytosine, 5- hydroxymethylcytosine, 5- formylcytosine and 5-carboxylcytosine.
[0276] In some embodiments, a sample can be isolated or obtained from a subject and transported to a site of sample analysis. The sample may be preserved and shipped at a desirable temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. A sample can be isolated or obtained from a subject at the site of the sample analysis.The subject can be a human, a mammal, an animal, a companion animal, a service animal, or a pet. The subject may have a cancer, precancer, infection, transplant rejection, or other disease or disorder related to changes in the immune system. The subject may not have cancer or a detectable cancer symptom. The subject may have been treated with one or more cancer therapy, e.g., any one or more of chemotherapies, antibodies, vaccines or biologies. The subject may be in remission. The subject may or may not be diagnosed of being susceptible to cancer or any cancer- associated genetic mutations / disorders.
[0277] In some embodiments, the sample comprises plasma. The volume of plasma obtained can depend on the desired read depth for sequenced regions.Exemplary volumes are 0.4-40 ml, 5-20 ml, 10-20 ml. For examples, the volume can be 0.5 mL, 1 mL, 5 mL 10 mL, 20 mL, 30 mL, or 40 mL. A volume of sampled plasma may be 5 to 20 mL.
[0278] In some embodiments, a sample is a nucleic acid composition. A sample can comprise various amount of nucleic acid that contains genome equivalents. For example, a sample of about 30 ng DNA can contain about 10,000 (104) haploid human genome equivalents and, in the case of cfDNA, about 200 billion (2xlOu) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents and, in the case of cfDNA, about 600 billion individual molecules.
[0279] A sample can comprise nucleic acids from different sources, e.g., from cells and cell-free of the same subject, from cells and cell-free of different subjects. A sample can comprise nucleic acids carrying mutations. For example, a sample can comprise DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations existing in germline DNA of a subject. Somatic mutations refer to mutations originating in somatic cells of a subject, e.g., cancer cells. A sample can comprise DNA carrying cancer-associated mutations (e.g., cancer- associated somatic mutations). A sample can comprise an epigenetic variant (i.e., a chemical or protein modification), wherein the epigenetic variant associated with the presence of a genetic variant such as a cancer-associated mutation. In some embodiments, thesample comprises an epigenetic variant associated with the presence of a genetic variant, wherein the sample does not comprise the genetic variant.
[0280] Exemplary amounts of cell-free nucleic acids in a sample before amplification range from about 1 fg to about 1 pg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can comprise obtaining 1 femtogram (fg) to 200 ng.
[0281] Cell -free DNA refers to DNA not contained within a cell at the time of its isolation from a subject. For example, cfDNA can be isolated from a sample as the DNA remaining in the sample after removing intact cells, without lysing the cells or otherwise extracting intracellular DNA. Cell- free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single- stranded, or a hybrid thereof. A cell-free nucleic acid can be released into bodily fluid through secretion or cell death processes, e.g., cellular necrosis and apoptosis. Some cell-free nucleic acids are released into bodily fluid from cancer cells e.g., circulating tumor DNA, (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA) In some embodiments, cell free nucleic acids are produced by tumor cells. In some embodiments, cell free nucleic acids are produced by a mixture of tumor cells and nontumor cells.
[0282] Cell -free nucleic acids have an exemplary size distribution of about 100- 500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of molecules, with a mode of about 168 nucleotides and a second minor peak in a range between 240 to 440 nucleotides.
[0283] In some embodiments, cell-free nucleic acids can be isolated from bodily fluids through a fractionation or partitioning step in which cell-free nucleic acids, as found in solution, are separated from intact cells and other non-soluble components of the bodily fluid. Partitioning may include techniques such as centrifugation or filtration. In some embodiments, cells in bodily fluids can be lysed and cell-free and cellular nucleic acids processed together. In some embodiments, after addition of buffers and wash steps, nucleic acids can be precipitated with an alcohol. In some embodiments, further clean up steps may be used such as silica based columns to remove contaminants or salts. In some embodiments, non-specific bulk carrier nucleic acids, such as C 1 DNA, DNA or protein for bisulfite sequencing, hybridization, and / or ligation, may be added throughout the reaction to optimize certain aspects of the procedure such as yield.
[0284] After such processing, samples can include various forms of nucleic acid including double stranded DNA, single stranded DNA and single stranded RNA. In some embodiments, single stranded DNA and RNA can be converted to double stranded forms so they are included in subsequent processing and analysis steps.Applications
[0285] The present methods can be used to diagnose presence of conditions, particularly cancer or precancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or any other somatic mutation detected in subject's blood if the treatment is successful as more cancers may die and shed nucleic acids. In other examples, this may not occur. In another example, certain treatment options may be correlated with profiles (e.g., genetic profiles) of cancers over time. This correlation may be useful in selecting a therapy. In some embodiments, hypermethylation variable epigenetic target regions are analyzed to determine whether they show hypermethylation characteristic of tumor cells or cells that do not ordinarily contributesignificantly to cfDNA and / or hypomethylation variable epigenetic target regions are analyzed to determine whether they show hypomethylation characteristic of tumor cells or cells that do not ordinarily contribute significantly to cfDNA.
[0286] In some embodiments, the present methods are used for screening for a cancer, or in a method for screening cancer. For example, in some embodiments, the sample can be from a subject who has not been previously diagnosed with cancer. In some embodiments, the subject may or may not have cancer. In some embodiments, the subject may or may not have an early-stage cancer. In some embodiments, the subject has one or more risk factors for cancer, such as tobacco use (e.g., smoking), being overweight or obese, having a high body mass index (BMI), being of advanced age, poor nutrition, high alcohol consumption, or a family history of cancer.
[0287] In some embodiments, the subject has used tobacco, e.g., for at least 1 , 5, 10, or 15 years. In some embodiments, the subject has a high BMI, e.g., a BMI of 25 or greater, 26 or greater, 27 or greater, 28 or greater, 29 or greater, or 30 or greater. In some embodiments, the subject is at least 40, 45, 50, 55, 60, 65, 70, 75, or 80 years old. In some embodiments, the subject has poor nutrition, e.g., high consumption of one or more of red meat and / or processed meat, trans fat, saturated fat, and refined sugars, and / or low consumption of fruits and vegetables, complex carbohydrates, and / or unsaturated fats. High and low consumption can be defined, e.g., as exceeding or falling below, respectively, recommendations in Dietary Guidelines for Americans 2020- 2025, available at www.dietaryguidelines.gov / sites / default / files / 2021- 03 / Dietary _Guidelines_for_Americans-2020-2025.pdf . In some embodiments, the subject has high alcohol consumption, e.g., at least three, four, or five drinks per day on average (where a drink is about one ounce or 30 mL of 80-proof hard liquor or the equivalent). In some embodiments, the subject has a family history of cancer, e.g., at least one, two, or three blood relatives were previously diagnosed with cancer. In some embodiments, the relatives are at least third-degree relatives (e.g., great-grandparent, great aunt or uncle, first cousin), at least second- degree relatives (e.g., grandparent, aunt or uncle, or halfsibling), or first-degree relatives (e.g., parent or full sibling).
[0288] Additionally, if a cancer is observed to be in remission after treatment, in some embodiments, the present methods can be used to monitor residual disease or recurrence of disease.
[0289] In some embodiments, the methods and systems disclosed herein may be used to identify customized or targeted therapies to treat a given disease or condition in patients based on the presence of one or more proteins of interest and / or classification of a nucleic acid variant as being of somatic or germline origin. In some embodiments, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, head and neck cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, gliomas, astrocytomas, breast carcinoma, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal carcinoma, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinomas, gastrointestinal stromal tumors (GISTs), endometrial carcinoma, endometrial stromal sarcomas, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder carcinomas, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinomas, Wilms tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, Lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphomas, nonHodgkin lymphoma, diffuse large B-cell lymphoma, Mantle cell lymphoma, T cell lymphomas, non- Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T cell lymphomas, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral cavity squamous cell carcinomas, osteosarcoma, ovarian carcinoma, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasms, acinar cell carcinomas. Prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine carcinomas, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma. Type and / orstage of cancer can be detected from genetic variations including mutations, rare mutations, indels, rearrangements, copy number variations, transversions, translocations, recombinations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0290] In some embodiments, methods described herein comprise identifying the presence of target regions and / or DNA produced by a tumor (or neoplastic cells, or cancer cells) or by precancer cells. In some embodiments, a method described herein comprises determining the level of target regions and / or identifying the presence of DNA produced by a tumor (or neoplastic cells, or cancer cells) or by precancer cells. In some embodiments, determining the level of target regions comprises determining either an increased level or decreased level of target regions, wherein the increased or decreased level of target regions is determined by comparing the level of target regions with a threshold level / value.
[0291] In some embodiments, genetic data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. In some embodiments, genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. In some embodiments, this information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. In some embodiments, the system and methods of this disclosure may be useful in determining disease progression.
[0292] In some embodiments, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of nucleic acids derived from the subject,wherein the genetic profile comprises a plurality of data resulting from sequencevariable and epigenetic analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.
[0293] In some embodiments, the present methods can be used to diagnose, prognose, monitor, or observe cancers, precancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non-invasive prenatal testing. In some embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor, or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.
[0294] Non-limiting examples of other genetic-based diseases, disorders, or conditions that are optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha- 1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cri du chat, Crohn's disease, cystic fibrosis, Dercum disease, down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay- Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson disease, or the like.
[0295] In some embodiments, methods described herein comprise detecting a presence or absence of a nucleic acid (e.g., DNA, such as cfDNA) originating or derivedfrom a tumor cell at a preselected timepoint following a previous cancer treatment of a subject previously diagnosed with cancer. In some embodiments, the method may further comprise determining a cancer recurrence score that is indicative of the presence or levels of DNA originating or derived from the tumor cell for the subject.
[0296] Where a cancer recurrence score is determined, it may further be used to determine a cancer recurrence status. The cancer recurrence status may be at risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. The cancer recurrence status may be at low or lower risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. In some embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status of either at risk for cancer recurrence or at low or lower risk for cancer recurrence.
[0297] In some embodiments, a cancer recurrence score is compared with a predetermined cancer recurrence threshold, and the subject is classified as a candidate for a subsequent cancer treatment when the cancer recurrence score is above the cancer recurrence threshold or not a candidate for therapy when the cancer recurrence score is below the cancer recurrence threshold. In particular embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for a subsequent cancer treatment or not a candidate for therapy.Methods of determining a risk of cancer recurrence in a subject and / or classifying a subject as being a candidate for a subsequent cancer treatment
[0298] In some embodiments, methods provided herein are used to determine a risk of cancer recurrence in a subject. In some embodiments, a method provided herein is a method of classifying a subject as being a candidate for a subsequent cancer treatment.
[0299] In some embodiments, such methods may comprise collecting a sample from the subject diagnosed with the cancer at one or more preselected timepoints following one or more previous cancer treatments to the subject. In some embodiments, the sample may comprise DNA, e.g., cfDNA. The DNA may be obtained from a tissue sample or a liquid sample.
[0300] In some embodiments, methods of determining a risk of cancer recurrence in a subject may comprise determining a cancer recurrence score that is indicative of the presence or absence, or amount, of type-specific target regions originating or derived from the tumor cell for the subject. The cancer recurrence score may further be used to determine a cancer recurrence status. The cancer recurrence status may be at risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. The cancer recurrence status may be at low or lower risk for cancer recurrence, e.g., when the cancer recurrence score is above a predetermined threshold. In particular embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status of either at risk for cancer recurrence or at low or lower risk for cancer recurrence.
[0301] In some embodiments, methods of classifying a subject as being a candidate for a subsequent cancer treatment may comprise comparing the cancer recurrence score of the subject with a predetermined cancer recurrence threshold, thereby classifying the subject as a candidate for the subsequent cancer treatment when the cancer recurrence score is above the cancer recurrence threshold or not a candidate for therapy when the cancer recurrence score is below the cancer recurrence threshold. In particular embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for a subsequent cancer treatment or not a candidate for therapy. In some embodiments, the subsequent cancer treatment comprises chemotherapy or administration of a therapeutic composition.
[0302] In some embodiments, such methods may comprise determining a disease-free survival (DFS) period for the subject based on the cancer recurrence score; for example, the DFS period may be 1 year, 2 years, 3, years, 4 years, 5 years, or 10 years.
[0303] In some embodiments, the set of sequence information comprises sequence-variable target region sequences and determining the cancer recurrence score may comprise determining at least a first subscore indicative of the levels of particular immune cell types, SNVs, insertions / deletions, CNVs, and / or fusions present in sequence-variable target region sequences.
[0304] In some embodiments, a number of mutations in the sequence-variable target regions chosen from 1 , 2, 3, 4, or 5 is sufficient for the first subscore to result in a cancer recurrence score classified as positive for cancer recurrence. In some embodiments, the number of mutations is chosen from 1 , 2, or 3.
[0305] In some embodiments, the set of sequence information comprises epigenetic target region sequences, and determining the cancer recurrence score comprises determining a second subscore indicative of the amount of molecules (obtained from the epigenetic target region sequences) that represent an epigenetic state different from DNA found in a corresponding sample from a healthy subject (e.g., cfDNA found in a blood sample from a healthy subject, or DNA found in a tissue sample from a healthy subject where the tissue sample is of the same type of tissue as was obtained from the test subject). These abnormal molecules (i.e. , molecules with an epigenetic state different from DNA found in a corresponding sample from a healthy subject) may be consistent with epigenetic changes associated with cancer, e.g., methylation of hypermethylation variable target regions and / or perturbed fragmentation of fragmentation variable target regions, where “perturbed” means different from DNA found in a corresponding sample from a healthy subject.
[0306] In some embodiments, a proportion of molecules corresponding to the hypermethylation variable target region set and / or fragmentation variable target region set that indicate hypermethylation in the hypermethylation variable target region set and / or abnormal fragmentation in the fragmentation variable target region set greater than or equal to a value in the range of 0.001 %-10% is sufficient for the second subscore to be classified as positive for cancer recurrence. The range may be 0.001 %- 1 %, 0.005%-l%, 0.01 %-5%, 0.01 %-2%, or 0.01 %-1 %.
[0307] In some embodiments, any of such methods may comprise determining a fraction of tumor DNA from the fraction of molecules in the set of sequence information that indicate one or more features indicative of origination from a tumor cell. In some embodiments, this may be done for molecules corresponding to some or all of the epigenetic target regions, e.g., including one or both of hypermethylation variable target regions and fragmentation variable target regions (hypermethylation of a hypermethylation variable target region and / or abnormal fragmentation of afragmentation variable target region may be considered indicative of origination from a tumor cell). In some embodiments, this may be done for molecules corresponding to sequence variable target regions, e.g., molecules comprising alterations consistent with cancer, such as SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA may be determined based on a combination of molecules corresponding to epigenetic target regions and molecules corresponding to sequence variable target regions.
[0308] In some embodiments, determination of a cancer recurrence score may be based at least in part on the fraction of tumor DNA, wherein a fraction of tumor DNA greater than a threshold in the range of 10 11 to 1 or 10 10 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, a fraction of tumor DNA greater than or equal to a threshold in the range of 1 CT10 to 1 CT9, 1 CT9 to 1 CT8, 1 CT8 to 1 CT7, 1 CT7 to 1 CT6, 1 CT6 to 1 CT5, 1 CT5 to 1 CT4, 10Ato 1 CT3, ICT3 to KG2, or 1 CT2 to 10_1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, the fraction of tumor DNA greater than a threshold of at least 107 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. A determination that a fraction of tumor DNA is greater than a threshold, such as a threshold corresponding to any of the foregoing embodiments, may be made based on a cumulative probability. For example, the sample was considered positive if the cumulative probability that the tumor fraction was greater than a threshold in any of the foregoing ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999. In some embodiments, the probability threshold is at least 0.95, such as 0.99.
[0309] In some embodiments, the set of sequence information comprises sequence-variable target region sequences and epigenetic target region sequences, and determining the cancer recurrence score comprises determining a first sub score indicative of the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in sequence-variable target region sequences and a second subscore indicative of the amount of abnormal molecules in epigenetic target region sequences, and combining the first and second subscores to provide the cancer recurrence score. In some embodiments, where the first and second subscores are combined, they may becombined by applying a threshold to each subscore independently (e.g., greater than a predetermined number of mutations (e.g., > 1 ) in sequence-variable target regions, and greater than a predetermined fraction of abnormal molecules (i.e. , molecules with an epigenetic state different from the DNA found in a corresponding sample from a healthy subject; e.g., tumor) in epigenetic target regions), or training a machine learning classifier to determine status based on a plurality of positive and negative training samples.
[0310] In some embodiments, a value for the combined score in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.
[0311] In some embodiments, where a cancer recurrence score is classified as positive for cancer recurrence, the cancer recurrence status of the subject may be at risk for cancer recurrence and / or the subject may be classified as a candidate for a subsequent cancer treatment.
[0312] In some embodiments, the cancer is any one of the types of cancer described elsewhere herein, e.g., colorectal cancer.Therapies and Related Administration
[0313] In certain embodiments, the methods disclosed herein relate to identifying and administering customized therapies to patients. In some embodiments, determination of the levels of particular nucleic acids facilitates selection of appropriate treatment. In some embodiments, the patient or subject has a given disease, disorder, or condition. In some embodiments, any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. In certain embodiments, the therapy administered to a subject comprises at least one chemotherapy drug. In some embodiments, the chemotherapy drug may comprise alkylating agents (for example, but not limited to, Chlorambucil, Cyclophosphamide, Cisplatin and Carboplatin), nitrosoureas (for example, but not limited to, Carmustine and Lomustine), anti-metabolites (for example, but not limited to, Fluorauracil, Methotrexate and Fludarabine), plant alkaloids and natural products (for example, but not limited to, Vincristine, Paclitaxel and Topotecan), anti- tumor antibiotics (for example, but notlimited to, Bleomycin, Doxorubicin and Mitoxantrone), hormonal agents (for example, but not limited to, Prednisone, Dexamethasone, Tamoxifen and Leuprolide) and biological response modifiers (for example, but not limited to, Herceptin and Avastin, Erbitux and Rituxan). In some embodiments, the chemotherapy administered to a subject may comprise FOLFOX or FOLFIRI. In certain embodiments, a therapy may be administered to a subject that comprises at least one PARP inhibitor. In certain embodiments, the PARP inhibitor may include OLAPARIB, TALAZOPARIB, RUCAPARIB, NIRAPARIB (trade name ZEJULA), among others. Typically, therapies include at least one immunotherapy (or an immunotherapeutic agent). Immunotherapy refers generally to methods of enhancing an immune response against a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer.
[0314] In some embodiments, therapy is customized based on the status of a nucleic acid variant as being of somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, customized therapies include at least one immunotherapy (or an immunotherapeutic agent). Immunotherapy refers generally to methods of enhancing an immune response against a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer.
[0315] In some embodiments, the immunotherapy or immunotherapeutic agents targets an immune checkpoint molecule. Certain tumors are able to evade the immune system by co-opting an immune checkpoint pathway. Thus, targeting immune checkpoints has emerged as an effective approach for countering a tumor’s ability to evade the immune system and activating anti-tumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012, 12:252-264, which is incorporated by reference herein in its entirety.
[0316] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule that reduces a signal involved in the T cell response to antigen. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (aka B7.1 ) or CD86 (aka B7.2) on antigen presenting cells. PD-1 isanother inhibitory checkpoint molecule that is expressed on T cells. PD-1 limits the activity of T cells in peripheral tissues during an inflammatory response. In addition, the ligand for PD-1 (PD-L1 or PD-L2) is commonly upregulated on the surface of many different tumors, resulting in the downregulation of anti tumor immune responses in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for PD-1 , such as PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for CTLA4, such as CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin like receptor (KIR), T cell membrane protein 3 (TIM3), galectin 9 (GAIN), or adenosine A2a receptor (A2aR).
[0317] Antagonists that target these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Accordingly, in certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the antagonist of the inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti- CTLA4, anti-PD-1 , anti- PD-LI, or anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal anti-PD-1 antibody. In some embodiments, the antibody is a monoclonal anti-PD- LI antibody. In certain embodiments, the monoclonal antibody is a combination of an anti- CTLA4 antibody and an anti-PD-1 antibody, an anti-CTLA4 antibody and an anti-PD-LI antibody, or an anti-PD-LI antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In certain embodiments, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In certain embodiments, the anti-PD-LI antibody is one or more of atezolizumab (Tecentriq®), avelumab (Bavencio®), or durvalumab (Imfinzi®).
[0318] In certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist (e.g. antibody) against CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other embodiments, the antagonist is a soluble version of the inhibitory immunecheckpoint molecule, such as a soluble fusion protein comprising the extracellular domain of the inhibitory immune checkpoint molecule and an Fc domain of an antibody. In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1 , PD-L1 , or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In some embodiments, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.
[0319] In certain embodiments, the immune checkpoint molecule is a costimulatory molecule that amplifies a signal involved in a T cell response to an antigen. For example, CD28 is a co stimulatory receptor expressed on T cells. When a T cell binds to antigen through its T cell receptor, CD28 binds to CD80 (aka B7.1 ) or CD86 (aka B7.2) on antigen-presenting cells to amplify T cell receptor signaling and promote T cell activation. Because CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 is able to counteract or regulate the co-stimulatory signaling mediated by CD28. In certain embodiments, the immune checkpoint molecule is a co stimulatory molecule selected from CD28, inducible T cell co-stimulator (ICOS), CD137, 0X40, or CD27. In other embodiments, the immune checkpoint molecule is a ligand of a costimulatory molecule, including, for example, CD80, CD86, B7RP1 , B7-H3, B7-H4, CD137L, OX40L, or CD70.
[0320] Agonists that target these co-stimulatory checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Accordingly, in certain embodiments, the immunotherapy or immunotherapeutic agent is an agonist of a co-stimulatory checkpoint molecule. In certain embodiments, the agonist of the co- stimulatory checkpoint molecule is an agonist antibody and preferably is a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti -0X40, or anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80, anti-CD86, anti-B7RPI, anti-B7-H3, anti-B7-H4, anti-CD137L, anti-OX40L, or anti-CD70 antibody.
[0321] In certain embodiments, the status of a nucleic acid variant from a sample from a subject as being of somatic or germline origin may be compared with adatabase of comparator results from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the subject and / or patients who are receiving, or who have received, the same therapy as the subject. A customized or targeted therapy (or therapies) may be identified when the nucleic variant and the comparator results satisfy certain classification criteria (e.g., are a substantial or an approximate match).
[0322] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing an immunotherapeutic agent are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by any method known in the art, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraauricular, which administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like.
[0323] Therapeutic options for treating specific genetic-based diseases, disorders, or conditions, other than cancer, are generally well-known to those of ordinary skill in the art and will be apparent given the particular disease, disorder, or condition under consideration.
[0324] In some embodiments, e.g., where genetic variants are detected, therapy is customized based on the status of a nucleic acid variant as being of somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, customized therapies include at least one immunotherapy (or an immunotherapeutic agent). Immunotherapy refers generally to methods of enhancing an immune response against a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer.
[0325] In certain embodiments, the status of a nucleic acid variant from a sample from a subject as being of somatic or germline origin may be compared with adatabase of comparator results from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the subject and / or patients who are receiving, or who have received, the same therapy as the subject. A customized or targeted therapy (or therapies) may be identified when the nucleic variant and the comparator results satisfy certain classification criteria (e.g., are a substantial or an approximate match).
[0326] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing an immunotherapeutic agent are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by methods such as, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraauricular, which administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like.Kits
[0327] Also provided are kits comprising the compositions as described herein. The kits can be for use in performing the methods as described herein. In some embodiments, a kit comprises a plurality of target-specific probes. In some embodiments, the plurality of target-specific probes comprises probes comprising a capture moiety that hybridize to target regions having a type-specific epigenetic variation and a copy number variation. In some embodiments, the kit comprises a solid support linked to a binding partner of the capture moiety. In some embodiments, the kit comprises adapters. In some embodiments, the kit comprises primers that anneal to an adapter. In some embodiments, the kit comprises additional elements elsewhere herein. In some embodiments, the kit comprises instructions for performing a method described herein.
Claims
CLAIMS1. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. hybridizing a probe to the single stranded DNA, wherein the probe a first member of a binding pair, to produce a probe-annealed DNA; c. enriching the probe-annealed DNA using a second member of the binding pair; d. ligating the probe-annealed DNA with a directional adapter to produce an adapter- ligated DNA; e. contacting the adapter-ligated DNA with a DNA polymerase and deoxynucleoside triphosphates to produce a double stranded primer-extended sample; and f. sequencing the primer-extended sample.
2. The method of claim 1, wherein the probe comprises a 5’ primer ligation block.
3. The method of claim 1 or claim 2, wherein the probe comprises a biotinylated probe.
4. The method of any one of claims 1-3, wherein the probe-annealed DNA is treated with an exonuclease.
5. The method of claim 4, wherein the exonuclease is a 3’ to 5’ exonuclease.
6. The method of any one of claims 1-5, further comprising adding a poly-A tail to the probe- annealed DNA.
7. The method of any one of claims 1-6, wherein the second member of the binding pair is streptavidin and the streptavidin is attached to magnetic beads.
8. The method of any one of claims 1-7, wherein the directional adapter comprises a motor protein.
9. The method of any one of claims 1-8, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
10. The method of any one of claims 1-9, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
11. The method of any one of claims 1-10, wherein the single stranded DNA is stabilized by single-stranded binding proteins.Probe-capture method 2:
12. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. hybridizing a probe to the single stranded DNA, wherein the probe a first member of a binding pair, to produce a probe-annealed DNA; c. enriching the probe-annealed DNA using a second member of the binding pair; d. ligating the probe-annealed DNA with a directional adapter to produce an adapter- ligated DNA; e. contacting the adapter-ligated DNA with a DNA polymerase and deoxynucleoside triphosphates to produce a double stranded primer-extended sample; and f. sequencing the primer-extended sample.
13. The method of claim 12, wherein the probe comprises a 5’ primer ligation block.
14. The method of claim 12 or claim 13, wherein the probe comprises a biotinylated probe.
15. The method of claim 12 or claim 13, wherein the probe further comprises a splint region that is not complementary to the single stranded DNA.
16. The method of any one of claims 12-5, wherein the probe- annealed DNA is treated with an exonuclease.
17. The method of claim 16, wherein the exonuclease is a 3’ to 5’ exonuclease.
18. The method of any one of claims 12-17, wherein the second member of the binding pair is streptavidin and the streptavidin is attached to magnetic beads.
19. The method of any one of claims 12-18, wherein the directional adapter comprises a motor protein.20 The method of any one of claims 12-19, wherein the directional adapter comprises a splint overhang that is complementary to the splint region.
21. The method of any one of claims 12-20, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
22. The method of any one of claims 12-21, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
23. The method of any one of claims 12-22, wherein the single stranded DNA is stabilized by single- stranded binding proteins.
24. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. hybridizing a probe to the single stranded DNA, wherein the probe a first member of a binding pair, to produce a probe-annealed DNA; c. enriching the probe-annealed DNA using a second member of the binding pair; d. ligating the probe-annealed DNA with a directional adapter to produce an adapter- ligated DNA; e. contacting the adapter-ligated DNA with a DNA polymerase and deoxynucleoside triphosphates to produce a double stranded primer-extended sample; and f. sequencing the primer-extended sample.
25. The method of claim 24, wherein the probe comprises a 5’ primer ligation block.
26. The method of claim 24 or claim 25, wherein the probe comprises a biotinylated probe.
27. The method of claim 24 or claim 25, wherein the probe further comprises a splint region that is not complementary to the single stranded DNA.
28. The method of claim 27, wherein the splint region further comprises a hybridized oligomer.
29. The method of claim 28, wherein the hybridized oligomer comprises digestible bases.
30. The method of any one of claims 24-29, wherein the probe-annealed DNA is treated with an exonuclease.
31. The method of claim 30, wherein the exonuclease is a 3’ to 5’ exonuclease.
32. The method of any one of claims 24-31, wherein the second member of the binding pair is streptavidin and the streptavidin is attached to magnetic beads.
33. The method of any one of claims 24-32, wherein the directional adapter comprises a motor protein.
34. The method of any one of claims 24-33, wherein the directional adapter comprises a splint overhang that is complementary to the splint region.
35. The method of any one of claims 24-34, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
36. The method of any one of claims 24-35, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
37. The method of any one of claims 24-36, wherein the single stranded DNA is stabilized by single-stranded binding proteins.
38. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. hybridizing a probe to the single stranded DNA, wherein the probe is a first member of a binding pair, and wherein the probe comprises a splint region that is not complementary to the single stranded DNA, wherein the splint region comprises a hybridized oligomer, to produce a probe-annealed DNA; c. ligating the hybridized oligomer to the 3’ end of the single stranded DNA; d. enriching the probe-annealed DNA using a second member of the binding pair; e. eluting the probe-annealed DNA; f. ligating the probe-annealed DNA with a directional adapter to produce an adapter- ligated DNA; and g. sequencing the adapter-ligated DNA.
39. The method of claim 38, wherein the probe comprises a 5’ primer ligation block.
40. The method of claim 38 or claim 39, wherein the probe comprises a biotinylated probe.
41. The method of any one of claims 38-40, wherein the probe-annealed DNA is treated with an exonuclease.
42. The method of claim 41, wherein the exonuclease is a 3’ to 5’ exonuclease.
43. The method of any one of claims 38-42, wherein the second member of the binding pair is streptavidin and the streptavidin is attached to magnetic beads.
44. The method of any one of claims 38-43, wherein the directional adapter comprises a motor protein.
45. The method of any one of claims 38-44, wherein the directional adapter comprises a splint overhang that complementary to the hybridized oligomer.
46. The method of any one of claims 38-45, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
47. The method of any one of claims 38-46, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
48. The method of any one of claims 38-47, wherein the single stranded DNA is stabilized by single-stranded binding proteins.
49. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. hybridizing a probe to the single stranded DNA, wherein the probe a first member of a binding pair, to produce a probe-annealed DNA; c. contacting the probe-annealed DNA with a DNA polymerase and deoxynucleoside triphosphates to produce a double stranded primer-extended sample; d. enriching the primer-extended sample using a second member of the binding pair; e. ligating the primer-extended sample with a directional adapter to produce an adapter- ligated DNA; f. contacting the adapter-ligated DNA with a DNA polymerase to produce an adapter- extended sample; and g. sequencing the adapter -extended sample.
50. The method of claim 49, wherein the probe comprises a 5’ primer ligation block.
51. The method of claim 49 or claim 50, wherein the probe comprises a biotinylated probe.
52. The method of any one of claims 49-51, wherein the probe-annealed DNA is treated with an exonuclease.
53. The method of claim 52, wherein the exonuclease is a 3’ to 5’ exonuclease.
54. The method of any one of claims 49-53, further comprising adding a poly-A tail to the primer-extended sample.
55. The method of any one of claims 49-54, wherein the second member of the binding pair is streptavidin and the streptavidin is attached to magnetic beads.
56. The method of any one of claims 49-55, wherein the directional adapter comprises a motor protein.
57. The method of any one of claims 49-56, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
58. The method of any one of claims 49-57, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
59. The method of any one of claims 49-58, wherein the single stranded DNA is stabilized by single-stranded binding proteins.
60. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. hybridizing a primer to the single stranded DNA to produce a primer-annealed DNA; c. ligating the primer-annealed DNA with a directional adapter to produce an adapter- ligated DNA; and d. sequencing the adapter-ligated DNA.
61. The method of claim 60, wherein the primer-annealed DNA is treated with an exonuclease.
62. The method of claim 61, wherein the exonuclease is a 3’ to 5’ exonuclease.
63. The method of any one of claims 60-62, wherein the directional adapter comprises a motor protein.
64. The method of any one of claims 60-63, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
65. The method of any one of claims 60-64, wherein the single stranded DNA is stabilized by single-stranded binding proteins.
66. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. ligating the single stranded DNA with a directional adapter, wherein the directional adapter comprises an overhang sequence, thereby producing adapter-ligated DNA; and c. sequencing the adapter-ligated DNA.
67. A method for sequencing a nucleic acid composition, comprising: a. denaturing the nucleic acid composition to produce a single stranded DNA; b. ligating the single stranded DNA with a directional adapter, wherein the directional adapter comprises an overhang sequence to produce an adapter-ligated DNA; c. contacting the adapter-ligated DNA with a DNA polymerase and deoxynucleoside triphosphates to produce a primer-extended sample; and d. sequencing the primer-extended sample.
68. The method of any one of claims 66-67, wherein the single stranded DNA is treated by endrepair.
69. The method of any one of claims 66-68, wherein the single stranded DNA is stabilized by single-stranded binding proteins.
70. The method of any one of claims 66-69, wherein the 3’ DNA ends are blocked.71 . The method of any one of claims 66-70, wherein the 5’ DNA ends are blocked.
72. A method for sequencing a single stranded DNA composition, comprising: a. adding a poly A tail to the single stranded DNA; b. annealing a poly-T oligonucleotide to the poly- A tail of the single stranded DNA to produce 3’ extended DNA; c. treating the 3’ extended DNA with an exonuclease; d. ligating the 3’ extended DNA with a directional adapter to produce an adapter-ligated DNA; and f. sequencing the adapter-ligated DNA.
73. The method of claim 72, wherein the exonuclease is a 3’ to 5’ exonuclease.
74. A method for sequencing a single stranded DNA composition, comprising: a. extending the single stranded DNA with a polymerase to produce a 3’ extended DNA; b. annealing a tail-complementary oligonucleotide to the 3’ extended DNA; c. annealing the 3’ extended DNA with a directional adapter, wherein the directional adapter comprises a region complementary to the 3’ extended DNA to produce an adapter-ligated DNA; and d. sequencing the adapter-ligated DNA.
75. The method of claim 74, further comprising adding a poly-A tail to the 3’ extended DNA.
76. The method of any one of claims 74-75, wherein the polymerase is a TdT polymerase.
77. The method of any one of claims 74-76, wherein the polymerase is a template-independent polymerase.
78. A method for sequencing a nucleic acid composition, comprising: a. ligating a double-stranded DNA with a 3’ nanopore adapter to produce 3’ nanopore adapter ligated DNA; b. denaturing the 3’ nanopore adapter ligated DNA to produce a single stranded nanopore adapter ligated DNA; c. hybridizing a probe to the single stranded nanopore adapter ligated DNA, wherein the probe is one part of a binding pair, to produce a probe-annealed DNA; c. enriching the probe-annealed DNA using the second part of the binding pair; d. contacting the probe-annealed DNA with a DNA polymerase and deoxynucleoside triphosphates to produce a double stranded primer-extended sample; e. loading a nanopore motor protein to the primer-extended sample; and f. sequencing the primer-extended sample.
79. The method of claim 78, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
80. The method of any one of claims 78-79, wherein the denaturing is preceded by end-repair of the nucleic acid composition.81 . The method of any one of claims 78-80, further comprising adding a poly-A tail to the nucleic acid composition.
81. A method for sequencing a nucleic acid composition, comprising: a. ligating a double-stranded DNA with an LP adapter to produce LP adapter ligated DNA; b. denaturing the LP adapter ligated DNA to produce a single stranded LP adapter ligated DNA; c. hybridizing a probe to the single stranded LP adapter ligated DNA, wherein the probe is the first member of a binding pair, to produce a probe-annealed DNA; c. enriching the probe-annealed DNA using a second member of the binding pair; d. hybridizing a universal primer to the LP adapter to produce primer adapted DNA; e. contacting the primer adapted DNA with a DNA polymerase and deoxynucleoside triphosphates to produce a double stranded primer-extended sample; f. ligating the primer-extended sample with a directional adapter to produce an adapter- ligated DNA; andg. sequencing the adapter-ligated DNA.
82. The method of claim 81, wherein the LP adapter comprises a sample index region.
83. The method of claim 81, wherein the LP adapter comprises a flanking constant region.
84. The method of claim 81, wherein the LP adapter comprises an internal base modification.
85. The method of claim 81, wherein the LP adapter comprises a 3’ overhang region.
86. The method of any one of claims 81-85, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
87. The method of claim 86, wherein the DNA polymerase is intolerant to the internal base modification.
88. The method of any one of claims 81-87, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
89. The method of any one of claims 81-87, further comprising adding a poly-A tail to the nucleic acid composition.
90. The method of any one of claims 81-87, wherein the second member of the binding pair is streptavidin and the streptavidin is attached to magnetic beads.
91. A method for sequencing a nucleic acid composition, comprising: a. ligating a double-stranded DNA with an LP adapter to produce LP adapter ligated DNA; b. denaturing the LP adapter ligated DNA to produce a single stranded LP adapter ligated DNA; c. hybridizing a probe to the single stranded LP adapter ligated DNA, wherein the probe is a first member of a binding pair, to produce a probe-annealed DNA; c. enriching the probe-annealed DNA using a second member of the binding pair; d. hybridizing a universal primer to the LP adapter to produce primer adapted DNA; e. ligating the primer adapted DNA with a directional adapter to produce an adapter- ligated DNA; and f. sequencing the adapter-ligated DNA.
92. The method of claim 91, wherein the LP adapter comprises a sample index region.
93. The method of claim 91, wherein the LP adapter comprises a flanking constant region.
94. The method of claim 91, wherein the LP adapter comprises a 3’ overhang region.
95. The method of any one of claims 91-94, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
96. The method of any one of claims 91-95, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
97. The method of any one of claims 91-96, further comprising adding a poly-A tail to the nucleic acid composition.
98. The method of any one of claims 91-97, wherein the second member of the binding pair is streptavidin and the streptavidin is attached to magnetic beads.
99. A method for sequencing a nucleic acid composition, comprising: a. ligating a double-stranded DNA with an LP adapter, wherein the LP adapter comprises a 3’ overhang region, to produce LP adapter ligated DNA; b. denaturing the LP adapter ligated DNA to produce a single stranded LP adapter ligated DNA; c. hybridizing a probe to the single stranded LP adapter ligated DNA, wherein the probe is the first member of a binding pair, to produce a probe-annealed DNA; c. enriching the probe-annealed DNA using a second member of the binding pair; d. hybridizing a universal primer to the LP adapter to produce primer adapted DNA; e. ligating the primer adapted DNA with a directional adapter to produce an adapter- ligated DNA; and f. sequencing the adapter-ligated DNA.
100. The method of claim 99, wherein the LP adapter comprises a sample index region.
101. The method of claim 99, wherein the LP adapter comprises a flanking constant region.
102. The method of any one of claims 99-101, wherein the DNA polymerase is a 5’ to 3’ exonuclease positive, strand displacement positive DNA polymerase.
103. The method of any one of claims 99-102, wherein the denaturing is preceded by end-repair of the nucleic acid composition.
104. The method of any one of claims 99-103, further comprising adding a poly-A tail to the nucleic acid composition.
105. The method of any one of claims 99-104, wherein the directional adapter comprises an overhang region that is complimentary to the 3’ overhang region of the LP adapter.
106. The method of any one of claims 1-105, wherein the nucleic acid composition comprises genomic DNA.
107. The method of any one of claims 1-106, wherein the nucleic acid composition comprises cell-free DNA.
108. The method of any one of claims 1-71 or 78-107, wherein the nucleic acid composition is enriched prior to denaturization.
109. The method of claim 108, wherein the enrichment comprises dead restriction enzyme enrichment, methylated DNA binding protein enrichment, dCas9-IP enrichment, or ChIP enrichment.
110. The method of any one of claims 1-109, wherein the single stranded DNA composition is produced by denaturing a nucleic acid composition.
Citation Information
Patent Citations
Methods for non-invasive assessment of genetic alterations
EP4235676A2
Methods for simultaneous molecular and sample barcoding
WO2023023402A2
Compositions and methods for synthesis and use of probes targeting nucleic acid rearrangements
WO2023056065A1