Compositions and methods for analyzing cell-free DNA in a methylation profiling assay
By partitioning and differentially treating cell-free DNA samples based on cytosine modification ratios, the method addresses the challenge of low concentration and heterogeneity in liquid biopsies, enhancing the detection of epigenetic changes for improved cancer diagnosis.
Patent Information
- Application Number
- JP2022519473
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-30
- Filing Date
- 2020-09-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-09-30
AI Technical Summary
Existing methods for analyzing cell-free DNA in liquid biopsies struggle to accurately detect nucleobase modifications due to low concentration and heterogeneity, limiting the ability to provide detailed information on epigenetic changes in cancer detection.
A method involving partitioning DNA samples into secondary samples with varying cytosine modification ratios, followed by procedures that differentially affect nucleobases, and sequencing to distinguish between modified and unmodified nucleobases, providing combined information on methylation levels and specific modifications.
Enhances the detection of epigenetic modifications in cell-free DNA, allowing for more accurate analysis of cancer-related DNA, improving early cancer detection through non-invasive liquid biopsies.
Smart Images

Figure 0007717057000008 
Figure 0007717057000009 
Figure 0007717057000010
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 908,569, filed Sep. 30, 2019, which is incorporated herein by reference in its entirety for all purposes.
Background Art
[0002] Field of the Invention The present disclosure provides compositions and methods related to analyzing DNA, such as cell - free DNA. In some embodiments, the cell - free DNA is DNA from a subject having or suspected of having cancer and / or the cell - free DNA comprises DNA from cancer cells. In some embodiments, the DNA is distributed into a first subsample and a second subsample, the first subsample contains DNA having a higher ratio of nucleotide modifications (e.g., cytosine modifications) than the second subsample, the first subsample is subjected to a procedure that affects a first nucleobase in the DNA to be different from a second nucleobase in the DNA of the first subsample, and the DNA is sequenced to distinguish the first nucleobase in the DNA of the first subsample from the second nucleobase.
Summary of the Invention
Means for Solving the Problems
[0003] Introduction and Overview Cancer is the cause of millions of deaths every year worldwide. Early detection of cancer can lead to improved outcomes because early - stage cancer tends to be more sensitive to treatment.
[0004] Inappropriately controlled cell proliferation is a characteristic of cancer, which is generally due to the accumulation of genetic and epigenetic changes, such as copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), cytosine modifications (e.g., 5-methylcytosine, 5-hydroxymethylcytosine, and other more oxidized forms), and epigenetic variations including the association of DNA with chromatin proteins and transcription factors.
[0005] Biopsy represents a conventional approach for detecting or diagnosing cancer, which involves extracting cells or tissues from a site that may be cancerous and analyzing them for relevant phenotypic and / or genotypic characteristics. Biopsy has the drawback of being invasive.
[0006] Cancer detection based on the analysis of body fluids, such as blood ("liquid biopsy"), is an interesting alternative based on the finding that DNA from cancer cells is released into body fluids. Liquid biopsy is non-invasive (sometimes only requiring blood sampling). However, due to the low concentration and heterogeneity of cell-free DNA, it has been a challenge to develop accurate and sensitive methods for analyzing liquid biopsy materials that provide detailed information on nucleobase modifications. Isolating and processing fractions of cell-free DNA that are useful for further analysis in liquid biopsy procedures is an important part of these methods. Therefore, there is a need for improved methods and compositions for analyzing cell-free DNA, for example, in liquid biopsy.
[0007] The present disclosure is based in part on the following recognition. It can be beneficial to analyze modifications of nucleobases (including, inter alia, methylation and / or hydroxymethylation of cytosine) along with other process steps, such as partitioning and sequencing based on the degree of methylation. For example, in an exemplary embodiment, a DNA sample (e.g., a cfDNA sample) is partitioned into a plurality of secondary samples having different amounts of cytosine methylation (e.g., based on binding to MBD (methyl-binding domain or methyl-binding protein) or an antibody specific for methylated cytosine), and then the secondary samples containing a high level of methylation are subjected to procedures that differentially affect different types of a given nucleobase (e.g., non-modified cytosine and methylated cytosine, or hydroxymethylated cytosine and methylated cytosine). Next, sequencing can be performed to identify the sequences in the first and second secondary samples and / or to identify the positions in the DNA from the first secondary sample where a particular species of nucleobase was present. Such methods according to the present disclosure can provide more information regarding epigenetic modifications in DNA, e.g., cfDNA, than existing approaches such as MeDIP-seq, MBD-seq, BS-seq, Ox-BS-seq, TAP-seq, ACE-seq, hmC-seal, and TAB-seq.See, for example, Schutsky, E.K. et al. Nondestructive, base-resolution sequencing of 5-hydroxymethylcytosine using a DNA deaminase. Nature Biotech, 2018; doi.10.1038 / nbt.4204 (ACE-Seq); Yu, Miao et al. Base-resolution analysis of 5-hydroxymethylcytosine in the Mammalian Genome. Cell, 2012; 149(6):1368-80 (TAB-Seq); Han, D. A highly sensitive and robust method for genome-wide 5hmC profiling of rare cell populations. Mol Cell. 2016; 63(4):711-719 (5hmC-Seal); Shen, S.Y. et al. Sensitive tumour detection and classification using plasma cell free DNA methylomes. Nature. 2018; 563(7732):579-583 (cfMeDIP); Nair, SS et al. Comparison of methyl-DNA immunoprecipitation (MeDIP) and methyl-CpG binding domain (MBD) protein capture for genome-wide DNA. Epigenetics. 2011; 6(1):34-44. Unlike such existing methods, the methods according to the present disclosure can provide a combination of a first modification by a partitioning step, e.g., information regarding methylation levels, and specific modifications by procedures differentially affecting different types of a given nucleobase and / or additional information regarding its location. Examples of such procedures include bisulfite, substituted borane, base-modifying enzymes, or various conversion or separation steps using modified base-specific antibodies that distinguish between different species of a class of nucleobases.In some embodiments, the methods described herein provide a combination of information regarding (i) the overall level of modification of a molecule (e.g., cytosine modification) (e.g., based on its distribution), and (ii) high-resolution information regarding the identity and / or location of specific modifications (e.g., based on the specific conversion of a particular modified or unmodified nucleotide, or on post-assignment sequencing that differentiates between particular types of modifications, as discussed in detail herein).
[0008] The method may further comprise capturing a set of two target regions from DNA. In some embodiments, the set of target regions includes a set of sequence-variable target regions and a set of epigenetic target regions. Each of these sets can provide information useful for determining the likelihood that a sample contains DNA from cancer cells. In some embodiments, the capture yield of the set of sequence-variable target regions is higher than the capture yield of the set of epigenetic target regions. The difference in capture yields can allow for deeper, and thus more accurate, sequencing of the set of sequence-variable target regions, and shallower but broader coverage of the set of epigenetic target regions, for example, during simultaneous sequencing, e.g., in the same sequencing cell or the same pool of material being sequenced.
[0009] Epigenetic target region sets can be analyzed in various ways. For example, if the acceptable confidence regarding the modification at a particular position is lower than the acceptable confidence regarding the accuracy in the set of array variable target regions (e.g., when the goal is to understand the different types of frequencies of modifications at various loci and it is not necessary to understand the exact position where the modification occurs), the analysis can use methods that do not rely on high accuracy in the sequencing of specific nucleotides within the target. Examples include determining the degree of modification, such as methylation, and / or the distribution and size of the fragments, which can indicate normal or abnormal chromatin structures in the cells from which the fragments were obtained. Such analysis can be performed by sequencing and requires less data (e.g., the number of sequence reads or the depth of sequencing coverage) than when determining the presence or absence of sequence variations such as base substitutions, insertions, or deletions.
[0010] This disclosure aims to meet the need for improved analysis of cell-free DNA and / or provide other advantages. Accordingly, the following exemplary embodiments are provided.
[0011] Embodiment 1 is a method for analyzing DNA in a sample, comprising: a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample includes DNA having nucleotide modifications (e.g., cytosine modifications) at a higher ratio than the second secondary sample; b) subjecting the first secondary sample to a procedure that affects the first nucleobase in the DNA of the first secondary sample to be different from the second nucleobase in the DNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and c) Sequencing the DNA in the first secondary sample and the DNA in the second secondary sample such that a first nucleobase in the DNA of the first secondary sample is distinguished from a second nucleobase which is a method comprising.
[0012] Embodiment 2 is the method according to Embodiment 1, wherein the DNA comprises cell-free DNA (cfDNA) obtained from a test subject.
[0013] Embodiment 3 is a method for analyzing DNA in a sample comprising cell-free DNA (cfDNA) obtained from a test subject, a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample comprises DNA having a higher ratio of nucleotide modification (e.g., cytosine modification) than the second secondary sample; b) subjecting the first secondary sample to a procedure that affects a first nucleobase in the DNA so as to be different from a second nucleobase in the DNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and c) sequencing at least the DNA in the first secondary sample such that a first nucleobase in the DNA of the first secondary sample is distinguished from a second nucleobase which is a method comprising.
[0014] Embodiment 4 is the method according to Embodiment 3, wherein step c) comprises sequencing at least the DNA in the first secondary sample.
[0015] Embodiment 5 is a method for analyzing a sample comprising cell-free DNA (cfDNA), a) A step of distributing a sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains cfDNA having a higher ratio of nucleotide modification (e.g., cytosine modification) than the second secondary sample; b) A step of subjecting the first secondary sample to a procedure that affects the first nucleobase in the cfDNA of the first secondary sample to be different from the second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; c) A step of capturing at least an epigenetic target region set of cfDNA from the first and second secondary samples, thereby providing the captured cfDNA; and d) A step of sequencing the captured cfDNA to distinguish the first nucleobase in the cfDNA from the first secondary sample from the second nucleobase A method comprising the above steps.
[0016] Embodiment 6 is a method for analyzing a sample containing cell-free DNA (cfDNA), a) A step of distributing a sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains cfDNA having a higher ratio of nucleotide modification (e.g., cytosine modification) than the second secondary sample; b) A step of subjecting the first secondary sample to a procedure that affects the first nucleobase in the cfDNA of the first secondary sample to be different from the second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; c) Capturing a set of multiple target regions of cfDNA from the first and second secondary samples, thereby providing the captured cfDNA, wherein the set of multiple target regions includes a set of sequence-variable target regions and a set of epigenetic target regions; and d) Sequencing the captured cfDNA to distinguish a first nucleobase in the captured cfDNA from the first secondary sample from a second nucleobase A method comprising:
[0017] Embodiment 7 is the method according to embodiment 6, wherein cfDNA molecules corresponding to the set of sequence-variable target regions are captured in the sample with a higher capture yield than cfDNA molecules corresponding to the set of epigenetic target regions.
[0018] Embodiment 8 is a method for isolating cell-free DNA (cfDNA) from a sample, comprising: a) Distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains cfDNA having a higher ratio of nucleotide modification (e.g., cytosine modification) than the second secondary sample; b) Subjecting the first secondary sample to a procedure that affects a first nucleobase in the cfDNA to be different from a second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; c) Contacting the cfDNA of the first and second secondary samples with a set of target-specific probes, wherein the set of target-specific probes includes a target-binding probe specific for the set of sequence-variable targets and a target-binding probe specific for the set of epigenetic targets, thereby forming a complex of the target-specific probe and the cfDNA; Separating the complex from cfDNA not bound to the target-specific probe, thereby providing captured cfDNA corresponding to a set of variable sequence targets and cfDNA corresponding to a set of epigenetic targets; and d) Sequencing the captured cfDNA to distinguish a first nucleobase in the cfDNA from a first secondary sample from a second nucleobase A method comprising.
[0019] Embodiment 9 is the method according to embodiment 8, wherein the target-specific probe set is configured to capture cfDNA corresponding to a set of variable sequence targets with a higher capture yield than cfDNA corresponding to a set of epigenetic targets.
[0020] Embodiment 10 is a method for identifying the presence of DNA produced by a tumor, comprising a) Collecting a cfDNA sample from a subject, b) Distributing the cfDNA sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample comprises captured cfDNA having a higher proportion of nucleotide modifications (e.g., cytosine modifications) than the second secondary sample; c) Subjecting the first secondary sample to a procedure that affects a first nucleobase in the cfDNA to be different from a second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first and second nucleobases have the same base pairing specificity; d) Capturing a set of a plurality of target regions from the cfDNA in the first and second secondary samples, thereby providing a sample comprising the captured cfDNA, wherein the set of a plurality of target regions comprises a set of variable sequence target regions and a set of epigenetic target regions; and e) sequencing the captured cfDNA in the first secondary sample and the captured cfDNA in the second secondary sample to distinguish a first nucleobase in the cfDNA of the first secondary sample from a second nucleobase which is a method comprising the same.
[0021] Embodiment 11 is the method according to Embodiment 10, wherein cfDNA molecules corresponding to a set of sequence-variable target regions are captured in a sample with a higher capture yield than cfDNA molecules corresponding to a set of epigenetic target regions.
[0022] Embodiment 12 is the method according to any one of Embodiments 6 to 11, comprising sequencing cfDNA molecules corresponding to a set of sequence-variable target regions to a higher sequencing depth than cfDNA molecules corresponding to a set of epigenetic target regions.
[0023] Embodiment 13 is the method according to Embodiment 12, wherein the captured cfDNA molecules of the sequence-variable target set are sequenced to a sequencing depth at least 2-fold higher than the captured cfDNA molecules of the epigenetic target region set.
[0024] Embodiment 14 is the method according to Embodiment 12, wherein the captured cfDNA molecules of the sequence-variable target set are sequenced to a sequencing depth at least 3-fold higher than the captured cfDNA molecules of the epigenetic target region set.
[0025] Embodiment 15 is the method according to Embodiment 12, wherein the captured cfDNA molecules of the sequence-variable target set are sequenced to a sequencing depth 4 to 10-fold higher than the captured cfDNA molecules of the epigenetic target region set.
[0026] Embodiment 16 is the method according to embodiment 12, wherein the captured cfDNA molecules of the set of variable targets are sequenced to a sequencing depth 4 to 100 times higher than the captured cfDNA molecules of the set of epigenetic target regions.
[0027] Embodiment 17 is the method according to any one of embodiments 6 to 16, wherein the captured cfDNA molecules of the set of variable targets and the captured cfDNA molecules of the set of epigenetic target regions are sequenced in the same sequencing cell.
[0028] Embodiment 18 is the method according to any one of the foregoing embodiments, wherein the DNA is amplified before sequencing, or the method includes a capture step and the DNA is amplified before the capture step.
[0029] Embodiment 19 further includes the step of ligating a barcode-containing adapter to the DNA before capture, and optionally the ligation step is performed before or simultaneously with the amplification, which is the method according to embodiments 5 to 18.
[0030] Embodiment 20 is the method according to any one of embodiments 5 to 19, wherein the set of epigenetic target regions includes a set of hypermethylated variable target regions.
[0031] Embodiment 21 is the method according to any one of embodiments 5 to 20, wherein the set of epigenetic target regions includes a set of hypomethylated variable target regions.
[0032] Embodiment 22 is the method according to embodiment 20 or 21, wherein the set of epigenetic target regions includes a set of methylation control target regions.
[0033] Embodiment 23 is the method according to any one of embodiments 5 to 22, wherein the set of epigenetic target regions includes a set of fragmented variable target regions.
[0034] Embodiment 24 is the method according to Embodiment 23, wherein the fragmented variable target region set includes a transcription start site region.
[0035] Embodiment 25 is the method according to Embodiment 23 or 24, wherein the fragmented variable target region set includes a CTCF binding region.
[0036] Embodiment 26 is the method according to any one of Embodiments 5 to 25, wherein the step of capturing a set of a plurality of target regions of cfDNA includes contacting the cfDNA with a target binding probe specific for the variable sequence target region set and a target binding probe specific for the epigenetic target region set.
[0037] Embodiment 27 is the method according to Embodiment 26, wherein the target binding probe specific for the variable sequence target region set is present at a higher concentration than the target binding probe specific for the epigenetic target region set.
[0038] Embodiment 28 is the method according to Embodiment 26, wherein the target binding probe specific for the variable sequence target region set is present at a concentration at least 2 times higher than the target binding probe specific for the epigenetic target region set.
[0039] Embodiment 29 is the method according to Embodiment 26, wherein the target binding probe specific for the variable sequence target region set is present at a concentration at least 4 or 5 times higher than the target binding probe specific for the epigenetic target region set.
[0040] Embodiment 30 is the method according to embodiment 26, wherein the target binding probe specific to the array-variable target region set is present at a concentration at least 10 times, 20 times, 30 times, 40 times, 50 times, 60 times, 70 times, 80 times, 90 times, or 100 times higher than the target binding probe specific to the epigenetic target region set, or the target binding probe specific to the array-variable target region set is present at a concentration in the range of 2 to 3, 3 to 4, 4 to 5, 5 to 7, 7 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 times the concentration of the target binding probe specific to the epigenetic target region set.
[0041] Embodiment 31 is the method according to any one of embodiments 26 to 30, wherein the target binding probe specific to the array-variable target region set has a higher target binding affinity than the target binding probe specific to the epigenetic target region set.
[0042] Embodiment 32 is the method according to any one of embodiments 6 to 32, wherein the footprint of the epigenetic target region set is at least twice as large as the size of the array-variable target region set.
[0043] Embodiment 33 is the method according to embodiment 32, wherein the footprint of the epigenetic target region set is at least 10 times as large as the size of the array-variable target region set.
[0044] Embodiment 34 is the method according to any one of embodiments 6 to 33, wherein the footprint of the array-variable target region set is at least 25 kB or 50 kB.
[0045] Embodiment 35 is the method according to any one of the foregoing embodiments, wherein the step of distributing the sample into a plurality of secondary samples includes the step of distributing based on the methylation level.
[0046] Embodiment 36 is the method according to embodiment 35, wherein the step of distributing includes the step of contacting the collected cfDNA with a methyl-binding reagent immobilized on a solid support.
[0047] Embodiment 37 is the method according to any one of the foregoing embodiments, including the step of differentially tagging a first secondary sample and a second secondary sample.
[0048] Embodiment 38 is the method according to embodiment 37, wherein the first secondary sample and the second secondary sample are differentially tagged before the first secondary sample is subjected to a procedure that affects a first nucleobase in the DNA so as to be different from a second nucleobase in the DNA of the first secondary sample.
[0049] Embodiment 39 is the method according to embodiment 37 or 38, wherein the first secondary sample and the second secondary sample are pooled after the first secondary sample is subjected to a procedure that affects a first nucleobase in the DNA so as to be different from a second nucleobase in the DNA of the first secondary sample.
[0050] Embodiment 40 is the method according to any one of embodiments 37 to 39, wherein the first secondary sample and the second secondary sample are sequenced in the same sequencing cell.
[0051] Embodiment 41 is the method according to any one of the foregoing embodiments, wherein the plurality of secondary samples includes a third secondary sample containing DNA having a nucleotide modification (e.g., cytosine modification) at a ratio higher than that of the second secondary sample but lower than that of the first secondary sample.
[0052] Embodiment 42 is the method according to embodiment 41, further including the step of differentially tagging the third secondary sample so as to be distinguishable from the first secondary sample and the second secondary sample.
[0053] Embodiment 43 is the method according to Embodiment 42, wherein the first secondary sample is subjected to a procedure that affects the first nucleobase in the DNA so as to be different from the second nucleobase in the DNA of the first secondary sample, and then the first, second, and third secondary samples are combined, and if necessary, the first, second, and third secondary samples are sequenced in the same sequencing cell.
[0054] Embodiment 44 is the method according to any one of the preceding embodiments, wherein the procedure to which the first secondary sample is subjected changes the base pairing specificity of the first nucleobase without substantially changing the base pairing specificity of the second nucleobase.
[0055] Embodiment 45 is the method according to any one of the preceding embodiments, wherein the first nucleobase is modified cytosine or unmodified cytosine, and the second nucleobase is modified cytosine or unmodified cytosine.
[0056] Embodiment 46 is the method according to any one of the preceding embodiments, wherein the first nucleobase contains unmodified cytosine (C).
[0057] Embodiment 47 is the method according to any one of the preceding embodiments, wherein the second nucleobase contains 5-methylcytosine (mC).
[0058] Embodiment 48 is the method according to any one of the preceding embodiments, wherein the procedure to which the first secondary sample is subjected includes bisulfite conversion.
[0059] Embodiment 49 is the method according to any one of Embodiments 1 to 46, wherein the first nucleobase contains mC.
[0060] Embodiment 50 is the method according to any one of the preceding embodiments, wherein the second nucleobase contains 5-hydroxymethylcytosine (hmC).
[0061] Embodiment 51 is the method according to Embodiment 50, wherein the procedure to which the first secondary sample is subjected includes protection of 5hmC.
[0062] Embodiment 52 is the method according to Embodiment 50, wherein the procedure for providing the first secondary sample includes Tet-assisted bisulfite conversion.
[0063] Embodiment 53 is the method according to Embodiment 50, wherein the procedure for providing the first secondary sample includes Tet-assisted conversion with a substituted borane reducing agent, which is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane as needed.
[0064] Embodiment 54 is the method according to Embodiment 53, wherein the substituted borane reducing agent is 2-picoline borane or borane pyridine.
[0065] Embodiment 55 is the method according to any one of Embodiments 49 to 51 or 53 to 54, wherein the second nucleobase contains C.
[0066] Embodiment 56 is the method according to any one of Embodiments 49 to 51 or 55, wherein the procedure for providing the first secondary sample includes protection of hmC and subsequent Tet-assisted conversion with a substituted borane reducing agent, which is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane as needed.
[0067] Embodiment 57 is the method according to Embodiment 56, wherein the substituted borane reducing agent is 2-picoline borane or borane pyridine.
[0068] Embodiment 58 is the method according to any one of Embodiments 46, 47, 49 to 51, or 55, wherein the procedure for providing the first secondary sample includes protection of hmC and subsequent deamination of mC and / or C.
[0069] Embodiment 59 is the method according to Embodiment 58, wherein the deamination of mC and / or C includes treatment with an AID / APOBEC family DNA deaminase enzyme.
[0070] Embodiment 60 is the method according to any one of Embodiments 51 or 55-59, wherein the protection of hmC includes glucosylation of hmC.
[0071] Embodiment 61 is the method according to any one of Embodiments 1-45, 47, 49, or 55, wherein the procedure for subjecting the first secondary sample includes chemical-assisted conversion with a substituted borane reducing agent, which is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane as needed.
[0072] Embodiment 62 is the method according to Embodiment 61, wherein the substituted borane reducing agent is 2-picoline borane or borane pyridine.
[0073] Embodiment 63 is the method according to any one of Embodiments 1-45, 47, 49, 55, or 61-62, wherein the first nucleobase contains hmC.
[0074] Embodiment 64 is the method according to any one of Embodiments 1-44, wherein the procedure for subjecting the first secondary sample includes separating DNA containing the first nucleobase from DNA not containing the first nucleobase originally.
[0075] Embodiment 65 is the method according to Embodiment 64, wherein the first nucleobase is hmC.
[0076] Embodiment 66 is the method according to Embodiment 64 or 65, wherein the step of separating DNA containing the first nucleobase from DNA not containing the first nucleobase originally includes labeling the first nucleobase.
[0077] Embodiment 67 is the method according to Embodiment 66, wherein the labeling step includes biotinylation.
[0078] Embodiment 68 is the method according to Embodiment 66 or 67, wherein the labeling step includes glucosylation.
[0079] Embodiment 69 is the method according to Embodiment 68, wherein the step of glucosylating is the step of attaching a glucosyl-azide moiety.
[0080] Embodiment 70 is the method according to Embodiment 66 or 68, wherein the step of labeling comprises a step of glucosylating before biotinylating and a subsequent step of attaching a biotin moiety to the glucosyl.
[0081] Embodiment 71 is the method according to Embodiment 70, wherein the step of attaching a biotin moiety to the glucosyl comprises Huisgen cycloaddition chemistry.
[0082] Embodiment 72 is the method according to any one of Embodiments 54 to 61, wherein the step of separating DNA containing a first nucleobase from DNA not containing the first nucleobase from the beginning comprises a step of binding the DNA containing the first nucleobase to a capture agent.
[0083] Embodiment 73 is the method according to Embodiment 72, wherein the capture agent comprises a biotin binder, and optionally the biotin binder comprises avidin or streptavidin.
[0084] Embodiment 74 is the method according to any one of Embodiments 64 to 73, which comprises a step of differentially tagging each of DNA containing a first nucleobase from the beginning, DNA not containing the first nucleobase from the beginning, and DNA of a second secondary sample.
[0085] Embodiment 75 is the method according to Embodiment 74, which comprises a step of pooling after differentially tagging DNA containing a first nucleobase from the beginning, DNA not containing the first nucleobase from the beginning, and DNA of a second secondary sample, and optionally DNA containing a first nucleobase from the beginning, DNA not containing the first nucleobase from the beginning, and DNA of a second secondary sample are sequenced in the same sequencing cell.
[0086] Embodiment 76 is the method according to any one of Embodiments 1 to 44, wherein the first nucleobase is a modified adenine or an unmodified adenine, and the second nucleobase is a modified adenine or an unmodified adenine.
[0087] Embodiment 77 is the method according to any one of Embodiments 1 to 44, wherein the first nucleobase is a modified guanine or an unmodified guanine, and the second nucleobase is a modified guanine or an unmodified guanine.
[0088] Embodiment 78 is the method according to any one of Embodiments 1 to 44, wherein the first nucleobase is a modified thymine or an unmodified thymine, and the second nucleobase is a modified thymine or an unmodified thymine.
[0089] Embodiment 79 is the method according to any one of the foregoing embodiments, wherein the second subpopulation is not subjected to a procedure that affects the first nucleobase so as to be different from the second nucleobase.
[0090] Embodiment 80 is a combination comprising a first and a second population of captured DNA, wherein the first population comprises or is derived from DNA having a higher ratio of nucleotide modification (e.g., cytosine modification) than the second population, the type of the first nucleobase originally present in the DNA before changing the base pairing specificity is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, the type of the first nucleobase and the second nucleobase originally present in the DNA before changing the base pairing specificity have the same base pairing specificity, and the second population does not contain the type of the first nucleobase originally present in the DNA having a changed base pairing specificity.
[0091] Embodiment 81 is the combination according to Embodiment 80, wherein the first population comprises sequence tags selected from a first set of one or more sequence tags, the second population comprises sequence tags selected from a second set of one or more sequence tags, and the second set of sequence tags is different from the first set of sequence tags.
[0092] Embodiment 82 is the combination described in Embodiment 81, wherein the array tag contains a barcode.
[0093] Embodiment 83 is the combination described in any one of Embodiments 80 - 82, wherein the cytosine modification is methylation.
[0094] Embodiment 84 is the combination described in any one of Embodiments 80 - 83, wherein the first nucleobase is modified or unmodified cytosine, and the second nucleobase is modified or unmodified cytosine.
[0095] Embodiment 85 is the combination described in any one of Embodiments 80 - 84, wherein the first nucleobase contains unmodified cytosine (C).
[0096] Embodiment 86 is the combination described in any one of Embodiments 80 - 85, wherein the second nucleobase contains one or both of 5 - methylcytosine (mC) and 5 - hydroxymethylcytosine (hmC).
[0097] Embodiment 87 is the combination described in any one of Embodiments 80 - 86, wherein the first population has been subjected to bisulfite conversion.
[0098] Embodiment 88 is the combination described in any one of Embodiments 80 - 86, wherein the first nucleobase contains mC.
[0099] Embodiment 89 is the combination described in any one of Embodiments 80 - 88, wherein the second nucleobase contains hmC.
[0100] Embodiment 90 is the combination described in any one of Embodiments 80 - 89, wherein the first population contains protected hmC.
[0101] Embodiment 91 is the combination described in Embodiment 84 or 90, wherein the first population has been subjected to Tet - assisted bisulfite conversion.
[0102] Embodiment 92 is the combination according to embodiment 84 or 90, wherein the first population is subjected to Tet-assisted conversion with a substituted borane reducing agent that is 2-picolinylborane, borane pyridine, tert-butylamine borane, or ammonia borane as needed.
[0103] Embodiment 93 is the combination according to embodiment 90, wherein the first population is subjected to Tet-assisted conversion with a substituted borane reducing agent that is 2-picolinylborane, borane pyridine, tert-butylamine borane, or ammonia borane as needed after being subjected to protection of hmC.
[0104] Embodiment 94 is the combination according to any one of embodiments 80-82, 86, 88-90, or 92-93, wherein the second nucleobase contains C.
[0105] Embodiment 95 is the combination according to embodiment 90, wherein the first population is subjected to protection of hmC and subsequent deamination of mC and / or C.
[0106] Embodiment 96 is the combination according to any one of embodiments 90-95, wherein the protected hmC contains glucosylated hmC.
[0107] Embodiment 97 is the combination according to any one of embodiments 80-83, wherein the first nucleobase contains hmC.
[0108] Embodiment 98 is the combination according to any one of embodiments 80-83 or 97, wherein the second nucleobase contains mC.
[0109] Embodiment 99 is the combination according to any one of embodiments 80-83 or 97-98, wherein the second nucleobase contains C.
[0110] Embodiment 100 is a combination according to any one of Embodiments 80 to 83 or 97 to 99, wherein the first population is subjected to chemical substance-assisted conversion with a substituted borane reducing agent that is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane as needed.
[0111] Embodiment 101 is a combination according to any one of Embodiments 80 to 83, wherein the first nucleobase is a modified adenine or an unmodified adenine, and the second nucleobase is a modified adenine or an unmodified adenine.
[0112] Embodiment 102 is a combination according to any one of Embodiments 80 to 83, wherein the first nucleobase is a modified guanine or an unmodified guanine, and the second nucleobase is a modified guanine or an unmodified guanine.
[0113] Embodiment 103 is a combination according to any one of Embodiments 80 to 83, wherein the first nucleobase is a modified thymine or an unmodified thymine, and the second nucleobase is a modified thymine or an unmodified thymine.
[0114] Embodiment 104 is a combination comprising a first and a second population of captured DNA, wherein the first population contains, or is derived from, DNA having a higher ratio of nucleotide modification (e.g., cytosine modification) than the second population; the first population comprises first and second subpopulations; the first subpopulation contains the first nucleobase at a higher ratio than the second subpopulation; the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, the first nucleobase and the second nucleobase have the same base pairing specificity; when the first nucleobase is a modified thymine or an unmodified thymine, the second nucleobase is a modified thymine or an unmodified thymine, and the second population does not contain the first nucleobase.
[0115] Embodiment 105 is the combination described in Embodiment 104, wherein the first nucleobase is a modified cytosine or an unmodified cytosine, and the second nucleobase is a modified cytosine or an unmodified cytosine.
[0116] Embodiment 106 is the combination described in Embodiment 105, wherein the first nucleobase is a protected modified cytosine.
[0117] Embodiment 107 is the combination described in Embodiment 105 or 106, wherein the first nucleobase is a derivative of hmC.
[0118] Embodiment 108 is the combination described in Embodiment 107, wherein the first nucleobase is glucosylated hmC.
[0119] Embodiment 109 is the combination described in Embodiment 107 or Embodiment 108, wherein the first nucleobase is biotinylated hmC.
[0120] Embodiment 110 is the combination described in any one of Embodiments 106 to 109, wherein the first nucleobase is a product of the Huisgen cycloaddition of an affinity label to β-6-azido-glucosyl-5-hydroxymethylcytosine.
[0121] Embodiment 111 is the combination described in Embodiment 104, wherein the first nucleobase is a modified adenine or an unmodified adenine, and the second nucleobase is a modified adenine or an unmodified adenine.
[0122] Embodiment 112 is the combination described in Embodiment 104, wherein the first nucleobase is a modified guanine or an unmodified guanine, and the second nucleobase is a modified guanine or an unmodified guanine.
[0123] Embodiment 113 is the combination described in Embodiment 104, wherein the first nucleobase is a modified thymine or an unmodified thymine, and the second nucleobase is a modified thymine or an unmodified thymine.
[0124] Embodiment 114 is a combination according to any one of Embodiments 104 to 113, wherein the first subpopulation contains a first array tag, the second subpopulation contains a second array tag different from the first array tag, the second population contains a third array tag different from the first and second array tags, and optionally, the first, second, and / or third tags are barcodes.
[0125] Embodiment 115 is a combination according to any one of Embodiments 80 to 114, wherein the captured DNA contains cfDNA.
[0126] Embodiment 116 is a combination according to any one of Embodiments 80 to 115, wherein the captured DNA contains a sequence-variable target region and an epigenetic target region, the concentration of the sequence-variable target region is higher than the concentration of the epigenetic target region, and the concentration is normalized with respect to the footprint sizes of the sequence-variable target region and the epigenetic target region.
[0127] Embodiment 117 is a combination according to Embodiment 116, wherein the concentration of the sequence-variable target region is at least twice higher than the concentration of the epigenetic target region.
[0128] Embodiment 118 is a combination according to Embodiment 116, wherein the concentration of the sequence-variable target region is at least four or five times higher than the concentration of the epigenetic target region.
[0129] Embodiment 119 is a combination according to any one of Embodiments 116 to 118, wherein the concentration is a mass / volume concentration normalized with respect to the footprint size of the target region.
[0130] Embodiment 120 is a combination according to any one of Embodiments 116 to 119, wherein the epigenetic target region contains one, two, three, or four of a hypermethylated variable target region; a hypomethylated variable target region; a transcription start site region; and a CTCF binding region; and optionally, the epigenetic target region further contains a methylation control target region.
[0131] Embodiment 121 is a combination according to any one of Embodiments 80-120, produced according to the method described in any one of Embodiments 1-79.
[0132] Embodiment 122 is a communication interface that receives, via a communication network, a plurality of sequence reads generated from sequencing of DNA in a first secondary sample and DNA in a second secondary sample according to the method described in any one of Embodiments 1-79 by a nucleic acid sequencer; and when executed by at least one electronic processor, (i) receiving, via the communication network, the sequence reads generated by the nucleic acid sequencer; and (ii) mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads a controller including or accessible to a computer-readable medium including non-transitory computer-executable instructions for performing a method including is a system including.
[0133] Embodiment 123 is a communication interface that receives, via a communication network, a plurality of sequence reads generated from sequencing of a combination of first and second populations of captured DNA according to any one of Embodiments 80-121 by a nucleic acid sequencer; and when executed by at least one electronic processor, (i) receiving, via the communication network, the sequence reads generated by the nucleic acid sequencer; and (ii) mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads a controller including or accessible to a computer-readable medium including non-transitory computer-executable instructions for performing a method including A system including
[0134] Embodiment 124 is a method executed by at least one electronic processor, (iii) Processing the mapped array reads corresponding to the array variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer The system according to Embodiment 122 or 123, further including
[0135] Embodiment 125 is the method according to any one of Embodiments 1 to 79, further including the step of determining the likelihood that the subject has cancer.
[0136] Embodiment 126 is a method in which sequencing generates a plurality of array reads, the method includes mapping the plurality of array reads to one or more reference arrays to generate mapped array reads, and processing the mapped array reads corresponding to the array variable target region set and the mapped array reads corresponding to the epigenetic target region set to determine the likelihood that the subject has cancer, and further including the steps described in the foregoing embodiments.
[0137] Embodiment 127 is the method according to any one of Embodiments 1 to 79, where the test subject has already been diagnosed with cancer and has received one or more previous cancer treatments, and cfDNA is obtained at one or more preselected time points after one or more previous cancer treatments as needed.
[0138] Embodiment 128 is the method according to the foregoing embodiments, further including the step of sequencing the captured set of cfDNA molecules, thereby producing a set of sequence information.
[0139] Embodiment 129 is the method described in the foregoing embodiments, in which the captured DNA molecules of the array-variable target region set are sequenced to a higher sequencing depth than the captured DNA sequences of the epigenetic target region set.
[0140] Embodiment 130 is the method described in Embodiment 128 or 129, further comprising the step of detecting the presence or absence of DNA originating from or derived from tumor cells at a preselected time point using a set of sequence information.
[0141] Embodiment 131 is the method described in the foregoing embodiments, further comprising the step of determining a cancer recurrence score indicating the presence or absence of DNA originating from or derived from the tumor cells of the test subject.
[0142] Embodiment 132 further comprises the step of determining the cancer recurrence status based on the cancer recurrence score. When it is determined that the cancer recurrence score is at or above a predetermined threshold, the cancer recurrence status of the test subject is determined to be at risk of cancer recurrence, or when the cancer recurrence score is below the predetermined threshold, the cancer recurrence status of the test subject is determined to be at low risk of cancer recurrence. This is the method described in the foregoing embodiments.
[0143] Embodiment 133 further comprises the step of comparing the cancer recurrence score of the test subject with a predetermined cancer recurrence threshold. When the cancer recurrence score is above the cancer recurrence threshold, the test subject is classified as a candidate for subsequent cancer treatment, or when the cancer recurrence score is below the cancer recurrence threshold, the test subject is classified as not being a candidate for subsequent cancer treatment. This is the method described in Embodiment 131 or 132.
[0144] Embodiment 134 is the method described in any one of Embodiments 131 to 133, in which the test subject is classified as being at risk of cancer recurrence and as a candidate for subsequent cancer treatment.
[0145] Embodiment 135 is the method according to any one of Embodiments 131, 133, or 134, wherein subsequent cancer treatment involves chemotherapy or administration of a therapeutic composition.
[0146] Embodiment 136 is the method according to any one of Embodiments 132 to 135, wherein the DNA originating from or derived from tumor cells is cell-free DNA.
[0147] Embodiment 137 is the method according to any one of Embodiments 132 to 135, wherein the DNA originating from or derived from tumor cells is obtained from a tissue sample.
[0148] Embodiment 138 is the method according to any one of Claims 129 to 137, further comprising determining the disease-free survival (DFS) period of a test subject based on a cancer recurrence score.
[0149] Embodiment 139 is the method according to Embodiment 138, wherein the DFS period is 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.
[0150] Embodiment 140 is the method according to any one of Embodiments 128 to 139, wherein the sequence information set includes a sequence variable target region sequence, and the step of determining a cancer recurrence score includes determining at least a first subscore indicating the amount of SNV, insertion / deletion, CNV, and / or fusion present in the sequence variable target region sequence.
[0151] Embodiment 141 is the method according to Embodiment 140, wherein the number of mutations in the sequence variable target region selected from 1, 2, 3, 4, or 5 is sufficient to result in a cancer recurrence score in which the first subscore is classified as positive for cancer recurrence, and optionally the number of mutations is selected from 1, 2, or 3.
[0152] Embodiment 142 is the method according to any one of Embodiments 128 to 141, wherein the array information set includes an epigenetic target region array, and the step of determining the cancer recurrence score includes the step of determining a second subscore indicating the amount of abnormal array reads in the epigenetic target region array.
[0153] Embodiment 143 is the method according to Embodiment 142, wherein the abnormal array reads include reads indicating methylation of a hypermethylation variable target array and / or reads indicating abnormal fragmentation in a fragmentation variable target region.
[0154] Embodiment 144 is the method according to Embodiment 143, wherein the ratio of reads corresponding to the hypermethylation variable target region set and / or the fragmentation variable target region, which indicates hypermethylation in the hypermethylation variable target region set and / or abnormal fragmentation in the fragmentation variable target region set, is greater than or equal to a value in the range of 0.001% to 10% is sufficient to classify the second subscore as positive for cancer recurrence.
[0155] Embodiment 145 is the method according to Embodiment 144, wherein the range is 0.001% to 1% or 0.005% to 1%.
[0156] Embodiment 146 is the method according to Embodiment 144, wherein the range is 0.01% to 5% or 0.01% to 2%.
[0157] Embodiment 147 is the method according to Embodiment 144, wherein the range is 0.01% to 1%.
[0158] Embodiment 148 is the method according to any one of Embodiments 128 to 147, further including the step of determining the proportion of tumor DNA from the proportion of reads in an array information set showing one or more characteristics indicating origin from tumor cells.
[0159] Embodiment 149 is the method according to embodiment 148, wherein one or more features indicating origin from tumor cells include one or more of a change in a sequence-variable target region, hypermethylation of a hypermethylation-variable target region, and abnormal fragmentation of a fragmentation-variable target region.
[0160] Embodiment 150 further includes the step of determining a cancer recurrence score based at least in part on the proportion of tumor DNA, and a proportion of tumor DNA greater than or equal to a default value in the range of 10 -11 ~1 or 10 -10 ~1 is sufficient to classify the cancer recurrence score as positive with respect to cancer recurrence, which is the method according to embodiment 148 or 149.
[0161] Embodiment 151 is the method according to embodiment 150, wherein a proportion of tumor DNA greater than or equal to a default value in the range of 10 -10 ~10 -9 、10 -9 ~10 -8 、10 -8 ~10 -7 、10 -7 ~10 -6 、10 -6 ~10 -5 、10 -5 ~10 -4 、10 -4 ~10 -3 、10 -3 ~10 -2 、or 10 -2 ~10 -1 is sufficient to classify the cancer recurrence score as positive with respect to cancer recurrence.
[0162] Embodiment 152 is the method according to embodiment 150 or 151, wherein the default value is in the range of 10 -8 ~10 -6 or is 10 -7 -7 .
[0163] Embodiment 153 is the method according to any one of Embodiments 149 to 152, wherein when the cumulative probability that the proportion of tumor DNA is greater than or equal to a predetermined value is at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999, the proportion of tumor DNA is determined to be greater than or equal to the predetermined value.
[0164] Embodiment 154 is the method according to Embodiment 153, wherein the cumulative probability is at least 0.95.
[0165] Embodiment 155 is the method according to Embodiment 153, wherein the cumulative probability is in the range of 0.98 to 0.995 or is 0.99.
[0166] Embodiment 156 is the method according to any one of Embodiments 128 to 155, wherein the sequence information set includes a sequence variable target region sequence and an epigenetic target region sequence, and the step of determining the cancer recurrence score includes determining a first subscore indicating the amount of SNV, insertion / deletion, CNV, and / or fusion present in the sequence variable target region sequence, and a second subscore indicating the amount of abnormal sequence reads in the epigenetic target region sequence, and combining the first and second subscores to provide the cancer recurrence score.
[0167] Embodiment 157 is the method according to Embodiment 156, wherein the step of combining the first and second subscores includes applying a threshold (e.g., greater than a predetermined number of mutations in the sequence variable target region (e.g., >1) and greater than a predetermined proportion of abnormal (e.g., tumor) reads in the epigenetic target region) independently to each subscore, or training a machine learning classifier to determine a state based on a plurality of positive and negative training samples.
[0168] Embodiment 158 is the method according to Embodiment 157, wherein a combined score value in the range of -4 to 2 or -3 to 1 is sufficient to classify the cancer recurrence score as positive regarding cancer recurrence.
[0169] Embodiment 159 is the method according to any one of Embodiments 127 to 158, wherein one or more preselected time points are selected from the group consisting of 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 1.5 years, 2 years, 3 years, 4 years, and 5 years after the administration of one or more previous cancer treatments.
[0170] Embodiment 160 is the method according to any one of Embodiments 127 to 159, wherein the cancer is colorectal cancer.
[0171] Embodiment 161 is the method according to any one of Embodiments 127 to 160, wherein one or more previous cancer treatments include surgery.
[0172] Embodiment 162 is the method according to any one of Embodiments 127 to 161, wherein one or more previous cancer treatments include the administration of a therapeutic composition.
[0173] Embodiment 163 is the method according to any one of Embodiments 127 to 162, wherein one or more previous cancer treatments include chemotherapy.
[0174] Embodiment 164 is the combination according to any one of Embodiments 80 to 120, wherein the altered base specificity is generated by chemical conversion.
[0175] Embodiment 165 is the combination according to Embodiment 164, wherein the chemical conversion is selected from the group consisting of (i) bisulfite conversion, (ii) Tet-assisted bisulfite conversion, (iii) Tet-assisted conversion with a substituted borane reducing agent, and (iv) protection of hmC, followed by Tet-assisted conversion with a substituted borane reducing agent. I. BRIEF DESCRIPTION OF THE DRAWINGS
BRIEF DESCRIPTION OF THE DRAWINGS
[0176]
Figure 1-1
Figure 1-2
[0177]
Figure 2-1
Figure 2-2
[0178]
Figure 3-1
Figure 3-2
[0179]
Figure 4
[0180]
Figure 5
BRIEF DESCRIPTION OF THE DRAWINGS
[0181] II. DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS Certain embodiments of the invention are described in detail. The invention is described in connection with such embodiments, but it is understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the scope of the invention as defined by the appended claims.
[0182] Before describing the teachings of the invention in detail, it is understood that the present disclosure is not limited to the specific compositions or process steps as they may vary. It should be noted that, as used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include the plural unless the context clearly indicates otherwise. Thus, for example, a reference to “a nucleic acid” includes a plurality of nucleic acids, a reference to “a cell” includes a plurality of cells, and the like.
[0183] A numerical range includes the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant figures and errors associated with the measurement. Similarly, the use of "comprise", "comprises", "comprising", "contain", "contains", "containing", "include", "includes", and "including" is intended to be non-limiting. It is understood that both the foregoing general description and the detailed description are exemplary and explanatory only and do not limit the present teachings.
[0184] Unless specifically stated otherwise in the above specification, embodiments in this specification that enumerate various components as "comprising" are also contemplated as "consisting of" or "consisting essentially of" the enumerated components, and embodiments in this specification that enumerate various components as "consisting of" are also contemplated as "comprising" or "consisting essentially of" the enumerated components, and embodiments in this specification that enumerate various components as "consisting essentially of" are also contemplated as "consisting of" or "comprising" the enumerated components (this interchangeability does not apply to the use of these terms in the claims).
[0185] The section headings used in this specification are for organizational purposes and are not to be construed as limiting the disclosed subject matter in any way. If any document or other material incorporated by reference conflicts with the explicit content of this specification, including definitions, this specification shall control.
[0186] A. Definitions "Cell-free DNA", "cfDNA molecule", or simply "cfDNA" includes DNA molecules that naturally exist in a subject in an extracellular form (e.g., in blood, serum, plasma, or other body fluids such as lymph, cerebrospinal fluid, urine, or sputum). CfDNA was originally present in one or more cells of a large and complex biological organism, such as a mammal, but has been released from the cell(s) into the fluid found in the organism and can be obtained by obtaining a sample of the fluid without performing an in vitro cell lysis step.
[0187] As used herein, a modification or other feature is "present at a higher rate" in a first subsample or population of nucleic acids than in a second subsample or population if the proportion of nucleotides having the modification or other feature is higher in the first subsample or population than in the second subsample. For example, if one tenth of the nucleotides in a first subsample are mC and one twelfth of the nucleotides in a second subsample are mC, the first subsample contains the 5-methylated cytosine modification at a higher rate than the second subsample.
[0188] As used herein, "without substantially altering the base pairing specificity of a given nucleobase" means that most of the molecules containing the nucleobase that can be sequenced do not have a change in the base pairing specificity of the second nucleobase compared to its base pairing specificity when present in the original isolated sample. In some embodiments, 75%, 90%, 95%, or 99% of the molecules containing the nucleobase that can be sequenced do not have a change in the base pairing specificity of the second nucleobase compared to its base pairing specificity when present in the original isolated sample.
[0189] As used herein, "base pairing specificity" refers to the standard DNA base (A, C, G, or T) with which a given base most preferentially forms a pair. Thus, for example, unmodified cytosine and 5-methylcytosine have the same base pairing specificity (i.e., specificity for G), while uracil and cytosine have different base pairing specificities, since uracil has a base pairing specificity for A, whereas cytosine has a base pairing specificity for G. The ability of uracil to form wobble base pairs with G is irrelevant since uracil most preferentially pairs with A among the four standard DNA bases nonetheless.
[0190] As used herein, a "combination" comprising a plurality of members refers to either a single composition or a set of proximate compositions containing the members in, for example, separate containers or compartments within a larger container, such as a multi-well plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other storage format.
[0191] The "capture yield" of a collection of probes for a given target set refers to the amount of nucleic acid corresponding to the target set that is captured by the collection of probes under typical conditions (e.g., amount compared to another target set or absolute amount). Exemplary typical capture conditions are to incubate the sample nucleic acid and the probes at 65° C. for 10-18 hours in a small reaction volume (about 20 μL) containing a stringent hybridization buffer. The capture yield can be expressed in absolute terms or, in the case of a collection of multiple probes, in relative terms. When comparing the capture yields of sets of multiple target regions, they are normalized with respect to the footprint size of the target region set (e.g., based per kilobase). Thus, for example, if the footprint sizes of the first and second target regions are 50 kb and 500 kb, respectively (as a normalization factor of 0.1), and the mass / volume concentration of the captured DNA corresponding to the first target region set is greater than 0.1 times the mass / volume concentration of the captured DNA corresponding to the second target region set, the DNA corresponding to the first target region set is captured with a higher yield than the DNA corresponding to the second target region set. As a further example, using the same footprint size, if the captured DNA corresponding to the first target region set has a mass / volume concentration that is 0.2 times the mass / volume concentration of the captured DNA corresponding to the second target region set, the DNA corresponding to the first target region set is captured with a capture yield that is 2 times higher than the DNA corresponding to the second target region set.
[0192] "Capturing" one or more target nucleic acids refers to preferentially isolating or separating one or more target nucleic acids from non-target nucleic acids.
[0193] A "captured set" of nucleic acids refers to the nucleic acids that have undergone capture.
[0194] A "target region set" or "set of target regions" refers to a plurality of genomic loci that are targeted for capture and / or targeted by a set of probes (e.g., through sequence complementarity).
[0195] "Corresponding to a target region set" means that a nucleic acid, such as cfDNA, originates from a locus in the target region set or specifically binds to one or more probes for the target region set.
[0196] "Specifically binds" in the context of a probe or other oligonucleotide and a target sequence means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a replica thereof to form a stable probe:target hybrid such that the formation of stable probe:non-target hybrids is minimized. Thus, the probe hybridizes to the target sequence or a replica thereof to a greater extent than to non-target sequences in order to enable capture or detection of the target sequence. Appropriate hybridization conditions are well known in the art and can be predicted based on sequence composition or determined by using conventional testing methods (see, e.g., §§ 1.90-1.91, 7.37-7.57, 9.47-9.51, and 11.47-11.57 of Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), which is incorporated herein by reference, particularly §§ 9.50-9.51, 11.12-11.13, 11.45-11.47, and 11.55-11.57).
[0197] "Set of sequence-variable target regions" refers to a set of target regions that can exhibit sequence changes such as nucleotide substitutions (i.e., single nucleotide variants), insertions, deletions, or gene fusions or translocations in neoplastic cells (e.g., tumor cells and cancer cells).
[0198] An "epigenetic target region set" refers to a set of target regions that can exhibit sequence-independent modifications in neoplastic cells (e.g., tumor cells and cancer cells), or sequence-independent changes in cfDNA from a subject with cancer compared to cfDNA from a healthy subject. Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. For the purposes of the present invention, loci that are prone to local amplifications and / or gene fusions related to neoplasia, tumors, or cancer may also be included in the epigenetic target region set, because the detection of copy number changes by sequencing or fusion sequences mapping to more than one locus in the reference genome can be detected at a relatively shallow sequencing depth, for example local amplifications and / or gene fusions, in a manner more similar to the detection of the exemplary epigenetic changes discussed above than the detection of nucleotide substitutions, insertions, or deletions, since the detection does not depend on the basecall accuracy at one or a few individual positions. In some embodiments, the epigenetic target region set includes one or more genomic regions in which the epigenetic state (e.g., methylation state) of cfDNA molecules in these regions is invariant in cancer, but the presence / amount in the blood indicates an increased abnormal presentation of cfDNA into the circulation from a particular tissue (e.g., the origin of the cancer).
[0199] The nucleic acid is "produced by the tumor" or, if it originates from tumor cells, is ctDNA or circulating tumor DNA. Tumor cells are neoplastic cells that originate from the tumor, whether they remain within the tumor or become detached from the tumor (e.g., in the case of metastatic cancer cells and circulating tumor cells).
[0200] The term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to cytosine at a CpG site (a cytosine-phosphate-guanine site, i.e., where guanine follows cytosine in the 5’→3’ direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to the addition of a methyl group to adenine, e.g., N 6 -methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the carbon at position 5 of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine to yield 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxyl cytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the carbon at position 3 of cytosine). In some embodiments, 3C methylation includes adding a methyl group to the 3C position of cytosine to produce 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, when DNA in a promoter region is methylated, gene transcription can be repressed. DNA methylation is extremely important for normal development, and abnormalities in methylation can interfere with epigenetic regulation. Interference with epigenetic regulation, e.g., repression, can cause diseases such as cancer. Methylation of a promoter in DNA can indicate cancer.
[0201] The term "hypermethylated" refers to an increase in the level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules within a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypermethylated DNA can include DNA molecules that contain at least 1 methylated residue, at least 2 methylated residues, at least 3 methylated residues, at least 5 methylated residues, or at least 10 methylated residues.
[0202] The term "hypomethylated" refers to a decrease in the level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules within a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules that contain 0 methylated residues, at most 1 methylated residue, at most 2 methylated residues, at most 3 methylated residues, at most 4 methylated residues, or at most 5 methylated residues.
[0203] The terms "or a combination thereof" and "or combinations thereof", as used herein, refer to any and all permutations and combinations of the terms listed prior to those terms. For example, "A, B, C, or a combination thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and also includes BA, CA, CB, ACB, CBA, BCA, BAC, or CAB if order is important in a particular context. Continuing with this example, combinations containing repeats of one or more items or terms, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc. are clearly included. One of ordinary skill in the art will understand that there is no limit to the number of items or terms in any combination, except where it is clear from the context that otherwise is the case.
[0204] "Or" is used in an inclusive sense, i.e., is equivalent to "and / or" unless the context requires otherwise.
[0205] B. Exemplary Methods 1. Distribution of a sample into a plurality of secondary samples; sample modalities; analysis of epigenetic features In certain embodiments described herein, different types of nucleic acid populations (e.g., hypermethylated and hypomethylated DNA in a sample, such as the captured cfDNA sets described herein) can be physically distributed based on one or more characteristics of the nucleic acids prior to further analysis, e.g., prior to differential modification or isolation of nucleobases, tagging, and / or sequencing. Using this approach, for example, it can be determined whether a particular sequence is hypermethylated or hypomethylated. In some embodiments, hypermethylated variable epigenetic target regions are analyzed to determine whether they exhibit hypermethylation characteristics of tumor cells, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit hypomethylation characteristics of tumor cells. Additionally, by distributing a heterogeneous nucleic acid population, rare signals may be increased, e.g., by enriching rare nucleic acid molecules that are more present in one fraction (or sub-fraction) of the population. For example, genetic variations that are present in hypermethylated DNA but are rare or absent in hypomethylated DNA can be more readily detected by distributing the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, a multidimensional analysis of a single locus of the genome or species of nucleic acid can be performed, thus achieving greater sensitivity.
[0206] In some examples, a heterogeneous nucleic acid sample is distributed into two or more fractions (e.g., at least 3, 4, 5, 6, or 7 fractions). In some embodiments, each fraction is differentially tagged. Next, the tagged fractions can be pooled together for collective sample preparation and / or sequencing. The distribute-tag-pool steps can occur more than once, and each distribution round occurs based on different characteristics (examples provided herein) and is tagged using differential tags that are distinct from other fractions and distribution means.
[0207] Examples of features that can be used for partitioning include the length of an array, methylation level, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins that bind to DNA. The resulting fractions can include one or more of the following nucleic acid types: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), short DNA fragments, and long DNA fragments. In some embodiments, partitioning based on cytosine modifications (e.g., cytosine methylation) or methylation is commonly performed and can be combined, if desired, with at least one additional partitioning step based on any of the aforementioned features or types of DNA. In some embodiments, a heterogeneous nucleic acid population is partitioned into nucleic acids having one or more epigenetic modifications and nucleic acids having no epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine vs. other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and the association and level of association with one or more proteins, such as histones. Alternatively or in addition, a heterogeneous nucleic acid population can be partitioned into nucleic acid molecules that associate with nucleosomes and nucleic acid molecules that lack nucleosomes. Alternatively or in addition, a heterogeneous nucleic acid population can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively or in addition, a heterogeneous nucleic acid population can be partitioned based on the length of the nucleic acid (e.g., molecules up to 160 bp and molecules having a length greater than 160 bp).
[0208] In some examples, each fraction (representing different nucleic acid types) is differentially labeled and the fractions are pooled together prior to sequencing. In other examples, different types are sequenced separately.
[0209] In some embodiments, different nucleic acid populations are distributed into two or more different fractions. Each fraction is representative of a different nucleic acid type, and the first fraction (also called the secondary sample) contains DNA with a higher ratio of cytosine modifications than the second secondary sample. Each fraction is tagged separately. The first secondary sample is subjected to a procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample, where the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first and second nucleobases have the same base pairing specificity. The tagged nucleic acids are pooled together prior to sequencing. Sequence reads are obtained and an analysis is performed that includes distinguishing, in silico, the first nucleobase from the second nucleobase in the DNA of the first secondary sample. Tags are used to sort reads from different fractions. Analyses for detecting genetic variants can be performed at the level of each fraction and at the level of the entire nucleic acid population. For example, the analysis can include in silico analysis to determine genetic variants, such as CNVs, SNVs, indels, fusions in the nucleic acids of each fraction. In some examples, the in silico analysis can include determining chromatin structure. For example, nucleosome positions in chromatin can be determined using the coverage of the sequence reads. High coverage may correlate with high nucleosome occupancy in a genomic region, while low coverage may correlate with low nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0210] The sample can contain modified nucleic acids with post-replication modifications to nucleotides and modifications to one or more proteins, typically non-covalent binding.
[0211] In one embodiment, the nucleic acid population is a population obtained from a serum, plasma, or blood sample from a subject suspected of having a neoplasm, tumor, or cancer, or a subject already diagnosed with a neoplasm, tumor, or cancer. The nucleic acid population includes nucleic acids having various methylation levels. Methylation can occur from any one or more post-replicative or post-transcriptional modifications. Post-replicative modifications include modifications of the nucleotide cytosine, particularly modifications at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0212] The affinity agent can be an antibody having a desired specificity, its natural binding partner or variant (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected to have specificity for a given target, for example by phage display.
[0213] Examples of capture moieties contemplated herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including proteins such as antibodies that preferentially bind to MeCP2 and 5-methylcytosine.
[0214] Similarly, the partitioning of different types of nucleic acids can be carried out using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.
[0215] For some affinity agents and modifications, binding to the agent can occur in an essentially all-or-none manner depending on whether the nucleic acid has the modification or not, although separation can be a matter of degree. In such cases, nucleic acids that are overrepresented in the modification bind to the agent to a greater extent than nucleic acids that are underrepresented in the modification. Alternatively, nucleic acids with the modification can bind in an all-or-none manner. However, various levels of modification can be eluted sequentially from the binder.
[0216] For example, in some embodiments, partitioning can be binary or based on the degree / level of modification. For example, all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequent partitioning can then involve eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and the bound fragments. As the salt concentration increases, fragments with a greater level of methylation are eluted.
[0217] In some examples, the final fractions are representative of nucleic acids with different degrees of modification (overrepresentation or underrepresentation of the modification). Overrepresentation and underrepresentation can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids of a sample is 2, nucleic acids containing more than 2 5-methylcytosine residues are overrepresented in this modification, and nucleic acids with 1 or zero 5-methylcytosine residues are underrepresented. The effect of affinity separation is to enrich nucleic acids that are overrepresented in the modification in the binding phase and nucleic acids that are underrepresented in the modification in the non-binding phase (i.e., in solution). Nucleic acids in the binding phase can be eluted prior to subsequent processing.
[0218] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be partitioned using sequential elution. For example, a hypomethylated fraction (e.g., no methylation) can be separated from a methylated fraction by contacting the nucleic acid population with MBD from the kit bound to magnetic beads. Beads are used to separate methylated nucleic acids from unmethylated nucleic acids. One or more elution steps are then performed sequentially to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, such as at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After eluting such methylated nucleic acids, magnetic separation is used again to separate high-level methylated nucleic acids from nucleic acids with low-level methylation. The elution and magnetic separation steps are repeated to generate various fractions such as a hypomethylated fraction (representative of no methylation), a methylated fraction (representative of low-level methylation), and a hypermethylated fraction (representative of high-level methylation).
[0219] In some methods, nucleic acids bound to the agent used for affinity separation are subjected to a washing step. The washing step washes away nucleic acids that weakly bind to the affinity agent. Such nucleic acids can be enriched in nucleic acids having a modification closer to the average or median (i.e., an intermediate between nucleic acids that remain bound to the solid phase and nucleic acids that do not bind to the solid phase when the agent is first contacted with the sample).
[0220] Affinity separation results in at least two, sometimes three or more fractions of nucleic acids with different degrees of modification. The fractions are still separated, but the nucleic acids of at least one fraction, usually two or three (or more) fractions, are usually linked to nucleic acid tags provided as components of adapters, and the nucleic acids in different fractions are tagged with different tags that distinguish the members of one fraction from the members of another fraction. The tags linked to the nucleic acid molecules of the same fraction can be the same or different from each other. However, if they are different from each other, the tags can have a common part of their code so as to identify the molecule to which it is bound as a molecule of a specific fraction.
[0221] For further details regarding the partitioning of nucleic acid samples based on features such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0222] In some embodiments, nucleic acid molecules can be fractionated into different fractions based on nucleic acid molecules bound to a specific protein or fragment thereof and nucleic acid molecules not bound to a specific protein or fragment thereof.
[0223] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzyme activity. Examples of proteins that can bind to DNA and serve as a basis for fractionation include, but are not limited to, Protein A and Protein G. Any suitable method can be used to fractionate nucleic acid molecules based on the region to which the protein is bound. Examples of methods used to fractionate nucleic acid molecules based on the region to which the protein is bound include, but are not limited to, SDS-PAGE, chromatin-immunoprecipitation (ChIP), heparin chromatography, and asymmetric flow field-flow fractionation (AF4).
[0224] In some embodiments, nucleic acid partitioning is performed by contacting the nucleic acid with the methyl-binding domain ("MBD") of a methyl-binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is linked via a biotin linker to paramagnetic beads, such as Dynabeads® M-280 streptavidin. Partitioning into fractions having different degrees of methylation can be performed by eluting the fractions by increasing the NaCl concentration.
[0225] Examples of MBPs contemplated herein include, but are not limited to, (a) MeCP2 is a protein that preferentially binds to 5-methyl-cytosine as compared to non-modified cytosine. (b) RPL26, PRP8, and DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethyl-cytosine as compared to non-modified cytosine. (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind to 5-formyl-cytosine as compared to non-modified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)). (d) An antibody specific for one or more methylated nucleotide bases is included.
[0226] Generally, elution is a function of the number of methylation sites per molecule, and molecules with more methylation elute under increased salt concentration. To elute DNA into separate populations based on the degree of methylation, a series of elution buffers with increasing NaCl concentration can be used. The salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three fractions. The molecules can be contacted with a solution of a first salt concentration containing molecules including a methyl binding domain, and the molecules can be bound to a capture moiety such as streptavidin. At the first salt concentration, a population of molecules binds to the MBD and a population remains unbound. The unbound population can be separated as the "low methylation" population. For example, the first fraction representing low-methylated DNA is the fraction that remains unbound at a low salt concentration, such as 100 mM or 160 mM. The second fraction representing intermediate-methylated DNA is eluted using an intermediate salt concentration, such as a concentration of 100 mM to 2000 mM. This is also separated from the sample. The third fraction representing high-methylated DNA is eluted using a high salt concentration, such as at least about 2000 mM.
[0227] a. Tagging of fractions In some embodiments, two or more fractions, e.g., each fraction, are differentially tagged. The tag or index can be a molecule such as a nucleic acid that contains information indicative of the identity of the molecule to which the tag associates. For example, the molecule can have a sample tag or sample index (to distinguish molecules in one sample from molecules in different samples), a fraction tag (to distinguish molecules in one fraction from molecules in different fractions), and / or a molecule tag / molecule barcode / barcode (to distinguish different molecules from each other (in both unique and non-unique tagging scenarios)). In certain embodiments, the tag can include one or a combination of barcodes. As used herein, the term "barcode" refers to, depending on the context, a nucleic acid molecule having a specific nucleotide sequence or the nucleotide sequence itself. The barcode can have, for example, between 10 and 100 nucleotides. The collection of barcodes can have degenerate sequences or sequences having a specific Hamming distance if desired for a particular purpose. Thus, for example, a molecule barcode can be composed of one barcode or a combination of two barcodes each attached to a different end of the molecule. Additionally or alternatively, different sets of molecule barcodes, molecule tags, or molecule indexes can be used for different fractions and / or samples such that the barcodes serve as molecule tags through their individual sequences and such that they serve to identify the corresponding fractions and / or samples based on the sets of which they are members.
[0228] Tags can be used to label fractions of individual polynucleotide populations to correlate that tag (or tags) with a particular fraction. Alternatively, tags can be used in embodiments of the invention that do not use the step of partitioning. In some embodiments, a single tag can be used to label a particular fraction. In some embodiments, multiple different tags can be used to label a particular fraction. In embodiments where multiple different tags are used to label a particular fraction, the set of tags used to label one fraction can be readily distinguished from the set of tags used to label other fractions. In some embodiments, the tag may have additional functionality, for example the tag can be used to index the source of a sample, or as a unique molecular identifier (which can be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations, for example as described in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS ONE,11: e0146638 (2016)), or as a non-unique molecular identifier as described, for example, in U.S. Patent No. 9,598,731. Similarly, in some embodiments, the tag may have additional functionality, for example the tag can be used to index the source of a sample, or as a non-unique molecular identifier (which can be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations).
[0229] In one embodiment, tagging of fractions includes tagging the molecules in each fraction with a fraction tag. The fractions are combined again (e.g., to reduce the number of sequencing runs required and avoid unnecessary costs), and after sequencing the molecules, the fraction tag identifies the source fraction. In another embodiment, different fractions are tagged with different sets of molecule tags, e.g., including barcode pairs. In this way, each molecule barcode is useful for indicating the source fraction and for distinguishing the molecules within the fraction. For example, a first set of 35 barcodes can be used to tag the molecules in a first fraction, and a second set of 35 barcodes can be used to tag the molecules in a second fraction.
[0230] In some embodiments, after dispensing and tagging with fraction tags, the molecules may be pooled for sequencing in a single run. In some embodiments, a sample tag is added to the molecules, e.g., in a step after addition and pooling of the fraction tags. The sample tag can facilitate pooling of material generated from multiple samples for sequencing in a single sequencing run.
[0231] Alternatively, in some embodiments, the fraction tag can be correlated with the sample as well as the fraction. As a simple example, a first tag can indicate the first fraction of a first sample, a second tag can indicate the second fraction of the first sample, a third tag can indicate the first fraction of a second sample, and a fourth tag can indicate the second fraction of the second sample.
[0232] Tags may bind to molecules that have already been assigned based on one or more characteristics, although the final tagged molecules in the library may no longer possess that characteristic. For example, single-stranded DNA molecules can be assigned and tagged, although the final tagged molecules in the library will probably be double-stranded. Similarly, DNA may be subjected to assignment based on different methylation levels, although in the final library, the tagged molecules derived from these molecules will probably not be methylated. Thus, tags bound to molecules in the library typically indicate the characteristics of the "parent molecule" from which the final tagged molecule is derived, and not necessarily the characteristics of the tagged molecule itself.
[0233] As an example, use barcodes 1, 2, 3, 4, etc. to tag and label the molecules in the first fraction; use barcodes A, B, C, D, etc. to tag and label the molecules in the second fraction; and use barcodes a, b, c, d, etc. to tag and label the molecules in the third fraction. The differentially tagged fractions can be pooled prior to sequencing. The differentially tagged fractions can be sequenced separately or, for example, together simultaneously in the same flow cell of an Illumina sequencer.
[0234] After sequencing, analysis of the reads to detect genetic variants can be performed at the level of each fraction as well as at the level of the total nucleic acid population. Tags are used to sort the reads from different fractions. The analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, length of genomic coordinates, coverage, and / or copy number. In some embodiments, high coverage may correlate with high nucleosome occupancy in a genomic region, while low coverage may correlate with low nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0235] b. Alternative methods for analysis of modified nucleic acids In some embodiments, adapters may be added to the nucleic acids after nucleic acid partitioning, and in other embodiments, adapters may be added to the nucleic acids before nucleic acid partitioning. In some such methods, a population of nucleic acids having modifications to varying degrees (e.g., 0, 1, 2, 3, 4, 5 or more methyl groups per nucleic acid molecule) is contacted with the adapter prior to fractionation of the population according to the degree of modification. The adapter binds to either one or both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a sufficient number of different tags such that the probability that two nucleic acids having the same starting and stopping points receive the same combination of tags is low, e.g., 95, 99, or 99.9%. The adapter may contain the same or different primer binding sites, whether having the same or different tags, but preferably the adapter contains the same primer binding site. After binding of the adapter, the nucleic acids are contacted with an agent (e.g., such an agent as already described) that preferentially binds to nucleic acids having the modification. The nucleic acids are partitioned from the binding to the agent into at least two secondary samples that differ in the degree to which the nucleic acids have the modification. For example, if the agent has an affinity for nucleic acids having the modification, nucleic acids in which the modification is overrepresented (compared to the median of the occurrences in the population) preferentially bind to the agent, while nucleic acids in which the modification is underrepresented do not bind to the agent or are more readily eluted from the agent. After partitioning, the first secondary sample is subjected to a procedure that affects the first nucleobase in the DNA such that it is different from the second nucleobase in the DNA of the first secondary sample, where the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. Next, the nucleic acids are amplified from a primer that binds to the primer binding site within the adapter. After amplification, the different fractions can be subjected to further processing steps, which typically include further amplification (e.g., clonal amplification), and sequence analysis, in parallel but separately. Next, the sequence data from the different fractions can be compared.
[0236] In another embodiment, the partitioning scheme can be implemented using the following exemplary procedure. The nucleic acid is ligated to both ends of a Y-shaped adapter that contains a primer binding site and a tag. The molecule is amplified. Next, the amplified molecule is fractionated by contacting it with an antibody that preferentially binds to 5-methylcytosine, producing two fractions. One fraction contains the original molecule lacking methylation and the amplified copies that have lost methylation. The other fraction contains the original DNA molecule having methylation. The fraction containing the original DNA molecule having methylation is subjected to a procedure that affects a first nucleobase in the DNA such that it is different from a second nucleobase in the DNA of the first subsample, where the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. Next, the two fractions are processed and sequenced separately by further amplifying the methylated fraction. Next, the sequence data of the two fractions can be compared. In this example, the tags are not used to distinguish between methylated and unmethylated DNA, but rather to distinguish between different molecules within those fractions so that it can be determined whether reads having the same start and stop points are based on the same or different molecules.
[0237] The present disclosure provides further methods for analyzing a nucleic acid population in which at least a portion of the nucleic acid comprises one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications already described. In these methods, after partitioning, an adapter comprising one or more cytosine residues modified at the 5C position, such as 5-methylcytosine, is contacted with a secondary sample of the nucleic acid. Preferably, all cytosine residues in such an adapter are also modified, or all such cytosines in the primer binding region of the adapter are modified. The adapter binds to both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a sufficient number of different tags such that the probability that two nucleic acids having the same starting and stopping points receive the same combination of tags is low, such as 95, 99, or 99.9%. The primer binding sites in such an adapter may be the same or different, but are preferably the same. After binding of the adapter, the nucleic acid is amplified from a primer that binds to the primer binding site of the adapter. The amplified nucleic acid is divided into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. The sequence data of the molecules in the first aliquot is thus determined regardless of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are subjected to a procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA, where the first nucleobase comprises cytosine modified at the 5 position and the second nucleobase comprises unmodified cytosine. This procedure can be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. Next, the nucleic acid subjected to the procedure is amplified by a primer that is specific for the original primer binding site of the adapter ligated to the nucleic acid. These nucleic acids retain cytosine at the primer binding site of the adapter, but the amplification products have undergone conversion to uracil in the bisulfite treatment and have lost methylation of these cytosine residues, such that only the nucleic acid molecules originally ligated to the adapter (different from their amplification products) are amplifiable. Thus, only the original molecules in the population that are at least partially methylated are amplified. After amplification, these nucleic acids are subjected to sequence analysis.Comparison of the sequences determined from the first and second aliquots can, inter alia, show which cytosines in the nucleic acid population have been subjected to methylation.
[0238] Such analysis can be carried out using the following exemplary procedure. After partitioning, the methylated DNA is ligated to a Y-shaped adapter at both ends containing primer binding sites and tags. The cytosines in the adapter are modified at the 5-position (e.g., 5-methylated). The modification of the adapter serves to protect the primer binding sites in subsequent conversion steps (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect modified cytosines but affects unmodified cytosines). After adapter ligation, the DNA molecules are amplified. The amplification products are split into two aliquots for sequencing with and without conversion. The aliquot not subjected to conversion can be subjected to sequence analysis with or without further processing. The other aliquot is subjected to a procedure that affects the first nucleobase in the DNA so as to be different from the second nucleobase in the DNA of the first secondary sample, the first nucleobase comprising cytosine modified at the 5-position and the second nucleobase comprising unmodified cytosine. This procedure can be bisulfite treatment or any other procedure that converts unmodified cytosine to uracil. Only the primer binding sites protected by the modification of cytosine can support amplification when contacting a primer specific for the original primer binding site. Thus, only the original molecules that are not copies of the first amplification are subjected to further amplification. Next, the further amplified molecules are subjected to sequence analysis. Next, the sequences from the two aliquots can be compared. Similar to the separation scheme above, the nucleic acid tags in the adapter are not used to distinguish between methylated and non-methylated DNA but are used to distinguish nucleic acid molecules within the same fraction.
[0239] 2. Subject the first secondary sample to a procedure that affects the first nucleobase in the DNA so as to be different from the second nucleobase in the DNA of the first secondary sample The method disclosed herein comprises subjecting a first secondary sample to a procedure that affects a first nucleobase in the DNA of the first secondary sample to be different from a second nucleobase in the DNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, when the first nucleobase is a modified adenine or an unmodified adenine, the second nucleobase is a modified adenine or an unmodified adenine; when the first nucleobase is a modified cytosine or an unmodified cytosine, the second nucleobase is a modified cytosine or an unmodified cytosine; when the first nucleobase is a modified guanine or an unmodified guanine, the second nucleobase is a modified guanine or an unmodified guanine; when the first nucleobase is a modified thymine or an unmodified thymine, the second nucleobase is a modified thymine or an unmodified thymine (modified and unmodified uracils are included within modified thymine for the purposes of this step).
[0240] In some embodiments, when the first nucleobase is a modified cytosine or an unmodified cytosine, the second nucleobase is a modified cytosine or an unmodified cytosine. For example, the first nucleobase may include unmodified cytosine (C), and the second nucleobase may include one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may include C, and the first nucleobase may include one or more of mC and hmC. For example, as shown in the above summary and the following description, other combinations are also possible, such as when one of the first and second nucleobases includes mC and the other includes hmC.
[0241] In some embodiments, the procedure for affecting the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample includes bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Thus, when using bisulfite conversion, the first nucleobase includes one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other types of cytosine affected by bisulfite, and the second nucleobase may include one or more of mC and hmC, e.g., one or more of mC and optionally hmC. Sequencing of bisulfite-treated DNA identifies positions that are read as cytosine as being mC or hmC positions. On the other hand, positions read as T are identified as T or a bisulfite-sensitive form of C, e.g., unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Thus, performing bisulfite conversion on the first secondary sample as described herein facilitates identifying positions containing mC or hmC using sequence reads obtained from the first secondary sample. For an exemplary description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068. An exemplary workflow for performing bisulfite conversion on a first secondary sample having high methylation is illustrated in FIG. 1.
[0242] In some embodiments, the procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample includes oxidative bisulfite (Ox - BS) conversion. This procedure first converts hmC to fC, which is bisulfite - sensitive, and then performs bisulfite conversion. Thus, when using oxidative bisulfite conversion, the first nucleobase includes one or more of unmodified cytosine, fC, caC, hmC, or other types of cytosine affected by bisulfite, and the second nucleobase includes mC. Sequencing of Ox - BS - converted DNA identifies positions read as cytosine as mC positions. On the other hand, positions read as T are identified as T, hmC, or bisulfite - sensitive forms of C, such as unmodified cytosine, fC, or hmC. Thus, performing Ox - BS conversion on the first secondary sample described herein facilitates identifying positions containing mC using sequence reads obtained from the first secondary sample. For an exemplary description of oxidative bisulfite conversion, see, for example, Booth et al., Science 2012; 336: 934 - 937.
[0243] In some embodiments, the procedure for affecting the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample includes Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from the conversion, mC is oxidized prior to the bisulfite treatment, such that the position originally occupied by mC is converted to U, and the position originally occupied by hmC remains as the protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, after protecting hmC using β-glucosyltransferase (to form 5-glucosylhydroxymethylcytosine (ghmC)), a TET protein such as mTet1 can be used to convert mC to caC, and then a bisulfite treatment can be used to convert C and caC to U, while ghmC remains unaffected. Thus, when using TAB conversion, the first nucleobase includes one or more of unmodified cytosine, fC, caC, mC, or other forms of cytosine affected by bisulfite, and the second nucleobase includes hmC. Sequencing of TAB-converted DNA identifies positions where the cytosine read is the hmC position. On the other hand, positions read as T are identified as T, mC, or a bisulfite-sensitive form of C, such as unmodified cytosine, fC, or caC. Thus, performing TAB conversion with respect to the first secondary sample described herein facilitates identifying positions containing hmC using sequence reads obtained from the first secondary sample.
[0244] In some embodiments, the procedure for affecting the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample includes Tet-assisted conversion with a substituted borane reducing agent, which is 2-picolyl borane, borane pyridine, tert-butylamine borane, or ammonia borane as needed. In the Tet-assisted pic-borane conversion by substituted borane reducing agent conversion, the TET protein is used to convert mC and hmC to caC without affecting unmodified C, and then caC and fC if present are similarly converted to dihydrouracil (DHU) by treatment with 2-picolyl borane (pic-borane) or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, without affecting unmodified C. See, for example, Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., Supplementary Figure 1 and Supplementary Note 7). DHU is read as T in sequencing. Thus, when using this type of conversion, the first nucleobase includes one or more of mC, fC, caC, or hmC, and the second nucleobase includes unmodified cytosine. Sequencing of the converted DNA identifies positions read as cytosine as unmodified C positions. On the other hand, positions read as T are identified as being T, mC, fC, caC, or hmC. Thus, performing TAP conversion on the first secondary sample as described herein facilitates identifying positions containing unmodified C using sequence reads obtained from the first secondary sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), as described in more detail above in Liu et al. 2019.
[0245] Alternatively, protection of hmC (e.g., using βGT) can be combined with Tet-assisted conversion by a substituted borane reducing agent. hmC can be protected as described above through glucosylation using βGT to form ghmC. Next, treatment with a TET protein, such as mTet1, converts mC to caC, but does not convert C or ghmC. caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, without similarly affecting unmodified C or ghmC. Thus, when using Tet-assisted conversion by a substituted borane reducing agent, the first nucleobase contains mC, and the second nucleobase contains one or more of unmodified cytosine or hmC, e.g., unmodified cytosine, and optionally hmC, fC, and / or caC. Sequencing of the converted DNA identifies positions read as cytosine as hmC or unmodified C positions. Positions read as T, on the other hand, are identified as T, fC, caC, or mC. Thus, performing TAPSβ conversion on the first secondary sample as described herein facilitates distinguishing positions containing one of unmodified C or hmC from positions containing mC using sequence reads obtained from the first secondary sample. For an exemplary illustration of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429.
[0246] In some embodiments, the procedure for affecting the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample includes chemical-assisted conversion with a substituted borane reducing agent that is, as needed, 2-picolyl borane, borane pyridine, tert-butylamine borane, or ammonia borane. In the chemical-assisted conversion with a substituted borane reducing agent, an oxidizing agent, such as potassium perruthenate (KRuO4) (also suitable for use in ox-BS conversion), is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, converts fC and caC to DHU, but does not affect mC or unmodified C. Thus, when using this type of conversion, the first nucleobase includes one or more of hmC, fC, and caC, and the second nucleobase includes one or more of unmodified cytosine or mC, such as one or more of unmodified cytosine and optionally mC. Sequencing of the converted DNA identifies positions read as cytosine as mC or unmodified C positions. On the other hand, positions read as T are identified as T, fC, caC, or hmC. Thus, performing this type of conversion on the first secondary sample as described herein facilitates distinguishing positions containing unmodified C or mC from positions containing hmC using sequence reads obtained from the first secondary sample. For an exemplary description of this type of conversion, see, for example, Liu et al., Nature Biotechnology 2019; 37:424-429.
[0247] In some embodiments, the procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample includes APOBEC-linked epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme, such as APOBEC3A (A3A), is used to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when using ACE conversion, the first nucleobase includes unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase includes hmC. Sequencing of ACE-converted DNA identifies positions read as cytosine as hmC, fC, or caC positions. On the other hand, positions read as T are identified as T, unmodified C, or mC. Thus, performing ACE conversion on the first secondary sample as described herein facilitates distinguishing positions containing hmC from positions containing mC or unmodified C using sequence reads obtained from the first secondary sample. For an exemplary description of ACE conversion, see, for example, Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0248] In some embodiments, procedures that affect the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample include enzymatic conversion of the first nucleobase, such as enzymatic conversion in EM-Seq. See, for example, Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, which are available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by deaminases (e.g., APOBEC3A), and then deaminases (e.g., APOBEC3A) can be used to deaminate non-modified cytosines and convert them into uracil.
[0249] In some embodiments, a procedure that affects a first nucleobase in DNA to be different from a second nucleobase in the DNA of a first secondary sample includes separating the DNA that originally contains the first nucleobase from DNA that does not originally contain the first nucleobase. In some such embodiments, the first nucleobase is hmC. The DNA that originally contains the first nucleobase can be separated from other DNA using a labeling procedure that includes a biotinylated position that originally contains the first nucleobase. In some embodiments, the first nucleobase is first derivatized by an azide-containing moiety, such as a glucosyl-azide-containing moiety. Next, the azide-containing moiety can serve as a reagent for binding biotin, for example, through Huisgen cycloaddition chemistry. Next, the DNA that now contains the biotinylated first nucleobase can be separated from the DNA that does not originally contain the first nucleobase using a biotin binder, such as avidin, NeutrAvidin (deglycosylated avidin having an isoelectric point of about 6.3), or streptavidin. An example of a procedure for separating the DNA that originally contains the first nucleobase from the DNA that does not originally contain the first nucleobase is hmC-seal, which labels hmC to form β-6-azido-glucosyl-5-hydroxymethylcytosine, then binds a biotin moiety through Huisgen cycloaddition, and then uses a biotin binder to separate the biotinylated DNA from other DNA. For an exemplary description of hmC-seal, see, for example, Han et al., Mol. Cell 2016; 63: 711-719. This approach is useful for identifying fragments that contain one or more hmC nucleobases.
[0250] In some embodiments, after such separation, the method further comprises differentially tagging each of the DNA that originally contains the first nucleobase, the DNA that does not originally contain the first nucleobase, and the DNA of the second secondary sample. After differential tagging, the method may further comprise pooling the DNA that originally contains the first nucleobase, the DNA that does not originally contain the first nucleobase, and the DNA of the second secondary sample. Next, the DNA that originally contains the first nucleobase, the DNA that does not originally contain the first nucleobase, and the DNA of the second secondary sample may be sequenced in the same sequencing cell while retaining the ability to resolve whether a given read is derived from a molecule of the DNA that originally contains the first nucleobase, the DNA that does not originally contain the first nucleobase, or the DNA of the second secondary sample using the differential tags.
[0251] In some embodiments, the first nucleobase is a modified adenine or an unmodified adenine, and the second nucleobase is a modified adenine or an unmodified adenine. In some embodiments, the modified adenine is 6 N-methyladenine (mA). In some embodiments, the modified adenine is 6 N-methyladenine (mA), 6 N-hydroxymethyladenine (hmA), or 6 one or more of N-formyladenine (fA).
[0252] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases, such as mA, from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific for mA are described in Sun et al., Bioessays 2015; 37:1155-62. Antibodies against various modified nucleic acid bases, such as thymine / uracil types including halogenated forms such as 5-bromouracil, are commercially available. Various modified bases can also be detected based on changes in their base pairing specificities. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read as G in sequencing. See, for example, U.S. Patent No. 8,486,630; Brown, Genomes, 2 nd Ed., John Wiley & Sons, Inc., New York, N.Y., 2002, chapter 14, "Mutation, Repair, and Recombination." See also.
[0253] 3. Enrichment / Capture Step; Amplification; Adapter; Barcode In some embodiments, the methods disclosed herein include capturing a set of one or more target regions of DNA, such as cfDNA. Capture can be performed using any suitable approach known in the art.
[0254] In some embodiments, the capturing step includes contacting the DNA to be captured with a target-specific probe set. The target-specific probe set can have any of the features described herein with respect to target-specific probe sets, including, but not limited to, those related to the above embodiments and the probes in the following sections. The capturing step can be performed on one or more secondary samples prepared during the methods disclosed herein. In some embodiments, the DNA is captured from at least a first secondary sample or a second secondary sample, such as from at least the first secondary sample and the second secondary sample. If the first secondary sample undergoes a separation step (e.g., a step of separating DNA originally containing a first nucleobase (e.g., hmC) from DNA not originally containing the first nucleobase, such as hmC-seal), the capturing step can be performed on the DNA originally containing the first nucleobase (e.g., hmC), the DNA not originally containing the first nucleobase, and / or the DNA of the second secondary sample, any one, any two, or all of them. In some embodiments, the secondary samples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.
[0255] The capturing step can generally be performed using conditions suitable for specific nucleic acid hybridization, which depends to some extent on the features of the probe, such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions, taking into account the general knowledge in the art regarding nucleic acid hybridization. In some embodiments, a complex of the target-specific probe and the DNA is formed.
[0256] In some embodiments, the methods described herein include capturing cfDNA obtained from a subject of interest with respect to a set of multiple target regions. The target regions include epigenetic target regions, which may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from tumors or healthy cells. The target regions also include sequence variable target regions, which may exhibit differences in sequence depending on whether they originate from tumors or healthy cells. The capturing step produces a set of captured cfDNA molecules, and the cfDNA molecules corresponding to the set of sequence variable target regions are captured at a higher capture yield in the set of captured cfDNA molecules than the cfDNA molecules corresponding to the set of epigenetic target regions. For further consideration of the capturing step, capture yield, and related aspects, see WO2020 / 160414, which is incorporated herein by reference for all purposes.
[0257] In some embodiments, the methods described herein include contacting the cfDNA obtained from a subject of interest with a set of target-specific probes, the set of target-specific probes being configured to capture cfDNA corresponding to a set of sequence variable target regions at a higher capture yield than cfDNA corresponding to a set of epigenetic target regions.
[0258] To analyze the array variable target regions with sufficient reliability or accuracy, a higher sequencing depth may be required than may be necessary to analyze the epigenetic target regions. Thus, it can be beneficial to capture cfDNA corresponding to the array variable target region set with a higher capture yield than cfDNA corresponding to the epigenetic target region set. The amount of data necessary to determine the fragmentation pattern (e.g., for testing perturbations of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in highly methylated and hypomethylated fractions) is generally less than the amount of data necessary to determine the presence or absence of sequence variations related to cancer. Capturing the target region sets at different yields can facilitate sequencing the target regions to different sequencing depths in the same sequencing run (e.g., using pooled mixtures and / or in the same sequencing cell).
[0259] In various embodiments, the method further includes, consistent with the discussion herein, sequencing the captured cfDNA to different sequencing depths, e.g., with respect to epigenetic and array variable target region sets.
[0260] In some embodiments, the complex of the target-specific probe and DNA is separated from DNA not bound to the target-specific probe. For example, if the target-specific probe is bound to a solid support covalently or non-covalently, a washing or aspiration step can be used to separate the unbound material. Alternatively, chromatography can be used if the complex has chromatographic properties distinct from the unbound material (e.g., if the probe contains a ligand that binds to a chromatographic resin).
[0261] As will be discussed in detail elsewhere in this specification, a target-specific probe set can include multiple sets such as probes for a set of array-variable target regions and probes for a set of epigenetic target regions. In some such embodiments, the capturing step is performed simultaneously in the same container for probes for a set of array-variable target regions and probes for a set of epigenetic target regions. For example, the probes for a set of array-variable target regions and the probes for a set of epigenetic target regions are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for a set of array-variable target regions is higher than the concentration of the probes for a set of epigenetic target regions.
[0262] Alternatively, the capturing step is performed using a set of array-variable target region probes in a first container and a set of epigenetic target region probes in a second container, or the contacting step is performed using a set of array-variable target region probes at a first time and in a first container and a set of epigenetic target region probes at a second time before or after the first time. This approach allows for the separate preparation of first and second compositions containing captured DNA corresponding to a set of array-variable target regions and captured DNA corresponding to a set of epigenetic target regions. The compositions can be separately processed, recombined in appropriate ratios if desired (e.g., for fractionation based on methylation as described elsewhere herein), and provided as materials for further processing and analysis such as sequencing.
[0263] In some embodiments, the DNA is amplified. In some embodiments, the amplification is performed before the capturing step. In some embodiments, the amplification is performed after the capturing step.
[0264] In some embodiments, the adapter is included in the DNA. This can be done simultaneously with the amplification procedure, for example, by providing the adapter at the 5' portion of the primer as described above. Alternatively, the adapter can be added by other approaches such as ligation.
[0265] In some embodiments, tags that can be or include barcodes are included in the DNA. The tags can facilitate identification of the origin of the nucleic acid. For example, barcodes can be used to identify the origin (e.g., the subject) from which the DNA is derived after pooling of multiple samples for parallel sequencing. This can be done simultaneously with the amplification procedure, for example, by providing the barcode at the 5' portion of the primer as described herein. In some embodiments, the adapter and the tag / barcode are provided by the same primer or primer set. For example, the barcode can be located at the 3' of the adapter and the 5' of the target hybridizing portion of the primer. Alternatively, the barcode can be added by other approaches, such as ligation, together with the adapter in the same ligation substrate as needed.
[0266] Additional details regarding amplification, tagging, and barcoding are discussed in the section "General Features of the Methods" below and can be practiced in any combination with the foregoing embodiments as well as the embodiments described in the Introduction and Summary sections to the extent practicable.
[0267] 4. Captured Set In some embodiments, a set of captured DNA (e.g., cfDNA) is provided. With respect to the methods of the present disclosure, a set of captured DNA can be provided, for example, by performing a capturing step after the distributing step described herein. The captured set can include DNA corresponding to a set of sequence-variable target regions, a set of epigenetic target regions, or a combination thereof. In some embodiments, when normalized with respect to the difference in the size (footprint size) of the targeted regions, the amount of captured sequence-variable target region DNA is greater than the amount of captured epigenetic target region DNA.
[0268] Alternatively, a first and a second captured set can be provided, each including DNA corresponding to a set of sequence-variable target regions and DNA corresponding to a set of epigenetic target regions, respectively. The first and second captured sets can be combined to provide a combined captured set.
[0269] In some embodiments, a captured set containing DNA corresponding to a variable array target region set and an epigenetic target region set includes the combinations of captured sets discussed above. For the DNA corresponding to the variable array target region set, it may be present at a higher concentration than the DNA corresponding to the epigenetic target region set, for example, at a concentration 1.1 - 1.2 times higher, 1.2 - 1.4 times higher, 1.4 - 1.6 times higher, 1.6 - 1.8 times higher, 1.8 - 2.0 times higher, 2.0 - 2.2 times higher, 2.2 - 2.4 times higher, 2.4 - 2.6 times higher, 2.6 - 2.8 times higher, 2.8 - 3.0 times higher, 3.0 - 3.5 times higher, 3.5 - 4.0 times higher, 4.0 - 4.5 times higher, 4.5 - 5.0 times higher, 5.0 - 5.5 times higher, 5.5 - 6.0 times higher, 6.0 - 6.5 times higher, 6.5 - 7.0 times higher, 7.0 - 7.5 times higher, 7.5 - 8.0 times higher, 8.0 - 8.5 times higher, 8.5 - 9.0 times higher, 9.0 - 9.5 times higher, 9.5 - 10.0 times higher, 10 - 11 times higher, 11 - 12 times higher, 12 - 13 times higher, 13 - 14 times higher, 14 - 15 times higher, 15 - 16 times higher, 16 - 17 times higher, 17 - 18 times higher, 18 - 19 times higher, 19 - 20 times higher, 20 - 30 times higher, 30 - 40 times higher, 40 - 50 times higher, 50 - 60 times higher, 60 - 70 times higher, 70 - 80 times higher, 80 - 90 times higher, 90 - 100 times higher, 10 - 20 times higher, 10 - 40 times higher, 10 - 50 times higher, 10 - 70 times higher, or 10 - 100 times higher. The degree of the difference in concentration explains the normalization with respect to the footprint size of the target region, as discussed in the Definitions section.
[0270] a. Epigenetic target region set An epigenetic target region set can include one or more types of target regions that can distinguish DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, such as non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. An epigenetic target region set can also include, for example, one or more control regions as described herein.
[0271] In some embodiments, the epigenetic target region set has a footprint of at least 100 kb, such as at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target region set has a footprint in the range of 100-1000 kb, such as 100-200 kb, 200-300 kb, 300-400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1,000 kb.
[0272] i. Hypermethylated variable target regions In some embodiments, the epigenetic target region set includes one or more hypermethylation variable target regions. Generally, a hypermethylation variable target region refers to a region where, for example, in a cfDNA sample, an increase in the observed methylation level indicates an increased likelihood that the sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells, such as tumor or cancer cells. For example, hypermethylation of the promoter of a tumor suppressor gene has been repeatedly observed. See, for example, Kang et al., Genome Biol. 18:53 (2017) and the references cited therein. In one example, a hypermethylation variable target region may not necessarily be different in terms of methylation in a cancerous tissue compared to DNA from the same type of healthy tissue, but may be different in methylation (e.g., have more methylation) compared to cfDNA that is typical in a healthy subject. For example, if the presence of cancer results in an increase in cell death, such as apoptosis of cells of the tissue type corresponding to the cancer, such cancer can be detected, at least in part, using such hypermethylation variable target regions. In some embodiments, a hypermethylation variable target region includes one or more genomic regions, and cfDNA molecules in those regions do not have a different methylation state in a cancerous subject compared to cfDNA from a healthy subject, but the presence / increased amount of hypermethylated cfDNA in those regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g., tumor shedding) circulating.
[0273] Extensive considerations regarding methylation variable target regions in colorectal cancer are provided in Lam et al., Biochim Biophys Acta. 1866:106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylated variable target regions based on tests of colorectal cancer (CRC) is provided in Table 1. Many of these genes are likely to be relevant to cancers other than colorectal cancer. For example, TP53 is a very important tumor suppressor, and it is widely recognized that inactivation based on hypermethylation of this gene can be a common tumorigenic mechanism.
Table 1-1
Table 1-2
[0274] In some embodiments, the hypermethylated variable target regions include a plurality of loci described in Table 1, such as at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci described in Table 1. For example, for each locus included as a target region, there may be one or more probes having a hybridization site that binds between the transcription start site of the gene and the stop codon (the last stop codon in the case of alternatively spliced genes), or in the promoter region of the gene. In some embodiments, one or more probes bind within 300 bp, such as within 200 or 100 bp, of the transcription start site of the gene in Table 1.
[0275] Methylation variable target regions in various types of lung cancer have been discussed in detail, for example, in Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77:453-74 (2015); Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017); Shi et al., BMC Genomics 18:901 (2017); Schneider et al., BMC Cancer. 11:102 (2011); Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016); Skvortsova et al., Br. J. Cancer. 94(10):1492-1495 (2006); Kim et al., Cancer Res. 61:3419-3424 (2001); Furonaka et al., Pathology International 55:303-309 (2005); Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014); Kim et al., Oncogene. 20:1765-70 (2001); Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003); Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005); Heller et al., Oncogene 25:959-968 (2006); Licchesi et al., Carcinogenesis. 29:895-904 (2008); Guo et al., Clin. Cancer Res. 10:7917-24 (2004); Palmisano et al., Cancer Res. 63:4620-4625 (2003); and Toyooka et al., Cancer Res. 61:4556-4560, (2001).
[0276] Table 2 provides an exemplary set of hypermethylated variable target regions based on lung cancer tests. Many of these genes may also be relevant to cancers other than lung cancer. For example, Casp8 (caspase 8) is an important enzyme in programmed cell death, and inactivation based on hypermethylation of this gene may be a common tumorigenic mechanism not limited to lung cancer. In addition, some genes appear in both Tables 1 and 2, indicating generality. [Table 2]
[0277] Any of the foregoing embodiments regarding the target regions identified in Table 2 may be combined with any of the above embodiments regarding the target regions identified in Table 1. In some embodiments, the hypermethylated variable target regions include a plurality of loci described in Table 1 or Table 2, such as at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci described in Table 1 or Table 2.
[0278] Additional hypermethylated target regions may be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a probabilistic method called CancerLocator using hypermethylated target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions may be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subsets of hypermethylated target regions that collectively show hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0279] ii. Hypomethylated variable target regions Global hypomethylation is a phenomenon commonly observed in various cancers. See, e.g., Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (review article describing findings on hypomethylation in colorectal cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular cancer, and cervical cancer). For example, regions such as repetitive elements, such as LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, as well as intergenic regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Thus, in some embodiments, the set of epigenetic target regions includes hypomethylation variable target regions, and the observed decrease in methylation level indicates an increased likelihood that the sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells, such as tumor cells or cancer cells. In one example, the hypomethylation variable target regions may include regions where the methylation status is not necessarily different in cancerous tissue compared to DNA from the same type of healthy tissue, but is different (e.g., less methylated) compared to cfDNA that is typical in healthy subjects. For example, if the presence of cancer results in an increase in cell death, such as apoptosis, of cells of the tissue type corresponding to the cancer, such cancer can be detected at least in part using such hypomethylation variable target regions. In some embodiments, the hypomethylation variable target regions include one or more genomic regions, and cfDNA molecules in those regions have a methylation status that is not different in cancerous subjects compared to cfDNA from healthy subjects, but the presence / increased amount of hypomethylated cfDNA in those regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g., tumor shedding) circulating.
[0280] In some embodiments, the hypomethylated variable target regions include repetitive elements and / or intergenic regions. In some embodiments, the repetitive elements include one, two, three, four, or five of LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.
[0281] Exemplary specific genomic regions showing cancer-related hypomethylation include nucleotides 8403565 - 8953708 and 151104701 - 151106035 of human chromosome 1. In some embodiments, the hypomethylated variable target regions overlap with one or both of these regions or include one or both of these regions.
[0282] iii. CTCF binding regions CTCF is a DNA-binding protein that contributes to chromatin organization and often co-localizes with cohesin. Perturbations of CTCF binding sites have been reported in a variety of different cancers. See, for example, Katainen et al., Nature Genetics, doi:10.1038 / ng.3335, published online on June 8, 2015; Guo et al., Nat. Commun. 9:1520 (2018). CTCF binding results in a recognizable pattern of cfDNA that can be detected by sequencing, for example, through fragment length analysis. Details regarding sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57 - 68 (2016); WO2018 / 009723; and US20170211143A1, each of which is incorporated herein by reference.
[0283] Thus, perturbations of CTCF binding result in variations in the fragmentation pattern of cfDNA. Therefore, CTCF binding sites represent one type of fragmentation variable target region.
[0284] There are many known CTCF binding sites. For example, refer to the CTCFBSDB (CTCF Binding Site Database) available at insulatordb.uthsc.edu / on the Internet; Cuddapah et al., Genome Res. 19:24-32 (2009); Martin et al., Nat. Struct. Mol. Biol. 18:708-14 (2011); Rhee et al., Cell. 147:1408-19 (2011), each of which is incorporated herein by reference. Exemplary CTCF binding sites are nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13.
[0285] Thus, in some embodiments, the set of epigenetic target regions includes CTCF binding regions. In some embodiments, the CTCF binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as the CTCF binding regions in the above or in CTCFBSDB or one or more of the papers by Cuddapah et al., Martin et al., or Rhee et al. cited above.
[0286] In some embodiments, at least a portion of the CTCF sites may or may not be methylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the CTCF binding sites.
[0287] iv. Transcription start site The transcription start site can also show perturbation in neoplastic cells. For example, nucleosome organization at various transcription start sites in healthy cells of the hematopoietic lineage substantially contributes to cfDNA in healthy individuals, but can be different from that at those transcription start sites in neoplastic cells. This results in different cfDNA patterns, which can be detected by sequencing, for example, as generally discussed in Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and US20170211143A1. In another example, the transcription start site is not necessarily epigenetically different in cancerous tissue compared to DNA from healthy tissue of the same type, but is epigenetically different (e.g., with respect to nucleosome architecture) compared to cfDNA that is typical in healthy subjects. For example, if the presence of cancer results in an increase in cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such cancer can be detected at least in part using such transcription start sites.
[0288] Thus, perturbation of the transcription start site also results in variation in the fragmentation pattern of cfDNA. Therefore, the transcription start site also represents one type of fragmentation variable target region.
[0289] Human transcription start sites are available from the DBTSS (DataBase of Human Transcription Start Sites) available at dbtss.hgc.jp on the Internet and are described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), which is incorporated herein by reference.
[0290] Thus, in some embodiments, the set of epigenetic target regions includes transcription start sites. In some embodiments, the transcription start sites include at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those described in DBTSS. In some embodiments, at least a portion of the transcription start sites may or may not be methylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the transcription start site.
[0291] v. Local amplification Local amplifications are somatic mutations, but these can be detected by sequencing based on read frequency in a manner similar to approaches for detecting certain epigenetic changes such as changes in methylation. Thus, regions that can show local amplification in cancer can be included in the set of epigenetic target regions, and they can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, the set of epigenetic target regions includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the above targets.
[0292] vi. Methylation control region It may be useful to include control regions to facilitate data verification. In some embodiments, the set of epigenetic target regions includes control regions that are expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is from cancer cells or normal cells. In some embodiments, the set of epigenetic target regions includes control hypomethylated regions that are expected to be hypomethylated in essentially all samples. In some embodiments, the set of epigenetic target regions includes control hypermethylated regions that are expected to be hypermethylated in essentially all samples.
[0293] b. Set of sequence-variable target regions In some embodiments, the set of sequence-variable target regions includes a plurality of regions known to undergo somatic mutations in cancer.
[0294] In some aspects, the set of sequence-variable target regions targets a plurality of different genes or genomic regions (a "panel") selected such that a predetermined proportion of subjects with cancer exhibit gene variants or tumor markers in one or more different genes or genomic regions in the panel. The panel can be selected to limit the sequencing region to a fixed number of base pairs. The panel can be selected to sequence a desired amount of DNA, for example, by adjusting the affinity and / or amount of the probe as described elsewhere herein. The panel can further be selected to achieve a desired depth of sequence reads. The panel can be selected to achieve a desired sequence read depth or sequence read coverage with respect to the amount of sequenced base pairs. The panel can be selected to achieve theoretical sensitivity, theoretical specificity, and / or theoretical accuracy with respect to the detection of one or more gene variants in a sample.
[0295] Probes for detecting panels of regions may include probes for detecting a genomic region of interest (hotspot region) as well as probes for detecting nucleosome recognition probes (e.g., KRAS codons 12 and 13), and may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation affected by nucleosome binding patterns and GC sequence composition. Regions used herein may also include non-hotspot regions optimized based on nucleosome position and GC models.
[0296] Examples of lists of genomic positions of interest can be found in Tables 3 and 4. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least a portion of at least 1, at least 2, or 3 of the indels in Table 3. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes in Table 4. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4.In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions of Table 4. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels of Table 4. Each of these genomic positions of interest can be identified as a backbone region or a hotspot region for a given panel. An example list of hot spot genomic positions of interest can be found in Table 5. In some embodiments, the set of array variable target regions used in the methods of the present disclosure includes at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes of Table 5. Each hot spot genomic region is described with several features including the associated gene, the chromosome on which it is present, the starting and stopping positions of the genome representing the locus, the length of the base pairs of the locus, the exons covered by the gene, and important features that the given genomic region of interest may capture (e.g., type of mutation).
Table 3
Table 4
Table 5-1
Table 5-2
[0297] In addition or alternatively, suitable target region sets are available from the literature. For example, Gale et al., PLoS One 13: e0194630 (2018), which is incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of a set of sequence-variable target regions. These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0298] In some embodiments, the set of sequence-variable target regions includes target regions from at least 10, 20, 30, or 35 cancer-related genes, such as the cancer-related genes described above.
[0299] 5. Subject In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having a neoplasm. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having a neoplasm. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject in a remission state from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm is a cancer, tumor, or neoplasm of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm is a cancer, tumor, or neoplasm of the lung. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm is a cancer, tumor, or neoplasm of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm is a cancer, tumor, or neoplasm of the breast. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm is a cancer, tumor, or neoplasm of the prostate. In any of the foregoing embodiments, the subject can be a human subject.
[0300] 6. Sequencing Sample nucleic acids adjacent to the adapter are generally subjected to sequencing, with or without prior amplification. Examples of sequencing methods include Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, Digital Gene Expression (Helicos), next-generation sequencing (NGS), single-molecule sequencing by synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. The sequencing reaction can be carried out in various sample processing units, which can include multiple lanes, multiple channels, multiple wells, or other means of processing multiple sets of samples substantially simultaneously. The sample processing unit can also include multiple sample chambers capable of processing multiple runs simultaneously.
[0301] The sequencing reaction can be performed on one or more types of nucleic acids, at least one of which is known to contain markers for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. In some embodiments, the sequence coverage of the genome can be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100%. In some embodiments, the sequencing reaction can provide sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the genome. The sequence coverage can be performed on at least 5, 10, 20, 70, 100, 200, or 500 different genes, or at most 5000, 2500, 1000, 500, or 100 different genes.
[0302] The simultaneous sequencing reaction can be performed using multiplex sequencing. In some examples, the cell-free nucleic acid can be sequenced by at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other examples, the cell-free nucleic acid can be sequenced by less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. The sequencing reactions can be performed sequentially or simultaneously. Subsequent data analysis can be performed on all or part of the sequencing reactions. In some examples, the data analysis can be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other examples, the data analysis can be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. An exemplary read depth is about 1000 to about 50000 reads per locus (base).
[0303] a. Differential depth of sequencing In some embodiments, the nucleic acids corresponding to the set of sequence variable target regions are sequenced to a higher sequencing depth than the nucleic acids corresponding to the set of epigenetic target regions. For example, the sequencing depth for the nucleic acids corresponding to the set of sequence variant target regions is at least 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 times higher than the sequencing depth for the nucleic acids corresponding to the set of epigenetic target regions, or 1.25 - 1.5, 1.5 - 1.75, 1.75 - 2, 2 - 2.25, 2.25 - 2.5, 2.5 - 2.75, 2.75 - 3, 3 - 3.5, 3.5 - 4, 4 - 4.5, 4.5 - 5, 5 - 5.5, 5.5 - 6, 6 - 7, 7 - 8, 8 - 9, 9 - 10, 10 - 11, 11 - 12, 13 - 14, 14 - 15 times, or 15 - 100 times higher. In some embodiments, the sequencing depth is at least 2 times higher. In some embodiments, the sequencing depth is at least 5 times higher. In some embodiments, the sequencing depth is at least 10 times higher. In some embodiments, the sequencing depth is 4 - 10 times higher. In some embodiments, the sequencing depth is 4 - 100 times higher. Each of these embodiments refers to the extent to which the nucleic acids corresponding to the set of sequence variable target regions are sequenced to a higher sequencing depth than the nucleic acids corresponding to the set of epigenetic target regions.
[0304] In some embodiments, the captured cfDNA corresponding to the set of sequence variable target regions and the captured cfDNA corresponding to the set of epigenetic target regions are sequenced simultaneously, for example, in the same sequencing cell (e.g., a flow cell of an Illumina sequencer) and / or in the same composition, and they may be a pooled composition obtained by recombining separately captured sets, or a composition obtained by capturing the cfDNA corresponding to the set of sequence variable target regions and the captured cfDNA corresponding to the set of epigenetic target regions in the same container.
[0305] 7. Analysis In some embodiments, the methods described herein include identifying the presence or absence of DNA produced by a tumor (or neoplastic or cancer cells).
[0306] The methods can be used to diagnose a condition in a subject, particularly the presence or absence of cancer, characterize the condition (e.g., stage cancer or determine cancer heterogeneity), monitor the response to treatment of the condition, and provide a prognostic risk of the occurrence or subsequent course of the condition. The present disclosure can also be useful in determining the effectiveness of a particular treatment option. When more cancer can be killed and shed DNA, if a treatment is successful, the successful treatment option can increase the amount of copy number variation or rare mutations detected in the subject's blood. In other instances, this may not occur. In another example, perhaps a particular treatment option can correlate with the cancer's genetic profile over time. This correlation can be useful in the selection of therapy.
[0307] Furthermore, if it is observed that the cancer is in a remission state after treatment, the methods can be used to monitor for residual disease or disease recurrence.
[0308] The types and numbers of cancers that can be detected can include blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, pharyngeal cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, homogeneous tumors, and the like. The type and / or stage of cancer can be detected by genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, segmental aneuploidy, ploidy, chromosomal instability, changes in chromosomal structure, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0309] Gene data can also be used to characterize specific forms of cancer. Cancer is often heterogeneous both in its composition and stage classification. Gene profile data can enable the characterization of specific subtypes of cancer that may be important in the diagnosis or treatment of that specific subtype. This information provides clues to the prognosis of a specific type of cancer to the subject or clinician, and enables the subject or clinician to adapt treatment options as the disease progresses. Some cancers may progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure may be useful in determining the progression of the disease.
[0310] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods may include, for example, generating a gene profile of extracellular polynucleotides derived from the subject, the gene profile including a plurality of data resulting from the analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition can be a condition that results in a heterogeneous genomic population. In the example of cancer, some tumors are known to contain tumor cells of different stages of cancer. In other examples, the heterogeneity can include multiple lesions of the disease. Again, in the example of cancer, there are multiple tumor lesions, and perhaps one or more of the lesions are the result of metastases that have spread from the primary site.
[0311] The method can be used to generate or profile a fingerprint or dataset that is the sum of gene information derived from different cells in a heterogeneous disease. This dataset can include the analysis of copy number variations, epigenetic variations, and mutations, either alone or in combination.
[0312] The methods can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods herein do not include the diagnosis, prognostication, or monitoring of a fetus and thus are not directed to non-invasive prenatal testing. In other embodiments, these methodologies can be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in an unborn subject whose DNA and other polynucleotides can circulate with maternal molecules.
[0313] An exemplary method for identifying molecular tags of an MBD-bead partitioning library through NGS includes the step of subjecting a first secondary sample to a procedure that affects a first nucleobase in the DNA such that it is different from a second nucleobase in the DNA of the first secondary sample, as follows: 1. Physical partitioning of an extracted DNA sample (e.g., plasma DNA extracted from a human sample that has been subjected to target capture as described herein, if desired) using a methyl-binding domain protein-bead purification kit, and storing all eluates from the process for downstream processing. 2. Parallel application of differential molecular tags and adapter sequences enabling NGS to each fraction. For example, fractions of hypermethylation, residual methylation (“washing”), and hypomethylation are ligated to NGS-adapters with molecular tags. 3. Subjecting the hypermethylated fraction to a procedure that affects a first nucleobase in the DNA such that it is different from a second nucleobase in the DNA, such as any of the procedures described herein. 4. Recombining all the molecular-tagged fractions and then amplifying using adapter-specific DNA primer sequences. 5. Capture / hybridization of the recombined and amplified total library targeting genomic regions of interest (e.g., cancer-specific gene variants and differentially methylated regions). 6. Re-amplification of the captured DNA library with sample tagging. Different samples are pooled and assayed multiplexly on an NGS instrument. 7. Bioinformatics analysis of NGS data and deconvolution of samples into differentially MBD-assigned molecules, where molecular tags are used to identify unique molecules. This analysis can generate information regarding relative 5-methylcytosine in genomic regions simultaneously with standard gene sequencing / variant detection.
[0314] In some embodiments of the methods described herein, including but not limited to the methods shown above, the molecular tag consists of nucleotides that do not change by a procedure that affects the first nucleobase in DNA so as to be different from a second nucleobase in DNA, such as any of the procedures described herein (e.g., if the procedure is bisulfite conversion or any other conversion that does not affect mC, mC along with A, T, and G; if the procedure is a conversion that does not affect hmC, hmC along with A, T, and G, etc.). In some embodiments of the methods described herein, including but not limited to the methods shown above, the molecular tag does not contain nucleotides that do not change by a procedure that affects the first nucleobase in DNA so as to be different from a second nucleobase in DNA, such as any of the procedures described herein (e.g., if the procedure is bisulfite conversion or any other conversion that affects C, the tag does not contain unmodified C; if the procedure is a conversion that affects mC, the tag does not contain mC; if the procedure is a conversion that affects hmC, the tag does not contain hmC, etc.).
[0315] Generally, the procedure that affects the first nucleobase in DNA so as to be different from the second nucleobase in DNA can be performed prior to the step of concurrently applying differential molecular tags and adapter sequences enabling NGS to each fraction. For example, this can be done if the procedure that affects the first nucleobase in DNA so as to be different from the second nucleobase in DNA is a separation, such as hmC-seal, in which case the separated populations can themselves be differentially tagged relative to each other. Such exemplary methods are as follows: 1. Physical partitioning of an extracted DNA sample (e.g., plasma DNA extracted from a human sample that has been subjected to target capture as described herein, if necessary) using a methyl-binding domain protein-bead purification kit, and preservation of all eluates from the process for downstream processing. 2. Subject the hypermethylated fraction to a procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA, such as any of the procedures described herein. 3. Parallel application of differential molecular tags and adapter sequences enabling NGS to each fraction. For example, ligate the hypermethylated fraction (or, where applicable, two or more sub-fractions of the hypermethylated fraction), the residual methylation ("wash") fraction, and the hypomethylated fraction to the NGS-adapter with the molecular tag. 4. Combine all the molecular-tagged fractions again and then amplify using an adapter-specific DNA primer sequence. 5. Capture / hybridization of the combined and amplified total library targeting genomic regions of interest (e.g., cancer-specific gene variants and differentially methylated regions). 6. Re-amplify the captured DNA library with sample tagging. Pool different samples for multiplexed assay on an NGS instrument. 7. Bioinformatics analysis of the NGS data where the molecular tag is used to identify unique molecules, and deconvolution of the samples to the molecules differentially distributed by MBD. This analysis can generate information regarding 5-methylcytosine relative to genomic regions simultaneously with standard gene sequencing / variant detection.
[0316] 8. Exemplary workflow An exemplary workflow for partitioning and library preparation is provided herein. In some embodiments, some or all of the features of the partitioning and library preparation workflow may be used in combination.
[0317] a. Partitioning In some embodiments, sample DNA (e.g., between 5 ng and 200 ng) is mixed with a methyl-binding domain (MBD) buffer and magnetic beads conjugated to an MBD protein and incubated overnight. Methylated DNA (hypermethylated DNA) binds to the MBD protein on the magnetic beads during this incubation. Unmethylated (hypomethylated DNA) or less methylated DNA (intermediate methylation) is washed away from the beads with a buffer containing an increased salt concentration. For example, one, two, or more fractions containing unmethylated, hypomethylated, and / or intermediate methylated DNA can be obtained by such washing. Finally, a buffer with a high salt concentration is used to elute the highly methylated DNA (hypermethylated DNA) from the MBD protein. In some embodiments, these washes result in three fractions of DNA with increasing methylation levels (a hypomethylated fraction, an intermediate methylated fraction, and a hypermethylated fraction).
[0318] In some embodiments, the three fractions of DNA are desalted and concentrated during the preparation of the enzymatic steps for library preparation.
[0319] b. Library preparation In some embodiments (e.g., after enrichment of DNA in a fraction), the distributed DNA is made ligatable, for example, by extending the overhangs at the ends of the DNA molecules, adding adenosine residues to the 3' ends of the fragments, and phosphorylating the 5' ends of each DNA fragment. DNA ligase and adapters are added to ligate an adapter to each distributed DNA molecule at each end. These adapters contain a fraction tag (e.g., a non-random, non-unique barcode) distinguishable from the fraction tags in the adapters used in other fractions. Before or after making the fragmented DNA ligatable and performing the ligation, the highly methylated fraction is subjected to a procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA, such as any of the procedures described herein. If the procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA further distributes the highly methylated fraction, adapter ligation must be performed after the procedure so that the sub-fractions of the highly methylated fraction can be differentially tagged. Next, three (or more) fractions are pooled together and amplified (e.g., by PCR using primers specific for the adapter).
[0320] After PCR, the amplified DNA may be purified and concentrated prior to enrichment. The amplified DNA is contacted with a collection of probes described herein that target the specific region of interest (which may be, for example, a biotinylated RNA probe). The mixture is incubated, for example, overnight in a salt buffer. The probes are captured (e.g., using streptavidin magnetic beads) and separated from the amplified DNA that was not captured, for example, by a series of salt washes, thereby enriching the sample. After enrichment, the enriched sample is amplified by PCR. In some embodiments, the PCR primers contain a sample tag, thereby incorporating the sample tag into the DNA molecule. In some embodiments, after pooling DNA from different samples, multiplex sequencing is performed, for example, using an Illumina NovaSeq sequencer.
[0321] C. Additional Features of a Particular Disclosed Method 1. Sample The sample may be any biological sample isolated from a subject. The sample may be a bodily sample. The sample may be a bodily tissue such as a solid tumor, known or suspected, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (leucocytes), endothelial cells, tissue biopsy material, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, fluid in the intercellular space, such as gingival sulcus exudate, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. The sample is preferably a body fluid, particularly blood and its fractions, as well as urine. The sample may be in its original form isolated from the subject, or further processed to remove or add components such as cells, or to enrich one component relative to another. Thus, a preferred body fluid for analysis is plasma or serum containing cell-free nucleic acids. The sample can be isolated or obtained from the subject and transported to the location of sample analysis. The sample can be stored and shipped at a desired temperature, such as room temperature, 4°C, -20°C, and / or -80°C. The sample can be isolated or obtained from the subject at the location of sample analysis. The subject may be a human, mammal, animal, companion animal, service animal, or pet. The subject may have cancer. The subject may not have cancer or detectable cancer symptoms. The subject may be treated with one or more cancer therapies, such as chemotherapy, antibodies, vaccines, or any one or more of biological agents. The subject may be in remission. The subject may or may not be diagnosed as being predisposed to cancer or any cancer-related genetic mutation / disorder. In some embodiments, the sample is a polynucleotide sample obtained from a tumor tissue biopsy.
[0322] The volume of plasma may depend on the desired read depth of the region to be sequenced. Exemplary volumes are 0.4 - 40 ml, 5 - 20 ml, 10 - 20 ml. For example, the volume may be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma collected may be 5 - 20 mL.
[0323] The sample can contain various amounts of nucleic acid containing genomic equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genomic equivalents, and in the case of cfDNA, can contain about 200 billion (2×10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA contains about 30,000 haploid human genomic equivalents, and in the case of cfDNA, can contain about 600 billion individual molecules.
[0324] The sample can contain nucleic acids from different sources, for example, from cells and cell-free of the same subject, from cells and cell-free of different subjects. The sample can contain nucleic acids having mutations. For example, the sample can contain DNA having germline mutations and / or somatic mutations. A germline mutation means a mutation present in the germline DNA of the subject. A somatic mutation means a mutation resulting from somatic cells of the subject, such as cancer cells. The sample can contain DNA having cancer-related mutations (e.g., cancer-related somatic mutations). The sample may contain epigenetic variants (i.e., chemical or protein modifications) that are associated with the presence of genetic variants such as cancer-related mutations. In some embodiments, a sample that does not contain genetic variants contains epigenetic variants associated with the presence of genetic variants.
[0325] Exemplary amounts of cell-free nucleic acids in a sample prior to amplification range from about 1 fg to about 1 μg, such as, for example, in the range of 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount may be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount may be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount may be 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or up to 200 ng of cell-free nucleic acid molecules. The method may include obtaining from the sample cell-free nucleic acid molecules from 1 femtogram (fg) to 200 ng.
[0326] Cell-free nucleic acids are nucleic acids that are not contained within cells or otherwise bound to cells, or in other words, nucleic acids that remain in a sample after intact cells have been removed. Cell-free nucleic acids include genomic DNA, mitochondrial DNA, siRNA, miRNA, circular RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or DNA, RNA, and hybrids thereof containing fragments of any of these. Cell-free nucleic acids may be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids via processes of secretion or cell death, such as necrosis and apoptosis of cells. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released from cancer cells into body fluids. Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, the cell-free nucleic acids are produced by tumor cells. In some embodiments, the cell-free nucleic acids are produced from a mixture of tumor cells and non-tumor cells.
[0327] Cell-free nucleic acids have an exemplary size distribution of about 100 to 500 nucleotides, with molecules of 110 to about 230 nucleotides corresponding to about 90% of the molecules, a mode of about 168 nucleotides, and a second minor peak in the range of 240 to 440 nucleotides.
[0328] Cell-free nucleic acids can be isolated from body fluids through fractionation or partitioning steps, in which the cell-free nucleic acids found in solution are separated from intact cells or other insoluble components of the body fluid. Partitioning can include techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid can be lysed and the nucleic acids of both cell-free and cells can be processed together. Generally, after the addition of buffer and washing steps, the nucleic acids can be precipitated with alcohol. Further purification steps such as silica-based columns to remove contaminants or salts may be used. In certain embodiments of this procedure, for example, to optimize the yield, non-specific bulk carrier nucleic acids such as C1 DNA, DNA, or protein for bisulfite sequencing, hybridization, and / or ligation can be added to the entire reaction.
[0329] After such treatment, the sample can contain nucleic acids in various forms including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA can be converted to double-stranded form and included in subsequent processing and analysis steps.
[0330] Double-stranded DNA molecules and single-stranded nucleic acid molecules converted to double-stranded DNA molecules in the sample can be ligated to the adapter at one or both ends. Typically, double-stranded molecules are blunt-ended by treatment with a polymerase containing 5'-3' polymerase and 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenow large fragment and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecules can be ligated to at least partially double-stranded adapters (such as Y-shaped or bell-shaped adapters). Alternatively, complementary nucleotides can be added to the blunt ends of the sample nucleic acid and the adapter to facilitate ligation. Both blunt-end ligation and sticky-end ligation are contemplated herein. In blunt-end ligation, both the nucleic acid molecule and the adapter tag have blunt ends. In sticky-end ligation, typically the nucleic acid molecule has an "A" overhang and the adapter has a "T" overhang.
[0331] 2. Amplification Sample nucleic acids adjacent to the adapter can be amplified by PCR and other amplification methods. Amplification is typically initiated by a primer that binds to a primer binding site in the adapter adjacent to the DNA molecule to be amplified. The amplification method may include cycles of denaturation, annealing, and extension due to thermal cycling, or may be isothermal as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication.
[0332] In some embodiments, the method performs dsDNA ligation using T-tail and C-tail adapters, which results in amplification of at least 50, 60, 70, or 80% of the double-stranded nucleic acid. Preferably, the method results in at least a 10, 15, or 20% increase in the amount or number of amplified molecules compared to a control method performed using only the T-tail adapter.
[0333] 3. Tags Tags containing barcodes can be incorporated into the adapter or otherwise attached to the adapter. The tags can be incorporated, among other methods, by ligation, overlap extension PCR.
[0334] a. Molecular tagging strategy In some embodiments, a nucleic acid molecule (derived from a sample of polynucleotides) can be tagged with a sample index and / or a molecular barcode (commonly referred to as a “tag”). The tag can be incorporated into an adapter by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or other methods such as overlapping-extension polymerase chain reaction (PCR), or ligated in other ways. Such an adapter may ultimately ligate to the target nucleic acid molecule. In other embodiments, one or more rounds of an amplification cycle (e.g., PCR amplification) are applied to introduce a sample index into the nucleic acid molecule, generally using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., multiple microwells in an array). The molecular barcode, and / or the sample index can be introduced simultaneously or in any sequential order. In some embodiments, the molecular barcode and / or the sample index are introduced before and / or after performing a sequence capture step. In some embodiments, only the molecular barcode is introduced prior to the probe capture step, and the sample index is introduced after performing the sequence capture step. In some embodiments, both the molecular barcode and the sample index are introduced prior to performing a probe-based capture step. In some embodiments, the sample index is introduced after performing the sequence capture step. In some embodiments, the molecular barcode is incorporated into nucleic acid molecules (e.g., cfDNA) in a sample through an adapter via ligation (e.g., blunt-end ligation or sticky-end ligation). In some embodiments, the sample index is incorporated into nucleic acid molecules (e.g., cfDNA molecules) in a sample through overlapping-extension polymerase chain reaction (PCR). Typically, a sequence capture protocol involves introducing single-stranded nucleic acid molecules complementary to targeted nucleic acid sequences (e.g., the coding sequences of genomic regions), where mutations in such regions are associated with cancer types.
[0335] In some embodiments, the tag may be located at one or both ends of the sample nucleic acid molecule. In some embodiments, the tag is an oligonucleotide of a predefined, or random or semi-random sequence. In some embodiments, the tag may be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide in length. The tag may be linked to the sample nucleic acid randomly or non-randomly.
[0336] In some embodiments, each sample is uniquely tagged by a sample index or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or a subsample is uniquely tagged by a molecular barcode or a combination of molecular barcodes. In other embodiments, multiple molecular barcodes may be used, such that the molecular barcodes are not necessarily unique to each other among the multiple molecular barcodes (e.g., non-unique molecular barcodes). In these embodiments, the molecular barcodes are generally attached to individual molecules (e.g., by ligation) to generate unique sequences that can be individually tracked in combination with the molecular barcode and the sequences to which it can be attached. Detection of non-unique molecular barcodes in combination with endogenous sequence information (e.g., start (departure) and / or stop (termination) genomic locations / positions corresponding to the sequence of the original nucleic acid molecule in the sample, start and stop genomic positions corresponding to the sequence of the original nucleic acid molecule in the sample, start (departure) and / or stop (termination) genomic locations / positions of sequence reads mapped to a reference sequence, start and stop genomic positions of sequence reads mapped to a reference sequence, secondary sequences of sequence reads at one or both ends, length of the sequence reads, and / or length of the original nucleic acid molecule in the sample) typically enables the assignment of a unique identity to a particular molecule. In some embodiments, the start region includes the position of the first base, the position of the first 2 bases, the position of the first 5 bases, the position of the first 10 bases, the position of the first 15 bases, the position of the first 20 bases, the position of the first 25 bases, the position of the first 30 bases, or at least the position of the first 30 bases at the 5' end of the sequencing read that aligns with the reference sequence. In some embodiments, the stop region includes the position of the last base, the position of the last 2 bases, the position of the last 5 bases, the position of the last 10 bases, the position of the last 15 bases, the position of the last 20 bases, the position of the last 25 bases, the position of the last 30 bases, or at least the position of the last 30 bases at the 3' end of the sequencing read that aligns with the reference sequence. The length of an individual sequence read or the number of its base pairs is also used, if necessary, to assign a unique identity to a given molecule.When described in this specification, a fragment from a single strand of a nucleic acid to which a unique identity has been assigned may enable subsequent identification of fragments from the parental strand and / or complementary strand thereby.
[0337] In certain embodiments, the number of different tags used to uniquely identify the number z of molecules in a class is 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z or 100 * either z (e.g., lower limit) and 100,000 * z, 10,000 * z, 1000 * z or 100 *It can be between any of z (e.g., the upper limit). In some embodiments, the molecular barcodes are introduced at an expected ratio for the molecules of a set of identifiers in the sample (e.g., a combination of unique or non-unique molecular barcodes). One example format uses from about 2 to about 1,000,000 different molecular barcode sequences, or from about 5 to about 150 different molecular barcode sequences, or from about 20 to about 50 different molecular barcode sequences that are ligated to both ends of the target molecule. Alternatively, from about 25 to about 1,000,000 different molecular barcode sequences may be used. For example, 20 - 50×20 - 50 molecular barcode sequences (i.e., one of 20 - 50 different molecular barcode sequences can be attached to each end of the target molecule) can be used. Such a number of identifiers is typically sufficient for different molecules having the same starting and stopping points to have a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of receiving different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same combination of molecular barcodes.
[0338] In some embodiments, the assignment of unique or non-unique molecular barcodes in the reaction is performed using, for example, the methods and systems described in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patents Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, only endogenous sequence information (e.g., starting and / or stopping positions, secondary sequences at one or both ends of the sequence, and / or length) may be used to identify different nucleic acid molecules in the sample.
[0339] 4. Bait Set; Capture Moiety As discussed above, the nucleic acids in a sample can be subjected to a capture step, in which molecules having a target sequence are captured for subsequent analysis. Target capture can involve the use of a bait set that includes baits of oligonucleotides labeled with a capture moiety such as biotin or other examples described below. The probes can have sequences that are selected to tile across a panel of regions, such as genes. In some embodiments, the bait set can have a higher or lower capture yield for a set of target regions, such as the capture yields of a set of sequence variable target regions and a set of epigenetic target regions, respectively, as discussed elsewhere herein. Such a bait set is combined with the sample under conditions that allow hybridization of the target molecules to the baits. Next, the captured molecules are isolated using a capture moiety, such as biotin capture moiety by streptavidin based beads. Such methods are further described, for example, in U.S. Patent No. 9,850,523, issued December 26, 2017, which is incorporated herein by reference.
[0340] Capture moieties include, without limitation, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically attractable particles. The extraction moiety can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, the capture moiety bound to the analyte is captured by a binding pair that binds to an isolatable moiety, such as a magnetically attractable particle or a large particle that can sediment by centrifugation. The capture moiety can be any type of molecule that allows affinity separation of a nucleic acid having the capture moiety from a nucleic acid lacking the capture moiety. Exemplary capture moieties are biotin, which allows affinity separation by binding to streptavidin that is linked or linkable to a solid phase, or oligonucleotides, which allow affinity separation through binding to a complementary oligonucleotide that is linked or linkable to a solid phase.
[0341] D. Collection of Target-Specific Probes In some embodiments, a collection of target-specific probes is used in the methods described herein. In some embodiments, the collection of target-specific probes includes target-binding probes specific for a set of sequence-variable target regions and target-binding probes specific for a set of epigenetic target regions. In some embodiments, the capture yield of the target-binding probes specific for the set of sequence-variable target regions is higher (e.g., at least two-fold higher) than the capture yield of the target-binding probes specific for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield for the set of sequence-variable target regions that is higher (e.g., at least two-fold higher) than its capture yield for the set of epigenetic target regions.
[0342] In some embodiments, the capture yield of the target-binding probes specific for the set of sequence-variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the capture yield of the target-binding probes specific for the set of epigenetic target regions. In some embodiments, the capture yield of the target-binding probes specific for the set of sequence-variable target regions is 1.25 to 1.5-fold, 1.5 to 1.75-fold, 1.75 to 2-fold, 2 to 2.25-fold, 2.25 to 2.5-fold, 2.5 to 2.75-fold, 2.75 to 3-fold, 3 to 3.5-fold, 3.5 to 4-fold, 4 to 4.5-fold, 4.5 to 5-fold, 5 to 5.5-fold, 5.5 to 6-fold, 6 to 7-fold, 7 to 8-fold, 8 to 9-fold, 9 to 10-fold, 10 to 11-fold, 11 to 12-fold, 13 to 14-fold, or 14 to 15-fold higher than the capture yield of the target-binding probes specific for the set of epigenetic target regions.
[0343] In some embodiments, the collection of target-specific probes is configured to have a capture yield specific to the set of variable sequence target regions that is at least 1.25 times, 1.5 times, 1.75 times, 2 times, 2.25 times, 2.5 times, 2.75 times, 3 times, 3.5 times, 4 times, 4.5 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 11 times, 12 times, 13 times, 14 times, or 15 times higher than its capture yield for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is 1.25 - 1.5 times, 1.5 - 1.75 times, 1.75 - 2 times, 2 - 2.25 times, 2.25 - 2.5 times, 2.5 - 2.75 times, 2.75 - 3 times, 3 - 3.5 times, 3.5 - 4 times, 4 - 4.5 times, 4.5 - 5 times, 5 - 5.5 times, 5.5 - 6 times, 6 - 7 times, 7 - 8 times, 8 - 9 times, 9 - 10 times, 10 - 11 times, 11 - 12 times, 13 - 14 times, or 14 - 15 times higher than its capture yield specific to the set of epigenetic target regions for the set of variable sequence target regions.
[0344] The collection of probes can be configured to provide a higher capture yield for the set of variable sequence target regions in various ways, including enrichment, various lengths and / or chemistries (such as those affecting affinity), and combinations thereof. Affinity can be modulated by adjusting the length of the probe and / or by including nucleotide modifications discussed below.
[0345] In some embodiments, target-specific probes that are specific for a set of array-variable target regions are present at a higher concentration than target-specific probes that are specific for a set of epigenetic target regions. In some embodiments, the concentration of target-binding probes that are specific for a set of array-variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the concentration of target-binding probes that are specific for a set of epigenetic target regions. In some embodiments, the concentration of target-binding probes that are specific for a set of array-variable target regions is 1.25 to 1.5-fold, 1.5 to 1.75-fold, 1.75 to 2-fold, 2 to 2.25-fold, 2.25 to 2.5-fold, 2.5 to 2.75-fold, 2.75 to 3-fold, 3 to 3.5-fold, 3.5 to 4-fold, 4 to 4.5-fold, 4.5 to 5-fold, 5 to 5.5-fold, 5.5 to 6-fold, 6 to 7-fold, 7 to 8-fold, 8 to 9-fold, 9 to 10-fold, 10 to 11-fold, 11 to 12-fold, 13 to 14-fold, or 14 to 15-fold higher than the concentration of target-binding probes that are specific for a set of epigenetic target regions. In such embodiments, the concentration can refer to the mass / volume concentration of the average of the individual probes within each set.
[0346] In some embodiments, target-specific probes that are specific for a set of array-variable target regions have a higher affinity for their target than target-specific probes that are specific for an epigenetic target region set. The affinity can be modulated by any method known to those of skill in the art, including using different probe chemistries. For example, modifications of certain nucleotides, such as cytosine 5-methylation (in the context of a particular sequence), introducing a heteroatom at the 2'-sugar position, and LNA nucleotides can increase the stability of double-stranded nucleic acids, and oligonucleotides having such modifications have been shown to have a relatively high affinity for their complementary sequences. See, for example, Severin et al., Nucleic Acids Res. 39: 8740-8751 (2011), Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997), U.S. Patent No. 9,738,894. Also, a long sequence length generally provides an increase in affinity. Other nucleotide modifications, such as substitution of guanine with the nucleobase hypoxanthine, reduce the affinity by reducing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, target-specific probes that are specific for a set of array-variable target regions have modifications that increase their affinity for their target. In some embodiments, alternatively or in addition, target-specific probes that are specific for an epigenetic target region set have modifications that decrease their affinity for their target. In some embodiments, target-specific probes that are specific for a set of array-variable target regions have a longer average length and / or a higher average melting temperature than target-specific probes that are specific for an epigenetic target region set. These embodiments can be combined with each other and / or with differences in concentration, as discussed above, to achieve a desired fold difference in capture yield, such as any of the fold differences or ranges thereof described above.
[0347] In some embodiments, the target-specific probe includes a capture moiety. The capture moiety can be any of the capture moieties described herein, such as biotin. In some embodiments, the target-specific probe is linked to a solid support, for example, by a covalent bond or a non-covalent bond such as an interaction of a binding pair of the capture moiety. In some embodiments, the solid support is a bead, such as a magnetic bead.
[0348] In some embodiments, the target-specific probe specific to a set of array-variable target regions and / or the target-specific probe specific to a set of epigenetic target regions is a probe comprising a capture moiety and a sequence selected to tile across a panel of regions such as genes of the bait set discussed above.
[0349] In some embodiments, the target-specific probe is provided as a single composition. The single composition can be a solution (liquid or frozen). Alternatively, the composition can be lyophilized.
[0350] Alternatively, the target-specific probe can be provided as a plurality of compositions, including a first composition comprising a probe specific to a set of epigenetic target regions and a second composition comprising a probe specific to a set of array-variable target regions. These probes can be mixed in appropriate ratios to provide a combined probe composition having any of the above-fold differences in concentration and / or capture yield. Alternatively, these probes can be used in separate capture procedures (e.g., with aliquots of the sample or sequentially with the same sample) to provide first and second compositions each comprising a captured epigenetic target region and an array-variable target region, respectively.
[0351] 1. Probe Specific to Epigenetic Target Region Probes for an epigenetic target region set can include one or more types of probes specific for one or more target regions that can distinguish DNA from neoplastic (e.g., tumor or cancer) cells from healthy cells, such as non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein, for example, in the above section regarding the captured set. Probes for an epigenetic target region set can also include, for example, probes for one or more control regions described herein.
[0352] In some embodiments, the probes of an epigenetic target region probe set have a footprint of at least 100 kb, such as at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the probes of an epigenetic target region set have a footprint in the range of 100 - 1000 kb, such as 100 - 200 kb, 200 - 300 kb, 300 - 400 kb, 400 - 500 kb, 500 - 600 kb, 600 - 700 kb, 700 - 800 kb, 800 - 900 kb, and 900 - 1,000 kb. In some embodiments, the probes of an epigenetic target region probe set have a footprint of less than 5 kb, at least 5 kb, such as at least 10, 20, or 50 kb.
[0353] a. Hypermethylation variable target region In some embodiments, the probes for the set of epigenetic target regions include probes specific for one or more hypermethylation variable target regions. The hypermethylation variable target region may be any of the target regions described above. For example, in some embodiments, the probes specific for the hypermethylation variable target region include probes specific for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the plurality of loci listed in Table 1, such as the loci listed in Table 1. In some embodiments, the probes specific for the hypermethylation variable target region include probes specific for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the plurality of loci listed in Table 2, such as the loci listed in Table 2. In some embodiments, the probes specific for the hypermethylation variable target region include probes specific for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the plurality of loci listed in Table 1 or Table 2, such as the loci listed in Table 1 or Table 2. In some embodiments, for each locus included as a target region, there may be one or more probes having a hybridization site that binds between the transcription start site of the gene and the stop codon (alternatively, the last stop codon for alternatively spliced genes). In some embodiments, this one or more probes bind within 300 bp of the listed position, such as within 200 or 100 bp. In some embodiments, the probe has a hybridization site that overlaps with the listed position. In some embodiments, the probes specific for the hypermethylation variable target region include one, two, three, four, or five probes specific for one, two, three, four, or five subsets of hypermethylated target regions that collectively show hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0354] b. Hypomethylation variable target region In some embodiments, the probes for the set of epigenetic target regions include probes specific for one or more hypomethylated variable target regions. The hypomethylated variable target regions can be any of the target regions described above. For example, probes specific for one or more hypomethylated variable target regions can include probes for repetitive elements, such as LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and regions such as satellite DNA, and intergenic regions that are normally methylated in healthy cells but show reduced methylation in tumor cells.
[0355] In some embodiments, the probes specific for hypomethylated variable target regions include probes specific for repetitive elements and / or intergenic regions. In some embodiments, the probes specific for repetitive elements include probes specific for one, two, three, four, or five of LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.
[0356] Exemplary probes specific for genomic regions showing cancer-related hypomethylation include probes specific for nucleotides 8403565 - 8953708 and / or 151104701 - 151106035 of human chromosome 1. In some embodiments, the probes specific for hypomethylated variable target regions include probes specific for regions that overlap with or include nucleotides 8403565 - 8953708 and / or 151104701 - 151106035 of human chromosome 1.
[0357] c.CTCF binding region In some embodiments, the probes for the set of epigenetic target regions include probes specific for CTCF binding regions. In some embodiments, the probes specific for CTCF binding regions include probes specific for at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 CTCF binding regions, such as the CTCF binding regions in one or more of the papers of Cuddapah et al., Martin et al., or Rhee et al. cited above or CTCFBSDB. In some embodiments, the probes for the set of epigenetic target regions include regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding sites.
[0358] d. Transcription start site In some embodiments, the probes for the set of epigenetic target regions include probes specific for transcription start sites. In some embodiments, the probes specific for transcription start sites include probes specific for at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 transcription start sites, such as the transcription start sites listed in DBTSS. In some embodiments, the probes for the set of epigenetic target regions include probes for sequences at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the transcription start sites.
[0359] e. Local amplification As described above, local amplifications are somatic mutations, but these can be detected by sequencing based on read frequency in a manner similar to approaches for detecting certain epigenetic changes such as changes in methylation. Thus, as discussed above, regions that can exhibit local amplifications in cancer can be included in the set of epigenetic target regions. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to local amplifications. In some embodiments, the probes specific to local amplifications include probes specific to one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, the probes specific to local amplifications include probes specific to one or more of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the above targets.
[0360] f. Control Region It may be useful to include control regions to facilitate validation of the data. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to control methylated regions that are expected to be methylated in essentially all samples. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to control hypomethylated regions that are expected to be hypomethylated in essentially all samples.
[0361] 2. Probes Specific to Variable Sequence Target Regions Probes for the array-variable target region set may include probes specific for a plurality of regions known to undergo somatic mutations in cancer. The probes may be specific for any array-variable target region set described herein. Exemplary array-variable target region sets are discussed in detail herein, for example, in the above section regarding the captured sets.
[0362] In some embodiments, the array-variable target region probe set has a footprint of at least 0.5 kb, such as at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, the epigenetic target region probe set has a footprint in the range of 0.5 to 100 kb, such as 0.5 to 2 kb, 2 to 10 kb, 10 to 20 kb, 20 to 30 kb, 30 to 40 kb, 40 to 50 kb, 50 to 60 kb, 60 to 70 kb, 70 to 80 kb, 80 to 90 kb, and 90 to 100 kb.
[0363] In some embodiments, the probes specific to the set of variable target regions include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or at least 70 of the genes in Table 3. In some embodiments, the probes specific to the set of variable target regions include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or at least 70 of the SNVs in Table 3. In some embodiments, the probes specific to the set of variable target regions include probes specific to at least 1, at least 2, at least 3, at least 4, at least 5, or at least 6 of the fusions in Table 3. In some embodiments, the probes specific to the set of variable target regions include probes specific to at least 1, at least 2, or at least 3 of the indels in Table 3. In some embodiments, the probes specific to the set of variable target regions include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or at least 73 of the genes in Table 4. In some embodiments, the probes specific to the set of variable target regions include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or at least 73 of the SNVs in Table 4. In some embodiments, the probes specific to the set of variable target regions include probes specific to at least 1, at least 2, at least 3, at least 4, at least 5, or at least 6 of the fusions in Table 4.In some embodiments, the probes specific for the set of array-variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. In some embodiments, the probes specific for the set of array-variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5.
[0364] In some embodiments, the probes specific for the set of array-variable target regions include probes specific for target regions from at least 10, 20, 30, or 35 cancer-related genes such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0365] E. A composition containing the captured DNA Provided herein is a combination comprising first and second populations of captured DNA. The first population may comprise or be derived from DNA having cytosine modifications at a higher rate than the second population. The first population may comprise the type of the first nucleobase originally present in DNA having altered base pairing specificity and a second nucleobase having no altered base pairing specificity, the type of the first nucleobase originally present in the DNA before the change in base pairing specificity is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the type of the first nucleobase and the second nucleobase originally present in the DNA before the change in base pairing specificity have the same base pairing specificity. The second population does not include the type of the first nucleobase originally present in DNA having altered base pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleobase is a modified cytosine or an unmodified cytosine, and the second nucleobase is a modified cytosine or an unmodified cytosine. The first and second nucleobases are either the nucleobases discussed in the summary of the present specification or the nucleobases discussed in connection with subjecting the first secondary sample to a procedure that affects the first nucleobase in the DNA of the first secondary sample to be different from the second nucleobase in the DNA of the first secondary sample.
[0366] In some embodiments, the first population comprises sequence tags selected from a first set of one or more sequence tags, the second population comprises sequence tags selected from a second set of one or more sequence tags, and the second set of sequence tags is different from the first set of sequence tags. The sequence tags may include barcodes.
[0367] In some embodiments, the first population comprises protected hmC, such as glucosylated hmC.
[0368] In some embodiments, the first population was subjected to any of the conversion procedures contemplated herein, such as bisulfite conversion, Ox-BS conversion, TAB conversion, ACE conversion, TAP conversion, TAPSβ conversion, or CAP conversion. In some embodiments, the first population was subjected to protection of hmC and subsequent deamination of mC and / or C.
[0369] In some embodiments of the combination, the first population comprises or is derived from DNA having a higher ratio of cytosine modification than the second population, the first population comprises first and second sub-populations, the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, the second population does not contain the first nucleobase. In some embodiments, the first nucleobase is modified or unmodified cytosine, the second nucleobase is modified or unmodified cytosine, and optionally, the modified cytosine is mC or hmC. In some embodiments, the first nucleobase is modified or unmodified adenine, the second nucleobase is modified or unmodified adenine, and optionally the modified adenine is mA.
[0370] In some embodiments, the first nucleobase (e.g., modified cytosine) is biotinylated. In some embodiments, the first nucleobase (e.g., modified cytosine) is the product of a Hu¨ sgen cycloaddition to β-6-azido-glucosyl-5-hydroxymethylcytosine containing an affinity label (e.g., biotin).
[0371] In any of the combinations described herein, the captured DNA can include cfDNA.
[0372] The captured DNA can have any of the features described herein with respect to the captured set, including, for example, a high concentration of DNA corresponding to a set of sequence-variable target regions (normalized with respect to the footprint size considered above) as compared to DNA corresponding to a set of epigenetic target regions. In some embodiments, the DNA of the captured set includes sequence tags, which may be added to the DNA as described herein. Generally, including sequence tags results in DNA molecules that are different from their naturally occurring untagged form.
[0373] The combination can further include a probe set or sequencing primer described herein, each of which can be different from its naturally occurring nucleic acid molecule. For example, the probe sets described herein can include a capture moiety, and the sequencing primers can include a label not found in nature.
[0374] F. Computer System The methods of the disclosure can be performed using or with the aid of a computer system. For example, such methods can include distributing a sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample includes DNA having a higher ratio of cytosine modification than the second secondary sample; subjecting the first secondary sample to a procedure that affects a first nucleobase in the DNA of the first secondary sample to be different from a second nucleobase in the DNA of the first secondary sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and sequencing the DNA in the first secondary sample and the DNA in the second secondary sample to distinguish the first nucleobase in the DNA of the first secondary sample from the second nucleobase.
[0375] FIG. 4 shows a computer system 401 programmed to execute the methods of the present disclosure or otherwise configured. The computer system 401 can regulate various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 401 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing, according to any of the methods disclosed herein, for example.
[0376] Computer system 401 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 405, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 401 also includes a memory or memory location 410 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 415 (e.g., hard disk), a communication interface 440 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 425, such as cache, other memory, data storage, and / or an electronic display adapter. Memory 410, storage unit 415, interface 420, and peripheral devices 425 communicate with CPU 405 through a communication network or bus (solid lines), such as a motherboard. Storage unit 415 may be a data storage unit (or data repository) for storing data. Computer system 401 can be operably coupled to computer network 430 with the aid of communication interface 420. Computer network 430 may be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. Computer network 430 may, in some examples, be a remote communication and / or data network. Computer network 430 may include one or more computer servers, which can enable distributed computing, such as cloud computing. Computer network 430 may, in some examples with the aid of computer system 0, execute a peer-to-peer network, which may enable devices to be coupled to computer system 401 and act as clients or servers.
[0377] The CPU 405 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 410. Examples of operations performed by the CPU 405 may include fetch, decode, execute, and writeback.
[0378] The storage unit 415 can store files, such as drivers, libraries, and saved programs. The storage unit 415 can store programs generated by the user and recorded sessions, as well as outputs related to the programs. The storage unit 415 can store user data, such as user preferences and user programs. Some examples of the computer system 401 may include one or more additional data storage units located external to the computer system 401, such as on a remote server that communicates with the computer system 401 through, for example, an intranet or the Internet. Data may be transferred from one location to another using, for example, a communication network or physical data transfer (e.g., using a hard drive, a thumb drive, or other data storage mechanisms).
[0379] The computer system 401 can communicate with one or more remote computer systems through the network 430. For embodiments, the computer system 401 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include personal computers (e.g., laptops), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), phones, smartphones (e.g., Apple® iPhone®, Android-enabled devices, Blackberry®), or personal digital assistants. The user can access the computer system 401 via the network 430.
[0380] The methods described in this specification can be executed by machine (e.g., computer processor) executable code stored in an electronic storage location of a computer system 401, such as memory 410 or electronic storage unit 415. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by processor 405. In some examples, the code is retrieved from storage unit 415 and stored in memory 410 for rapid access by processor 405. In some situations, electronic storage unit 415 can be excluded and machine executable instructions are stored in memory 410.
[0381] In one aspect, the present disclosure is a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform at least part of a method comprising: collecting cfDNA from a subject; capturing a set of a plurality of target regions from the cfDNA, wherein the set of a plurality of target regions includes a set of sequence-variable target regions and a set of epigenetic target regions, and a set of captured cfDNA molecules is produced; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the set of sequence-variable target regions are sequenced to a higher sequencing depth than the captured cfDNA molecules of the set of epigenetic target regions; obtaining a plurality of sequence reads generated by a nucleic acid sequencer from the step of sequencing the captured cfDNA molecules; mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the set of sequence-variable target regions and the set of epigenetic target regions to determine the likelihood that the subject has cancer.
[0382] The code can be precompiled and configured for use on a machine having a processor adapted to execute the code, or can be compiled at runtime. The code can be provided in a programming language selected to enable execution of the code as precompiled or while compiling.
[0383] Aspects of the systems and methods provided herein, such as computer system 401, can be embodied during programming. Various aspects of the technology can typically be thought of as a "product" or "manufactured article" in the form of machine (or processor) executable code and / or associated data embodied in or included in a type of machine-readable medium. Machine executable code can be stored in an electronic memory unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Memory" type media includes any or all of a computer's tangible memory, a processor or the like, or associated modules, such as various semiconductor memories, tape drives, disk drives, and the like, which can provide non-transitory storage at any time for software programming.
[0384] All or part of the software may sometimes communicate over the Internet or other various remote communication networks. Such communication may enable, for example, the loading of software from one computer or processor to another, such as from a management server or host computer to an application server's computer platform. Thus, another type of medium that may have software elements includes light, electrical, and electromagnetic waves such as those used over various air links through wired and optical terrestrial communication networks across physical interfaces between local devices. Physical elements that carry such waves, such as wired or wireless links, optical links, or the like, may also be considered media having software. As used herein, the term "computer or machine-readable medium" and the like, unless limited to non-transitory tangible "memory" media, means any medium that contributes to providing instructions for execution to a processor.
[0385] Accordingly, machine-readable media, such as computer-executable code, can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, any computer storage device such as optical or magnetic disks like those used to execute databases shown in the figures. Volatile storage media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables, copper wire, and fiber optics (including wires that comprise a bus within a computer system). Carrier wave transmission media can take the form of electrical or electromagnetic signals, or acoustic or optical waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punch cards, paper tape, any other physical storage media with patterns of holes, RAM, ROM, PROM, and EPROM, FLASH (registered trademark)-EPROM, any other memory chip or cartridge, carrier wave transporting data or instructions, cables or links transporting such carrier waves, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0386] Computer system 401 can include, or be in communication with, an electronic display that includes, for example, a user interface (UI) for providing one or more results of a sample analysis. Examples of UIs include, without limitation, graphical user interfaces (GUIs) and web-based user interfaces.
[0387] Further details regarding computer systems and networks, databases, and computer program products can be found, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety.
[0388] G. Applications 1. Cancer and Other Diseases The method can be used to diagnose a condition in a subject, particularly the presence of cancer, to characterize the condition (e.g., stage classify cancer or determine cancer heterogeneity), to monitor the response to treatment of the condition, and to achieve a prognostic risk of progression of the condition or subsequent course of the condition. The present disclosure may also be useful in determining the effectiveness of a particular treatment option. If a treatment option is successful, more cancer may die and excrete DNA, so that the amount of copy number variation or rare mutations detected in the blood of the subject may increase if the treatment is successful. In other instances, this may not occur. In another instance, perhaps a particular treatment option may correlate with the genetic profile of cancer over time. This correlation may be useful in the selection of therapy.
[0389] Furthermore, if it is observed that the cancer is in a remission state after treatment, the method can be used to monitor for residual disease or recurrence of the disease.
[0390] In some embodiments, the methods and systems disclosed herein can be used to identify customized or targeted therapies for treating a given disease or condition in a patient based on classifying nucleic acid variants as being of somatic origin or germline origin. Typically, the disease being considered is a type of cancer. Non-limiting examples of such cancers include cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, squamous cell carcinoma of the cervix, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatocarcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, germinoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B cell lymphoma, non-Hodgkin lymphoma, diffuse large B cell lymphoma, mantle cell lymphoma, T cell lymphoma, non-Hodgkin lymphoma, precursor T lymphoblastic lymphoma / leukemia, peripheral T cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostatic adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, stomach cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma. The type and / or stage of cancer can be detected by genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, segmental aneuploidy, ploidy, chromosomal instability, changes in chromosomal structure, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0391] Gene data can also be used to characterize specific forms of cancer. Cancer is often heterogeneous in composition and staging. Gene profile data can enable the characterization of specific subtypes of cancer that may be important in the diagnosis or treatment of those specific subtypes. This information provides clues to the prognosis of a specific type of cancer to the subject or clinician and enables the subject or clinician to adapt treatment options according to the progression of the disease. Some cancers may progress to be more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure may be useful in determining the progression of the disease.
[0392] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of abnormal conditions in a subject. Such methods include, for example, generating a gene profile of extracellular polynucleotides derived from the subject, where the gene profile includes multiple data resulting from the analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition can be a condition that results in a heterogeneous genomic population. In the example of cancer, some tumors are known to contain tumor cells of different cancer stages. In other examples, the heterogeneity can include multiple disease lesions. Again, in the example of cancer, there are multiple tumor lesions, and perhaps one or more of the lesions are the result of metastases that have spread from the primary site.
[0393] The method can be used to generate or profile a fingerprint or dataset that is the sum of gene information derived from different cells in a heterogeneous disease. This dataset can include, alone or in combination, the analysis of copy number variations, epigenetic variations, and mutations.
[0394] The methods can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods of the present specification do not include the diagnosis, prognosis, or monitoring of a fetus and thus are not related to non-invasive prenatal testing. In other embodiments, these methodologies can be used in a pregnant subject to diagnose, prognose, monitor, or observe cancer or other diseases in an unborn subject whose DNA and other polynucleotides can circulate with the maternal molecules.
[0395] Non-limiting examples of other genetic diseases, disorders, or conditions that may be evaluated using the methods and systems disclosed herein as needed include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cat-eye syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibroma, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson's disease, or the like.
[0396] In some embodiments, the methods described herein include detecting the presence or absence of DNA originating from or derived from tumor cells at a preselected time point following a previous cancer treatment of a subject already diagnosed with cancer using a set of sequence information obtained as described herein. The method may further include determining a cancer recurrence score indicative of the presence or absence of DNA originating from or derived from tumor cells for the test subject.
[0397] If a cancer recurrence score is determined, this score can be further used to determine the cancer recurrence status. The cancer recurrence status, for example, has a risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. The cancer recurrence status, for example, has a low or lower risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. In certain embodiments, a cancer recurrence score equal to the predetermined threshold can result in a cancer recurrence status having a risk of cancer recurrence or a low or lower risk of cancer recurrence.
[0398] In some embodiments, the cancer recurrence score is compared to a predetermined cancer recurrence threshold. If the cancer recurrence score is above the cancer recurrence threshold, the subject is classified as a candidate for subsequent cancer treatment, or if the cancer recurrence score is below the cancer recurrence threshold, the subject is classified as not being a candidate for treatment. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold can result in a classification of being a candidate for subsequent cancer treatment or not being a candidate for treatment.
[0399] The methods discussed above may further include any suitable feature(s) described elsewhere in this specification, including sections related to methods for determining the risk of cancer recurrence in a subject and / or methods for classifying a subject as a candidate for subsequent cancer treatment.
[0400] 2. Method for Determining the Risk of Cancer Recurrence in a Subject and / or Method for Classifying a Subject as a Candidate for Subsequent Cancer Treatment In some embodiments, the method provided herein is a method for determining the risk of cancer recurrence in a subject. In some embodiments, the method provided herein is a method for classifying a subject as a candidate for subsequent cancer treatment.
[0401] Any such method may include the step of collecting DNA (e.g., originating from or derived from tumor cells) from a test subject diagnosed with cancer at one or more preselected time points after one or more previous cancer treatments of the test subject. The subject may be any of the subjects described herein. The DNA may be cfDNA. The DNA is obtained from a tissue sample.
[0402] Any such method may include the step of capturing a plurality of sets of target regions from the DNA from the subject, the plurality of sets of target regions including a set of sequence-variable target regions and a set of epigenetic target regions, and producing a set of captured DNA molecules. The capturing step may be performed according to any of the embodiments described elsewhere herein.
[0403] In any of such methods, the previous cancer treatment may include surgery, administration of a therapeutic composition, and / or chemotherapy.
[0404] Any such method includes the step of sequencing the captured DNA molecules, thereby producing a set of sequence information. The captured DNA molecules of the set of sequence-variable target regions may be sequenced to a higher sequencing depth than the captured DNA molecules of the set of epigenetic target regions.
[0405] Any such method may include the step of using the set of sequence information to detect the presence or absence of DNA originating from or derived from tumor cells at a preselected time point. Detection of the presence or absence of DNA originating from or derived from tumor cells may be performed according to any of those embodiments described elsewhere herein.
[0406] A method for determining the risk of cancer recurrence in a subject may include determining a cancer recurrence score indicative of the presence, absence, or amount of DNA originating from or derived from tumor cells in the subject. The cancer recurrence score may be further used to determine the cancer recurrence status. The cancer recurrence status may, for example, indicate a risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. The cancer recurrence status may, for example, indicate a low or lower risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. In certain embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status indicating a risk of cancer recurrence or a low or lower risk of cancer recurrence.
[0407] A method for classifying a subject as a candidate for subsequent cancer treatment includes comparing the cancer recurrence score of the subject to a predetermined cancer recurrence threshold, and if the cancer recurrence score is above the cancer recurrence threshold, classifying the subject as a candidate for subsequent cancer treatment or, if the cancer recurrence score is below the cancer recurrence threshold, classifying the subject as not a candidate for treatment. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in a classification of the subject as a candidate for subsequent cancer treatment or not a candidate for treatment. In some embodiments, the subsequent cancer treatment includes administration of chemotherapy or a therapeutic composition.
[0408] Any such method may include determining the disease-free survival (DFS) period of the subject based on the cancer recurrence score, where, for example, the DFS period may be 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.
[0409] In some embodiments, the sequence information set includes a sequence variable target region sequence, and the step of determining the cancer recurrence score may include determining at least a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence variable target region sequence.
[0410] In some embodiments, the number of mutations in the array-variable target region selected from 1, 2, 3, 4, or 5 is sufficient to yield a cancer recurrence score in which the first subscore is classified as positive for cancer recurrence. In some embodiments, the number of mutations is selected from 1, 2, or 3.
[0411] In some embodiments, the array information set includes an epigenetic target region array, and the step of determining the cancer recurrence score includes determining a second subscore indicative of the amount of molecules (e.g., obtained from the epigenetic target region array) representing an epigenetic state different from that of DNA found in a corresponding sample from a healthy subject (e.g., cfDNA found in a blood sample from a healthy subject, or DNA found in a tissue sample from a healthy subject if the tissue sample is of the same tissue type as that obtained from the test subject). These abnormal molecules (i.e., molecules having an epigenetic state different from that of DNA found in a corresponding sample from a healthy subject) may coincide with epigenetic changes associated with cancer, such as methylation of hypermethylated variable target regions and / or perturbed fragmentation of fragmented variable target regions, where "perturbed" means different from that of DNA found in a corresponding sample from a healthy subject.
[0412] In some embodiments, it is sufficient for the ratio of molecules corresponding to the hypermethylated variable target region set and / or the fragmented variable target region set, which indicates hypermethylation in the hypermethylated variable target region set and / or abnormal fragmentation in the fragmented variable target region set, to be greater than or equal to a value in the range of 0.001% to 10% to classify the second subscore as positive for cancer recurrence. This range can be 0.001% to 1%, 0.005% to 1%, 0.01% to 5%, 0.01% to 2%, or 0.01% to 1%.
[0413] In some embodiments, any such method may include determining the proportion of tumor DNA from the proportion of molecules in a set of sequence information that exhibits one or more characteristics indicative of originating from tumor cells. This can be done for molecules corresponding to some or all of an epigenetic target region, for example, including one or both of a hypermethylated variable target region and a fragmented variable target region (hypermethylation of the hypermethylated variable target region and / or abnormal fragmentation of the fragmented variable target region may be considered indicative of originating from tumor cells). This can be done for molecules corresponding to a sequence variable target region, for example, molecules containing changes consistent with cancer such as SNVs, indels, CNVs, and / or fusions. The proportion of tumor DNA can be determined based on a combination of molecules corresponding to the epigenetic target region and molecules corresponding to the sequence variable target region.
[0414] Determination of the cancer recurrence score may be at least partially based on the proportion of tumor DNA, and 10 -11 ~1 or 10 -10 ~1. A proportion of tumor DNA greater than a threshold in the range of 10 -10 ~10 -9 、10 -9 ~10 -8 、10 -8 ~10 -7 、10 -7 ~10 -6 、10 -6 ~10 -5 、10 -5 ~10 -4 、10 -4 ~10 -3 、10 -3 ~10 -2 、or 10 -2 ~10 -1 or equal to is sufficient to classify the cancer recurrence score as positive with respect to cancer recurrence. In some embodiments, at least 10 -7The proportion of tumor DNA greater than the threshold value is sufficient to classify the cancer recurrence score as positive regarding cancer recurrence. The determination that the proportion of tumor DNA is greater than the threshold value, for example, the threshold value corresponding to any of the above embodiments, can be made based on the cumulative probability. For example, if the cumulative probability that the tumor proportion is greater than any of the threshold values in the above range exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999, the sample is considered positive. In some embodiments, the probability threshold is at least 0.95, for example, 0.99.
[0415] In some embodiments, the sequence information set includes a sequence variable target region sequence and an epigenetic target region sequence, and the step of determining the cancer recurrence score includes determining a first subscore indicating the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence variable target region sequence and a second subscore indicating the amount of abnormal molecules in the epigenetic target region sequence, and combining the first and second subscores to provide a cancer recurrence score. When combining the first and second subscores, a threshold value (for example, greater than a predetermined number of mutations in the sequence variable target region (e.g., >1), and greater than a predetermined proportion of abnormal molecules (i.e., molecules having an epigenetic state different from the DNA found in corresponding samples from healthy subjects; e.g., tumors) in the epigenetic target region) is applied independently to each subscore, or these can be combined by training a machine learning classifier to determine the state based on a plurality of positive and negative training samples.
[0416] In some embodiments, if the value of the combined score is in the range of -4 to 2 or -3 to 1, it is sufficient to classify the cancer recurrence score as positive regarding cancer recurrence.
[0417] In any embodiment where the cancer recurrence score is classified as positive for cancer recurrence, the subject's cancer recurrence status is at risk of cancer recurrence and / or the subject may be classified as a candidate for subsequent cancer treatment.
[0418] In some embodiments, the cancer is any one of the cancer types described elsewhere herein, such as colorectal cancer.
[0419] 3. Treatment and Related Administration In certain embodiments, the methods disclosed herein relate to identifying and administering customized therapies to patients given the status of nucleic acid variants of somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, the customized therapy includes at least one immunotherapy (or immunotherapy agent). Immunotherapy generally refers to methods of enhancing the immune response to a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing the T cell response to a tumor or cancer.
[0420] In certain embodiments, the status of nucleic acid variants in a sample from a subject of somatic or germline origin is compared to a database of comparator results from a reference population, and a customized or targeted therapy for that subject is identified. Typically, the reference population includes patients having the same cancer or disease type as the test subject and / or patients who are receiving or have received the same therapy as the test subject. If the nucleic acid variant and comparator results meet certain classification criteria (e.g., are substantially or approximately identical), a customized or targeted therapy (one or more therapies) may be identified.
[0421] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may be administered by methods such as buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intratympanic, etc., and the administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like.
[0422] Preferred embodiments of the invention are shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited by the specific examples provided herein. The invention is described with reference to the foregoing specification, but the description and explanation of the embodiments herein are not meant to be construed in a limiting sense. Many variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Further, it should be understood that all aspects of the invention are not limited to the specific depictions, configurations, or relative ratios described herein which depend on various conditions and variables. It should be understood that various alternative options may be employed in practicing the invention with respect to the disclosed embodiments herein. Accordingly, it is intended that the present disclosure should encompass any such options, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0423] The foregoing disclosure has been described in some detail for purposes of illustration and understanding, but it will be apparent to those skilled in the art that various changes in form and detail can be made without departing from the true scope of the disclosure and that the disclosure can be practiced within the scope of the appended claims. For example, all features, steps, elements, or other aspects of the methods, systems, computer-readable media, and / or components can be used in various combinations.
[0424] H. Kit Kits are also provided that include the compositions described herein. The kits can be useful for practicing the methods described herein. In some embodiments, the kit includes a first reagent for dispensing a sample into a plurality of secondary samples described herein, such as any of the dispensing reagents described elsewhere herein. In some embodiments, the kit includes a second reagent for subjecting a first secondary sample to a procedure that affects a first nucleobase in the DNA of the first secondary sample to be different from a second nucleobase in the DNA of the first secondary sample, where the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity (e.g., any of the reagents described elsewhere herein for converting a nucleobase such as cytosine or methylated cytosine to a different nucleobase). The kit can include the first and second reagents and additional elements discussed below and / or elsewhere herein.
[0425] The kit may further include a plurality of oligonucleotide probes that selectively hybridize to at least 5, 6, 7, 8, 9, 10, 20, 30, 40, or all genes selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSFIR, CTNNBl, ERBB4, EZH2, FGFRl, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID 1 A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRKl. The number of genes to which the oligonucleotide probes can selectively hybridize can vary. For example, the number of genes can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or 54. The kit can include a container containing a plurality of oligonucleotide probes and instructions for performing any of the methods described herein.
[0426] Oligonucleotide probes can selectively hybridize to exonic regions of genes, such as at least 5 genes. In some examples, oligonucleotide probes can selectively hybridize to at least 30 exons of genes, such as at least 5 genes. In some examples, multiple probes can selectively hybridize to each of at least 30 exons. Probes that hybridize to each exon can have sequences that overlap with at least one other probe. In some embodiments, oligoprobes can selectively hybridize to non-coding regions of the genes disclosed herein, such as intronic regions of genes. Oligoprobes can also selectively hybridize to regions of genes that include both exonic and intronic regions of the genes disclosed herein.
[0427] Any number of exons can be targeted by oligonucleotide probes. For example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 400, 500, 600, 700, 800, 900, 1,000 or more exons can be targeted.
[0428] The kit can include at least 4, 5, 6, 7, or 8 different library adapters having distinct molecular barcodes and identical sample barcodes. The library adapter may not be a sequencing adapter. For example, the library adapter does not include an array that enables flow cell alignment or hairpin loop formation for sequencing. Different variants and combinations of molecular barcodes and sample barcodes are described throughout and are applicable to the kit. Further, in some examples, the adapter is not a sequencing adapter. Further, the adapter provided in the kit may also include a sequencing adapter. The sequencing adapter may include an array that hybridizes to one or more sequencing primers. The sequencing adapter may further include an array that hybridizes to a solid support, such as a flow cell array. For example, the sequencing adapter can be a flow cell adapter. The sequencing adapter can be bound to one or both ends of the polynucleotide fragment. In some examples, the kit can include at least 8 different library adapters having distinct molecular barcodes and identical sample barcodes. The library adapter may not be a sequencing adapter. The kit may further include a sequencing adapter having a first array that selectively hybridizes to the library adapter and a second array that selectively hybridizes to the flow cell array. In another example, the sequencing adapter can be in the shape of a hairpin. For example, the hairpin-shaped adapter may include a complementary double-stranded portion and a loop portion, and the double-stranded portion can be bound (e.g., ligated) to the double-stranded polynucleotide. The hairpin-shaped sequencing adapter can be bound to both ends of the polynucleotide fragment to generate a circular molecule, which can be sequenced multiple times.The sequencing adapter can be 10 to 100 or more bases from end to end. The sequencing adapter can contain 20 to 30, 20 to 40, 30 to 50, 30 to 60, 40 to 60, 40 to 70, 50 to 60, 50 to 70 bases from end to end. In certain examples, the sequencing adapter can contain 20 to 30 bases from end to end. In another example, the sequencing adapter can contain 50 to 60 bases from end to end. The sequencing adapter can contain one or more barcodes. For example, the sequencing adapter can contain a sample barcode. The sample barcode can contain a predefined sequence. The sample barcode can be used to identify the source of the polynucleotide.The sample barcode can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or more (or of any length described throughout this specification) nucleobases, for example, at least 8 bases. The barcode can be a continuous or discontinuous sequence as described above.
[0429] The library adapter can be blunt-ended and Y-shaped and is less than 40 nucleobases in length or equal thereto. Other variant forms can be found throughout this specification and are applicable to the kit.
[0430] All patents, patent applications, websites, other publications and documents, accession numbers, and the like cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each separate item were specifically and individually indicated to be incorporated by reference. Where different versions of sequences are associated with different accession numbers, the version associated with the accession number as of the effective filing date of this application is meant. The effective filing date means either the actual filing date or, if applicable, the earlier of the filing date of the priority application that mentions the accession number. Similarly, where different versions of publications, websites, or the like are published at different times, the version published closest to the effective filing date of this application is meant, unless otherwise indicated.
Examples
[0431] III. Examples (Example 1: Analysis of cfDNA for Detecting the Presence / Absence of Tumors) A patient sample set is analyzed by a blood-based NGS assay at Guardant Health (Redwood City, CA, USA) to detect the presence / absence of cancer. cfDNA is extracted from the plasma of these patients. Next, the cfDNA of the patient samples is combined with magnetic beads conjugated to methyl-binding domain (MBD) buffer and MBD protein and incubated overnight. Methylated cfDNA, if present in the cfDNA sample, binds to the MBD protein during this incubation. Unmethylated, or less methylated, DNA is washed away from the beads with a buffer containing an increasing concentration of salt. Finally, a high-salt buffer is used to wash away highly methylated DNA from the MBD protein. These washes result in three fractions of cfDNA with increasing methylation (a low-methylation fraction, a residual methylation fraction, and a high-methylation fraction). In this example, cfDNA molecules in the high-methylation fraction are subjected to a procedure called hmC-Seal / 5hmC-Seal (Han, D. A highly sensitive and robust method for genome-wide 5hmC profiling of rare cell populations. Mol Cell. 2016; 63(4):711-719), which labels hmC residues (in cfDNA molecules of the high-methylation fraction), forms β-6-azido-glucosyl-5-hydroxymethylcytosine, then attaches a biotin moiety through a Huisgen cycloaddition, and then uses a biotin binder to separate the biotinylated DNA from other DNA, thereby separating cfDNA molecules with hmC residues in the high-methylation fraction into a fourth fraction, the hmC fraction. After separation, the cfDNA molecules in the four fractions are purified to remove salts and concentrated in the preparation for the enzymatic steps of library preparation.
[0432] After enrichment of cfDNA in the partition, the overhangs at the ends of the distributed cfDNA are extended, and adenosine residues are added to the 3' ends of the cfDNA fragments by polymerase during the extension. Phosphorylate the 5' end of each fragment. These modifications render the distributed cfDNA ligatable. Add DNA ligase and adapters to ligate an adapter to each end of each distributed cfDNA molecule. These adapters contain non-unique molecular barcodes and ligate an adapter having a non-unique molecular barcode distinguishable from the barcode in the adapters used in other fractions to each fraction. After ligation, pool the four fractions together and amplify by PCR.
[0433] After PCR, wash and concentrate the amplified DNA prior to enrichment. Once concentrated, combine the amplified DNA with a biotinylated RNA probe containing a salt buffer as well as probes for a set of variable target regions and probes for epigenetic target regions, and incubate this mixture overnight. The probe for the set of variable target regions has a footprint of approximately 50 kb, and the probe for the epigenetic target regions has a footprint of approximately 500 kb. The probe for the set of variable target regions includes oligonucleotides targeting at least a subset of the genes identified in Tables 3-5, and the probe for the epigenetic target regions includes oligonucleotides targeting those selected from hypermethylated variable target regions, hypomethylated variable target regions, CTCF binding target regions, transcription start site target regions, local amplification target regions, and methylation control regions.
[0434] The biotinylated RNA probe (hybridized to DNA) is captured by streptavidin magnetic beads and separated from the amplified DNA that is not captured by a series of salt-based washes, thereby enriching the sample. After enrichment, an aliquot of the enriched sample is sequenced using an Illumina NovaSeq sequencer. Next, the sequence reads generated by the sequencer are analyzed using bioinformatics tools / algorithms. Molecular barcodes are used to identify unique molecules and for deconvolution of the sample into molecules differentially distributed by MBD. The method described in this example provides information regarding the overall methylation level of molecules based on their fraction (i.e., methylated cytosine residues), and can also provide higher-resolution information regarding the identity and / or location of the type of methylated cytosine (i.e., mC or hmC) based on further partitioning of the highly methylated fraction to yield the hmC fraction. The sequence of the variable target region is analyzed by detecting genomic changes such as SNVs, insertions, deletions, and fusions that can be called with sufficient support to distinguish true tumor variants from technical errors (e.g., with respect to PCR errors, sequencing errors). The sequence of the epigenetic target region is analyzed independently to detect cfDNA molecules methylated in regions shown to be differentially methylated in cancer compared to normal cells. Finally, the results of both analyses are combined to generate a final tumor presence / absence call.
[0435] (Example 2: Analysis of methylation at single-base resolution in cfDNA samples from healthy subjects and subjects with early-stage colorectal cancer) cfDNA samples from healthy subjects and subjects with early-stage colorectal cancer were analyzed as follows. The cfDNA was partitioned using MBD to provide a hypermethylated fraction, an intermediate fraction, and a hypomethylated fraction. To the DNA partitioned in each fraction, an adapter (the adapter contains methylated cytosine to protect against EM-seq conversion and to maintain compatibility of the NGS library preparation) was ligated and subjected to an EM-seq conversion procedure, whereby non-modified cytosine undergoes deamination but mC and hmC do not. After such deamination, the fractions were prepared for sequencing and subjected to whole-genome sequencing. Each fraction was sequenced separately, but in an alternative procedure, the fractions can be differentially tagged (e.g., before EM-seq conversion after partitioning, or before further preparation for sequencing after partitioning and EM-seq conversion), pooled, and processed and sequenced in parallel.
[0436] Sequence data from hypermethylated variable target regions were isolated by bioinformatics, but in an alternative procedure, the target regions can be enriched in vitro before sequencing. Base-by-base methylation of hypermethylated variable target regions was quantified as shown in FIG. 5, which shows the number of methylated CpGs per molecule in the hypermethylated variable target regions from the hypermethylated fraction. The x-axis shows the total number of CpGs per molecule, and the points along the diagonal represent molecules with methylation for each CpG. Thus, it was possible to analyze methylation at single-base resolution and quantify methylation per base and partial molecular methylation of the MBD partitioning material. Samples from subjects with colorectal cancer showed significantly higher overall methylation in these regions than samples from healthy subjects. In certain embodiments, for example, the following items are provided. (Item 1) A method for analyzing DNA in a sample, comprising: a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains DNA having a higher ratio of cytosine modification than the second secondary sample; b) subjecting the first secondary sample to a procedure that affects a first nucleobase in the DNA of the first secondary sample so as to be different from a second nucleobase in the DNA of the first secondary sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and c) sequencing the DNA in the first secondary sample and the DNA in the second secondary sample so as to distinguish the first nucleobase in the DNA of the first secondary sample from the second nucleobase A method comprising the steps of. (Item 2) The method according to item 1, wherein the DNA comprises cell-free DNA (cfDNA) obtained from a test subject. (Item 3) A method for analyzing DNA in a sample containing cell-free DNA (cfDNA) obtained from a test subject in the sample, comprising: a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains DNA having a higher ratio of cytosine modification than the second secondary sample; b) subjecting the first secondary sample to a procedure that affects a first nucleobase in the DNA of the first secondary sample so as to be different from a second nucleobase in the DNA of the first secondary sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and c) sequencing at least the DNA in the first secondary sample so as to distinguish the first nucleobase in the DNA of the first secondary sample from the second nucleobase A method comprising the steps of. (Item 4) The method according to item 3, wherein step c) comprises at least sequencing the DNA in the first secondary sample. (Item 5) A method for analyzing a sample containing cell-free DNA (cfDNA), comprising: a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains cfDNA having a higher ratio of cytosine modification than the second secondary sample; b) subjecting the first secondary sample to a procedure that affects a first nucleobase in the cfDNA of the first secondary sample to be different from a second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; c) capturing at least a set of epigenetic target regions of cfDNA from the first secondary sample and the second secondary sample, thereby providing the captured cfDNA; and d) sequencing the captured cfDNA to distinguish the first nucleobase in the cfDNA of the first secondary sample from the second nucleobase. (Item 6) A method for analyzing a sample containing cell-free DNA (cfDNA), comprising: a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains cfDNA having a higher ratio of cytosine modification than the second secondary sample; b) subjecting the first secondary sample to a procedure that affects a first nucleobase in the cfDNA of the first secondary sample to be different from a second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; c) capturing a set of a plurality of target regions of cfDNA from the first secondary sample and the second secondary sample, thereby providing the captured cfDNA, a step in which the plurality of target region sets includes a variable sequence target region set and an epigenetic target region set; and d) a method comprising sequencing the captured cfDNA to distinguish the first nucleobase in the captured cfDNA of the first secondary sample from the second nucleobase. (Item 7) The method according to item 6, wherein cfDNA molecules corresponding to the variable sequence target region set are captured with a higher capture yield in the sample than cfDNA molecules corresponding to the epigenetic target region set. (Item 8) A method for isolating cell-free DNA (cfDNA) from a sample, a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains cfDNA having a higher ratio of cytosine modification than the second secondary sample; b) subjecting the first secondary sample to a procedure that affects the first nucleobase in the cfDNA so as to be different from the second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; c) contacting the cfDNA of the first secondary sample and the second secondary sample with a target-specific probe set, wherein the target-specific probe set includes a target-binding probe specific for a variable sequence target set and a target-binding probe specific for an epigenetic target set, thereby forming a complex of the target-specific probe and cfDNA; separating the complex from cfDNA that did not bind to the target-specific probe, thereby providing captured cfDNA corresponding to the variable sequence target set and cfDNA corresponding to the epigenetic target set; and d) sequencing the captured cfDNA to distinguish the first nucleobase in the cfDNA of the first secondary sample from the second nucleobase comprising a method. (Item 9) The method according to item 8, wherein the target-specific probe set is configured to capture cfDNA corresponding to the set of variable target regions with a higher capture yield than cfDNA corresponding to the set of epigenetic target regions. (Item 10) A method for identifying the presence of DNA produced by a tumor, comprising: a) collecting a cfDNA sample from a subject; b) distributing the cfDNA sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample includes captured cfDNA having a higher ratio of cytosine modification than the second secondary sample; c) subjecting the first secondary sample to a procedure that affects a first nucleobase in the cfDNA so as to be different from a second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, the second nucleobase is a modified nucleobase or an unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; d) capturing a set of a plurality of target regions from the cfDNA of the first secondary sample and the second secondary sample, thereby providing a sample containing the captured cfDNA, wherein the set of the plurality of target regions includes a set of variable target regions and a set of epigenetic target regions; and e) sequencing the captured cfDNA in the first secondary sample and the captured cfDNA in the second secondary sample to distinguish the first nucleobase in the cfDNA of the first secondary sample from the second nucleobase. The method comprising. (Item 11) The method according to item 10, wherein cfDNA molecules corresponding to the set of variable target regions are captured with a higher capture yield than cfDNA molecules corresponding to the set of epigenetic target regions in the sample. (Item 12) The method according to any one of items 6 to 11, comprising sequencing the cfDNA molecules corresponding to the set of variable target regions to a higher sequencing depth than the cfDNA molecules corresponding to the set of epigenetic target regions. (Item 13) The method according to item 12, wherein the captured cfDNA molecules of the set of sequence-variable targets are sequenced to a sequencing depth that is at least 2-fold higher than the captured cfDNA molecules of the set of epigenetic target regions. (Item 14) The method according to item 12, wherein the captured cfDNA molecules of the set of sequence-variable targets are sequenced to a sequencing depth that is at least 3-fold higher than the captured cfDNA molecules of the set of epigenetic target regions. (Item 15) The method according to item 12, wherein the captured cfDNA molecules of the set of sequence-variable targets are sequenced to a sequencing depth that is 4 to 10-fold higher than the captured cfDNA molecules of the set of epigenetic target regions. (Item 16) The method according to item 12, wherein the captured cfDNA molecules of the set of sequence-variable targets are sequenced to a sequencing depth that is 4 to 100-fold higher than the captured cfDNA molecules of the set of epigenetic target regions. (Item 17) The method according to any one of items 6 to 16, wherein the captured cfDNA molecules of the set of sequence-variable targets and the captured cfDNA molecules of the set of epigenetic target regions are sequenced in the same sequencing cell. (Item 18) The method according to any one of the preceding items, wherein the DNA is amplified before sequencing, or the method includes a capture step and the DNA is amplified before the capture step. (Item 19) The method according to items 5 to 18, further including the step of ligating a barcode-containing adapter to the DNA before capture, and optionally, the ligation step is performed before or simultaneously with amplification. (Item 20) The method according to any one of items 5 to 19, wherein the set of epigenetic target regions includes a set of hypermethylation variable target regions. (Item 21) The method according to any one of items 5 to 20, wherein the set of epigenetic target regions includes a set of hypomethylation variable target regions. (Item 22) The method according to item 20 or 21, wherein the set of epigenetic target regions includes a set of methylation control target regions. (Item 23) The method according to any one of items 5 to 22, wherein the set of epigenetic target regions includes a set of fragmented variable target regions. (Item 24) The method according to item 23, wherein the fragmented variable target region set includes a transcription start site region. (Item 25) The method according to item 23 or 24, wherein the fragmented variable target region set includes a CTCF binding region. (Item 26) The method according to any one of items 5 to 25, wherein the step of capturing the set of the plurality of target regions of cfDNA includes contacting the cfDNA with a target binding probe specific for the variable sequence target region set and a target binding probe specific for the epigenetic target region set. (Item 27) The method according to item 26, wherein the target binding probe specific for the variable sequence target region set is present at a higher concentration than the target binding probe specific for the epigenetic target region set. (Item 28) The method according to item 26, wherein the target binding probe specific for the variable sequence target region set is present at a concentration at least 2 times higher than the target binding probe specific for the epigenetic target region set. (Item 29) The method according to item 26, wherein the target binding probe specific for the variable sequence target region set is present at a concentration at least 4 times or 5 times higher than the target binding probe specific for the epigenetic target region set. (Item 30) The method according to item 26, wherein the target binding probe specific for the variable sequence target region set is present at a concentration at least 10 times, 20 times, 30 times, 40 times, 50 times, 60 times, 70 times, 80 times, 90 times, or 100 times higher than the target binding probe specific for the epigenetic target region set, or the target binding probe specific for the variable sequence target region set is present at a concentration in the range of 2 to 3, 3 to 4, 4 to 5, 5 to 7, 7 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 times the concentration of the target binding probe specific for the epigenetic target region set. (Item 31) The method according to any one of items 26 to 30, wherein the target binding probe specific for the variable sequence target region set has a higher target binding affinity than the target binding probe specific for the epigenetic target region set. (Item 32) The method according to any one of items 6 to 32, wherein the footprint of the epigenetic target region set is at least twice as large as the size of the array variable target region set. (Item 33) The method according to item 32, wherein the footprint of the epigenetic target region set is at least ten times as large as the size of the array variable target region set. (Item 34) The method according to any one of items 6 to 33, wherein the footprint of the array variable target region set is at least 25 kb or 50 kb. (Item 35) (a) The step of distributing the sample into a plurality of secondary samples includes distribution based on the methylation level; and / or (b) the step of distributing the sample into a plurality of secondary samples includes distribution based on binding to a protein, and optionally the protein is a methylated protein, an acetylated protein, an unmethylated protein, a non-acetylated protein; and / or optionally the protein is a histone. The method according to any one of the preceding items. (Item 36) (a) The step of distributing includes contacting the collected cfDNA with a methyl-binding reagent immobilized on a solid support; and / or (b) the step of distributing includes contacting the collected cfDNA with a binding reagent that is specific for the protein and immobilized on a solid support. The method according to item 35. (Item 37) The method according to any one of the preceding items, including the step of differentially tagging the first secondary sample and the second secondary sample. (Item 38) The method according to item 37, wherein the first secondary sample and the second secondary sample are differentially tagged before the first secondary sample is subjected to a procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample. (Item 39) The method according to item 37 or 38, wherein the first secondary sample and the second secondary sample are pooled after the first secondary sample is subjected to a procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample. (Item 40) The method according to any one of items 37 to 39, wherein the first secondary sample and the second secondary sample are sequenced in the same sequencing cell. (Item 41) The method according to any one of the preceding items, wherein the plurality of secondary samples includes a third secondary sample containing DNA having cytosine modification at a ratio higher than that of the second secondary sample but lower than that of the first secondary sample. (Item 42) The method according to item 41, further comprising the step of differentially tagging the third secondary sample such that the first secondary sample and the second secondary sample are distinguishable. (Item 43) After subjecting the first secondary sample to a procedure that affects the first nucleobase in the DNA so as to be different from the second nucleobase in the DNA of the first secondary sample, the first secondary sample, the second secondary sample, and the third secondary sample are combined, and if necessary, the first secondary sample, the second secondary sample, and the third secondary sample are sequenced in the same sequencing cell. The method according to item 42. (Item 44) The method according to any one of the preceding items, wherein the procedure to which the first secondary sample is subjected changes the base pairing specificity of the first nucleobase without substantially changing the base pairing specificity of the second nucleobase. (Item 45) The method according to any one of the preceding items, wherein the first nucleobase is modified cytosine or unmodified cytosine, and the second nucleobase is modified cytosine or unmodified cytosine. (Item 46) The method according to any one of the preceding items, wherein the first nucleobase includes unmodified cytosine (C). (Item 47) The method according to any one of the preceding items, wherein the second nucleobase includes 5-methylcytosine (mC). (Item 48) The method according to any one of the preceding items, wherein the procedure to which the first secondary sample is subjected includes bisulfite conversion. (Item 49) The method according to any one of items 1 to 46, wherein the first nucleobase includes mC. (Item 50) The method according to any one of the preceding items, wherein the second nucleobase includes 5-hydroxymethylcytosine (hmC). (Item 51) The method according to item 50, wherein the procedure to which the first secondary sample is subjected includes protection of 5hmC. (Item 52) The method according to item 50, wherein the procedure to which the first secondary sample is subjected includes Tet-assisted bisulfite conversion. (Item 53) The method according to item 50, wherein the procedure for providing the first secondary sample includes Tet-assisted conversion with a substituted borane reducing agent that is, if necessary, 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. (Item 54) The method according to item 53, wherein the substituted borane reducing agent is 2-picoline borane or borane pyridine. (Item 55) The method according to any one of items 49 to 51 or 53 to 54, wherein the second nucleobase contains C. (Item 56) The method according to any one of items 49 to 51 or 55, wherein the procedure for providing the first secondary sample includes protection of hmC and subsequent Tet-assisted conversion with a substituted borane reducing agent that is, if necessary, 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. (Item 57) The method according to item 56, wherein the substituted borane reducing agent is 2-picoline borane or borane pyridine. (Item 58) The method according to any one of items 46, 47, 49 to 51 or 55, wherein the procedure for providing the first secondary sample includes protection of hmC and subsequent deamination of mC and / or C. (Item 59) The method according to item 58, wherein the deamination of mC and / or C includes treatment with an AID / APOBEC family DNA deaminase enzyme. (Item 60) The method according to any one of items 51 or 55 to 59, wherein the protection of hmC includes glucosylation of hmC. (Item 61) The method according to any one of items 1 to 45, 47, 49, or 55, wherein the procedure for providing the first secondary sample includes chemical-assisted conversion with a substituted borane reducing agent that is, if necessary, 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. (Item 62) The method according to item 61, wherein the substituted borane reducing agent is 2-picoline borane or borane pyridine. (Item 63) The method according to any one of items 1 to 45, 47, 49, 55, or 61 to 62, wherein the first nucleobase contains hmC. (Item 64) The method according to any one of items 1 to 44, wherein the procedure for providing the first secondary sample includes separating the DNA originally containing the first nucleobase from the DNA not originally containing the first nucleobase. (Item 65) The method according to item 64, wherein the first nucleobase is hmC. (Item 66) The method according to item 64 or 65, wherein the step of separating the DNA containing the first nucleobase from the DNA not containing the first nucleobase from the beginning includes the step of labeling the first nucleobase. (Item 67) The method according to item 66, wherein the labeling step includes biotinylation. (Item 68) The method according to item 66 or 67, wherein the labeling step includes glucosylation. (Item 69) The method according to item 68, wherein the glucosylation step binds a glucosyl-azide moiety. (Item 70) The method according to item 66 or 68, wherein the labeling step includes a step of glucosylation before biotinylation and a subsequent step of binding a biotin moiety to the glucosyl. (Item 71) The method according to item 70, wherein the step of binding the biotin moiety to the glucosyl includes Huisgen cycloaddition chemistry. (Item 72) The method according to any one of items 54 to 61, wherein the step of separating the DNA containing the first nucleobase from the DNA not containing the first nucleobase from the beginning includes the step of binding the DNA containing the first nucleobase to a capture agent. (Item 73) The method according to item 72, wherein the capture agent includes a biotin binder, and optionally the biotin binder includes avidin or streptavidin. (Item 74) The method according to any one of items 64 to 73, including the step of differentially tagging each of the DNA containing the first nucleobase from the beginning, the DNA not containing the first nucleobase from the beginning, and the DNA of the second secondary sample. (Item 75) The method according to item 74, including the step of pooling each of the DNA containing the first nucleobase from the beginning, the DNA not containing the first nucleobase from the beginning, and the DNA of the second secondary sample after differentially tagging, and optionally the DNA containing the first nucleobase from the beginning, the DNA not containing the first nucleobase from the beginning, and the DNA of the second secondary sample are sequenced in the same sequencing cell. (Item 76) The method according to any one of items 1 to 44, wherein the first nucleobase is modified adenine or unmodified adenine, and the second nucleobase is modified adenine or unmodified adenine. (Item 77) The method according to any one of items 1 to 44, wherein the first nucleobase is a modified guanine or an unmodified guanine, and the second nucleobase is a modified guanine or an unmodified guanine. (Item 78) The method according to any one of items 1 to 44, wherein the first nucleobase is a modified thymine or an unmodified thymine, and the second nucleobase is a modified thymine or an unmodified thymine. (Item 79) The method according to any one of the preceding items, wherein the second subpopulation is not subjected to a procedure that affects the first nucleobase so as to be different from the second nucleobase. (Item 80) A combination comprising a first population and a second population of captured DNA, wherein the first population comprises or is derived from DNA having a higher ratio of cytosine modification than the second population, the first population comprising a first type of nucleobase originally present in DNA having altered base pairing specificity and a second nucleobase having no altered base pairing specificity, the type of the first nucleobase originally present in the DNA before the change in base pairing specificity being a modified nucleobase or an unmodified nucleobase, the second nucleobase being a modified nucleobase or an unmodified nucleobase different from the first nucleobase, the type of the first nucleobase and the second nucleobase originally present in the DNA before the change in base pairing specificity having the same base pairing specificity, and the second population not containing the type of the first nucleobase originally present in the DNA having altered base pairing specificity. (Item 81) The combination according to item 80, wherein the first population comprises sequence tags selected from a first set of one or more sequence tags, the second population comprises sequence tags selected from a second set of one or more sequence tags, and the second set of sequence tags is different from the first set of sequence tags. (Item 82) The combination according to item 81, wherein the sequence tag comprises a barcode. (Item 83) The combination according to any one of items 80 to 82, wherein the cytosine modification is methylation. (Item 84) The combination according to...
Claims
Claim 1 A method for isolating cell-free DNA (cfDNA) from a sample, comprising: a) distributing the sample into a plurality of secondary samples including a first secondary sample and a second secondary sample, wherein the first secondary sample contains cfDNA having a higher ratio of cytosine modification than the second secondary sample, and the step of distributing the sample into a plurality of secondary samples includes distribution based on methylation level; b) subjecting the first secondary sample to a procedure that affects a first nucleobase in the cfDNA of the first secondary sample to be different from a second nucleobase in the cfDNA of the first secondary sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, the first nucleobase and the second nucleobase have the same base pairing specificity, the first nucleobase is modified or unmodified cytosine, and the second nucleobase is modified or unmodified cytosine; c) contacting the cfDNA of the first secondary sample and the second secondary sample with a target-specific probe set, wherein the target-specific probe set includes a target-binding probe specific for a variable sequence target set and a target-binding probe specific for an epigenetic target set, the target-specific probe set is configured to capture cfDNA corresponding to the variable sequence target set with a higher capture yield than cfDNA corresponding to the epigenetic target set, thereby forming a complex of the target-specific probe and cfDNA; d) separating the complex from cfDNA that did not bind to the target-specific probe, thereby providing the captured cfDNA corresponding to the variable sequence target set and the cfDNA corresponding to the epigenetic target set; and e) sequencing the captured cfDNA to distinguish the first nucleobase in the cfDNA of the first secondary sample from the second nucleobase A method comprising the steps of. Claim 2 The method according to claim 1, comprising the step of sequencing the cfDNA molecules corresponding to the array-variable target region set to a sequencing depth higher than that of the cfDNA molecules corresponding to the epigenetic target region set.
3. The method according to claim 2, wherein the captured cfDNA molecules of the array-variable target set are sequenced to a sequencing depth 4 to 100 times higher than that of the captured cfDNA molecules of the epigenetic target region set.
4. The method according to any one of claims 1 to 3, further comprising the step of ligating a barcode-containing adapter to the DNA before capture.
5. a) The epigenetic target region set includes a hypermethylation variable target region set and / or the epigenetic target region set includes a hypomethylation variable target region set; and / or b) The epigenetic target region set includes a fragmentation variable target region set, The method according to any one of claims 1 to 4.
6. The method according to any one of claims 1 to 5, wherein the step of distributing the sample into a plurality of secondary samples includes distribution based on binding to a protein.
7. The method according to any one of claims 1 to 6, comprising the step of differentially tagging the first secondary sample and the second secondary sample.
8. The first secondary sample and the second secondary sample are pooled after subjecting the first secondary sample to a procedure that affects the first nucleobase in the DNA to be different from the second nucleobase in the DNA of the first secondary sample, and / or the first secondary sample and the second secondary sample are sequenced in the same sequencing cell. The method according to claim 7.
9. The method according to any one of claims 1 to 8, wherein the plurality of secondary samples includes a third secondary sample containing DNA having a cytosine modification at a higher ratio than the second secondary sample but a lower ratio than the first secondary sample.
10. a) The first nucleobase includes unmodified cytosine (C); b) The second nucleobase includes 5-methylcytosine (mC); and / or c) The procedure to which the first secondary sample is subjected includes bisulfite conversion. The method according to any one of claims 1 to 9.
11. a) The first nucleobase includes mC; b) the second nucleobase comprises 5-hydroxymethylcytosine (hmC); and / or c) the procedure to which the first secondary sample is subjected comprises protection of 5hmC, or the procedure to which the first secondary sample is subjected comprises Tet-assisted bisulfite conversion,[[]] The method according to any one of claims 1 to 9.[[]]
12. [[]] The method according to any one of claims 1 to 11, further comprising the step of determining the likelihood that the subject has cancer.[[]]
Citation Information
Patent Citations
Methods of attaching adapters to sample nucleic acids
US20180305738A1
Methods for multi-resolution analysis of cell-free nucleic acids
US9850523B1
Methods and systems for analyzing nucleic acid molecules
WO2018119452A2