Compositions and methods for analyzing DNA using partition and methylation-dependent nucleases
By distributing cell-free DNA samples based on cytosine modification levels and using methylation-dependent nucleases, the method improves cancer detection and monitoring by analyzing methylation patterns in liquid biopsies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2026-02-27
- Publication Date
- 2026-05-19
AI Technical Summary
Current methods for analyzing cell-free DNA in liquid biopsies struggle to accurately detect nucleobase modifications due to low concentration and heterogeneity, making it difficult to provide detailed information about cancer detection.
A method involving the distribution of a DNA sample into partial samples with varying cytosine modification proportions, followed by the use of methylation-dependent and methylation-sensitive nucleases to degrade nonspecific DNA, and subsequent capture and sequencing of epigenetic target regions to analyze methylation patterns.
Enhances the detection of cancer by providing a more comprehensive assessment of tumor state through improved analysis of cell-free DNA, enabling early detection and monitoring of cancer recurrence.
Smart Images

Figure 2026083210000010 
Figure 2026083210000011 
Figure 2026083210000012
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority based on U.S. Provisional Patent Application No. 63 / 086,000, filed on September 30, 2020, and U.S. Provisional Patent Application No. 63 / 105,183, filed on October 23, 2020, each of which is hereby incorporated by reference in its entirety for all purposes.
[0002] Field of the Invention The present disclosure provides compositions and methods related to analyzing DNA, such as cell - free DNA. In some embodiments, the cell - free DNA is derived from a subject having or suspected of having cancer, and / or the cell - free DNA contains DNA derived from cancer cells. In some embodiments, the DNA is distributed into a first partial sample and a second partial sample, where the first partial sample contains DNA having a higher proportion of nucleotide modifications (e.g., cytosine modifications) than the second partial sample, and the second partial sample is contacted with a methylation - dependent nuclease.
Background Art
[0003] Introduction and Summary Cancer is responsible for millions of deaths annually worldwide. Early detection of cancer may lead to improved outcomes as early - stage cancers tend to be more sensitive to treatment.
[0004] Inappropriately regulated cell growth is a prominent feature of cancer that generally results from genetic and epigenetic changes, such as copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions, and / or deletions (indels), epigenetic variations including modifications of cytosine (e.g., 5 - methylcytosine, 5 - hydroxymethylcytosine, and other more oxidized forms), and the accumulation of associations of DNA with chromatin proteins and transcription factors.
[0005] Biopsy represents a conventional approach for detecting or diagnosing cancer by extracting cells or tissues from a potential cancer site and analyzing them for relevant phenotypic and / or genotypic characteristics. Biopsy has the drawback of being invasive.
[0006] The detection of cancer based on the analysis of body fluids ("liquid biopsy"), for example, blood, is an interesting alternative based on the observation that DNA derived from cancer cells is released into body fluids. Liquid biopsy is non-invasive (only blood collection may be required). Current methods for cancer diagnostic assays of cell-free nucleic acids (e.g., cell-free DNA or cell-free RNA) may focus on the detection of tumor-related somatic variants, including single nucleotide variants (SNVs), copy number variations (CNVs), fusions, and indels (i.e., insertions or deletions), all of which are the mainstream targets of liquid biopsy. There is increasing evidence that non-sequence modifications such as methylation status and fragmentome signals in cell-free DNA can provide information regarding the origin and disease level of cell-free DNA. Non-sequence modifications of cell-free DNA, when combined with the calling of somatic mutations, can provide a more comprehensive assessment of the tumor state than what is available with either approach alone. However, considering the low concentration and heterogeneity of cell-free DNA, it has been difficult to develop accurate and sensitive methods for analyzing liquid biopsy materials that provide detailed information regarding nucleobase modifications.
[0007] Isolating and processing fractions of cell-free DNA useful for further analysis in liquid biopsy procedures is an important part of these methods. Thus, for example, in liquid biopsy, improved methods and compositions for analyzing cell-free DNA are needed. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0008] This disclosure aims to satisfy the need for improved analysis of cell-free DNA and / or to provide other advantages. Accordingly, the following exemplary embodiments are provided.
[0009] Embodiment 1 is a method for analyzing DNA in a sample, a) A step of distributing a sample into a plurality of partial samples, including a first partial sample and a second partial sample, wherein the first partial sample contains DNA having a cytosine modification in a larger proportion than the second partial sample. b) A step of contacting a second partial sample with a methylation-dependent nuclease to degrade the nonspecifically distributed DNA in the second partial sample to produce a treated second partial sample; optionally, a step of contacting a first partial sample with a methylation-sensitive endonuclease to degrade the nonspecifically distributed DNA in the first partial sample to produce a treated first partial sample; c) capturing a first set of target regions including epigenetic target regions from a first partial sample or at least a portion of a treated first partial sample, and capturing a second set of target regions including epigenetic target regions from at least a portion of a treated second partial sample. This method includes [something].
[0010] Embodiment 2 is a method for analyzing DNA in a sample, a) A step of capturing a first set of target regions, including epigenetic target regions, from a sample, b) A step of distributing a target region set into a plurality of partial samples, including a first partial sample and a second partial sample, wherein the first partial sample contains DNA having a cytosine modification in a larger proportion than the second partial sample. c) A step of contacting a second partial sample with a methylation-dependent nuclease to degrade the nonspecifically distributed DNA in the second partial sample and produce a treated second partial sample; optionally, a step of contacting a first partial sample with a methylation-sensitive endonuclease to degrade the nonspecifically distributed DNA in the first partial sample and produce a treated first partial sample. This method includes [something].
[0011] Embodiment 3 is the method of Embodiment 1, further comprising the step of quantifying epigenetic target regions captured from or present in one or more of the first partial sample, the treated first partial sample, or the treated second partial sample.
[0012] Embodiment 4 is the method according to Embodiment 2, wherein the quantification step includes quantitative amplification, and if necessary, the quantitative amplification is quantitative PCR.
[0013] Embodiment 5 is a method according to any one of the prior embodiments, further comprising the step of sequencing DNA in a first target region set and a second target region set or in a treated second partial sample.
[0014] Embodiment 6 is the method of the preceding embodiment, wherein the DNA in the treated second partial sample and the DNA in the treated first partial sample are sequenced.
[0015] Embodiment 7 is a method according to any one of the prior embodiments, wherein the epigenetic target region comprises a set of low-methylation variable target regions.
[0016] Embodiment 8 is the method of the preceding embodiment, wherein the set of low-methylation variable target regions includes regions in at least one type of tissue that have a lower degree of methylation than that in cell-free DNA derived from a healthy subject.
[0017] Embodiment 9 is a method of the prior embodiments, further comprising the step of determining the presence, absence, or likelihood of cancer based at least partially on the arrangement or quantity of regions in a set of low-methylation variable target regions.
[0018] Embodiment 10 is a method according to any one of Embodiments 7 to 9, further comprising the step of quantifying tumor DNA in a sample based at least partially on the sequence or number of regions in a set of low-methylation variable target regions.
[0019] Embodiment 11 is a method for analyzing DNA in a sample, a) A step of distributing a sample into a plurality of partial samples, including a first partial sample and a second partial sample, wherein the first partial sample contains DNA having a cytosine modification in a larger proportion than the second partial sample. b) A step of contacting a second partial sample with a methylation-dependent nuclease to degrade the nonspecifically distributed DNA in the second partial sample to produce a treated second partial sample; optionally, a step of contacting a first partial sample with a methylation-sensitive endonuclease to degrade the nonspecifically distributed DNA in the first partial sample to produce a treated first partial sample; c) A step of capturing a first set of target regions, including epigenetic target regions, from a first partial sample or at least a portion of a treated first partial sample. This method includes [something].
[0020] Embodiment 12 is the method of Embodiment 11, further comprising the step of quantifying epigenetic target regions captured from one or more of the first partial samples or treated first partial samples, or present in the treated second partial sample.
[0021] Embodiment 13 is the method according to Embodiment 12, wherein the quantification step includes quantitative amplification, and if necessary, the quantitative amplification is quantitative PCR.
[0022] Embodiment 14 is the method according to any one of Embodiments 11 to 13, further comprising the step of sequencing DNA in a first target region set and DNA derived from a second partial sample.
[0023] Embodiment 15 is a method according to any one of the preceding embodiments, wherein the DNA includes DNA obtained from a body fluid, and the body fluid is optionally plasma, urine, lymph, or cerebrospinal fluid.
[0024] Embodiment 16 is a method according to any one of the prior embodiments, wherein the DNA includes cell-free DNA (cfDNA) obtained from the test subject.
[0025] Embodiment 17 is a method according to any one of the prior embodiments, wherein the cytosine modification is methylation.
[0026] Embodiment 18 is a method according to any one of the prior embodiments, wherein the cytosine modification is methylation at the 5-position of cytosine.
[0027] Embodiment 19 is a method according to any one of the preceding embodiments, wherein the first partial sample is brought into contact with a methylation-sensitive endonuclease.
[0028] Embodiment 20 is a method according to the preceding embodiment, wherein a methylation-sensitive endonuclease cleaves a non-methylated CpG sequence.
[0029] Embodiment 21 is the method according to any one of the prior embodiments, wherein the methylation-sensitive endonuclease is one or more of the following: AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, and SnaBI.
[0030] Embodiment 22 is the method of the preceding embodiment, wherein the methylation-sensitive endonuclease is one or more of BstUI, HpaII, Hin6I, HhaI, or AccII, and optionally the methylation-sensitive endonuclease is (i) BstUI and HpaII, (ii) BstUI, HpaII, and Hin6I, or (iii) HhaI and AccII.
[0031] Embodiment 23 is a method according to any one of the preceding embodiments, wherein a methylation-dependent endonuclease cleaves a methylated CpG sequence.
[0032] Embodiment 24 is the method according to any one of the prior embodiments, wherein the methylation-dependent endonuclease is one or more of MspJI, LpnPI, FspEI, or McrBC.
[0033] Embodiment 25 is a method according to any one of the prior embodiments, wherein the first target region set includes a highly methylated variable target region set.
[0034] Embodiment 26 is the method of the preceding embodiment, wherein the set of highly methylated variable target regions includes regions having a higher degree of methylation in at least one type of tissue than the degree of methylation in cell-free DNA derived from a healthy subject.
[0035] Embodiment 27 is the method of Embodiment 25 or 26, further comprising the step of determining the presence, absence, or likelihood of cancer based at least partially on the arrangement or number of regions in a set of highly methylated variable target regions.
[0036] Embodiment 28 is a method according to any one of Embodiments 25-27, further comprising the step of quantifying tumor DNA in a sample based at least partially on the sequence or number of regions in a set of low-methylation variable target regions.
[0037] Embodiment 29 is a method according to any one of the prior embodiments, wherein the epigenetic target region includes a set of methylation control target regions.
[0038] Embodiment 30 is a method according to any one of the prior embodiments, wherein the first and / or second set of epigenetic target regions includes a set of fragmentation-variable target regions.
[0039] Embodiment 31 is the method according to the preceding embodiment, wherein the fragmentation variable target region set includes a transcription start site region.
[0040] Embodiment 32 is the method according to Embodiment 30 or 31, wherein the fragmentation variable target region set includes a CTCF binding region.
[0041] Embodiment 33 is a method according to any one of the prior embodiments, wherein the first set of target regions further comprises sequence-variable target regions.
[0042] Embodiment 34 is a method according to any one of the prior embodiments, wherein the second set of target regions further comprises sequence-variable target regions.
[0043] Embodiment 35 is the method according to Embodiment 33 or 34, wherein DNA molecules corresponding to sequence-variable target region sets are captured with a higher capture yield than DNA molecules corresponding to epigenetic target region sets.
[0044] Embodiment 36 is a method according to any one of the preceding embodiments, wherein the capture step includes contacting the DNA to be captured with a set of target-specific probes, thereby forming a complex of the target-specific probes and the DNA.
[0045] Embodiment 37 is a method of the prior embodiments, further comprising the capture step of separating the complex from DNA not bound to the target-specific probe, thereby obtaining the captured DNA.
[0046] Embodiment 38 is the method of Embodiment 36 or 37, wherein the set of target-specific probes is configured to capture DNA corresponding to a sequence-variable target region set with a higher capture yield than DNA corresponding to an epigenetic target region set.
[0047] Embodiment 39 is a method according to any one of Embodiments 33 to 38, further comprising the step of sequencing a DNA molecule corresponding to a sequence variable target region set to a sequencing depth greater than that of the DNA molecule corresponding to the epigenetic target region set.
[0048] Embodiment 40 is a method according to one of the preceding embodiments, wherein the DNA is amplified before the sequencing step or before the capture step.
[0049] Embodiment 41 is a method according to any one of the preceding embodiments, further comprising the step of ligating the adapter to DNA before capture, wherein the ligation step occurs, if necessary, before or concurrently with amplification.
[0050] Embodiment 42 is a method according to the prior embodiment, wherein the adapter is a barcode-containing adapter.
[0051] Embodiment 43 is the method of Embodiment 41 or 42, wherein the adapter is ligated to DNA before the second partial sample is brought into contact with the methylation-dependent nuclease.
[0052] Embodiment 44 is a method according to the prior embodiments, wherein the adapter is resistant to digestion by methylation-dependent nucleases.
[0053] Embodiment 45 is the method of the prior embodiment, wherein the adapter is non-methylated.
[0054] Embodiment 46 is a method according to any one of the prior embodiments, wherein the step of distributing the sample into a plurality of partial samples includes distributing them based on methylation levels.
[0055] Embodiment 47 is a method of the preceding embodiment, wherein the dispensing step includes contacting the collected cfDNA with a methyl-binding reagent immobilized on a solid support.
[0056] Embodiment 48 is a method according to any one of the preceding embodiments, comprising the step of differentially tagging a first partial sample and a second partial sample, a treated first partial sample and a second partial sample, a first partial sample and a treated second partial sample, or a treated first partial sample and a treated second partial sample.
[0057] Embodiment 49 is the method of the preceding embodiment, wherein a treated first partial sample and a second partial sample, a first partial sample and a treated second partial sample, or a treated first partial sample and a treated second partial sample are pooled after the first partial sample has been contacted with a methylation-sensitive endonuclease and / or the second partial sample has been contacted with a methylation-dependent nuclease.
[0058] Embodiment 50 is a method according to any one of Embodiments 47 to 49, wherein DNA derived from a treated first partial sample and a second partial sample, a first partial sample and a treated second partial sample, or a treated first partial sample and a treated second partial sample is sequenced in the same sequencing cell.
[0059] Embodiment 51 is a method according to any one of the prior embodiments, comprising a third partial sample in which multiple partial samples contain DNA having cytosine modifications at a higher rate than the second partial sample but at a lower rate than the first partial sample.
[0060] Embodiment 52 is a method of the prior embodiment, further comprising the step of differentially tagging a third partial sample.
[0061] Embodiment 53 is the method of the preceding embodiment, wherein the first, second, and third partial samples are combined after the first partial sample has been contacted with a methylation-sensitive endonuclease and / or the second partial sample has been contacted with a methylation-dependent nuclease, and optionally the DNA derived from the first, second, and third partial samples is sequenced in the same sequencing cell.
[0062] Embodiment 54 is the method according to any one of Embodiments 51 to 53, wherein a third partial sample is brought into contact with a methylation-sensitive endonuclease.
[0063] Embodiment 55 is the method according to any one of Embodiments 52 to 53, wherein a third partial sample is combined with a first partial sample, and the combined first and third partial samples are brought into contact with a methylation-sensitive endonuclease.
[0064] Embodiment 56 is the method of the preceding embodiment, wherein the combined first and third partial samples are brought into contact with a methylation-sensitive endonuclease, the second partial sample is brought into contact with a methylation-dependent nuclease, and then further combined with the second partial sample, and if necessary, the DNA derived from the first, second, and third partial samples is sequenced in the same sequencing cell.
[0065] Embodiment 57 is a method according to any one of the preceding embodiments, wherein a methylation-dependent nuclease is thermally inactivated after degrading nonspecifically distributed DNA.
[0066] Embodiment 58 is a method according to any one of the preceding embodiments, wherein a first partial sample is subjected to a procedure in which a first nucleic acid base in the DNA of the first partial sample is affected differently from a second nucleic acid base in the DNA, the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base pairing specificity.
[0067] Embodiment 59 is a method of the preceding embodiment, wherein the procedure for providing the first partial sample alters the base pairing specificity of the first nucleic acid base without substantially altering the base pairing specificity of the second nucleic acid base.
[0068] Embodiment 60 is the method according to Embodiment 58 or 59, wherein the first nucleic acid base is modified or unmodified cytosine, and the second nucleic acid base is modified or unmodified cytosine.
[0069] Embodiment 61 is the method according to any one of Embodiments 58 to 60, wherein the first nucleic acid base comprises unmodified cytosine (C).
[0070] Embodiment 62 is the method according to any one of Embodiments 58 to 61, wherein the second nucleic acid base comprises 5-methylcytosine (mC).
[0071] Embodiment 63 is the method according to any one of Embodiments 58 to 62, wherein the procedure for providing the first partial sample includes bisulfite conversion.
[0072] Embodiment 64 is the method according to any one of Embodiments 58 to 60, wherein the first nucleic acid base comprises mC.
[0073] Embodiment 65 is the method according to any one of Embodiments 58 to 62, wherein the second nucleic acid base comprises 5-hydroxymethylcytosine (hmC).
[0074] Embodiment 66 is the method of Embodiment 62, wherein the procedure for providing the first partial sample includes protection at 5hmC.
[0075] Embodiment 67 is the method of Embodiment 65, wherein the procedure for providing the first partial sample includes Tet-assisted bisulfite conversion.
[0076] Embodiment 68 is the method of Embodiment 65, wherein the procedure for providing the first partial sample includes a Tet-assisted conversion using a substituted borane reducing agent, the substituted borane reducing agent being, if necessary, 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane.
[0077] Embodiment 69 is the method according to Embodiment 68, wherein the substituted borane reducing agent is 2-picoline borane or borampyridine.
[0078] Embodiment 70 is a method according to any one of Embodiments 58-60, 64-66, or 68-69, wherein the second nucleic acid base contains C.
[0079] Embodiment 71 is the method according to any one of Embodiments 64-66 or 70, wherein the procedure to which the first partial sample is provided includes protection of hmC, followed by Tet-assisted conversion with a substituted borane reducing agent, wherein the substituted borane reducing agent is optionally 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane.
[0080] Embodiment 72 is the method according to Embodiment 71, wherein the substituted borane reducing agent is 2-picoline borane or borampyridine.
[0081] Embodiment 73 is the method according to any one of Embodiments 61, 62, 64-66, or 70, wherein the procedure to which the first partial sample is provided includes protection of hmC, followed by deamination of mC and / or C.
[0082] Embodiment 74 is the method of Embodiment 73, wherein the deamination of mC and / or C includes treatment with an AID / APOBEC family DNA deaminase enzyme.
[0083] Embodiment 75 is a method according to any one of Embodiments 66 or 70-74, wherein the protection of hmC includes glucosylation of hmC.
[0084] Embodiment 76 is the method according to any one of Embodiments 58-60, 62, 64, or 70, wherein the procedure for providing the first partial sample includes a chemical assisted transformation using a substituted borane reducing agent, and optionally the substituted borane reducing agent is 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane.
[0085] Embodiment 77 is the method according to Embodiment 76, wherein the substituted borane reducing agent is 2-picolineborane or borampyridine.
[0086] Embodiment 78 is a method according to any one of Embodiments 58-60, 62, 64, 70, or 76-77, wherein the first nucleic acid base comprises hmC.
[0087] Embodiment 79 is a method according to any one of the preceding embodiments, wherein the DNA of a first partial sample and the DNA of a second partial sample are differentially tagged, and after differential tagging, a portion of the DNA derived from the second partial sample is added to the first partial sample or the treated first partial sample or at least a portion thereof, thereby forming a pool from which sequence variable target regions and epigenetic target regions are captured.
[0088] Embodiment 80 is the method according to the prior embodiments, wherein the pool contains less than or equal to about 45%, less than or equal to 40%, less than or equal to 35%, less than or equal to 30%, less than or equal to 25%, less than or equal to 20%, less than or equal to 15%, less than or equal to 10%, or less than or equal to 5% of the DNA of the second partial sample.
[0089] Embodiment 81 is the method according to the prior embodiments, wherein the pool contains about 70–90%, about 75–85%, or about 80% of the DNA of the second partial sample.
[0090] Embodiment 82 is the method according to any one of Embodiments 79 to 81, wherein the pool contains substantially all of the DNA of the first partial sample.
[0091] Embodiment 83 is the method according to any one of Embodiments 79 to 82, wherein the pool contains substantially all of the DNA of the first partial sample or the treated first partial sample.
[0092] Embodiment 84 is a method according to any one of Embodiments 79 to 83, wherein the first target region set is captured from at least a portion of the first partial sample or the treated first partial sample after the pool has been formed.
[0093] Embodiment 85 is a method according to any one of the prior embodiments, further comprising the step of determining the likelihood that the subject has cancer.
[0094] Embodiment 86 is a method of a prior embodiment, wherein the sequencing step is to generate a plurality of sequencing read data, and the method further includes the step of mapping the plurality of sequence read data to one or more reference sequences to generate mapped sequence read data, and the step of processing the mapped sequence read data corresponding to a sequence variable target region set and an epigenetic target region set to determine the likelihood that the subject has cancer.
[0095] Embodiment 87 is a method according to any one of Embodiments 1 to 85, wherein the subject of test has been previously diagnosed with cancer and has undergone one or more prior cancer treatments, and cfDNA is acquired, if necessary, at one or more pre-selected time points after one or more prior cancer treatments, and a captured set of cfDNA molecules is sequenced, thereby producing a set of sequence information.
[0096] Embodiment 88 is a method of the preceding embodiment, further comprising the step of using a set of sequence information to detect the presence or absence of DNA originating from or derived from tumor cells at a pre-selected time point.
[0097] Embodiment 89 is the method of the preceding embodiments, further comprising the step of determining a cancer recurrence score for a test subject indicating the presence or absence of DNA originating from or derived from tumor cells, and optionally further comprising the step of determining a cancer recurrence status based on the cancer recurrence score, wherein the cancer recurrence status of the test subject is determined to be at risk of cancer recurrence if the cancer recurrence score is determined to be at or above a predetermined threshold, or the cancer recurrence status of the test subject is determined to be at low risk of cancer recurrence if the cancer recurrence score is below a predetermined threshold.
[0098] Embodiment 90 is a method of the preceding embodiment, further comprising the step of comparing the cancer recurrence score of a test subject with a predetermined cancer recurrence threshold, wherein the test subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or as not a candidate for subsequent cancer treatment if the cancer recurrence score is below the cancer recurrence threshold. I. Brief Description of the Drawings [Brief explanation of the drawing]
[0099] [Figure 1] Figure 1A is a schematic diagram of a methylation-dependent nuclease (e.g., methylation-dependent restriction enzyme (MDRE)) that digests / cuts DNA when the restriction enzyme (RE) recognition site contains methylated nucleotides, but does not cut DNA when the restriction enzyme (RE) recognition site contains unmethylated nucleotides. Figure 1B is a schematic diagram of a methylation-sensitive nuclease (e.g., methylation-sensitive restriction enzyme (MSRE)) that digests / cuts DNA when the restriction enzyme (RE) recognition site contains unmethylated nucleotides, but does not cut DNA when the restriction enzyme (RE) recognition site contains methylated nucleotides.
[0100] [Figure 2] Figure 2 is a flowchart illustrating a method for determining the methylation state of nucleic acid molecules in a polynucleotide sample obtained from a target, according to an embodiment of the present disclosure.
[0101] [Figure 3] Figure 3 is a flowchart illustrating a method for determining the methylation state of nucleic acid molecules in a polynucleotide sample obtained from a target, according to an embodiment of the present disclosure.
[0102] [Figure 4-1] Figure 4 is a schematic diagram of a method for detecting the presence or absence of cancer in a subject, according to a particular embodiment of the present disclosure. [Figure 4-2]Figure 4 is a schematic diagram of a method for detecting the presence or absence of cancer in a subject, according to a particular embodiment of the present disclosure.
[0103] [Figure 5] A schematic diagram of an example of a system suitable for use according to some embodiments of this disclosure.
[0104] [Figure 6] Figure 6 shows the number of molecules in three partitions with and without MSRE treatment in normal and diluted CRC samples.
[0105] [Figure 7] Figure 7 shows the CpG methylation quantification results obtained as described in Example 2 for three samples from subjects with early-stage colorectal cancer ("early-stage CRC") and three healthy subjects ("normal"). For the early-stage CRC plot, MAF represents the mutant allele fraction.
[0106] [Figure 8-1] Figures 8A-8D show the number of positive and negative control molecules with FspEI palindromic sites for the enzyme and buffer conditions described, as in Example 4. Figures 8A and 8C correspond to the first donor, and Figures 8B and 8D correspond to the second donor. The data points are distributed along the horizontal axis for readability. [Figure 8-2] Figures 8A-8D show the number of positive and negative control molecules with FspEI palindromic sites for the enzyme and buffer conditions described, as in Example 4. Figures 8A and 8C correspond to the first donor, and Figures 8B and 8D correspond to the second donor. The data points are distributed along the horizontal axis for readability.
[0107] [Figure 9-1] Figures 9A-D show the digestion efficiency and the number of positive control molecules, as described in Example 4. [Figure 9-2] Figures 9A-D show the digestion efficiency and the number of positive control molecules, as described in Example 4.
[0108] [Figure 10-1] Figures 10A–J show the number of low-methylated variable target regions ("low VTR") molecules (10A–E) or the low VTR / negative control molecule ratio (10F–J) for the conditions described, as in Example 5. The data points are distributed along the horizontal axis for readability. Triangles, circles, plus signs, and squares indicate that the origin of the normal cfDNA was the first, second, third, or fourth person out of four healthy donors, respectively. [Figure 10-2] Figures 10A–J show the number of low-methylated variable target regions ("low VTR") molecules (10A–E) or the low VTR / negative control molecule ratio (10F–J) for the conditions described, as in Example 5. The data points are distributed along the horizontal axis for readability. Triangles, circles, plus signs, and squares indicate that the origin of the normal cfDNA was the first, second, third, or fourth person out of four healthy donors, respectively. [Modes for carrying out the invention]
[0109] II. Detailed Description of a Specific Embodiment Herein, a detailed reference is made to certain embodiments of the present invention. While the present invention is described in relation to such embodiments, it should be understood that they are not intended to limit the invention to those embodiments. On the contrary, the present invention is intended to cover all alternative forms, modifications, and equivalents, which may be included within the present invention as defined by the appended claims.
[0110] Before describing this instruction in detail, it should be understood that this disclosure is not limited to any particular composition or process step, and that it may vary. When used herein and in the appended claims, the singular forms "one (a)", "one (an)", etc. The noun "the" refers to plural objects unless otherwise explicitly indicated by the context. It should be noted that, for example, a reference to "a nucleic acid" includes multiple nucleic acids, and a reference to "a cell" includes multiple cells. .
[0111] Numerical ranges include the number that defines that range. Measured values and measurable values are understood to be approximations, taking into account the significant figures and errors associated with the measurement. Also, "comprise," "comprises," "comprising," "contain," "contains," "containing," and "include" are used. The use of "includes" and "including" is restrictive. This is not intended to be so. Please understand that the general and detailed explanations above are merely illustrative and descriptive, and do not limit the teaching.
[0112] Unless specifically referred to in the above specification, embodiments in this specification that list various components “including” are also intended to “consist of” or “essentially consist of” the listed components, and embodiments in this specification that list various components “including” are also intended to “consist of” or “essentially consist of” the listed components, and embodiments in this specification that list various components “essentially consist of” are also intended to “consist of” or “consist of” the listed components (this interchangeability does not apply to the use of these terms in the claims).
[0113] The section headings used herein are for structural purposes only and should not be construed as limiting the subject matter disclosed. In the event of any conflict between any document or other material incorporated by reference and any express content of this specification, including definitions, this specification shall prevail. A.Definition
[0114] "Cell-free DNA," "cfDNA molecules," or simply "cfDNA" refers to DNA molecules that are naturally present in extracellular form in a subject (e.g., in blood, serum, plasma, or other bodily fluids, e.g., lymph, cerebrospinal fluid, urine, or sputum). cfDNA originally existed in large, complex organisms, such as the cells (one or more) of mammals, but undergoes release from cells into fluids found in organisms, and can be obtained from fluid samples without the need for an in vitro cell lysis step.
[0115] As used herein, “cellular nucleic acid” means a nucleic acid located within one or more cells from which the nucleic acid originates, at least at the time the sample is taken or collected from the subject, even if the nucleic acid is later removed as part of a given analytical process (e.g., by cell lysis).
[0116] As used herein, a modification or other feature is present in a “higher proportion” in the first sample or population of nucleic acids than in the second sample or population if the fraction of nucleotides having the modification or other feature is higher in the first sample or population than in the second population. For example, if one-tenth of the nucleotides in the first sample are mC and one-twentieth of the nucleotides in the second sample are mC, then the first sample contains a higher proportion of the cytosine 5-methylation modification than the second sample.
[0117] As used herein, “without substantially altering the base-pairing specificity” of a given nucleic acid base means that the majority of molecules containing that nucleic acid base that can be sequenced have no alteration to the base-pairing specificity of the second nucleic acid base compared to the base-pairing specificity it had when it was originally isolated in the sample. In some embodiments, 75%, 90%, 95%, or 99% of the molecules containing that nucleic acid base that can be sequenced have no alteration to the base-pairing specificity of the second nucleic acid base compared to the base-pairing specificity it had when it was originally isolated in the sample.
[0118] As used herein, “base pairing specificity” refers to the standard DNA base (A, C, G, or T) to which a given base most preferentially pairs. Therefore, for example, unmodified cytosine and 5-methylcytosine have the same base pairing specificity (i.e., specificity to G), but uracil and cytosine have different base pairing specificities, as uracil has base pairing specificity to A while cytosine has base pairing specificity to G. Since uracil preferentially pairs to A among the four standard DNA bases in any case, uracil’s ability to form unstable pairs with G is irrelevant.
[0119] As used herein, “combination” including multiple members means either a single composition including members, or a set of compositions located in separate containers or compartments in a nearby larger container, such as a multiwell plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other storage configuration.
[0120] The “capture yield” of a probe collection for a given target set refers to the amount of nucleic acid corresponding to the target set captured by the probe collection under typical conditions (e.g., an amount compared to another target set, or an absolute amount). An exemplary typical capture condition is incubation of the sample nucleic acid and probe in a small reaction volume (approximately 20 μL) containing stringent hybridization buffer at 65°C for 10–18 hours. Capture yields can be expressed as absolute values or, for collections of multiple probes, as relative values. When comparing capture yields of multiple target region sets, they are normalized with respect to the footprint size of the target region set (e.g., based on kilobases). Therefore, for example, if the footprint sizes of the first and second target regions are 50kb and 500kb, respectively (normalization factor 0.1), then DNA corresponding to the first set of target regions will be captured in a higher yield than DNA corresponding to the second set of target regions if the mass per unit volume of captured DNA corresponding to the first set of target regions is greater than 0.1 times the mass per unit volume of captured DNA corresponding to the second set of target regions. As a further example, using the same footprint size, DNA corresponding to the first set of target regions will be captured in twice the capture yield than DNA corresponding to the second set of target regions if the mass per unit volume of captured DNA corresponding to the first set of target regions is 0.2 times the mass per unit volume of captured DNA corresponding to the second set of target regions.
[0121] "Capturing" one or more target nucleic acids means preferentially isolating or separating one or more target nucleic acids from non-target nucleic acids.
[0122] A "captured set" of nucleic acids refers to the nucleic acids that have been captured.
[0123] A “target region set” or “set of target regions” refers to a set of genomic loci that are targeted for capture and / or targeted by a set of probes (e.g., through sequence complementarity).
[0124] "Corresponding to a target region set" means that the nucleic acid, for example, cfDNA, originates from a locus in the target region set or specifically binds to one or more probes targeting the target region set.
[0125] In the context of probes or other oligonucleotides and target sequences, "specifically binding" means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a replica thereof to form a stable probe:target hybrid, while simultaneously minimizing the formation of stable probe:non-target hybrids. Therefore, the probe hybridizes to the target sequence or a replica thereof to a sufficiently higher degree than the non-target sequence, enabling the capture or detection of the target sequence. Appropriate hybridization conditions are well known in the art and can be predicted based on the sequence composition, or determined by using conventional test methods (see, for example, Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, in particular §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, which are incorporated herein by reference).
[0126] A "set of sequence-variable target regions" refers to a set of target regions in neoplastic cells (e.g., tumor cells and cancer cells) that may exhibit sequence changes, such as nucleotide substitutions (i.e., single-nucleotide mutations), insertions, deletions, or gene fusions or transpositions.
[0127] The “epigenetic target region set” refers to a set of target regions in neoplastic cells (e.g., tumor cells and cancer cells) that may exhibit sequence-independent changes, or that may exhibit sequence-independent changes in cfDNA from cancerous subjects compared to cfDNA from healthy subjects. Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. For the purposes of this, loci susceptible to local amplification and / or gene fusion associated with neoplasms, tumors, or cancer may also be included in the epigenetic target region set because the detection of copy number changes by sequencing or fusion sequences mapping to more than one locus in the reference genome tends to be more similar to the detection of the exemplary epigenetic changes discussed above than the detection of nucleotide substitutions, insertions, or deletions, in that, for example, the detection of local amplification and / or gene fusion can be detected at a relatively shallow sequencing depth because it does not depend on the accuracy of base calls at one or a few individual locations.
[0128] Nucleic acids, when originating from tumor cells, are "produced by the tumor" or produced by ctDNA or circulating tumor DNA. Tumor cells are neoplastic cells originating from a tumor, whether they remain within the tumor or are isolated from it (for example, as in metastatic cancer cells and circulating tumor cells).
[0129] The terms "methylation" or "DNA methylation" refer to the addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to cytosine at a CpG site (cytosine-phosphate-guanine site, i.e., guanine following cytosine in the 5'→3' direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to the addition of a methyl group to adenine, e.g., N 6-This refers to methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine to produce 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (caryboxylcytosine) (5-caC). No. In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of cytosine to produce 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription may be repressed. DNA methylation is crucial for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruption, e.g., repression, in epigenetic regulation can lead to disease, e.g., cancer. Promoter methylation in DNA can indicate cancer.
[0130] The term "hypermethylation" refers to an increased level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules in a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypermethylated DNA may include DNA molecules containing at least one methylated residue, at least two methylated residues, at least three methylated residues, at least five methylated residues, or at least ten methylated residues.
[0131] The term "hypomethylation" refers to a reduced level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules in a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA may include DNA molecules containing 0 methylated residues, at most 1 methylated residue, at most 2 methylated residues, at most 3 methylated residues, at most 4 methylated residues, or at most 5 methylated residues.
[0132] The term "methylation-dependent nuclease" refers to a nuclease that preferentially cleaves methylated DNA compared to unmethylated DNA. For example, a methylation-dependent nuclease may cleave at or near a recognition site, such as a restriction site, in a manner dependent on the methylation of at least one nucleic acid base within the recognition sequence, e.g., cytosine. In some embodiments, the nucleic acid lysis activity of a methylation-dependent nuclease is at least 10, 20, 50, or 100 times higher than that of an unmethylated control at a methylated recognition site in a standard nucleic acid lysis assay. Methylation-dependent nucleases include methylation-dependent restriction enzymes.
[0133] As used herein, “methylation-dependent restriction enzyme” or “MDRE” refers to a restriction enzyme that depends on the methylation of DNA (e.g., cytosine methylation), meaning that the presence or absence of a methyl group in a nucleotide base alters the proportion of target DNA cleaved by the enzyme. In some embodiments, a methylation-dependent restriction enzyme will not cleave DNA if a particular nucleotide base is unmethylated in the recognition sequence. For example, MspJI is a methylation-dependent restriction enzyme with the recognition sequence “mCNNR(N9)” and will not cleave DNA if methylated cytosine (mC) is absent in the recognition sequence.
[0134] The term "methylation-sensitive nuclease" refers to a nuclease that preferentially cleaves unmethylated DNA compared to methylated DNA. For example, a methylation-sensitive nuclease may cleave at or near a recognition sequence, such as a restriction site, in a manner dependent on the absence of methylation of at least one nucleic acid base in the recognition sequence, e.g., cytosine. In some embodiments, the nucleic acid lysis activity of a methylation-sensitive nuclease is at least 10, 20, 50, or 100 times higher than that of a methylated control for unmethylated recognition sequences in a standard nucleic acid lysis assay. Methylation-sensitive nucleases include methylation-sensitive restriction enzymes.
[0135] As used herein, “methylation-sensitive restriction enzyme” or “MSRE” refers to a restriction enzyme that is sensitive to the methylation status of DNA (e.g., cytosine methylation), i.e., the presence or absence of a methyl group in a nucleotide base alters the proportion of target DNA cleaved by the enzyme. In some embodiments, a methylation-sensitive restriction enzyme will not cleave DNA if a particular nucleotide base is methylated in the recognition sequence. For example, HpaII is a methylation-sensitive restriction enzyme having the recognition sequence “CCGG,” and will not cleave DNA if the second cytosine in the recognition sequence is methylated.
[0136] As used herein, “digestion efficiency” or “cleavage efficiency” refers to the efficiency of restriction enzyme digestion. Digestion efficiency can be calculated based on the number of control molecules observed during restriction enzyme digestion and the number of control molecules observed in the absence of restriction enzyme digestion. MSRE digestion efficiency can be calculated as follows: efficiency = 1 - (negative control molecules) [MSRE] Number of negative control molecules [Mock] (Number of). MDRE digestion efficiency can be calculated as follows: Efficiency = 1 - (Positive control molecule) [MDRE] Number of positive control molecules [Mock] (The number).
[0137] As used herein, “methylation state” can refer to the presence or absence of a methyl group on a DNA base (e.g., cytosine) at a specific genomic location within a nucleic acid molecule. It can also refer to the degree of methylation in a nucleic acid sequence (e.g., highly methylated, hypomethylated, intermediately methylated, or unmethylated nucleic acid molecules). The methylation state can also refer to the number of methylated nucleotides in a particular nucleic acid molecule.
[0138] As used herein, “mutation” refers to a mutation from a known reference sequence, including, for example, single nucleotide variants (SNVs) and mutations such as insertions or deletions (indels). Mutations may be germline or somatic mutations. In some embodiments, the reference sequence for comparative purposes is the wild-type genome sequence of the species under consideration, typically the human genome, which provides the test sample.
[0139] As used herein, the terms “neoplasm” and “tumor” are interchangeable. They refer to the abnormal growth of cells in an object. Neoplasms or tumors can be benign, possibly malignant, or malignant. Malignant tumors are called cancer or cancerous tumors.
[0140] As used herein, “next-generation sequencing” or “NGS” refers to sequencing technologies that have increased processing power compared to conventional Sanger and capillary electrophoresis-based approaches, for example, the ability to generate hundreds of thousands of relatively small sequence read data at once. Some examples of next-generation sequencing technologies, but not limited to these, include sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. In some embodiments, next-generation sequencing involves the use of instruments capable of sequencing single molecules. Examples of commercially available instruments for performing next-generation sequencing, but not limited to these, include NextSeq, HiSeq, NovaSeq, MiSeq, Ion PGM, and Ion GeneStudio S5.
[0141] As used herein, “nucleic acid tag” refers to a short nucleic acid (e.g., less than approximately 500 nucleotides, less than approximately 100 nucleotides, less than approximately 500 nucleotides, or less than approximately 10 nucleotides in length) used to identify nucleic acids from different samples (e.g., representing sample indices), nucleic acids from different distributions (e.g., representing distribution tags), or different nucleic acid molecules in the same sample undergoing different types or different processing (e.g., representing molecular barcodes). Nucleic acid tags comprise a predetermined, fixed, non-random, random, or semi-random oligonucleotide sequence. Such nucleic acid tags may be used to label different nucleic acid molecules or different nucleic acid samples or partial samples. Nucleic acid tags may be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags may have the same or diverse lengths as needed. Nucleic acid tags may also comprise a double-stranded molecule with one or more blunt ends, a 5' or 3' single-stranded region (e.g., an overhang), and / or one or more other single-stranded regions at other locations within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the sample, morphology, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples, including nucleic acids with different molecular barcodes and / or sample indices, by detecting (e.g., reading) the nucleic acid tag, which then deconvolutes the nucleic acid. Nucleic acid tags may also be called identifiers (e.g., molecular identifiers, sample identifiers). Furthermore, or or, nucleic acid tags can be used as molecular identifiers (e.g., to identify between different molecules or between amplicons of different parent molecules in the same sample or partial sample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample, or non-uniquely tagging such molecules.In non-unique tagging applications, each nucleic acid molecule may be tagged using a limited number of tags (i.e., molecular barcodes) so that different molecules can be identified based on their intrinsic sequence information (e.g., start and / or end positions to which they map to a selected reference genome, subsequences of one or both ends of the sequence, and / or sequence length) in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used so that the probability of any two molecules having the same intrinsic sequence information (e.g., start and / or end positions, subsequences of one or both ends of the sequence, and / or length), and similarly the probability of them having the same molecular barcode, is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1%).
[0142] As used herein, “distributing” means physically separating or fractionating a mixture of nucleic acid molecules in a sample based on the characteristics of the nucleic acid molecules. Distributing can be the physical distribution of molecules. Distributing may include separating nucleic acid molecules into groups or sets based on the level of epigenetic characteristics (e.g., relating to methylation). For example, nucleic acid molecules can be distributed based on the level of methylation of the nucleic acid molecules. In some embodiments, methods and systems used for distribution can be found in PCT Patent Application No. PCT / US2017 / 068329, which is incorporated herein by reference in its entirety.
[0143] As used herein, “distribution set” or “distribution” refers to a set or group of nucleic acid molecules distributed to a binder based on the differential binding affinity of the nucleic acid molecules or proteins associated with nucleic acid molecules to a binder. A distribution set may also be referred to as a partial sample. The binder preferentially binds to nucleic acid molecules containing nucleotides having epigenetic modifications. For example, if the epigenetic modification is methylation, the binder may be a methyl-binding domain (MBD) protein. In some embodiments, a distribution set may include nucleic acid molecules belonging to a particular level or degree of epigenetic feature (e.g., methylation). For example, nucleic acid molecules may be distributed to three sets: one set of highly methylated nucleic acid molecules (first partial sample, high distribution, high distribution set, or high-methylated distribution set), a second set of low-methylated nucleic acid molecules (second partial sample, low distribution, low distribution set, or low-methylated distribution set), and a third set of intermediate-methylated nucleic acid molecules (third partial sample, intermediate distribution set, intermediate-methylated distribution set, residual distribution, or residual distribution set). In another example, nucleic acid molecules may be distributed based on the number of methylated nucleotides, with one distribution set having nucleic acid molecules with 9 methylated nucleotides and another distribution set having unmethylated nucleic acid molecules (0 methylated nucleotides).
[0144] As used herein, “polynucleotide,” “nucleic acid,” “nucleic acid molecule,” or “oligonucleotide” refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogues) joined by nucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a small number of monomeric units, e.g., 3-4 to several hundred monomeric units. When a polynucleotide is represented by a sequence of letters, e.g., “ATGCCTG,” the nucleotides are always in the 5'→3' order from left to right, and in the case of DNA, unless otherwise specified, “A” represents deoxyadenosine, “C” represents deoxycytidine, “G” represents deoxyguanosine, and “T” represents deoxythymidine. The letters A, C, G, and T may be used to refer to the base itself, a nucleoside, or a nucleotide containing a base.
[0145] As used herein, “processing” refers to a set of steps used to generate a library of nucleic acids suitable for sequencing. The set of steps may include, but is not limited to, a distributing step, an end repair step, a sequencing adapter addition, a tagging step, and / or PCR amplification of nucleic acids.
[0146] As used herein, “quantitative measurement” refers to absolute or relative measurement. A quantitative measurement may be, but is not limited to, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation, or quantile), or a degree or relative quantity (e.g., high, medium, and low). A quantitative measurement may be a ratio of two quantitative measurements. A quantitative measurement may be a linear combination of quantitative measurements. A quantitative measurement may be a normalized measurement.
[0147] As used herein, “reference sequence” refers to a known sequence used for comparison with a sequence determined experimentally. For example, a known sequence may be an entire genome, a chromosome, or any segment thereof. A reference sequence may be aligned with a single contiguous sequence of the genome, or a chromosome, or a chromosomal arm, or it may include discontinuous segments aligned with different regions of the genome or chromosome. Examples of reference sequences include, for example, the human genome, e.g., HG19 and HG38.
[0148] As used herein, "restriction enzyme" is an enzyme that recognizes and cleaves DNA at or near a specific recognition site.
[0149] As used herein, “sample” means any sample that can be analyzed by the methods and / or systems disclosed herein.
[0150] As used herein, “sequencing” refers to any of several technologies used to determine the sequence (e.g., identity and order of monomeric units) of a biomolecule, such as nucleic acids, such as DNA or RNA. Examples of sequencing methods, but not limited to, include, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, sungardideoxy termination sequencing, whole-genome sequencing, hybridization sequencing, pyrosequencing, double-strand sequencing, cycle sequencing, single-nucleotide extension sequencing, solid-phase sequencing, high-throughput sequencing, ultra-parallel signature sequencing, emulsion PCR, cold-PCR (COLD-PCR), multiplex PCR, reversible dye-terminator sequencing, paired-end sequencing, and near-term sequencing. Examples of sequencing methods include exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, synthesis sequencing, real-time sequencing, reverse terminater sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD® sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing may be performed using a gene analyzer, among others commercially available from companies such as Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific.
[0151] As used herein, “sequence information” in the context of nucleic acid polymers means the order and identity of monomeric units (e.g., nucleotides) in that polymer.
[0152] As used herein, “sequence-variable target region set” refers to a set of target regions in neoplastic cells (e.g., tumor cells and cancer cells) that may exhibit sequence changes, such as nucleotide substitutions, insertions, deletions, or gene fusions or transpositions.
[0153] As used herein, the terms “somatic mutation” and “somatic variant” are interchangeable. They refer to genomic mutations that occur after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0154] As used herein, "specifically binds" in the context of a probe or other oligonucleotide and a target sequence means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a replica thereof to form a stable probe:target hybrid, while simultaneously minimizing the formation of stable probe:non-target hybrids. Thus, the probe hybridizes to the target sequence or a replica thereof to a sufficiently greater extent than to the non-target sequence, enabling the capture or detection of the target sequence. Appropriate hybridization conditions are well known in the art and can be predicted based on the sequence composition, or can be determined by using conventional test methods (see, for example, Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, in particular §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, which are incorporated herein by reference).
[0155] As used herein, “Subject” refers to an animal, e.g., a mammalian species (e.g., human), or a bird species (e.g., birds), or another organism, e.g., a plant. More specifically, a subject may be a vertebrate, e.g., a mammal, e.g., a mouse, primate, monkey, or human. Animals include farm animals (e.g., beef cattle, dairy cows, poultry, horses, pigs, etc.), sports animals, and companion animals (e.g., pets or service animals). A subject may be a healthy individual, an individual with or suspected to have a disease or a predisposition to a disease, or an individual in need of or suspected to need treatment. The terms “individual” or “patient” are intended to be interchangeable with “subject.” For example, a subject may be an individual who has been diagnosed with cancer, is scheduled to undergo cancer treatment, and / or has undergone at least one cancer treatment. A subject may be in remission from cancer. As another example, a subject may be an individual who has been diagnosed with an autoimmune disease. As another example, the subjects may be female individuals who have been diagnosed with or are suspected of having a disease, such as cancer or an autoimmune disease, and who are pregnant or planning to become pregnant.
[0156] As used herein, “target region set” or “set of target regions” or “target region” or “target target region” or “target region” or “target genomic region” refers to multiple genomic loci or multiple genomic regions that are targeted for capture and / or targeted by a set of probes (e.g., through sequence complementarity).
[0157] As used herein, “tumor fraction” refers to the proportion of cfDNA molecules originating from tumor cells in a given sample or sample-region pair.
[0158] The terms “or any combination thereof” (singular and plural) as used herein refer to all permutations and combinations of the terms listed before them. For example, “A, B, C, or any combination thereof” is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and is also intended to include BA, CA, CB, ACB, CBA, BCA, BAC, or CAB, where the order is important in the particular context. Continuing this example, combinations containing repetitions of one or more items or terms, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc. A person skilled in the art will understand that, unless particularly evident from the context, there is typically no limit to the number of items or terms in any combination.
[0159] "Or" is used in a comprehensive sense, meaning it is equivalent to "and / or" unless the context specifically requires otherwise. B. Exemplary Methods 1. Overview
[0160] Cancer formation and progression can arise from both genetic modification and epigenetic features of deoxyribonucleic acid (DNA). This disclosure provides methods and systems for analyzing DNA, such as cell-free DNA (cfDNA). This disclosure also provides methods and systems for reducing the signal-to-noise ratio in methylation partition assays.
[0161] While we do not wish to be bound by any particular theory, cells in or around cancer or neoplasms may shed more DNA than cells of the same tissue type in healthy subjects. Therefore, the tissue distribution of a particular DNA sample, such as cfDNA, may change during carcinogenesis. For this reason, for example, an increase in the level of hypermethylated variable target regions, which show lower methylation in healthy cfDNA than in at least one other tissue type, may be an indicator of the presence of cancer (or recurrence, depending on the subject's history). Similarly, an increase in the level of hypomethylated variable target regions in a sample may be an indicator of the presence of cancer (or recurrence, depending on the subject's history).
[0162] Furthermore, cancer can be indicated by non-sequence alterations such as methylation. Examples of methylation changes in cancer include localized increases in DNA methylation at CpG islands in TSSs of genes involved in the regulation of normal growth, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation may be associated with an abnormal loss of transcriptional ability of the genes involved and occurs at least as frequently as point mutations and deletions that cause altered gene expression.
[0163] Therefore, DNA methylation profiling can be used to detect abnormal methylation in the DNA of a sample. DNA is normally hypermethylated or hypomethylated in a given sample type (e.g., cfDNA derived from bloodstream), but may correspond to certain genomic regions ("differentially methylated regions" or "DMRs") that indicate a degree of abnormal methylation correlated with neoplasms or cancer, based on the degree of methylation in the genome, for example, due to an abnormally increased tissue contribution to that sample type (e.g., due to increased shedding of DNA around or in neoplasms or cancers), and / or altered during development or disrupted by disease, e.g., cancer or any cancer-related disease.
[0164] In some embodiments, DNA methylation involves the addition of a methyl group to a cytosine residue at a CpG site (cytosine-phosphate-guanine site (i.e., guanine after cytosine in the 5'→3' direction of the nucleic acid sequence)). In some embodiments, DNA methylation is, for example, N 6- This involves the addition of a methyl group to an adenine residue, as in the case of methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the 6-carbocyclic ring of cytosine). In some embodiments, 5-methylation involves the addition of a methyl group to the 5C position of a cytosine residue to produce 5-methylcytosine (m5c or 5-mC or 5mC). In some embodiments, methylation includes derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC or 5hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the 6-carbocyclic ring of cytosine residue). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of a cytosine residue to produce 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if the DNA in a promoter region is methylated, gene transcription may be repressed. DNA methylation is crucial for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, such as repression, can lead to diseases such as cancer. Methylation of promoters in DNA can indicate cancer.
[0165] Methylation profiling can involve determining methylation patterns across different regions of the genome. For example, after distributing and sequencing molecules based on their degree of methylation (e.g., the relative number of methylated nucleotides per molecule), the sequences of molecules in different distributions can be mapped to a reference genome. This can reveal regions of the genome that are more or less highly methylated compared to other regions. In this way, genomic regions may exhibit different degrees of methylation compared to individual molecules.
[0166] Combining signals obtained from methylation profiling with signals obtained from somatic mutations (e.g., SNVs, indels, CNVs, and gene fusions) facilitates cancer detection.
[0167] Nucleic acid molecules in a sample can be fractionated or distributed based on the methylation status of the nucleic acid molecules. Distributing nucleic acid molecules in a sample can increase the detection of rare signals. For example, genetic variations that are present in highly methylated DNA but are not very (or not at all) present in low-methylated DNA can be more easily detected by distributing the sample into highly methylated and low-methylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of single molecules can be performed, and therefore higher sensitivity can be achieved. Distribution may involve physically distributing nucleic acid molecules into subsets or groups based on the presence or absence of one or more methylated nucleotides. A sample can be fractionated or distributed into one or more distribution sets based on features indicating differential gene expression or disease status. A sample can be fractionated based on features or combinations that result in signal differences between normal and diseased states during the analysis of nucleic acids, e.g., cell-free DNA ("cfDNA"), non-cfDNA, tumor DNA, circulating tumor DNA ("ctDNA"), and cell-free nucleic acids ("cfNA").
[0168] The distribution procedure may result in incomplete sorting of DNA molecules between partial samples. For example, a small portion of the molecules in the second partial sample may be highly modified (e.g., highly methylated), and / or a small portion of the molecules in the first partial sample may be unmodified or only slightly modified (e.g., unmethylated or mostly unmethylated). The highly modified molecules in the second partial sample and the unmodified or only slightly modified molecules in the first partial sample are considered to be distributed nonspecifically. The method described herein includes steps that can reduce the technical noise originating from nonspecifically distributed DNA, for example, by decomposing it and / or by transforming certain bases so that the nonspecifically distributed DNA can be identified after sequencing. Thus, the method described herein may provide improved sensitivity and / or streamlined analysis.
[0169] Figure 2 illustrates an embodiment of method 200 as an example of a method for determining the methylation status of nucleic acid molecules in a sample obtained from a subject. In 202, a polynucleotide sample is obtained from the subject. In some embodiments, the sample is a DNA sample obtained from a tumor tissue biopsy. In some embodiments, the sample is a cell-free DNA (cfDNA) sample obtained from blood. In 204, the polynucleotide sample is distributed into at least two distribution sets (partial samples). In some embodiments, the distribution involves distributing nucleic acid molecules based on the differential binding affinity of the polynucleotide to a binder that preferentially binds to the polynucleotide containing methylated nucleotides.
[0170] In 206, nucleic acid molecules in at least one partition set are digested with at least one methylation-dependent nuclease, such as a methylation-dependent restriction enzyme (MDRE). In some embodiments, nucleic acids in at least one partition set are digested with at least two MDREs. In some embodiments, two MDREs are used to digest nucleic acid molecules in at least one partition set. Any or a combination of MDREs described elsewhere in this specification may be used.
[0171] In some embodiments, at least one adapter is attached to at least one end of the nucleic acid molecule (i.e., the 5' and / or 3' ends of the DNA molecule) before restriction digestion in MDRE. In other embodiments, at least one adapter is attached to at least one end of the nucleic acid molecule after digestion but before enrichment in 208. In some embodiments, the adapter is resistant to digestion by methylation-dependent nucleases or restriction enzymes due to, for example, the presence of a nucleotide or a suitable nucleotide analog (e.g., a ligation variant, e.g., a nucleotide analog having a phosphorothioate).
[0172] In 208, after MDRE digestion, nucleic acid molecules in one or more distribution sets can be enriched with respect to a genomic region of interest. Alternatively, the enrichment step can be performed before the distribution step. In some embodiments, the genomic region of interest may include differentially methylated regions (e.g., a set of highly methylated variable target regions and / or a set of lowly methylated variable target regions) for cancer detection. In 210, at least a subset of the enriched molecules are sequenced by a next-generation sequencer. In 212, the sequencing read data generated by the sequencer is then analyzed using bioinformatics tools / algorithms to determine the number of molecules in one or more distribution sets, and this is then used to determine the methylation status of nucleic acid molecules in at least one distribution set at one or more loci. In some embodiments, one or more loci may include multiple loci. In some embodiments, one or more loci may include one or more genomic regions. In some embodiments, a genomic region may be a gene promoter region. In some embodiments, the nucleic acid molecules can be amplified by PCR amplification before sequencing. In some embodiments, the primers used in amplification may include at least one sample index.
[0173] Figure 3 illustrates an embodiment of a method 300 for detecting the presence or absence of cancer in a subject according to embodiments of the present disclosure. In 302, a polynucleotide sample is obtained from the subject. In some embodiments, the polynucleotide sample is a DNA sample obtained from a tumor tissue biopsy. In some embodiments, the polynucleotide sample is a cell-free DNA (cfDNA) sample obtained from blood. In 304, the polynucleotide sample is distributed into at least two distribution sets. In some embodiments, the distribution involves distributing nucleic acid molecules based on the differential binding affinity of the polynucleotide to a binder that preferentially binds to the polynucleotide, including methylated nucleotides. Examples of binders include, but are not limited to, methyl-binding domains (MBDs) and methyl-binding proteins (MBPs), which are discussed in detail elsewhere in this specification.
[0174] In 306, nucleic acid molecules in one or more distribution sets are attached to an adapter, the adapter comprising at least one tag attached to at least one end of the nucleic acid molecule (i.e., the 5' and 3' ends of the DNA molecule). In some embodiments, the adapter is resistant to digestion by methylation-dependent restriction enzymes. In some embodiments, the adapter comprises an unmethylated nucleotide. In some embodiments, the adapter comprises one or more nucleotide analogs resistant to methylation-dependent restriction enzymes. In some embodiments, the adapter comprises a nucleotide sequence not recognized by methylation-dependent restriction enzymes. In some embodiments, the tag may be provided as a component of the adapter. In some embodiments, the tag comprises a molecular barcode (i.e., a molecular identifier). In some embodiments, the tag attached to a nucleic acid molecule in one distribution set is different from the tag attached to a nucleic acid molecule in the other distribution set. In some embodiments, one distribution set is tagged differentially from the other distribution set. Differential tagging of distribution sets helps to keep track of nucleic acid molecules belonging to a particular distribution set. Nucleic acid molecules in different distribution sets accept different tags that allow them to distinguish members of one distribution set from others. Tags attached to nucleic acid molecules in the same distribution set may be the same or different. However, if they are different, the tags may have a common sequence portion to identify the molecule to which they are attached as belonging to a particular distribution set. For example, if molecules in a sample are distributed into two distribution sets, P1 and P2, molecules in P1 may be tagged with A1, A2, A3, etc., and molecules in P2 may be tagged with B1, B2, B3, etc. Such a tagging system makes it possible to distinguish between distribution sets and molecules within them. In some embodiments, the tags include distribution tags (i.e., distribution identifiers). In such embodiments, nucleic acid molecules in a distribution set accept the same distribution tag, which is different from the distribution tag attached to nucleic acid molecules in the other distribution set.
[0175] In 308, nucleic acid molecules in at least one partition set are digested with at least one methylation-dependent nuclease or methylation-dependent restriction enzyme (MDRE). In some embodiments, nucleic acids in at least one partition set are digested with at least two MDREs. In some embodiments, two MDREs are used to digest nucleic acid molecules in at least one partition set. The MDREs may be any MDREs or combinations thereof as described elsewhere in this specification.
[0176] In 310, after MDRE digestion, nucleic acid molecules in one or more distribution sets can be enriched with respect to a genomic region of interest. Alternatively, the enrichment step can be performed before the distribution step. In some embodiments, the genomic region of interest may include differentially methylated regions for cancer detection. In 312, at least a subset of the enriched molecules are sequenced by a next-generation sequencer. In 314, the sequencing read data generated by the sequencer is then analyzed using a bioinformatics tool / algorithm to determine the number of molecules in one or more distribution sets, and this is then used to determine the methylation status of nucleic acid molecules in at least one distribution set at one or more loci. In some embodiments, one or more loci may include multiple loci. In some embodiments, one or more loci may include one or more genomic regions. In some embodiments, a genomic region may be a gene promoter region. In some embodiments, the nucleic acid molecules can be amplified by PCR amplification before sequencing. In some embodiments, the primers used in amplification may include at least one sample index.
[0177] In some embodiments, the method may further include the step of detecting the presence or absence of cancer in a subject based, for example, on the methylation status of nucleic acid molecules in at least one distribution set at one or more loci. In some embodiments, the method further includes the step of determining the level of DNA derived from tumor cells in a polynucleotide sample.
[0178] Figure 4 illustrates an exemplary workflow for detecting the presence or absence of cancer according to a particular embodiment of the present disclosure, for example, starting with a cfDNA sample, wherein the cfDNA is isolated from a blood sample, and the cfDNA sample includes cfDNA molecules belonging to a highly methylated variable target region (high DMR), cfDNA molecules belonging to a low methylated variable target region (low DMR), and cfDNA molecules belonging to a non-methylated control region. The cfDNA is distributed into a low-methylated partial sample and a highly methylated partial sample using a methyl-binding domain protein (MBD), each distribution set is subjected to molecular barcoding to identifyly tag the DNA derived from the partial sample, the low-distribution set is digested with one or more MDREs to cleave methylated cfDNA molecules at the RE recognition site, and optionally the highly-distribution set is digested with one or more MSREs to cleave non-methylated cfDNA molecules at the RE recognition site, and then the distribution sets (including the low-distribution set digested with MDREs) are pooled, captured, amplified, and sequenced. 2. Distribute the sample into multiple partial samples; sample configuration
[0179] In certain embodiments described herein, populations of different forms of nucleic acids (e.g., hypermethylated and hypomethylated DNA in a sample, e.g., cfDNA) may be physically partitioned based on one or more characteristics of the nucleic acids before further analysis, e.g., contact with nucleases, differential modification or isolation of nucleic acid bases, tagging, and / or sequencing. Using this approach, for example, it is possible to determine whether a particular sequence is hypermethylated or hypomethylated. Furthermore, by partitioning heterogeneous nucleic acid populations, rare signals can be increased, for example, by enriching rare nucleic acid molecules that are more abundant in one fraction (or partition) of the population. For example, genetic variations that are present in hypermethylated DNA but less (or none) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of a single locus of nucleic acid genome or species can be performed, and thus higher sensitivity can be achieved.
[0180] In some examples, heterogeneous nucleic acid samples are distributed into two or more partitions (e.g., at least three, four, five, six, or seven partitions). The sample partitions are also referred to herein as partial samples. In some embodiments, each partition is differentially tagged. The tagged partitions can then be pooled together for collective sample preparation and / or sequencing. The partition-tagging-pooling steps may be performed more than once, with each partition being based on different characteristics (examples are provided herein) and tagged using differential tags that identify them with other partitions and distributing means.
[0181] Examples of features that can be used for distribution include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting distribution may include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, distribution based on cytosine modifications (e.g., cytosine methylation) or methylation is commonly performed and, if necessary, combined with at least one additional distribution step which may be based on any of the aforementioned DNA features or forms. In some embodiments, a heterogeneous population of nucleic acids is distributed into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine versus other types of methylation, e.g., adenine methylation and / or cytosine hydroxymethylation), as well as association with one or more proteins, e.g., histones, and the level of association. Alternatively, the heterogeneous nucleic acid population may be distributed between nucleic acid molecules associated with nucleosomes and nucleic acid molecules lacking nucleosomes. Alternatively, the heterogeneous nucleic acid population may be distributed between single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, the heterogeneous nucleic acid population may be distributed based on the length of the nucleic acid (e.g., molecules up to 160 bp and molecules longer than 160 bp).
[0182] In some cases, each partition (representing a different nucleic acid morphology) is differentially labeled, and the partitions are pooled together before sequencing. In other cases, the different morphologies are sequenced separately.
[0183] In some embodiments, a population of different nucleic acids is divided into two or more different partitions. Each partition represents a different nucleic acid morphology, with the first partition (also referred to as a partial sample) containing DNA with a higher proportion of cytosine modifications than the second partial sample. Each partition is explicitly tagged. The first partial sample is subjected to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA, where the first nucleic acid base is modified or unmodified, and the second nucleic acid base is a modified or unmodified nucleic acid base different from the first, and the first and second nucleic acid bases have the same base-pairing specificity. The tagged nucleic acids are pooled together before sequencing. Sequence reading data is obtained and analyzed in silico, for example, to distinguish the first nucleic acid base from the second nucleic acid base in the DNA of the first partial sample. The tags are used to sort the reading data from different partitions. Analysis to detect genetic variants can be performed at the partition level as well as at the whole nucleic acid population level. For example, the analysis may include in silico analysis to determine genetic variants in nucleic acids during each distribution, such as CNVs, SNVs, indels, and fusions. In some cases, the in silico analysis may include determining chromatin structure. For example, the coverage of sequence reading data can be used to determine the positioning of nucleosomes in chromatin. Higher coverage may correlate with higher nucleosome occupancy in genomic regions, while lower coverage may correlate with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0184] The samples may contain nucleic acids that differ in their modification, including post-replication modifications to nucleotides, and their binding to one or more proteins, usually by non-covalent bonds.
[0185] In one embodiment, the nucleic acid population is obtained from serum, plasma, or blood samples from a subject suspected of having a neoplasm, tumor, or cancer, or a subject previously diagnosed with a neoplasm, tumor, or cancer. The nucleic acid population includes nucleic acids having varying levels of methylation. Methylation may result from any one or more post-replication or transcriptional modifications. Post-replication modifications include modifications of nucleotide cytosines, particularly those at the 5-position of the nucleic acid base, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0186] The affinity agent may be an antibody with desired specificity, a natural binding partner or its variant (Bock et al., Nat Biotech 28: 1106-1114 (2010), Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected to have specificity for a given target, for example by phage display.
[0187] Examples of capture portions intended herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, which include proteins such as MeCP2, MBDs such as MBD2, and antibodies that preferentially bind to 5-methylcytosine. When methylated DNA is immunoprecipitated using an antibody, the methylated DNA can be recovered in single-strand form. In such embodiments, a second strand can be synthesized. A highly methylated (and optionally intermediate-methylated) partial sample can then be contacted with a methylation-sensitive nuclease that does not cleave semi-methylated DNA, such as HpaII, BstUI, or Hin6i. Alternatively, a less methylated (and optionally intermediate-methylated) partial sample can then be contacted with a methylation-dependent nuclease that cleaves semi-methylated DNA.
[0188] Similarly, the step of distributing different forms of nucleic acids can be carried out using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.
[0189] With respect to certain affinity agents and modifications, binding to the agonist may occur in an all-or-nothing manner, depending on whether the nucleic acid has the modification, while separation may be degree-dependent. In such cases, nucleic acids with overexpression of the modification bind to the agonist to a higher degree than nucleic acids with underexpression of the modification. Alternatively, nucleic acids with the modification may bind in an all-or-nothing manner. However, varying levels of modification may then be sequentially eluted from the binder.
[0190] For example, in some embodiments, the distribution may be binary or based on the degree / level of modification. For instance, all methylated fragments can be distributed from the unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Further distribution may then involve eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and the bound fragments. As the salt concentration increases, fragments with higher levels of methylation are eluted.
[0191] In some cases, the final allocation will represent nucleic acids with varying degrees of modification (overexpression or underexpression of the modification). Overexpression and underexpression can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is 2, nucleic acids containing more than 2 5-methylcytosine residues are overexpressing this modification, while nucleic acids with 1 or 0 5-methylcytosine residues are underexpressing it. The effect of affinity separation is to enrich nucleic acids that are overexpressing the modification in the conjugated phase and nucleic acids that are underexpressing the modification in the unconjugated phase (i.e., in solution). Nucleic acids in the conjugated phase can be eluted before subsequent processing.
[0192] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be partitioned using sequential elution. For example, low-methylation partitioning (no methylation) can be separated from methylation partitioning by contacting the nucleic acid population with MBD from the kit attached to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at salt concentrations of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After eluting such methylated nucleic acids, magnetic separation is used again to separate the nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps themselves can be repeated to create various partitions, such as a low-methylation partition (representing no methylation), a methylated partition (representing low levels of methylation), and a high-methylation partition (representing high levels of methylation).
[0193] In some methods, nucleic acids bound to the activator used for affinity separation are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity activator. Such nucleic acids can be enriched with nucleic acids that have modifications to a degree close to the mean or median (i.e., an intermediate value between nucleic acids that remained bound to the solid phase when the sample was first brought into contact with the activator and nucleic acids that were not bound to the solid phase).
[0194] Affinity segregation results in at least two, sometimes three or more, distributions of nucleic acids having different degrees of modification. The distributions, while still distinct, involve the attachment of nucleic acids from at least one distribution, and typically two or three (or more) distributions, to nucleic acid tags, usually provided as components of an adapter, where nucleic acids in different distributions receive different tags that identify one distribution member as another. Tags attached to nucleic acid molecules in the same distribution may be identical or different from each other. However, if different, the tags may have a common coding portion to identify the molecule to which they are attached as belonging to a particular distribution.
[0195] For further details regarding the distribution of nucleic acid samples based on characteristics such as methylation, please refer to WO2018 / 119452, which is incorporated herein by reference.
[0196] In some embodiments, nucleic acid molecules may be allocated to different distributions based on which nucleic acid molecules are bound to a particular protein or fragment thereof and which are not bound to that particular protein or fragment thereof.
[0197] Nucleic acid molecules can be partitioned based on DNA-protein binding. Protein-DNA complexes can be partitioned based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that can bind to DNA and function as a criterion for fractionation include, but are not limited to, protein A and protein G. Nucleic acid molecules can be partitioned based on protein-binding regions using any preferred method. Examples of methods used to partition nucleic acid molecules based on protein-binding regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric flow-field separation (AF4).
[0198] In some embodiments, the nucleic acid distribution step is performed by contacting the nucleic acid with the methylation-binding domain ("MBD") of a methylation-binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 streptavidin, via a biotin linker. Distribution to fractions with different degrees of methylation can be performed by eluting the fractions by increasing the NaCl concentration.
[0199] Examples of MBPs intended herein include, but are not limited to, the following: (a) MeCP2 and MBD2 are proteins that preferentially bind to 5-methylcytosine rather than unmodified cytosine. (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethylcytosine rather than unmodified cytosine. (c)FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind to 5-formyl-cytosine more than unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)). (d) An antibody specific to one or more methylated nucleotide bases.
[0200] Generally, elution is a function of the number of methylation sites per molecule, with higher salt concentrations eluting molecules that have more methylation. A series of elution buffers with progressively increasing NaCl concentrations can be used to elute DNA into separate populations, also based on the degree of methylation. The salt concentration can range from about 100 nm to about 2500 mM NaCl. In one embodiment, the process yields three partitions. Molecules are brought into contact with a solution containing molecules with methyl-binding domains at a first salt concentration, allowing the molecules to attach to a capture site, such as streptavidin. At the first salt concentration, some molecular populations bind to the MBD, while others remain unbound. The unbound population can be separated as a "low-methylated" population. For example, the first partition, representing a low-methylated form of DNA, remains unbound at low salt concentrations, e.g., 100 mM or 160 mM. The second partition, representing intermediate methylated DNA, is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM. This is also separated from the sample. The third partition, representing the highly methylated form of DNA, is eluted using a high salt concentration, e.g., at least about 2000 mM. a. Tagging of partitions
[0201] In some embodiments, two or more distributions, for example, each distribution, are differentially tagged. The tag or index may be a molecule, such as a nucleic acid, containing information that characterizes the molecule to which the tag associates. The tag may allow for the distinction of the molecule from which the sequence reading data originates. For example, a molecule may have a sample tag or sample index (thereby distinguishing a molecule in one sample from one in a different sample), a distribution tag (thereby distinguishing a molecule in one distribution from one in a different distribution), or a molecular tag / molecular barcode / barcode (thereby distinguishing different molecules from one another (in both unique and non-unique tagging scenarios)). In certain embodiments, a tag may include one barcode or a combination of barcodes. As used herein, the term “barcode” refers, depending on the context, to a nucleic acid molecule having a particular nucleotide sequence, or the nucleotide sequence itself. A barcode may have, for example, between 10 and 100 nucleotides. A collection of barcodes may have a degenerate sequence as desired for a specific purpose. It may have sequences that have a certain Hamming distance. For example, a molecular barcode may consist of one barcode or a combination of two barcodes, each attached to a different end of the molecule. Furthermore, or for different distributions and / or samples, different sets of molecular barcodes, molecular tags, or molecular indices may be used so that the barcodes function as molecular tags through their individual sequences and also function to identify the corresponding distributions and / or samples based on the set to which they are members. Tags containing barcodes may be incorporated into an adapter or otherwise ligated to an adapter. Tags may be incorporated by ligation, overlap extension PCR, among other methods.
[0202] Tagging strategies can be divided into unique tagging strategies and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample have different tags, so that the read data can be assigned to the original molecules based on the tag information alone. Tags used in such a method may be referred to as "unique tags". In non-unique tagging, different molecules in the same sample may have the same tag, so that in addition to the tag information, other information is used to assign the sequence read data to the original molecules. Such information can include start and end coordinates, coordinates to which the molecule is mapped, start or end coordinates alone, etc. Tags used in such a method may be referred to as "non-unique tags". Thus, it is not necessarily the case that all molecules in a sample are uniquely tagged. It is sufficient to uniquely tag the molecules contained within an identifiable class within the sample. Thus, molecules within different identifiable families can have the same tag without losing information regarding the identity of the tagged molecules.
[0203] In certain embodiments of non-unique tagging, the number of different tags used can be sufficient if there is a very high probability (e.g., at least 99%, at least 99.9%, at least 99.99%, or at least 99.999%) that all molecules in a particular group have different tags. Note that when using barcodes as tags, and the barcodes are attached, for example, randomly to both ends of the molecule, the combinations of barcodes can together constitute a tag. This number is a function of the number of molecules clearly included in the call. For example, a class can be all molecules that map to the same start-end position on a reference genome. A class can be all molecules that map across a particular locus, e.g., a particular base or a particular region (e.g., up to 100 bases or a gene or an exon of a gene). In certain embodiments, the number of different tags used to uniquely identify the number z of molecules within a class is 2 * z, 3 * z, 4 * z, 5* z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z, or 100 * One of z (for example, the lower limit) and 100,000 * z, 10,000 * z, 1000 * z, or 100 * It can be between any of z (for example, the upper limit).
[0204] For example, in a sample of approximately 5 ng to 30 ng of cell-free DNA, it is predicted that approximately 3,000 molecules will be mapped to specific nucleotide coordinates, and that between 3 and 10 molecules with arbitrary start coordinates will share the same end coordinate. Therefore, approximately 50 to 50,000 different tags (e.g., between 6 and 220 barcode combinations) may be sufficient to uniquely tag all such molecules. To uniquely tag all 3,000 molecules that map to the entire nucleotide coordinate system, approximately 1 million to 20 million different tags would be required.
[0205] Generally, the assignment of unique or non-unique tag barcodes in a reaction follows the methods and systems described in U.S. Patent Applications No. 20010053519, 20030152490, 20110160078, and U.S. Patents No. 6,582,908, 7,537,898, and 9,598,731.
[0206] The tags can be attached to the sample nucleic acid randomly or non-randomly.
[0207] In some embodiments, tagged nucleic acids are sequenced after being loaded into a microwell plate. The microwell plate may have 96, 384, or 1536 microwells. In some cases, they are introduced in a ratio of uniquely tagged microwells to the expected number of uniquely tagged microwells. For example, unique tags may be loaded so that more than approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 10000, 10000, 100000, 100000, 100000, 500000, 1,000000, 10,000000, 5000000, or 1,000000000 unique tags are loaded per genome sample. In some cases, unique tags may be loaded so that approximately fewer than 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 10000, 50000, 50000, 50000, 100000, 100000, 500000, 1000000, 10000000, 10000000, 5000000, or 1000000000 unique tags are loaded per genome sample.In some cases, the average number of unique tags loaded per sample genome is approximately less than or more than 1, less than or more than 2, less than or more than 3, less than or more than 4, less than or more than 5, less than or more than 6, less than or more than 7, less than or more than 8, less than or more than 9, less than or more than 10, less than or more than 20, less than or more than 50, less than or more than 100. , less than or more than 500, less than or more than 1,000, less than or more than 5,000, less than or more than 10,000, less than or more than 50,000, less than or more than 100,000, less than or more than 500,000, less than or more than 1,000,000, less than or more than 10,000,000, less than or more than 50,000,000, or less than or more than 1,000,000,000 unique tags.
[0208] A preferred format uses 20 to 50 different tags (e.g., barcodes) ligated to both ends of the target nucleic acid. For example, ligating 35 different tags (e.g., barcodes) to both ends of a target molecule creates 35 × 35 permutations, which is equal to 1225 for 35 tags. Such a number of tags is sufficient for different molecules with the same start and end points to accept different combinations of tags with a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%). Other barcode combinations include any number between 10 and 500, e.g., approximately 15 × 15, approximately 35 × 35, approximately 75 × 75, approximately 100 × 100, approximately 250 × 250, and approximately 500 × 500.
[0209] In some cases, the unique tag may be an oligonucleotide of a predetermined, random, or semi-random sequence. In other cases, multiple barcodes may be used such that the barcodes are not necessarily unique to one another. In this example, the barcodes may be ligated to individual molecules such that the combination of the barcode and the sequence it can be ligated to forms a unique sequence that can be tracked individually. As described herein, the detection of a non-unique barcode combined with the sequence data of the beginning and end portions of the sequence reading data may enable the assignment of unique identity to a particular molecule. The length or number of base pairs of individual sequence reading data may also be used to assign unique identity to such molecules. As described herein, a fragment derived from a single strand of nucleic acid to which unique identity has been assigned may thereby enable the identification of a subsequent fragment derived from the parent strand.
[0210] Tags can be used to label individual polynucleotide population distributions in order to correlate tags(one or more) with specific distributions. Alternatively, tags can be used in embodiments of the present invention that do not involve a distribution step. In some embodiments, a single tag can be used to label a specific distribution. In some embodiments, multiple different tags can be used to label a specific distribution. In embodiments where multiple tags are used to label a specific distribution, the set of tags used to label one distribution can be easily distinguished from the set of tags used to label other distributions. In some embodiments, tags may have additional functions, for example, being used to index the origin of a sample or as a unique molecular identifier (which distinguishes sequencing errors from mutations, as seen, for example, in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011) and Kou et al., PLoS ONE, 11: e0146638 (2016)). They can be used as (which can be used to improve the quality of sequencing data) or as non-unique molecular identifiers, for example, as described in U.S. Patent No. 9,598,731. Similarly, in some embodiments, tags may have additional functions, for example, they can be used to index the origin of a sample or as non-unique molecular identifiers (which can be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations).
[0211] In one embodiment, distribution tagging involves tagging the molecules in each distribution with distribution tags. After the distributions are recombined (e.g., to reduce the number of sequencing runs required and avoid unnecessary costs) and the molecules are sequenced, the distribution tags identify the distribution of origin. In another embodiment, different distributions are tagged with different sets of molecular tags, for example, consisting of pairs of barcodes. Thus, each molecular barcode is useful for identifying the distribution of origin, as well as the molecules within the distribution. For example, a first set of 35 barcodes can be used to tag the molecules in a first distribution, while a second set of 35 barcodes can be used to tag the molecules in a second distribution.
[0212] In some embodiments, after distribution and tagging with distribution tags, molecules may be pooled for single-pass sequencing. In some embodiments, sample tags are attached to molecules, for example, in a step after the attachment of distribution tags and pooling. Sample tags can facilitate pooling materials generated from multiple samples for sequencing in single-pass sequencing.
[0213] Alternatively, in some embodiments, the distribution tags may be correlated with the sample and the distribution. As a simple example, a first tag may indicate a first distribution of a first sample, a second tag may indicate a second distribution of a first sample, a third tag may indicate a first distribution of a second sample, and a fourth tag may indicate a second distribution of a second sample.
[0214] Tags can be attached to already distributed molecules based on one or more characteristics, but the final tagged molecules in the library may no longer possess those characteristics. For example, single-stranded DNA molecules may be distributed and tagged, but the final tagged molecules in the library are likely to be double-stranded. Similarly, DNA may be subjected to distribution based on different levels of methylation, but the tagged molecules derived from these molecules in the final library are likely to be unmethylated. Therefore, tags attached to molecules in the library typically represent characteristics of the "parent molecule" from which the final tagged molecule originates, rather than characteristics of the tagged molecule itself.
[0215] For example, molecules in the first distribution may be tagged and labeled using barcodes 1, 2, 3, 4, etc., molecules in the second distribution may be tagged and labeled using barcodes A, B, C, D, etc., molecules in the third distribution may be tagged and labeled using barcodes a, b, c, d, etc. Differentially tagged distributions may be pooled before sequencing. Differentially tagged distributions may be sequenced separately or simultaneously together in the same flow cell of an Illumina sequencer, for example.
[0216] After sequencing, analysis of read data to detect genetic variants can be performed at the distribution level as well as at the whole nucleic acid population level. Tags are used to sort read data from different distributions. Analysis may include in silico analysis to determine genetic and epigenetic variants (one or more of the following, such as methylation and chromatin structure) using sequence information, genomic coordinate length, coverage, and / or copy number. In some embodiments, higher coverage may correlate with higher nucleosome occupancy in genomic regions, while lower coverage may correlate with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0217] In some embodiments, an adapter is used that does not contain a sequence recognized by the nuclease used in the method and / or is resistant to cleavage due to the presence of, for example, nucleotide modifications, e.g., ligation modifications (e.g., phosphorothioates). In some embodiments, a tag is used that does not contain a sequence recognized by the nuclease used in the method and / or is resistant to cleavage due to the presence of, for example, nucleotide modifications, e.g., ligation modifications (e.g., phosphorothioates). When both one or more methylation-dependent restriction enzymes and one or more methylation-sensitive restriction enzymes are used, the adapter and / or tag may be methylation-deficient and may lack a recognition sequence for one or more methylation-sensitive restriction enzymes so that they do not become substrates for cleavage by any of the restriction enzymes used. b. Alternative methods for modified nucleic acid analysis
[0218] In some embodiments, the adapter may be attached to the nucleic acid after the nucleic acid has been distributed, and in other embodiments, the adapter may be attached to the nucleic acid before it has been distributed. In some such methods, a population of nucleic acids having different degrees of modification (e.g., 0, 1, 2, 3, 4, 5, or more methyl groups per nucleic acid molecule) is brought into contact with the adapter before fractionation of the population, depending on the degree of modification. The adapter attaches to one or both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a number of different tags sufficient to result in a low probability, e.g., 95, 99, or 99.9%, that two nucleic acids having the same start and end points accept the same tag combination. The adapter may contain the same or different primer binding sites, regardless of whether they have the same or different tags, but preferably the adapter contains the same primer binding sites. After the attachment of the adapter, the nucleic acid is brought into contact with an activator that preferentially binds to the modified nucleic acid (e.g., such activators described previously). Nucleic acids are distributed from binding to an agonist into at least two partial samples with different degrees of modification. For example, if the agonist has affinity for modified nucleic acids, nucleic acids with overexpression of the modification (compared to the median occurrence in the population) will preferentially bind to the agonist, while nucleic acids with underexpression of the modification will not bind or will elute more easily from the agonist. After distribution, the first partial sample is subjected to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA, where the first nucleic acid base is modified or unmodified, and the second nucleic acid base is a modified or unmodified nucleic acid base different from the first, and the first and second nucleic acid bases have the same base-pairing specificity. The nucleic acids are then amplified from a primer that binds to the primer-binding site in the adapter. After amplification, the different distributions may then be subjected to further processing steps, which typically include further (e.g., clonal) amplification and sequence analysis, in parallel but separately. Sequence data from different distributions can then be compared.
[0219] In another embodiment, the distribution scheme can be carried out using the following exemplary procedure: Nucleic acids are ligated to both ends of a Y-shaped adapter containing a primer binding site and a tag. The molecules are amplified. The amplified molecules are then fractionated by contact with an antibody that preferentially binds to 5-methylcytosine to produce two partitions. One partition contains the original molecule lacking methylation and an amplified copy that has lost methylation. The other partition contains the original DNA molecule with methylation. The partition containing the original DNA molecule with methylation is subjected to a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA, where the first nucleic acid base is a modified or unmodified nucleic acid base, and the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity. The two partitions are then processed and sequenced separately from further amplification of the methylation partition. The sequence data of the two distributions can then be compared. In this example, the tags are not used to distinguish between methylated and unmethylated DNA, but rather to distinguish between different molecules within these distributions so that read data with the same start and end points can be determined as to whether they are based on the same molecule or different molecules.
[0220] This disclosure further provides methods for analyzing a population of nucleic acids in which at least a portion of the nucleic acid contains one or more modified cytosine residues, such as 5-methylcytosine, and any of the other modifications described above. In these methods, after distribution, a portion of the nucleic acid sample is brought into contact with an adapter containing one or more cytosine residues modified at the 5C position, e.g., 5-methylcytosine. Preferably, all cytosine residues in such an adapter are also modified, or all such cytosines in the primer-binding region of the adapter are modified. The adapter is attached to both ends of the nucleic acid molecule in the population. Preferably, the adapter contains a number of different tags sufficient to result in a low probability, e.g., 95, 99, or 99.9%, that two nucleic acids having the same start and end points accept the same tag combination. The primer-binding sites in such an adapter may be the same or different, but are preferably the same. After attachment of the adapter, the nucleic acid is amplified from a primer that binds to the primer-binding site of the adapter. The amplified nucleic acid is divided into a first aliquot and a second aliquot. The first aliquot is assayed for sequence data with or without further processing. The sequence data of the molecule in the first aliquot is therefore determined regardless of the initial methylation state of the nucleic acid molecule. The nucleic acid molecule in the second aliquot is subjected to a procedure that affects the first nucleic acid base in the DNA differently than the second nucleic acid base in the DNA, where the first nucleic acid base contains cytosine modified at position 5 and the second nucleic acid base contains unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. The nucleic acids subjected to the procedure are then amplified with primers to the original primer binding sites of adapters ligated to the nucleic acids. These nucleic acids retain cytosine at the primer binding sites of the adapters, but the amplified products have lost methylation of these cytosine residues due to conversion to uracil in bisulfite treatment; therefore, only the nucleic acid molecule originally ligated to the adapter (different from its amplified product) is amplified here.Therefore, only the original molecules within the population that are at least partially methylated undergo amplification. After amplification, these nucleic acids are subjected to sequence analysis. A comparison of the sequences determined from the first and second aliquots can, among other things, indicate which cytosines within the nucleic acid population were subjected to methylation.
[0221] Such analysis can be performed using the following exemplary procedure. After distribution, the methylated DNA is ligated to both ends of a Y-shaped adapter containing a primer binding site and a tag. The cytosine in the adapter is modified at position 5 (e.g., 5-methylation). The modification of the adapter serves to protect the primer binding site in subsequent conversion steps (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but affects the unmodified cytosine). After the adapter is attached, the DNA molecule is amplified. The amplified product is divided into two aliquots for sequencing with and without conversion. The aliquot not subjected to conversion can be subjected to sequence analysis with or without further processing. The other aliquot is subjected to a procedure that affects a first nucleic acid base in the DNA differently from a second nucleic acid base in the DNA, where the first nucleic acid base contains cytosine modified at position 5 and the second nucleic acid base contains unmodified cytosine. This procedure may involve bisulfite treatment or another procedure for converting unmodified cytosine to uracil. When contacted with a primer specific to the original primer binding site, only the primer binding site protected by the cytosine modification can aid in amplification. Therefore, only the original molecule, and not copies derived from the first amplification, is subjected to further amplification. The further amplified molecule is then subjected to sequence analysis. The sequences can then be compared from the two aliquots. As in the separation scheme considered above, the nucleic acid tag in the adapter is used to identify nucleic acid molecules within the same distribution, not to distinguish between methylated and unmethylated DNA. 3. Contact the partial sample with a methylation-dependent or methylation-sensitive nuclease.
[0222] In some embodiments, partial samples (for example, first, second, or third partial samples prepared by partitioning the sample based on the level of cytosine modification, e.g., methylation, e.g., 5-methylation, as described herein) are brought into contact with a methylation-dependent nuclease or a methylation-sensitive nuclease. Unless otherwise indicated, if the partitioning is based on cytosine modification, the first partial sample is a partial sample having a higher level of modification, the second partial sample is a partial sample having a lower level of modification, and, if present, the third partial sample has an intermediate level of modification between the first and second partial samples.
[0223] As discussed above, the distribution procedure may result in incomplete sorting of DNA molecules between partial samples. The selection of methylation-dependent nucleases or methylation-sensitive nucleases may be carried out to degrade nonspecifically distributed DNA. For example, a second partial sample may be contacted with a methylation-dependent nuclease, such as a methylation-dependent restriction enzyme. This can degrade the nonspecifically distributed DNA (e.g., methylated DNA) in the second partial sample to produce a treated second partial sample. Alternatively, the first partial sample may be contacted with a methylation-sensitive endonuclease, such as a methylation-sensitive restriction enzyme, thereby degrading the nonspecifically distributed DNA in the first partial sample to produce a treated first partial sample. Degradation of nonspecifically distributed DNA in one or both of the first or second partial samples is proposed as an improvement in the performance of methods that rely on the precise distribution of DNA based on cytosine modification, for example, to detect the presence of abnormally modified DNA in a sample, to determine the tissue from which the DNA originates, and / or to determine whether a subject has cancer. For example, such degradation may result in improved sensitivity and / or simplify downstream analysis. In general, when the nonspecifically distributed DNA is highly methylated, such as in hypomethylated distribution, methylation-dependent nucleases, e.g., methylation-dependent restriction enzymes, should be used. In contrast, when the nonspecifically distributed DNA is hypomethylated, such as in highly methylated distribution, methylation-sensitive nucleases, e.g., methylation-sensitive restriction enzymes, should be used. Methylation-dependent nucleases, e.g., methylation-dependent restriction enzymes, preferentially cleave methylated DNA compared to unmethylated DNA, while methylation-sensitive nucleases, e.g., methylation-sensitive restriction enzymes, preferentially cleave unmethylated DNA compared to methylated DNA.
[0224] When contacting a partial sample with a nuclease, one or more nucleases may be used. In some embodiments, the partial sample is contacted with multiple nucleases. The partial sample may be contacted with nucleases sequentially or simultaneously. The simultaneous use of nucleases may be beneficial to avoid unnecessary sample handling when the nucleases are active under similar conditions (e.g., buffer composition). By contacting a second partial sample with one or more methylation-dependent restriction enzymes, nonspecifically distributed highly methylated DNA can be more completely degraded. Similarly, by contacting a first partial sample with one or more methylation-sensitive restriction enzymes, nonspecifically distributed low-methylated and / or unmethylated DNA can be more completely degraded.
[0225] In some embodiments, the methylation-dependent nuclease includes one or more of MspJI, LpnPI, FspEI, or McrBC. In some embodiments, at least two methylation-dependent nucleases are used. In some embodiments, at least three methylation-dependent nucleases are used. In some embodiments, the methylation-dependent nuclease includes FspEI. In some embodiments, the methylation-dependent nuclease includes, for example, FspEI and MspJI, used sequentially.
[0226] In some embodiments, the methylation-sensitive nucleases include one or more of AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, and SnaBI. In some embodiments, at least two methylation-sensitive nucleases are used. In some embodiments, at least three methylation-sensitive nucleases are used. In some embodiments, the methylation-sensitive nucleases include BstUI and HpaII. In some embodiments, the two methylation-sensitive nucleases include HhaI and AccII. In some embodiments, the methylation-sensitive nucleases include BstUI, HpaII, and Hin6I.
[0227] In some embodiments, FspEI is used to digest nucleic acid molecules in at least one subsample (e.g., a low-methylation partition). In some embodiments, BstUI, HpaII, and Hin6I are used to digest nucleic acid molecules in at least one subsample (e.g., a high-methylation partition), and FspEI is used to digest nucleic acid molecules in at least one other subsample (e.g., a low-methylation partition). In embodiments including an intermediate methylation partition, the nucleic acid molecules therein may be digested with a methylation-sensitive nuclease or a methylation-dependent nuclease. In some embodiments, the nucleic acid molecules in the intermediate methylation partition are digested with the same nuclease as the high-methylation partition. For example, the intermediate methylation partition may be pooled together with the high-methylation partition, and then the pooled partition may be subjected to digestion. In some embodiments, the nucleic acid molecules in the intermediate methylation partition are digested with the same nuclease as the low-methylation partition. For example, the intermediate methylation partition may be pooled together with the low-methylation partition, and then the pooled partition may be subjected to digestion.
[0228] In some embodiments, a partial sample is brought into contact with a nuclease as described above after the step of tagging or attaching adapters to both ends of the DNA. The tags or adapters may be resistant to cleavage by nucleases using one of the approaches described above. In this approach, cleavage can prevent nonspecifically distributed molecules from being brought to analysis because the cleavage product lacks tags or adapters at both ends.
[0229] Alternatively, the tagging or adapter attachment step can be performed after nuclease cleavage, as described above. The cleaved molecules can then be identified in the sequence reading data based on the presence of ends (attachment sites to the tag or adapter) corresponding to the nuclease recognition sites. Processing molecules in this manner can also provide information from the observation of cleaved molecules, such as somatic mutations. When analyzing low molecular weight DNA, such as cfDNA, by tagging or adapter attachment after contacting a partial sample with a nuclease, it may be desirable to remove high molecular weight DNA (e.g., contaminating genomic DNA) from the sample before the contact step. Also, since denaturation can interfere with subsequent ligation steps, it may be desirable to avoid DNA denaturation by using nucleases that can be thermally inactivated at relatively low temperatures (e.g., 65°C or below, or 60°C or below).
[0230] When a sample is divided into three partial samples, including a third partial sample containing an intermediate methylation molecule, the third partial sample is, in some embodiments, brought into contact with a methylation-sensitive nuclease. Such a step may have any of the features described elsewhere in this specification relating to the contact step and may be performed before or after the tagging or adapter attachment step, as considered above. In some embodiments, the first and third partial samples are combined before contact with the methylation-sensitive nuclease. Such a step may have any of the features described elsewhere in this specification relating to the contact step and may be performed before or after the tagging or adapter attachment step, as considered above. In some embodiments, the first and third partial samples are differentially tagged before combination.
[0231] Alternatively, if the sample is divided into three partial samples, including a third partial sample containing an intermediate methylated molecule, the third partial sample is, in some embodiments, brought into contact with a methylation-dependent nuclease. Such a step may have any of the features described elsewhere in this specification relating to the contact step and may be performed before or after the tagging or adapter attachment step, as considered above. In some embodiments, the second and third partial samples are combined before contact with the methylation-dependent nuclease. Such a step may have any of the features described elsewhere in this specification relating to the contact step and may be performed before or after the tagging or adapter attachment step, as considered above. In some embodiments, the second and third partial samples are differentially tagged before combination.
[0232] In some embodiments, the DNA is purified after contact with a nuclease, for example, using SPRI beads. Such purification may be performed after thermal inactivation of the nuclease. Alternatively, the purification may be omitted, and so, for example, subsequent steps such as amplification can be performed on the partial sample containing the thermally inactivated nuclease. In another embodiment, the contact step may be performed in the presence of a purification reagent, for example, SPRI beads, so as to minimize loss associated with tube transfer, for example. After cleavage and thermal inactivation, the SPRI beads may be reused for cleanup by adding a molecular crowding reagent (e.g., PEG) and salts. 4. The first partial sample is subjected to a procedure that affects the first nucleic acid base in the DNA of the first partial sample differently from the second nucleic acid base in the DNA.
[0233] The methods disclosed herein may include the step of subjecting a first partial sample to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity (but, for example, the second partial sample is contacted with a methylation-dependent nuclease according to any embodiment described elsewhere herein). In some embodiments, if the first nucleic acid base is modified or unmodified adenine, then the second nucleic acid base is modified or unmodified adenine; if the first nucleic acid base is modified or unmodified cytosine, then the second nucleic acid base is modified or unmodified cytosine; if the first nucleic acid base is modified or unmodified guanine, then the second nucleic acid base is modified or unmodified guanine; and if the first nucleic acid base is modified or unmodified thymine, then the second nucleic acid base is modified or unmodified thymine (modified or unmodified uracil is incorporated within modified thymine for the purposes of this step). Using such a procedure, nucleotides in a partial sample can be identified that have a particular modification, for example, methylation or lack thereof.
[0234] In some embodiments, the first nucleic acid base is modified or unmodified cytosine, and the second nucleic acid base is modified or unmodified cytosine. For example, the first nucleic acid base may include unmodified cytosine (C), and the second nucleic acid base may include one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleic acid base may include C, and the first nucleic acid base may include one or more of mC and hmC. Other combinations are also possible, for example, as shown in the above summary of the invention and the following discussion, such as when one of the first and second nucleic acid bases includes mC and the other includes hmC.
[0235] In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA involves bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Therefore, when bisulfite conversion is used, the first nucleic acid base includes one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleic acid base may include one or more of mC and hmC, e.g., mC and optionally hmC. Sequencing of bisulfite-treated DNA identifies the position read as cytosine as an mC or hmC position. On the other hand, positions read as T are identified as bisulfite-sensitive forms of T or C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Performing bisulfite conversion on the first partial sample as described herein thus facilitates the identification of mC or hmC-containing positions using sequence reading data obtained from the first partial sample. For an illustrative description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068.
[0236] In some embodiments, a procedure that affects a first nucleic acid base in a first partial sample DNA differently from a second nucleic acid base in the DNA involves oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to the bisulfite-sensitive fC, and then performs the bisulfite conversion. Thus, when oxidative bisulfite conversion is used, the first nucleic acid base includes one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms affected by bisulfite, and the second nucleic acid base includes mC. Sequencing of DNA converted with Ox-BS identifies positions read as cytosine as mC positions. On the other hand, positions read as T are identified as T, hmC, or bisulfite-sensitive forms of C, e.g., unmodified cytosine, fC, or hmC. Performing the Ox-BS conversion on the first partial sample as described herein facilitates the identification of mC-containing locations using sequence reading data obtained from the first partial sample. For an illustrative description of oxidative bisulfite conversion, see, for example, Booth et al., Science. See 2012; 336: 934-937.
[0237] In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA involves Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion, and mC is oxidized prior to bisulfite treatment, resulting in the position originally occupied by mC being converted to U, while the position originally occupied by hmC remains as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, hmC can be protected using β-glucosyltransferase (forming 5-glucosylhydroxymethylcytosine (ghmC)), then mC can be converted to caC using a TET protein, e.g., mTet1, and then C and caC can be converted to U using bisulfite treatment, while ghmC remains unaffected. Therefore, when TAB conversion is used, the first nucleic acid base contains one or more of the following: unmodified cytosine, fC, caC, mC, or other cytosine forms affected by bisulfites, and the second nucleic acid base contains hmC. Sequencing of TAB-converted DNA identifies positions read as cytosine as hmC positions. On the other hand, positions read as T are identified as bisulfite-sensitive forms of T, mC, or C, e.g., unmodified cytosine, fC, or caC. Performing TAB conversion on a first partial sample as described herein thus facilitates the identification of hmC-containing positions using sequence reading data obtained from the first partial sample.
[0238] In some embodiments, a procedure that affects the first nucleic acid base in the DNA of a first partial sample differently from the second nucleic acid base in the DNA includes Tet-assisted conversion with a substituted borane reducing agent, which, if necessary, is 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane. In Tet-assisted pic-borane conversion with a substituted borane reducing agent, the TET protein is used to convert mC and hmC to caC without affecting the unmodified C. caC and fC, if present, are then converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or other substituted borane reducing agents, e.g., borampyridine, tert-butylamine borane, or ammonia borane, also without affecting the unmodified C. See, for example, Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., Supplementary Figure 1 and Supplementary Explanation 7). DHU is sequenced In the sequence, it is read as T. Therefore, when this type of conversion is used, the first nucleic acid base contains one or more of mC, fC, caC, or hmC, and the second nucleic acid base contains unmodified cytosine. Sequencing of the converted DNA identifies the position read as cytosine as the unmodified C position. On the other hand, the position read as T is identified as T, mC, fC, caC, or hmC. Performing the TAP conversion on the first partial sample as described herein thus facilitates the identification of the unmodified C-containing position using the sequence reading data obtained from the first partial sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), which is described in more detail in Liu et al. 2019 (above).
[0239] Alternatively, the protection of hmC (e.g., using βGT) can be combined with Tet-assisted conversion using a substituted borane reducing agent. hmC can be protected as described above through glucosylation using βGT to form ghmC. Treatment with a TET protein, e.g., mTet1, then converts mC to caC, but not C or ghmC. caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent, e.g., borampyridine, tert-butylamine borane, or ammonia borane, also without affecting unmodified C or ghmC. Thus, when Tet-assisted conversion using a substituted borane reducing agent is used, the first nucleic acid base contains mC, and the second nucleic acid base contains one or more of unmodified cytosine or hmC, e.g., unmodified cytosine, and optionally hmC, fC, and / or caC. Sequencing of the converted DNA identifies locations read as cytosine as hmC or unmodified C locations, while locations read as T are identified as T, fC, caC, or mC. Performing TAPSβ conversion on the first partial sample as described herein thus facilitates the identification of locations containing unmodified C or hmC as locations containing mC using sequence reading data obtained from the first partial sample. For an illustrative description of this type of conversion, see, for example, Liu et al., Nature Biotechnology 2019; 37:424-429.
[0240] In some embodiments, the procedure for affecting a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes a chemically assisted conversion using a substituted borane reducing agent, which may optionally be 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane. In the chemically assisted conversion using a substituted borane reducing agent, an oxidizing agent, such as potassium perruthenate (KRuO4) (also suitable for use in ox-BS conversion), is used to specifically oxidize hmC to fC. Treatment with pic-borane or other substituted borane reducing agents, such as borampyridine, tert-butylamine borane, or ammonia borane, converts fC and caC to DHU but does not affect mC or unmodified C. Therefore, when this type of conversion is used, the first nucleic acid base contains one or more of hmC, fC, and caC, and the second nucleic acid base contains one or more of unmodified cytosine or mC, e.g., unmodified cytosine and optionally mC. Sequencing of the converted DNA identifies positions read as cytosine as mC or unmodified C positions. On the other hand, positions read as T are identified as T, fC, caC, or hmC. Performing this type of conversion on a first partial sample as described herein therefore facilitates the identification of positions containing unmodified C or mC as positions containing hmC using sequence reading data obtained from the first partial sample. For an illustrative description of this type of conversion, see, for example, Liu et al., Nature Biotechnology 2019; 37:424-429.
[0241] In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA involves APOBEC coupling epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme, e.g., APOBEC3A (A3A), is used to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when ACE conversion is used, the first nucleic acid base contains unmodified C and / or mC (e.g., unmodified C and mC as appropriate), and the second nucleic acid base contains hmC. Sequencing of the ACE-converted DNA identifies positions read as cytosine as hmC, fC, or caC positions. On the other hand, positions read as T are identified as T, unmodified C, or mC. Performing ACE conversion on the first partial sample as described herein thus facilitates the identification of hmC-containing locations from mC- or unmodified C-containing locations using sequence reading data obtained from the first partial sample. For an illustrative description of ACE conversion, see, for example, Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0242] In some embodiments, the procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes, for example, the enzymatic conversion of the first nucleic acid base, as found in EM-Seq. For example, see Vaisvila R, et al. (2019) EM-seq: Detection, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. of DNA methylation at single base resolution from picograms of DNA. See bioRxiv; DOI: 10.1101 / 2019.12.20.884692. For example, TET2 and Using T4-βGT, 5mC and 5hmC can be converted into substrates that cannot be deaminated by deaminase (e.g., APOBEC3A), and then the unmodified cytosine can be deaminated and converted to uracil using deaminase (e.g., APOBEC3A).
[0243] In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes separating the DNA that originally contained the first nucleic acid base from the DNA that did not originally contain the first nucleic acid base. In some such embodiments, the first nucleic acid base is hmC. The DNA that originally contained the first nucleic acid base can be separated from the other DNA using a labeling procedure that includes biotinylating the position that originally contained the first nucleic acid base. In some embodiments, the first nucleic acid base is first derivatized with an azide-containing moiety, for example, a glucosyl-azide-containing moiety. The azide-containing moiety can then function as a reagent for biotin deposition, for example, through a hysgen cycloaddition reaction. Next, the DNA originally containing the first nucleic acid base, which has been biotinylated, can be separated from the DNA originally not containing the first nucleic acid base using a biotin binder, such as avidin, neutraavidin (a deglycosylated avidin with an isoelectric point of approximately 6.3), or streptavidin. An example of a procedure for separating DNA originally containing the first nucleic acid base from DNA originally not containing the first nucleic acid base is the hmC seal, which involves labeling hmC to form β-6-azido-glucosyl-5-hydroxymethylcytosine, then attaching the biotin moiety via hysgen cycloaddition, and subsequently separating the biotinylated DNA from other DNA using a biotin binder. For an illustrative description of the hmC seal, see, for example, Han et al. See Mol. Cell 2016; 63: 711-719. This approach is useful for identifying fragments containing one or more hmC nucleic acid bases.
[0244] In some embodiments, following such separation, the method further includes the step of differentially tagging each of the DNAs originally containing the first nucleic acid base from the DNA originally not containing the first nucleic acid base and the DNA of the second partial sample. After differential tagging, the method further includes the step of pooling the DNA originally containing the first nucleic acid base, the DNA originally not containing the first nucleic acid base, and the DNA of the second partial sample. The DNA originally containing the first nucleic acid base, the DNA originally not containing the first nucleic acid base, and the DNA of the second partial sample can then be sequenced in the same sequencing cell, but the differential tagging is used to retain the ability to determine whether a given read data originates from a molecule of the DNA originally containing the first nucleic acid base, the DNA originally not containing the first nucleic acid base, or the DNA of the second partial sample.
[0245] In some embodiments, the first nucleic acid base is modified or unmodified adenine, and the second nucleic acid base is modified or unmodified adenine. In some embodiments, the modified adenine is N 6 -methyl adenine (mA). In some embodiments, the modified adenine is N 6 -Methyladenine (mA), N 6 -Hydroxymethyladenine (hmA), or N 6 - One or more of the formyladenine (fA) compounds.
[0246] Using techniques including methylated DNA immunoprecipitation (MeDIP), DNA containing modified bases, such as mA, can be isolated from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific to mA are described in Sun et al., Bioessays 2015; 37:1155-62. Various modified nucleic acid bases, such as halogenated forms, e.g., 5-br Antibodies against thymine / uracil forms, including romouracil, are commercially available. Various modified bases can also be detected based on changes in their base-pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read as G in sequencing. See, for example, U.S. Patent No. 8,486,630, Brown, Genomes, 2 nd See Ed., John Wiley & Sons, Inc., New York, NY, 2002, chapter 14, "Mutation, Repair, and Recombination". 5. Enrichment / Capture Step; Amplification, Adapter, Barcode
[0247] In some embodiments, the methods disclosed herein include the step of capturing one or more sets of target regions of DNA, for example, cfDNA. Capture may be carried out using any preferred approach known in the art.
[0248] In some embodiments, the capturing step comprises contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes can have any of the features described herein with respect to sets of target-specific probes, including but not limited to those in the embodiments described above and in the sections related to the probes below. The capturing step can be performed on one or more sub-samples prepared during the methods disclosed herein. In some embodiments, the DNA is captured from at least a first sub-sample or a second sub-sample, e.g., from at least a first sub-sample and a second sub-sample. If the first sub-sample undergoes a separation step (e.g., separating DNA that originally contains a first nucleobase (e.g., hmC) from DNA that originally does not contain the first nucleobase, e.g., hmC-seal), the capturing step can be performed on any one, any two, or all of the DNA that originally contains a first nucleobase (e.g., hmC), the DNA that originally does not contain the first nucleobase, and the second sub-sample. In some embodiments, the sub-samples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.
[0249] The capturing step can generally be performed using conditions suitable for specific nucleic acid hybridization, which depends to some extent on the features of the probes, such as length, base composition, etc. Those skilled in the art are familiar with appropriate conditions, taking into account general knowledge in the art regarding nucleic acid hybridization. In some embodiments, a complex of target-specific probes and DNA is formed.
[0250] In some embodiments, the methods described herein include capturing cfDNA obtained from a subject with respect to a set of multiple target regions. The target regions include epigenetic target regions, which may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from tumors or healthy cells. The target regions also include sequence-variable target regions, which may exhibit differences in sequence depending on whether they originate from tumors or healthy cells. The capturing step produces a captured set of cfDNA molecules, and cfDNA molecules corresponding to the set of sequence-variable target regions are captured at a higher capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the set of epigenetic target regions. For further discussion of the capturing step, capture yield, and related aspects, see WO2020 / 160414, which is incorporated herein by reference for all purposes.
[0251] In some embodiments, the methods described herein include contacting cfDNA obtained from a subject with a set of target-specific probes, the set of target-specific probes being configured to capture cfDNA corresponding to the set of sequence-variable target regions at a higher capture yield than cfDNA corresponding to the set of epigenetic target regions.
[0252] To analyze sequence-variable target regions with sufficient confidence or accuracy, sequencing to a greater depth may be required than may be necessary to analyze epigenetic target regions. Therefore, capturing cfDNA corresponding to a set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to a set of epigenetic target regions may be beneficial. The amount of data required to determine fragmentation patterns (e.g., to test for disruption of transcription start sites or CTCF binding sites) or the abundance of fragments (e.g., in high-methylation and low-methylation distributions) is generally less than the amount of data required to determine the presence or absence of cancer-associated sequence mutations. Capturing target region sets with different yields may facilitate sequencing target regions to different depths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).
[0253] In various embodiments, the method further includes the step of sequencing the captured cfDNA to varying degrees of sequencing depth with respect to, for example, epigenetic and sequence-variable target region sets, consistent with the discussion herein.
[0254] In some embodiments, the target-specific probe-DNA complex is separated from the DNA that is not bound to the target-specific probe. For example, if the target-specific probe is bound to a solid support by covalent or non-covalent bonds, the unbound material can be separated using a washing or aspiration step. Alternatively, if the complex has chromatographic properties distinct from the unbound material (for example, if the probe contains a ligand that binds to a chromatography resin), chromatography can be used.
[0255] As will be discussed in detail elsewhere herein, a set of target-specific probes may include multiple sets, such as probes for sequence-variable target regions and probes for epigenetic target regions. In some such embodiments, the capture step is performed simultaneously for the probes for sequence-variable target regions and the probes for epigenetic target regions in the same container, for example, the probes for sequence-variable target regions and the probes for epigenetic target regions are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for sequence-variable target regions is higher than the concentration of the probes for epigenetic target regions.
[0256] Alternatively, the capture step may be performed for the sequence variable target region probe set in a first container and for the epigenetic target region probe set in a second container, or contact may be performed for the sequence variable target region probe set in the first time and in the first container, and for the epigenetic target region probe set in the second time, either before or after the first time. This approach allows for the preparation of separate first and second compositions containing captured DNA corresponding to the sequence variable target region set and captured DNA corresponding to the epigenetic target region set. The compositions can be processed individually as desired (e.g., to fractionate based on methylation as described elsewhere herein) and recombined in appropriate proportions to provide material for further processing and analysis such as sequencing.
[0257] In some embodiments, the DNA is amplified. In some embodiments, the amplification is performed before the capture step. In some embodiments, the amplification is performed after the capture step.
[0258] In some embodiments, the adapter is contained within the DNA. This can be done concurrently with the amplification procedure, for example, by providing the adapter to the 5' portion of the primer, as described above. Alternatively, the adapter may be added by other approaches, such as ligation.
[0259] In some embodiments, a tag, which may be or may contain a barcode, is included in the DNA. The tag can facilitate the identification of the nucleic acid's origin. For example, after pooling multiple samples for parallel sequencing, the barcode may be used to identify the origin (e.g., target) from which the DNA originates. This can be done concurrently with the amplification procedure, for example, by providing the barcode on the 5' portion of the primer, as described above. In some embodiments, the adapter and tag / barcode are provided by the same primer or primer set. For example, the barcode may be positioned at the 3' of the adapter and the 5' of the portion that hybridizes to the target of the primer. Alternatively, the barcode may be added together with the adapter on the same ligation substrate as needed, by other approaches, such as ligation.
[0260] Further details regarding amplification, tagging, and barcodes are discussed in the following section, “General Features of the Method,” which can be combined to a practical degree with any of the embodiments described above as well as those described in the Introduction and Summary sections. 6. Captured Set
[0261] In some embodiments, a captured set of DNA (e.g., cfDNA) is provided. With respect to the disclosed method, the captured set of DNA may be provided, for example, by performing a capture step after a distribution step as described herein. The captured set may include DNA corresponding to a sequence variable target region set, an epigenetic target region set, or a combination thereof.
[0262] In some embodiments, the first set of target regions is captured from a first partial sample containing at least epigenetic target regions. The epigenetic target regions captured from the first partial sample may include hypermethylated variable target regions. In some embodiments, the hypermethylated variable target regions are CpG-containing regions in cfDNA derived from a healthy subject that are unmethylated or have low methylation (e.g., below-average methylation compared to bulk cfDNA). In some embodiments, the hypermethylated variable target regions are regions in healthy cfDNA that exhibit lower methylation than those in at least one other tissue type. While we do not wish to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the tissue distribution of cfDNA origin may change during carcinogenesis. Thus, an increase in the level of hypermethylated variable target regions in the first partial sample may indicate the presence of cancer (or recurrence, depending on the subject's history).
[0263] In some embodiments, a second set of target regions is captured from a second partial sample containing at least an epigenetic target region. The epigenetic target region may include a hypomethylated variable target region. In some embodiments, the hypomethylated variable target region is a CpG-containing region that is methylated or has high methylation (e.g., above-average methylation compared to bulk cfDNA) in cfDNA derived from a healthy subject. In some embodiments, the hypomethylated variable target region is a region in healthy cfDNA that exhibits higher methylation than in at least one other tissue type. While we do not wish to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the tissue distribution of cfDNA origin may change during carcinogenesis. Thus, an increase in the level of hypomethylated variable target regions in the second partial sample may indicate the presence of cancer (or recurrence, depending on the subject's history).
[0264] In some embodiments, the quantity of captured sequence-variable target region DNA is greater than the quantity of captured epigenetic target region DNA, when normalized with respect to the difference in the size (footprint size) of the targeted region.
[0265] Alternatively, a first and second captured set may be provided, each containing DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set. The first and second captured sets may be combined to provide a combined captured set.
[0266] In some embodiments, where the captured set includes DNA corresponding to the sequence variable target region set and the epigenetic target region set, the DNA corresponding to the sequence variable target region set is present in a higher concentration than the DNA corresponding to the epigenetic target region set, for example, 1.1 to 1.2 times higher, 1.2 to 1.4 times higher, 1.4 to 1.6 times higher, 1.6 to 1.8 times higher, 1.8 to 2.0 times higher, 2.0 to 2.2 times higher, 2.2 to 2.4 times higher, 2.4 to 2.6 times higher, 2.6 to 2.8 times higher, 2.8 to 3.0 times higher, 3.0 to 3.5 times higher, 3.5 to 4.0, 4.0 to 4.5 times higher, 4.5 to 5.0 times higher, 5.0 to 5.5 times higher, 5 0.5 to 6.0 times higher concentration, 6.0 to 6.5 times higher concentration, 6.5 to 7.0 times higher, 7.0 to 7.5 times higher concentration, 7.5 to 8.0 times higher concentration, 8.0 to 8.5 times higher concentration, 8.5 to 9.0 times higher concentration, 9.0 to 9.5 times higher concentration, 9.5 to 10.0 times higher concentration, 10 to 11 times higher concentration, 11 to 12 times higher concentration, 12 to 13 times higher concentration, 13 to 14 times higher concentration, 14 times The concentration can be approximately 15 times higher, 15 to 16 times higher, 16 to 17 times higher, 17 to 18 times higher, 18 to 19 times higher, 19 to 20 times higher, 20 to 30 times higher, 30 to 40 times higher, 40 to 50 times higher, 50 to 60 times higher, 60 to 70 times higher, 70 to 80 times higher, 80 to 90 times higher, or 90 to 100 times higher. The degree of the concentration difference explains the normalization with respect to the footprint size of the target area, as discussed in the definition section. a. Epigenetic target region set
[0267] The epigenetic target region set may include one or more types of target regions that are likely to distinguish DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The epigenetic target region set may also include one or more control regions, as described herein, for example.
[0268] In some embodiments, the epigenetic target region set has a footprint of at least 100 kbp, for example, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the epigenetic target region set has a footprint in the range of 100 to 20 Mbp, for example, 100 to 200 kbp, 200 to 300 kbp, 300 to 400 kbp, 400 to 500 kbp, 500 to 600 kbp, 600 to 700 kbp, 700 to 800 kbp, 800 to 900 kbp, 900 to 1,000 kbp, 1 to 1.5 Mbp, 1.5 to 2 Mbp, 2 to 3 Mbp, 3 to 4 Mbp, 4 to 5 Mbp, 5 to 6 Mbp, 6 to 7 Mbp, 7 to 8 Mbp, 8 to 9 Mbp, 9 to 10 Mbp, or 10 to 20 Mbp. In some embodiments, the epigenetic target region set has a footprint of at least 20 Mbp. i. Highly methylated variable target regions
[0269] In some embodiments, the epigenetic target region set includes one or more hypermethylation variable target regions. Generally, a hypermethylation variable target region refers to a region where, for example, an increase in the level of methylation observed in a cfDNA sample indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by neoplastic cells, such as tumor cells or cancer cells. For example, hypermethylation of the promoter of a tumor suppressor gene has been repeatedly observed. See, for example, Kang et al., Genome Biol. 18:53 (2017) and the references cited therein. In another example, as discussed above, a hypermethylation variable target region may include a region that is not necessarily different in methylation compared to DNA from the same type of healthy tissue in a cancerous tissue, but is different in methylation (e.g., has more methylation) compared to typical cfDNA in a healthy subject. For example, if the presence of cancer results in an increase in cell death, such as apoptosis of cells of the tissue type corresponding to the cancer, such cancer can be detected, at least in part, using such hypermethylation variable target regions.
[0270] An extensive discussion of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866:106-20 (2016). These include VIM , SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylation variable target regions based on colorectal cancer (CRC) studies is provided in Table 1. Many of these genes are likely to be relevant to cancers other than colorectal cancer. For example, TP53 is widely recognized as a very important tumor suppressor, and inactivation based on hypermethylation of this gene can be a common carcinogenic mechanism.
Table 1
[0271] In some embodiments, the highly methylated variable target region comprises multiple loci listed in Table 1, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. For example, with respect to each locus included as a target region, there may be one or more probes having a hybridization site that binds between the transcription start site and the stop codon (the last stop codon of a gene that is alternatively spliced) of the gene, or to the promoter region of the gene. In some embodiments, one or more probes bind within 300 bp, for example, 200 or 100 bp, of the transcription start site of the gene in Table 1.
[0272] Methylation-variable target regions in various types of lung cancer are discussed in, for example, Ooki et al., Clin. Cancer Res. 23:7141-52 (2017), Belinksy, Annu. Rev. Physiol. 77:453-74 (2015), Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017), Shi et al., BMC Genomics 18:901 (2017), Schneider et al., BMC Cancer. 11:102 (2011), Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016), Skvortsova et al., Br. J. Cancer. 94(10):1492-1495 (2006), Kim et al., Cancer Res. 61:3419-3424 (2001), Furonaka et al., Pathology International 55:303-309 (2005), Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014), Kim et al., Oncogene. 20:1765-70 (2001), Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003), Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005), Heller et al., Oncogene 25:959-968 (2006), Licchesi et al., Carcinogenesis. 29:895-904 (2008), Guo et al., Clin. Cancer Res. This topic is discussed in detail in 10:7917-24 (2004), Palmisano et al., Cancer Res. 63:4620-4625 (2003), and Toyooka et al., Cancer Res. 61:4556-4560 (2001).
[0273] An exemplary set of hypermethylated variable target regions based on lung cancer research is provided in Table 2. Many of these genes are likely to be relevant to cancers other than lung cancer; for example, Casp8 (caspase 8) is a key enzyme in programmed cell death, and inactivation based on hypermethylation of this gene may be a general carcinogenic mechanism not limited to lung cancer. Furthermore, several genes appear in both Tables 1 and 2, demonstrating their generality. [Table 2-1] [Table 2-2]
[0274] Any of the embodiments described above relating to the target regions identified in Table 2 may be combined with any of the embodiments described above relating to the target regions identified in Table 1. In some embodiments, the highly methylated variable target region comprises a plurality of loci listed in Table 1 or Table 2, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2.
[0275] Further hypermethylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a stochastic method called CancerLocator using hypermethylated target regions derived from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions may be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0276] In some embodiments, when different epigenetic target regions are captured from the first and second partial samples, the epigenetic target regions captured from the first partial sample include highly methylated variable target regions. ii. Low-methylation variable target regions
[0277] Overall hypomethylation is a phenomenon commonly observed in various cancers. For example, see Hon et al., Genome Res. 22:246-258 (2012) (breast cancer), Ehrlich, Epigenomics. 1:239-259 (2009) (Colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and See review articles that mention the observation of hypomethylation in cervical cancer. For example, regions such as repeating elements, e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA, as well as intergenetic regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Therefore, in some embodiments, the epigenetic target region set may include hypomethylated variable target regions in which the observed decrease in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by neoplastic cells, e.g., tumor cells or cancer cells. In another example, as discussed above, hypomethylated variable target regions may include regions in cancerous tissue that are methylated differently (e.g., less unmethylated) compared to typical cfDNA in healthy subjects, but not necessarily differently methylated compared to DNA from the same type of healthy tissue. For example, if the presence of cancer leads to increased cell death, such as increased apoptosis of cells in tissue types corresponding to cancer, then such cancer can be detected, at least partially, using such hypomethylated variable target regions.
[0278] In some embodiments, the low-methylation variable target region includes repeat elements and / or intergenetic regions. In some embodiments, the repeat elements include one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and / or satellite DNA.
[0279] Exemplary specific genomic regions that exhibit cancer-related hypomethylation include nucleotides 8403565–8953708 and 151104701–151106035 on human chromosome 1. In some embodiments, the hypomethylation variable target region overlaps with or includes one or both of these regions.
[0280] In some embodiments, when different epigenetic target regions are captured from the first and second partial samples, the epigenetic target region captured from the second partial sample includes a low-methylation variable target region. iii.CTCF binding region
[0281] CTCF is a DNA-binding protein that contributes to chromatin structure and often co-localizes with cohesin. Disruptions of CTCF binding sites have been reported in various different cancers. For example, see Katainen et al., Nature Genetics, doi:10.1038 / ng.3335, published online 8 June 2015, and Guo et al., Nat. Commun. 9:1520 (2018). Please refer to the following. CTCF binding results in a recognizable pattern in cfDNA that can be detected by sequencing, for example, through fragment length analysis. For details on sequencing-based fragment length analysis, see Snyder et al., Cell 164:57-68 (2016). These are provided in WO2018 / 009723 and US20170211143A1, each of which is incorporated herein by reference.
[0282] Therefore, disruption of CTCF binding leads to variations in the fragmentation pattern of cfDNA. Thus, the CTCF binding site represents one type of variable fragmentation target region.
[0283] Numerous known CTCF binding sites exist. For example, each of them is incorporated by reference in the CTCFBSDB (CTCF Binding Site Database), which is available on the internet at insulatordb.uthsc.edu / , and Cuddapah et al., Genome Res. 19:24-32. (2009), Martin et al., Nat. Struct. Mol. Biol. 18:708-14 (2011), Rhee See et al., Cell. 147:1408-19 (2011). Exemplary CTCF binding sites are nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13.
[0284] Therefore, in some embodiments, the epigenetic target region set includes CTCF-binding regions. In some embodiments, the CTCF-binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF-binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF-binding regions, such as those described above or listed in CTCFBSDB or in one or more of the papers by Cuddapah et al., Martin et al., or Rhee et al.
[0285] In some embodiments, at least some of the CTCF sites may be methylated or unmethylated, where the methylation status correlates with whether the cell is cancerous or not. In some embodiments, the epigenetic target region set includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, and at least 1000 bp upstream and downstream of the CTCF binding site. iv. Transcription initiation site
[0286] Transcription start sites can also exhibit disruption in neoplastic cells. For example, the nucleosome composition at various transcription start sites in healthy hematopoietic cells (which substantially contributes to cfDNA in healthy individuals) may differ in those transcription start sites in neoplastic cells. This is documented in Snyder et al., Cell 164:57-68 (2016), WO2018 / 009723, and US20170211143A1. As is generally considered, this results in different cfDNA patterns that can be detected by sequencing. In another example, transcription start sites in cancerous tissue are not necessarily epigenetically different from DNA derived from healthy tissue of the same type, but they are epigenetically different from cfDNA that is typical in healthy subjects (e.g., with respect to nucleosome composition). For example, if the presence of cancer results in increased cell death, e.g., apoptosis of cells in the tissue type corresponding to cancer, then such cancer can be detected, at least partially, using such transcription start sites.
[0287] Therefore, disruption of the transcription start site also leads to variations in the fragmentation pattern of cfDNA. Thus, the transcription start site also represents one type of fragmentation-variable target region.
[0288] Human transcription start sites are available via DBTSS, which can be accessed online at btss.hgc.jp. Available from the (Database of Human Transcription Start Sites) and described in Yamashita et al., Nucleic Acids Res. 34 (Database issue): D86-D89 (2006), which is incorporated herein by reference.
[0289] Therefore, in some embodiments, the epigenetic target region set includes transcription start sites. In some embodiments, the transcription start sites include at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those listed in DBTSS. In some embodiments, at least some of the transcription start sites may be methylated or unmethylated, where the methylation status correlates with whether the cell is a cancer cell or not. In some embodiments, the epigenetic target region set includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, and at least 1000 bp upstream and downstream of the transcription start site. v. Local amplification
[0290] Local amplifications are somatic mutations, but they can be detected by sequencing based on the frequency of read data in a manner similar to approaches for detecting certain epigenetic changes, such as changes in methylation. Therefore, regions that may exhibit local amplification in cancer may be included in a set of epigenetic target regions, which may include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, the set of epigenetic target regions includes at least two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, or eighteen of the aforementioned targets. vi. Methylation control region
[0291] Including control regions can be useful to facilitate data validation. In some embodiments, the epigenetic target region set includes a control region that is expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA originates from cancer cells or normal cells. In some embodiments, the epigenetic target region set includes a control hypomethylated region that is expected to be hypomethylated in essentially all samples. In some embodiments, the epigenetic target region set includes a control hypermethylated region that is expected to be hypermethylated in essentially all samples. b. Variable array target region set
[0292] In some embodiments, the sequence-variable target region set includes multiple regions known to undergo somatic mutation in cancer.
[0293] In some embodiments, the sequence-variable target region set targets several different genes or genomic regions ("Panel") selected such that a determined proportion of subjects with cancer exhibit genetic variants or tumor markers in one or more different genes or genomic regions within the Panel. The Panel may be selected to limit the region for sequencing to a fixed number of base pairs. The Panel may be selected to sequence a desired amount of DNA, for example, by adjusting the affinity and / or amount of the probe, as described elsewhere in this Spec. The Panel may further be selected to achieve a desired sequence read data depth. The Panel may be selected to achieve a desired sequence read data depth or sequence read data coverage for the amount of base pairs being sequenced. The Panel may be selected to achieve a theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for detecting one or more genetic variants in a sample.
[0294] Probes for detecting panels of regions include those for detecting target genomic regions (hotspot regions), as well as nucleosome recognition probes (e.g., KRAS codons 12 and 13), which may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations influenced by nucleosome binding patterns and GC sequence composition. Regions used herein may also include non-hotspot regions optimized based on nucleosome location and GC model.
[0295] Examples of lists of target genome locations can be found in Tables 3 and 4. In some embodiments, the sequence-variable target region set used in the methods of the present disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the sequence-variable target region set used in the methods of the present disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the sequence-variable target region set used in the methods of the present disclosure includes at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 3. In some embodiments, the sequence-variable target region set used in the methods of the present disclosure includes at least one, at least two, or at least a portion of three of the indels in Table 3. In some embodiments, the sequence-variable target region set used in the methods of the present disclosure includes at least five, at least ten, at least fifteen, at least twenty, at least twenty-five, at least thirty, at least thirty-five, at least forty, at least forty-five, at least fifty, at least fifty-five, at least sixty, at least sixty-sixty, at least seventy, or at least seventy-sixty of the genes in Table 4. In some embodiments, the sequence-variable target region set used in the method of the present disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4.In some embodiments, the sequence-variable target region set used in the methods of the present disclosure includes at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 4. In some embodiments, the sequence-variable target region set used in the methods of the present disclosure includes at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, or eighteen of the indels in Table 4. Each of these target genome locations may be identified as a skeletal region or hotspot region of a given panel. An example list of target hotspot genome locations can be found in Table 5. In some embodiments, the set of sequence-variable target regions used in the methods of this disclosure includes at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, at least eighteen, at least nineteen, or at least twenty of the genes listed in Table 5. Each hotspot genomic region is enumerated with several features, including the associated gene, the chromosome in which it resides, the genomic start and end positions representing the gene locus, the length of the gene locus in base pairs, the exons covered by the gene, and important features that a given genomic region of interest may attempt to capture (e.g., the type of mutation). [Table 3] [Table 4-1] [Table 4-2] [Table 5-1] [Table 5-2] [Table 5-3]
[0296] Furthermore, or alternatively, suitable sets of target regions are available from the literature. For example, Gale et al., PLoS One 13: e0194630 (2018), incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of a set of sequence-variable target regions. These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0297] In some embodiments, the sequence-variable target region set includes target regions from at least 10, 20, 30, or 35 cancer-related genes, such as the cancer-related genes listed above.
[0298] In some embodiments, the sequence-variable target region set has a footprint of at least 50 kbp, for example, at least 100 kbp, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the sequence-variable target region set has a footprint in the range of 100 to 2000 kbp, for example, 100 to 200 kbp, 200 to 300 kbp, 300 to 400 kbp, 400 to 500 kbp, 500 to 600 kbp, 600 to 700 kbp, 700 to 800 kbp, 800 to 900 kbp, 900 to 1,000 kbp, 1 to 1.5 Mbp, or 1.5 to 2 Mbp. In some embodiments, the sequence-variable target region set has a footprint of at least 2 Mbp. 7. Target
[0299] In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having a neoplasm. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having a neoplasm. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject who has achieved remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the embodiments described above, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, may be of the lung. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the breast. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the prostate. In any of the embodiments described above, the subject may be a human subject. 8.Quantitative
[0300] In some embodiments, epigenetic target regions captured from one or more of the first partial sample, the treated first partial sample, or the treated second partial sample are quantified. For example, a low-methylation variable target region may be quantified in the treated second partial sample, and / or a high-methylation variable target region may be quantified in the first partial sample or the treated first partial sample. Quantification may be performed by any suitable technique, e.g., quantitative amplification, e.g., quantitative PCR. In some embodiments, quantification is based on sequencing data (e.g., the number of sequencing reads or the number of unique molecules being sequenced).
[0301] The quantification of epigenetic target regions as discussed above can be used to determine the presence, absence, or likelihood of cancer in a subject. For example, the determination of the presence or absence of cancer may be based at least in part on whether the amount of hypermethylated variable target regions in a first partial sample or a treated first partial sample, and / or the amount of hypomethylated variable target regions in a treated second partial sample, exceeds a predetermined threshold. In some embodiments, such amounts can be used together with other data collected from the sample, e.g., the presence of mutations and / or other epigenetic features described elsewhere in this specification, e.g., disruption of transcription start sites and / or CTCF binding sites. 9. A pool of DNA derived from the first and second partial samples or parts thereof.
[0302] In some embodiments, the method includes the step of preparing a pool containing at least a portion of the DNA from a second partial sample (also referred to as a low-methylation partition) and at least a portion of the DNA from a first partial sample (also referred to as a high-methylation partition). Target regions, for example, including epigenetic target regions and / or sequence-variable target regions, may be captured from the pool. The step of capturing a set of target regions from at least a portion of the partial samples, as described elsewhere in this specification, encompasses a capture step performed on a pool containing DNA derived from the first and second partial samples. The step of amplifying the DNA in the pool may be performed before capturing the target regions from the pool. The capture step may have any of the features described elsewhere in this specification.
[0303] Epigenetic target regions may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from tumors or healthy cells, or from which tissue type they originate, as discussed elsewhere in this specification. Sequence-variable target regions may exhibit differences in sequence depending on whether they originate from tumors or healthy cells.
[0304] Analysis of epigenetic target regions derived from low methylation distributions may, in some applications, provide less information than analysis of sequence-variable target regions derived from high and low methylation distributions, as well as epigenetic target regions derived from high methylation distributions. Therefore, in methods for capturing sequence-variable target regions and epigenetic target regions, the latter may be captured to a lower degree than one or more of the sequence-variable target regions derived from high and low methylation distributions, as well as epigenetic target regions derived from high methylation distributions. For example, sequence-variable target regions may be captured from portions of low methylation distributions that are not pooled with high methylation distributions, and the pool may be prepared using a portion (e.g., most, substantially all, or all) of DNA derived from high methylation distributions and no or a portion (e.g., a small amount) of DNA derived from low methylation distributions. Such an approach can reduce or eliminate sequencing of epigenetic target regions derived from low methylation distributions, thereby reducing the amount of sequencing data sufficient for further analysis.
[0305] In some embodiments, including a small portion of the low-methylated DNA in the pool makes it easier, for example, relatively easier to quantify one or more epigenetic features (e.g., methylation or other epigenetic features discussed elsewhere herein).
[0306] In some embodiments, the pool includes, for example, less than about 50% of the low-methylated partition DNA, for example, less than or equal to about 45%, less than or equal to about 40%, less than or equal to 35%, less than or equal to 30%, less than or equal to 25%, less than or equal to 20%, less than or equal to 15%, less than or equal to 10%, or less than or equal to 5%. In some embodiments, the pool includes about 5% to 25% of the low-methylated partition DNA. In some embodiments, the pool includes about 10% to 20% of the low-methylated partition DNA. In some embodiments, the pool includes about 10% of the low-methylated partition DNA. In some embodiments, the pool includes about 15% of the low-methylated partition DNA. In some embodiments, the pool includes about 20% of the low-methylated partition DNA.
[0307] In some embodiments, the pool contains a portion of the highly methylated partition, which may be at least about 50% of the DNA in the highly methylated partition. For example, the pool may contain at least about 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the DNA in the highly methylated partition. In some embodiments, the pool contains 50-55%, 55-60%, 60-65%, 65-70%, 70-75%, 75-80%, 80-85%, 85-90%, 90-95%, or 95-100% of the DNA in the highly methylated partition. In some embodiments, the second pool contains all or substantially all of the highly methylated partition.
[0308] In some embodiments, the method includes the step of preparing a first pool containing at least a portion of the low-methylated partition DNA. In some embodiments, the method includes the step of preparing a second pool containing at least a portion of the high-methylated partition DNA. In some embodiments, the first pool further contains a portion of the high-methylated partition DNA. In some embodiments, the second pool further contains a portion of the low-methylated partition DNA. In some embodiments, the first pool contains the majority of the low-methylated partition DNA and optionally a minority of the high-methylated partition DNA. In some embodiments, the second pool contains the majority of the high-methylated partition DNA and a minority of the low-methylated partition DNA. In some embodiments including intermediate methylated partitions, the second pool contains at least a portion of the intermediate methylated partition DNA, for example, the majority of the intermediate methylated partition DNA. In some embodiments, the first pool contains the majority of the low-methylated partition DNA, and the second pool contains the majority of the high-methylated partition DNA and the majority of the intermediate methylated partition DNA.
[0309] In some embodiments, the method includes the step of capturing at least a first set of target regions from a first pool, for example, the first pool being as described in any of the embodiments described above. In some embodiments, the first set includes sequence-variable target regions. In some embodiments, the first set includes hypomethylation-variable target regions and / or fragmentation-variable target regions. In some embodiments, the first set includes sequence-variable target regions and fragmentation-variable target regions. In some embodiments, the first set includes sequence-variable target regions, hypomethylation-variable target regions, and fragmentation-variable target regions. The step of amplifying the DNA in the first pool may be performed before this capture step. In some embodiments, the step of capturing a first set of target regions from the first pool includes contacting the DNA in the first pool with a first set of target-specific probes. In some embodiments, the first set of target-specific probes includes target-binding probes specific to sequence-variable target regions. In some embodiments, the first set of target-specific probes includes target-binding probes specific to sequence-variable target regions, hypomethylation-variable target regions, and / or fragmentation-variable target regions.
[0310] In some embodiments, the method includes the step of capturing a second set or multiple sets of target regions from a second pool, for example, the first pool being as described in any of the embodiments described above. In some embodiments, the second multiple includes epigenetic target regions, e.g., highly methylated variable target regions and / or fragmentation variable target regions. In some embodiments, the second multiple includes sequence variable target regions and epigenetic target regions, e.g., highly methylated variable target regions and / or fragmentation variable target regions. The step of amplifying the DNA in the second pool may be performed before this capture step. In some embodiments, the step of capturing a second multiple set of target regions from the second pool includes contacting the DNA in the first pool with a second set of target-specific probes, the second set of target-specific probes including target-binding probes specific to sequence variable target regions and target-binding probes specific to epigenetic target regions. In some embodiments, the first set of target regions and the second set of target regions are not identical. For example, the first set of target regions may include one or more target regions that are not present in the second set of target regions. Alternatively, or additionally, the second set of target regions may include one or more target regions that are not present in the first set of target regions. In some embodiments, at least one hypermethylation-variable target region is captured from the second pool rather than from the first pool. In some embodiments, multiple hypermethylation-variable target regions are captured from the second pool rather than from the first pool. In some embodiments, the first set of target regions includes sequence-variable target regions, and / or the second set of target regions includes epigenetic target regions. In some embodiments, the first set of target regions includes sequence-variable target regions and fragmentation-variable target regions, and the second set of target regions includes epigenetic target regions, e.g., hypermethylation-variable target regions and fragmentation-variable target regions.In some embodiments, a first set of target regions includes sequence-variable target regions, fragmentation-variable target regions, and low-methylation-variable target regions, and a second set of target regions includes epigenetic target regions, such as high-methylation-variable target regions, and fragmentation-variable target regions.
[0311] In some embodiments, the first pool comprises the majority of the low-methylated DNA and a portion (e.g., about half) of the high-methylated DNA, and the second pool comprises a portion (e.g., about half) of the high-methylated DNA. In some such embodiments, the first set of target regions comprises sequence-variable target regions, and / or the second set of target regions comprises epigenetic target regions. The sequence-variable target regions and / or epigenetic target regions may be as described in any of the embodiments described elsewhere in this specification. 10. Sequencing
[0312] In general, sample nucleic acids adjacent to an adapter can be subjected to sequencing, regardless of whether or not they have been previously amplified. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, synthesis sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, digital gene expression (Helicos), next-generation sequencing (NGS), synthesis single-molecule sequencing (SMSS) (Helicos), large-scale parallel sequencing, cloned single-molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions can be carried out in various sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means for processing multiple sample sets substantially simultaneously. Sample processing units may also include multiple sample chambers that allow for the simultaneous processing of multiple runs.
[0313] In some embodiments, the sequencing step is performed on a library containing a set of captured target regions, which may include any of the target region sets described herein. In some embodiments, the sequencing step is performed on a library containing partial samples (e.g., whole genome partial samples) that have not undergone capture / enrichment. For example, target regions may be captured from a first partial sample and a second sample and then sequenced, or target regions may be captured from a first partial sample and combined with a second partial sample after processing such as contacting and tagging, or target regions may be captured from a second partial sample and combined with a first partial sample after processing such as contacting and tagging, or both first and second partial samples may be processed and combined without undergoing capture / enrichment.
[0314] The sequencing reaction can be performed on one or more forms of nucleic acids, at least one of which is known to contain markers for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragments present in the sample. In some embodiments, genome sequence coverage may be less than 5%, less than 10%, less than 15%, less than 20%, less than 25%, less than 30%, less than 40%, less than 50%, less than 60%, less than 70%, less than 80%, less than 90%, less than 95%, less than 99%, less than 99.9%, or less than 100%. In some embodiments, the sequencing reaction may provide sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the genome. Sequence coverage can be performed on at least 5, 10, 20, 70, 100, 200, or 500 different genes, or at most 5000, 2500, 1000, 500, or 100 different genes.
[0315] Multiple sequencing may be used to perform simultaneous sequencing reactions. In some cases, cell-free nucleic acids can be sequenced in at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. In other cases, cell-free nucleic acids can be sequenced in fewer than 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. The sequencing reactions may be performed sequentially or simultaneously. Subsequent data analysis may be performed on all or part of the sequencing reactions. In some cases, data analysis may be performed on at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions. In other cases, data analysis may be performed on fewer than 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions. An example of a read data depth is 1,000 to 50,000 reads per locus (base). a. Differential depth of sequencing
[0316] In some embodiments, nucleic acids corresponding to sequence-variable target region sets are sequenced to a greater sequencing depth than nucleic acids corresponding to epigenetic target region sets. For example, the sequencing depth of nucleic acids corresponding to sequence-variant target region sets is at least 1.25 times, 1.5 times, 1.75 times, 2 times, 2.25 times, 2.5 times, 2.75 times, 3 times, 3.5 times, 4 times, 4.5 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 11 times, 12 times, 13 times, 14 times, or 15 times greater than the sequencing depth of nucleic acids corresponding to epigenetic target region sets, or 1.25 times to 1. It can be 5 times, 1.5 to 1.75 times, 1.75 to 2 times, 2 to 2.25 times, 2.25 to 2.5 times, 2.5 to 2.75 times, 2.75 to 3 times, 3 to 3.5 times, 3.5 to 4 times, 4 to 4.5 times, 4.5 to 5 times, 5 to 5.5 times, 5.5 to 6 times, 6 to 7 times, 7 to 8 times, 8 to 9 times, 9 to 10 times, 10 to 11 times, 11 to 12 times, 13 to 14 times, 14 to 15 times, or 15 to 100 times larger. In some embodiments, the sequencing depth is at least 2 times larger. In some embodiments, the sequencing depth is at least 5 times larger. In some embodiments, the sequencing depth is at least 10 times larger. In some embodiments, the sequencing depth is 4 to 10 times greater. In some embodiments, the sequencing depth is 4 to 100 times greater. Each of these embodiments refers to the extent to which the nucleic acids corresponding to the sequence variable target region set are sequenced to a greater sequencing depth than the nucleic acids corresponding to the epigenetic target region set.
[0317] In some embodiments, captured cfDNA corresponding to a sequence variable target region set and captured cfDNA corresponding to an epigenetic target region set are sequenced simultaneously, for example, in the same sequencing cell (e.g., a flow cell of an Illumina sequencer), and / or in the same composition, which may be a pooled composition resulting from rearranging separately captured sets, or a composition obtained by capturing cfDNA corresponding to a sequence variable target region set and captured cfDNA corresponding to an epigenetic target region set in the same container. 11.Analysis
[0318] In some embodiments, the methods described herein include the step of identifying the presence of DNA produced by a tumor (or neoplastic cells or cancer cells).
[0319] This method can be used to diagnose a condition in a subject, particularly the presence of cancer; to characterize the condition (e.g., to stage cancer or determine cancer heterogeneity); to monitor the response to treatment of the condition; and to determine the risk of developing the condition or the prognosis of the subsequent course of the condition. This disclosure may also be useful in determining the effectiveness of a particular treatment option. In a successful treatment option, if the treatment is successful, more cancer cells may be killed and DNA may be lost, which may increase the amount of copy number variations or rare mutations detected in the subject's blood. In other cases, this may not occur. In another case, perhaps a particular treatment option may be correlated with the genetic profile of the cancer over time. This correlation may be useful in selecting a treatment.
[0320] In addition, if the cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.
[0321] The types and number of cancers that can be detected include blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, and homogeneous tumors. Cancer types and / or stages can be detected from genetic mutations, including mutations, rare mutations, indels, copy number variations, base transpositions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, alterations in chromosomal structure, gene fusions, chromosome fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0322] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profiling data may enable the characterization of specific cancer subtypes, which may be important for the diagnosis or treatment of that particular subtype. This information may also provide subjects or practitioners with clues about the prognosis of a particular cancer type, enabling either the subjects or practitioners to adapt treatment options as the disease progresses. Some cancers may become more invasive and genetically unstable as they progress. Other cancers may remain benign, inactive, or dormant. The systems and methods of this disclosure may be useful in determining disease progression.
[0323] Furthermore, the methods of this disclosure may be used to characterize heterogeneity of an abnormal condition in a subject. Such methods may include, for example, a step of generating a gene profile of extracellular polynucleotides derived from the subject, where the gene profile includes multiple data obtained by analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition may result in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells at different stages of cancer. In other examples, the heterogeneity may constitute multiple disease lesions. Furthermore, in the example of cancer, there may be multiple tumor lesions, possibly one or more lesions resulting from metastasis spreading from the primary site.
[0324] This method can be used to generate a profile, fingerprint, or set of data that aggregates genetic information from different cells in heterogeneous diseases. This set of data may include, individually or in combination, analyses of copy number variations, epigenetic variations, and mutations.
[0325] This method can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods described herein do not involve diagnosing, prognosing, or monitoring a fetus, and therefore do not apply to non-invasive prenatal testing. In other embodiments, these methodologies may be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in prenatal subjects where DNA and other polynucleotides can co-circulate with maternal molecules.
[0326] An exemplary method for molecular tag identification of a library distributed by MBD beads via NGS is as follows, which includes the step of subjecting a first partial sample to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA: 1. Using a methyl-binding domain protein-bead purification kit, the extracted DNA sample (e.g., extracted blood plasma DNA from a human sample, which is subjected to target capture as described herein, if necessary) is physically distributed, and all elutes from the process are stored for downstream processing. 2. Parallel application of differential molecular tags and adapter sequences to each partition performed by NGS. For example, ligating high methylation, residual methylation ("wash"), and low methylation partitions with NGS-adapters having molecular tags. 3. Subject the DNA to a procedure that affects the first nucleic acid base in the DNA differently from the second nucleic acid base in the DNA, such as any of the procedures described herein. 4. All molecularly tagged partitions are recombined and then amplified using adapter-specific DNA primer sequences. 5. Capture / hybridization of the entire recombinant and amplified library targeting the desired genomic region (e.g., cancer-specific genetic variants and differentially methylated regions). 6. Re-amplify the captured DNA library and add sample tags. Pool different samples and perform multiple assays using an NGS instrument. 7. Bioinformatics analysis of NGS data using molecular tags to identify unique molecules, and deconvolution of samples into differentially MBD-distributed molecules. This analysis allows for obtaining relative 5-methylcytosine information for genomic regions simultaneously with standard gene sequencing / mutant detection.
[0327] In some embodiments of the methods described herein, including but not limited to those described above, the molecular tag consists of nucleotides that are not modified by a procedure that affects a first nucleic acid base in the DNA differently than a second nucleic acid base in the DNA, e.g., any of those described herein (e.g., if the procedure is a bisulfite conversion or any other conversion that does not affect mC, mC together with A, T, and G; if the procedure is a conversion that does not affect hmC, hmC together with A, T, and G, etc.). In some embodiments of the methods described herein, including but not limited to those described above, the molecular tag does not contain nucleotides that are modified by a procedure that affects a first nucleic acid base in the DNA differently than a second nucleic acid base in the DNA, e.g., any of those described herein (e.g., if the procedure is a bisulfite conversion or any other conversion that affects C, the tag does not contain an unmodified C; if the procedure is a conversion that affects mC, the tag does not contain mC; if the procedure is a conversion that affects hmC, the tag does not contain hmC, etc.).
[0328] In general, procedures that affect a first nucleic acid base in DNA differently from those affecting a second nucleic acid base in DNA may instead be performed before the step of parallel application of the adapter sequence to each distribution, which is carried out by differential molecular tagging and NGS. For example, this may be performed when the procedure that affects a first nucleic acid base in DNA differently from those affecting a second nucleic acid base in DNA is a separation, e.g., hmC-seal, in which case the separated populations themselves may be differentially tagged to one another. An exemplary such method is as follows: 1. Physically distribute the extracted DNA sample (e.g., extracted blood plasma DNA derived from a human sample, which is, if necessary, subjected to target capture as described herein) using a methyl-binding domain protein-bead purification kit, and store all elutes from the process for downstream processing. 2. Subject the DNA to a procedure that affects the first nucleic acid base in the DNA differently from the second nucleic acid base in the DNA, such as any of the procedures described herein. 3. Parallel application of differential molecular tags and adapter sequences to each partition performed by NGS. For example, ligating highly methylated partitions (or, where applicable, two or more partial partitions of highly methylated partitions), residual methylation ("wash") partitions, and low methylation partitions with NGS-adapters having molecular tags. 4. All molecularly tagged partitions are recombined and then amplified using adapter-specific DNA primer sequences. 5. Capture / hybridization of the entire recombinant and amplified library targeting the desired genomic region (e.g., cancer-specific genetic variants and differentially methylated regions). 6. Re-amplify the captured DNA library and add sample tags. Pool different samples and perform multiple assays using an NGS instrument. 7. Bioinformatics analysis of NGS data, including the identification of unique molecules using molecular tags, and deconvolution of samples into differentially MBD-distributed molecules. This analysis allows for the simultaneous acquisition of relative 5-methylcytosine information for genomic regions, along with standard gene sequencing / mutant detection. 12. Exemplary Workflow
[0329] Exemplary workflows for distribution and library preparation are provided herein. In some embodiments, some or all features of the distribution and library preparation workflows may be used in combination. a. Distribution
[0330] In some embodiments, sample DNA (e.g., between 5 and 200 ng) is mixed with a methyl-binding domain (MBD) buffer, and magnetic beads are conjugated with the MBD protein and incubated overnight. Methylated DNA (highly methylated DNA) binds to the MBD protein on the magnetic beads during this incubation. Unmethylated (lowly methylated DNA) or less unmethylated (intermediately methylated) DNA is washed away from the beads with a buffer containing progressively increasing concentrations of salt. For example, one, two, or more fractions containing unmethylated, lowly methylated, and / or intermediately methylated DNA may be obtained from such washings. Finally, highly salted buffer is used to elute highly methylated DNA (highly methylated DNA) from the MBD protein. In some embodiments, these washings yield three partitions of DNA with progressively increasing methylation levels (lowly methylated partition, intermediately methylated partition, and highly methylated partition).
[0331] In some embodiments, the three partitions of DNA are desalted and concentrated during the enzymatic preparation step of library preparation. b. Library preparation
[0332] In some embodiments (for example, after enriching the DNA during distribution), the distributed DNA is made ligable by, for example, extending the terminal overhangs of the DNA molecules, adding adenosine residues to the 3' ends of the fragments, and phosphorylating the 5' ends of each DNA fragment. DNA ligase and adapters are added to ligate each distributed DNA molecule with the adapter at each end. These adapters contain distribution tags (e.g., non-random, non-unique barcodes) that are distinguishable from the distribution tags of adapters used in other distributions. Either before or after making the distributed DNA ligable and performing ligation, at least one partial sample (e.g., a low-methylated distribution, or, where applicable, a low-methylated distribution and an intermediate-methylated distribution) is digested with a methylation-dependent nuclease (e.g., a methylation-dependent restriction enzyme, e.g., FspEI). If necessary, the hypermethylated partition may be digested with a methylation-sensitive nuclease, such as a methylation-sensitive restriction enzyme (e.g., one or more of HpaII, BstUI, and Hin6i, or each of them). If necessary, the hypermethylated partition may be subjected to a procedure that affects the first nucleic acid base in the DNA differently from the second nucleic acid base in the DNA, such as any of those described herein. If the hypermethylated partition is further partitioned by a procedure that affects the first nucleic acid base in the DNA differently from the second nucleic acid base in the DNA, adapter ligation should be performed after this procedure so that the partial partitions of the hypermethylated partition can be differentiated and tagged. The three (or more) partitions are then pooled together and amplified (e.g., by PCR, e.g., using primers specific to the adapter).
[0333] After PCR, the amplified DNA may be cleaned and enriched before enrichment. The amplified DNA is contacted with a collection of probes described herein (which may be, for example, biotinylated RNA probes) that target a specific region of interest. The mixture is incubated, for example, overnight in a salt buffer. The probes are captured (for example, using streptavidin magnetic beads) and separated from the uncaptured amplified DNA by, for example, a series of salt washes, thereby enriching the sample. After enrichment, the enriched sample is amplified by PCR. In some embodiments, the PCR primers contain a sample tag, thereby incorporating the sample tag into the DNA molecule. In some embodiments, DNA from different samples is pooled together and then multiplexed using, for example, an Illumina NovaSeq sequencer. C. Further features of a particular method of disclosure 1. Sample
[0334] The sample may be any biological sample isolated from the subject. The sample may be a body sample. Samples may include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsy, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid. Examples of fluids in the intercellular spaces include gingival crevicular fluid, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions, as well as urine. The sample may be in a form originally isolated from the subject, or it may have been subjected to further processing to remove or add components, e.g., cells, or to enrich one component with another. Therefore, preferred body fluids for analysis are plasma or serum containing cell-free nucleic acids. The sample can be isolated or obtained from the subject and transported to the site of sample analysis. The sample can be stored or shipped at a desired temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. The sample can be isolated or obtained from the subject at the site of sample analysis. The subject may be a human, mammal, animal, companion animal, service animal, or pet. The subject may have cancer. Participants may not have cancer or detectable symptoms of cancer. Participants may be treated with one or more cancer treatments, such as chemotherapy, antibodies, vaccines, or biological therapies. Participants may be in remission. Participants may or may not be diagnosed as susceptible to cancer or any cancer-related gene mutation / disorder.
[0335] The volume of plasma may depend on the desired reading data depth of the region being sequenced. Exemplary volumes are 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, the volume could be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma sampled may be 5–20 mL.
[0336] The sample may contain varying amounts of nucleic acids, including genome equivalents. For example, a sample of about 30 ng of DNA may contain about 10,000 (10 4 It contains ) haploid human genome equivalents, and in the case of cfDNA, it contains approximately 200 billion (2 × 10⁻¹⁶) 11It may contain ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, it may contain about 600 billion individual molecules.
[0337] The sample may contain nucleic acids from different origins, e.g., from cells and cell-free cells of the same subject, or from cells and cell-free cells of different subjects. The sample may contain nucleic acids with mutations. For example, the sample may contain DNA with germline mutations and / or somatic mutations. Germline mutations refer to mutations present in the germline DNA of the subject. Somatic mutations refer to mutations originating from somatic cells of the subject, e.g., cancer cells. The sample may contain DNA with cancer-associated mutations (e.g., cancer-associated somatic mutations). The sample may contain epigenetic variants (i.e., chemical or protein modifications), where the epigenetic variant is associated with the presence of genetic variants, e.g., cancer-associated mutations. In some embodiments, the sample contains epigenetic variants associated with the presence of genetic variants, where the sample does not contain genetic variants.
[0338] Exemplary amounts of cell-free nucleic acids in the sample before amplification range from approximately 1 fg to approximately 1 μg, for example, 1 pg to 200 ng, 1 ng to 100 ng, and 10 ng to 1000 ng. For example, the amount may be up to approximately 600 ng, up to approximately 500 ng, up to approximately 400 ng, up to approximately 300 ng, up to approximately 200 ng, up to approximately 100 ng, up to approximately 50 ng, or up to approximately 20 ng of cell-free nucleic acid molecules. The amount may be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity may be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method may include obtaining 1 femtogram (fg) to 200 ng.
[0339] Cell-free DNA refers to DNA that is not contained within cells at the time of its isolation from the subject. For example, cfDNA can be isolated from a sample as DNA remaining in the sample after intact cells have been removed, without lysing cells or otherwise extracting intracellular DNA. Cell-free nucleic acids include DNA, RNA, and their hybrids, and include genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, nucleolar small RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids through secretion or cell death processes, such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free embryonic DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.
[0340] Cell-free nucleic acids exhibit an exemplary size distribution of approximately 100–500 nucleotides, with molecules of 110–230 nucleotides accounting for approximately 90% of the molecules, the mode being approximately 168 nucleotides, and a second minor peak in the range of 240–440 nucleotides.
[0341] Cell-free nucleic acids can be isolated from body fluids by fractionation or partitioning steps, which separate the cell-free nucleic acids found in the solution from intact cells and other insoluble components of the body fluid. Partitioning steps may include techniques such as centrifugation or filtration. Alternatively, cells in the body fluid may be lysed, and the cell-free nucleic acids and cellular nucleic acids may be processed together. Generally, after buffer addition and washing steps, the nucleic acids can be precipitated with alcohol. Further washing steps, such as using a silica-based column, may be used to remove impurities or salts. Nonspecific bulk carrier nucleic acids, such as C1 DNA, DNA, or proteins, for bisulfite sequencing, hybridization, and / or ligation may be added throughout the reaction, for example, to optimize yield, in certain aspects of the procedure.
[0342] Following such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA may be converted to double-stranded form so that they can be included in subsequent processing and analysis steps.
[0343] Double-stranded DNA molecules in a sample and single-stranded nucleic acid molecules to be converted into double-stranded DNA molecules can be ligated to an adapter at one or both ends. Typically, double-stranded molecules are blunt-ended by treatment with a polymerase having 5'-3' polymerase and 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenourage fragments and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecules can be ligated with an adapter that is at least partially double-stranded (e.g., a Y-shaped or bell-shaped adapter). Alternatively, complementary nucleotides may be added to the blunt ends of the sample nucleic acid and the adapter to facilitate ligation. Both blunt-end ligation and adherent-end ligation are contemplated herein. In blunt-end ligation, both the nucleic acid molecule and the adapter tag have blunt ends. In attached end ligation, typically the nucleic acid molecule has an "A" overhang, and the adapter has a "T" overhang. 2. Amplification
[0344] The sample nucleic acid adjacent to the adapter can be amplified by PCR and other amplification methods. Amplification is typically primed by the binding of a primer to a primer binding site in the adapter adjacent to the DNA molecule to be amplified. The amplification method may include cycles of denaturation, annealing, and extension as a result of thermocycling, or it may be isothermal, as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and replication methods based on auto-persistent sequences.
[0345] In some embodiments, the method performs dsDNA ligation using T-tail and C-tail adapters, resulting in amplification of at least 50, 60, 70, or 80% of the double-stranded nucleic acid before ligation to the adapter. Preferably, the method increases the amount or number of amplified molecules by at least 10, 15, or 20% compared to a control method performed using a T-tail adapter alone. 3. Bait set; capture section
[0346] As discussed above, nucleic acids in a sample can be subjected to a capture step, where molecules having a target sequence are captured for subsequent analysis. Target capture may involve the use of a bait set containing an oligonucleotide bait labeled with a capture moiety, e.g., biotin or other examples described below. Probes may have sequences selected to tile across a panel of regions, such as genes. In some embodiments, the bait set may have higher and lower capture yields with respect to sets of target regions, e.g., sequence-variable target region sets and epigenetic target region sets, as discussed elsewhere in this specification. Such a bait set is combined with the sample under conditions that allow hybridization of the target molecule with the bait. The captured molecules are then isolated using a capture moiety, e.g., a biotin capture moiety with streptavidin based on beads. Such methods are further described, for example, in U.S. Patent No. 9,850,523 issued December 26, 2017, which is incorporated herein by reference.
[0347] The capture portion may include, but is not limited to, biotin, avidin, streptavidin, nucleic acids containing a specific nucleotide sequence, haptens recognized by antibodies, and magnetically adsorbable particles. The extraction portion may be a binding pair, e.g., a biotin / streptavidin or hapten / antibody member. In some embodiments, the capture portion attached to the sample is captured by its binding pair attached to an isolateable portion, e.g., a magnetically adsorbable particle or a larger particle that can be settled by centrifugation. The capture portion may be any type of molecule that enables affinity separation of nucleic acids having the capture portion from nucleic acids lacking the capture portion. Exemplary capture portions include biotin enabling affinity separation by binding to streptavidin that is linked to or can be linked to a solid phase, or oligonucleotides enabling affinity separation through binding to complementary oligonucleotides that are linked to or can be linked to a solid phase. D. Collection of target-specific probes
[0348] In some embodiments, a collection of target-specific probes is used in the methods described herein. In some embodiments, the collection of target-specific probes includes target-binding probes specific to a set of sequence-variable target regions and target-binding probes specific to a set of epigenetic target regions. In some embodiments, the capture yield of the target-binding probes specific to the set of sequence-variable target regions is higher (e.g., at least twice as high) than the capture yield of the target-binding probes specific to the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific to a set of sequence-variable target regions that is higher (e.g., at least twice as high) than its capture yield specific to a set of epigenetic target regions.
[0349] In some embodiments, the capture yield of target-binding probes specific to sequence-variable target region sets is at least 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 times higher than the capture yield of target-binding probes specific to epigenetic target region sets. In some embodiments, the capture yield of target-binding probes specific to sequence-variable target region sets is 1.25 to 1.5, 1.5 to 1.75, 1.75 to 2, 2 to 2.25, 2.25 to 2.5, 2.5 to 2.75, 2.75 to 3, 3 to 3.5, 3.5 to 4, 4 to 4.5, 4.5 to 5, 5 to 5.5, 5.5 to 6, 6 to 7, 7 to 8, 8 to 9, 9 to 10, 10 to 11, 11 to 12, 13 to 14, or 14 to 15 times higher than the capture yield of target-binding probes specific to epigenetic target region sets.
[0350] In some embodiments, the collection of target-specific probes is configured to have a capture yield specific to a sequence-variable target region set that is at least 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 times higher than its capture yield to the epigenetic target region set. In some embodiments, the collection of target-specific probes is configured to have a capture yield for sequence-variable target regions that is 1.25 to 1.5, 1.5 to 1.75, 1.75 to 2, 2 to 2.25, 2.25 to 2.5, 2.5 to 2.75, 2.75 to 3, 3 to 3.5, 3.5 to 4, 4 to 4.5, 4.5 to 5, 5 to 5.5, 5.5 to 6, 6 to 7, 7 to 8, 8 to 9, 9 to 10, 10 to 11, 11 to 12, 13 to 14, or 14 to 15 times higher than its capture yield for epigenetic target regions.
[0351] Probe collections can be configured to provide higher capture yields for sequence-variable target region sets in various ways, including concentration, varying lengths, and / or chemistry (e.g., affecting affinity), as well as combinations thereof. Affinity can be modulated by adjusting probe length and / or including nucleotide modifications, as discussed below.
[0352] In some embodiments, target-specific probes specific to sequence-variable target region sets are present at higher concentrations than target-specific probes specific to epigenetic target region sets. In some embodiments, the concentration of target-binding probes specific to sequence-variable target region sets is at least 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 times higher than the concentration of target-binding probes specific to epigenetic target region sets. In some embodiments, the concentration of a target-binding probe specific to a sequence-variable target region set is 1.25 to 1.5, 1.5 to 1.75, 1.75 to 2, 2 to 2.25, 2.25 to 2.5, 2.5 to 2.75, 2.75 to 3, 3 to 3.5, 3.5 to 4, 4 to 4.5, 4.5 to 5, 5 to 5.5, 5.5 to 6, 6 to 7, 7 to 8, 8 to 9, 9 to 10, 10 to 11, 11 to 12, 13 to 14, or 14 to 15 times higher than the concentration of a target-binding probe specific to an epigenetic target region set. In such embodiments, the concentration may refer to the average mass per volume concentration of individual probes in each set.
[0353] In some embodiments, target-specific probes specific to sequence-variable target region sets have a higher affinity for those targets than target-specific probes specific to epigenetic target region sets. Affinity can be modulated in any way known to those skilled in the art, including by using different probe chemistry. Certain nucleotide modifications, such as cytosine 5-methylation (in the context of a particular sequence), modifications providing a heteroatom at the 2' position of a sugar, and LNA nucleotides can increase the stability of double-stranded nucleic acids, and oligonucleotides having such modifications have been shown to have relatively high affinity for their complementary sequences. See, for example, Severin et al., Nucleic Acids Res. 39: 8740-8751 (2011); Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997); U.S. Patent No. 9,738,894. Similarly, longer sequence lengths generally provide increased affinity. Other nucleotide modifications, such as substituting the nucleic acid base hypoxanthine for guanine, reduce affinity by decreasing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, target-specific probes specific to sequence-variable target region sets have modifications that increase their affinity for their target. In some embodiments, or / or further, target-specific probes specific to epigenetic target region sets have modifications that decrease their affinity for their target. In some embodiments, target-specific probes specific to sequence-variable target region sets have a longer average length and / or a higher average melting temperature than target-specific probes specific to epigenetic target region sets. These embodiments may be combined with each other and / or with the concentration differences considered above to achieve a desired multiplier difference in capture yield, such as any multiplier difference or range described above.
[0354] In some embodiments, the target-specific probe includes a capture portion. The capture portion may be any of the capture molecules described herein, for example, biotin. In some embodiments, the target-specific probe is covalently or noncovalently linked to a solid support, for example, through the interaction of binding pairs of the capture portion. In some embodiments, the solid support is a bead, for example, a magnetic bead.
[0355] In some embodiments, target-specific probes specific to a sequence-variable target region set and / or target-specific probes specific to an epigenetic target region set are probes containing sequences selected to tile across a bait set, e.g., a panel of regions such as capture regions and genes, as discussed above.
[0356] In some embodiments, the target-specific probe is provided in a single composition. This single composition may be a solution (liquid or frozen), or it may be a lyophilized product.
[0357] Alternatively, target-specific probes may be provided as a plurality of compositions, for example, comprising a first composition containing a probe specific to an epigenetic target region set, and a second composition containing a probe specific to a sequence-variable target region set. These probes may be mixed in appropriate proportions to provide combined probe compositions having either of the aforementioned multiplier differences in concentration and / or capture yield. Alternatively, they may be used in separate capture procedures (e.g., using aliquots of the sample, or sequentially with the same sample) to provide the first and second compositions containing the captured epigenetic target regions and sequence-variable target regions, respectively. 1. Probes specific to epigenetic target regions
[0358] A set of epigenetic target region probes may include probes specific to one or more types of target regions that are likely to distinguish DNA from neoplastic (e.g., tumor or cancer) cells from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein, for example, in the section above concerning captured sets. A set of epigenetic target region probes may also include probes for one or more control regions, for example, as described herein.
[0359] In some embodiments, the probes of the epigenetic target region set have a footprint of at least 100 kbp, for example, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the epigenetic target region set has a footprint in the range of 100 to 20 Mbp, for example, 100 to 200 kbp, 200 to 300 kbp, 300 to 400 kbp, 400 to 500 kbp, 500 to 600 kbp, 600 to 700 kbp, 700 to 800 kbp, 800 to 900 kbp, 900 to 1,000 kbp, 1 to 1.5 Mbp, 1.5 to 2 Mbp, 2 to 3 Mbp, 3 to 4 Mbp, 4 to 5 Mbp, 5 to 6 Mbp, 6 to 7 Mbp, 7 to 8 Mbp, 8 to 9 Mbp, 9 to 10 Mbp, or 10 to 20 Mbp. In some embodiments, the epigenetic target region set has a footprint of at least 20 Mbp. a. Highly methylated variable target region
[0360] In some embodiments, the probes for the epigenetic target region set include probes specific to one or more highly methylated variable target regions. Highly methylated variable target regions may also be referred to herein as highly methylated DMRs (differentially methylated regions). Highly methylated variable target regions can be any of those described above. For example, in some embodiments, the probes specific to highly methylated variable target regions include probes specific to a plurality of loci listed in Table 1, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, the probes specific to highly methylated variable target regions include probes specific to a plurality of loci listed in Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. In some embodiments, probes specific to highly methylated variable target regions include probes specific to a plurality of loci listed in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2. In some embodiments, for each locus included as a target region, there may be one or more probes having a hybridization site that binds between the transcription start site and the stop codon (or the last stop codon in the case of a gene that is alternatively spliced). In some embodiments, one or more probes bind within 300 bp, e.g., 200 or 100 bp, of the listed locations. In some embodiments, the probes have hybridization sites that overlap with the locations listed above. In some embodiments, probes specific to hypermethylated target regions include probes specific to one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer. b. Low-methylation variable target regions
[0361] In some embodiments, a set of epigenetic target region probes includes probes specific to one or more hypomethylated variable target regions. Hypomethylated variable target regions may also be referred to herein as hypomethylated DMRs (differentially methylated regions). Hypomethylated variable target regions can be any of those described above. For example, probes specific to one or more hypomethylated variable target regions may include probes for repeating elements, such as LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA, where intergeneric regions normally methylated in healthy cells may exhibit reduced methylation in tumor cells.
[0362] In some embodiments, probes specific to low-methylation variable target regions include probes specific to repeat elements and / or intergenetic regions. In some embodiments, probes specific to repeat elements include probes specific to one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and / or satellite DNA.
[0363] Exemplary probes specific to genomic regions exhibiting cancer-related hypomethylation include probes specific to human chromosome 1 nucleotides 8403565–8953708 and / or 151104701–151106035. In some embodiments, probes specific to hypomethylation variable target regions include probes specific to regions overlapping with or containing human chromosome 1 nucleotides 8403565–8953708 and / or 151104701–151106035. c.CTCF binding region
[0364] In some embodiments, the probes in the epigenetic target region set include probes specific to CTCF binding regions. In some embodiments, the probes specific to CTCF binding regions include probes specific to at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those mentioned above, or one or more of the CTCFBSDB, or the CTCF binding regions in the papers by Cuddapah et al., Martin et al., or Rhee et al. cited above. In some embodiments, the probes in the epigenetic target region set include regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding site. d. Transcription initiation site
[0365] In some embodiments, the probes of the epigenetic target region set include probes specific to transcription start sites. In some embodiments, the probes specific to transcription start sites include probes specific to at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those listed in DBTSS. In some embodiments, the probes of the epigenetic target region set include probes for sequences at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the transcription start site. e. Local amplification
[0366] As described above, local amplification is somatic mutation, but they can be detected by sequencing based on the frequency of readout data in a manner similar to approaches for detecting certain epigenetic changes, such as changes in methylation. Therefore, regions that may exhibit local amplification in cancer can be included in the epigenetic target region set as discussed above. In some embodiments, probes specific to the epigenetic target region set include probes specific to local amplification. In some embodiments, probes specific to local amplification include probes specific to one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, a probe specific to local amplification includes a probe specific to one or more of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the aforementioned targets. f. Control area
[0367] Including control regions can be useful to facilitate data validation. In some embodiments, probes specific to a set of epigenetic target regions include probes specific to control methylated regions that are expected to be methylated in essentially all samples. In some embodiments, probes specific to a set of epigenetic target regions include probes specific to control hypomethylated regions that are expected to be hypomethylated in essentially all samples. 2. Probes specific to sequence-variable target regions
[0368] Probes for sequence-variable target region sets may include probes specific to multiple regions known to undergo somatic mutation in cancer. Probes may be specific to any sequence-variable target region set described herein. Exemplary sequence-variable target region sets are discussed in detail herein, for example, in the section above relating to captured sets.
[0369] In some embodiments, the sequence-variable target region probe set has a footprint of at least 0.5kb, for example, at least 1kb, at least 2kb, at least 5kb, at least 10kb, at least 20kb, at least 30kb, or at least 40kb. In some embodiments, the epigenetic target region probe set has a footprint in the range of 0.5 to 100kb, for example, 0.5 to 2kb, 2 to 10kb, 10 to 20kb, 20 to 30kb, 30 to 40kb, 40 to 50kb, 50 to 60kb, 60 to 70kb, 70 to 80kb, 80 to 90kb, and 90 to 100kb. In some embodiments, the sequence-variable target region probe set has a footprint of at least 50kbp, for example, at least 100kbp, at least 200kbp, at least 300kbp, or at least 400kbp. In some embodiments, the sequence-variable target region probe set has a footprint in the range of 100 to 2000 kbp, for example, 100 to 200 kbp, 200 to 300 kbp, 300 to 400 kbp, 400 to 500 kbp, 500 to 600 kbp, 600 to 700 kbp, 700 to 800 kbp, 800 to 900 kbp, 900 to 1,000 kbp, 1 to 1.5 Mbp, or 1.5 to 2 Mbp. In some embodiments, the sequence-variable target region set has a footprint of at least 2 Mbp.
[0370] In some embodiments, the probes specific to the sequence variable target region set include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the probes specific to the sequence variable target region set include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the probes specific to the sequence variable target region set include probes specific to at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, probes specific to a set of sequence-variable target regions include probes specific to at least one, at least two, or at least a portion of three of the indels in Table 3. In some embodiments, probes specific to a set of sequence-variable target regions include probes specific to at least five, at least ten, at least fifteen, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes in Table 4. In some embodiments, probes specific to a set of sequence-variable target regions include probes specific to at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4.In some embodiments, the probes specific to the sequence-variable target region set include probes specific to at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 4. In some embodiments, the probes specific to the sequence-variable target region set include probes specific to at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, or eighteen of the indels in Table 4. In some embodiments, the probes specific to the sequence variable target region set include probes specific to at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, at least eighteen, at least nineteen, or at least twenty of the genes listed in Table 5.
[0371] In some embodiments, probes specific to a sequence-variable target region set include probes specific to target regions from at least 10, 20, 30, or 35 cancer-related genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1. E. Compositions containing captured DNA
[0372] Provided herein are combinations comprising first and second populations of DNA, wherein the second population comprises a fragment of DNA having a terminal or attached tag or adapter at the recognition site of at least one methylation-dependent nuclease, which may be any one or any combination of the methylation-dependent nucleases described herein. In some embodiments, the first and second populations are differentially tagged. The first population may contain or be derived from DNA having cytosine modifications at a higher rate than the second population. The first population may contain a first nucleic acid base morphology that was originally present in the DNA having the altered base-pairing specificity and a second nucleic acid base that does not have the altered base-pairing specificity, wherein the first nucleic acid base morphology that was originally present in the DNA before the alteration of base-pairing specificity is altered or unaltered nucleic acid base, and the second nucleic acid base is a different altered or unaltered nucleic acid base from the first nucleic acid base, and the first nucleic acid base morphology that was originally present in the DNA before the alteration of base-pairing specificity and the second nucleic acid base have the same base-pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleic acid base is modified or unmodified cytosine, and the second nucleic acid base is modified or unmodified cytosine. The first and second nucleic acid bases may be any of those considered herein in the outline of the invention or in relation to subjecting a first partial sample to a procedure that affects the first nucleic acid base in the DNA of the first partial sample differently from the second nucleic acid base in the DNA. In some embodiments, the first population includes a DNA fragment having a terminal or attached tag or adapter at the recognition site of at least one methylation-sensitive nuclease, which may be any one or any combination of the methylation-sensitive nucleases described herein.
[0373] In some embodiments, the first group includes array tags selected from a first set of one or more array tags, and the second group includes array tags selected from a second set of one or more array tags, wherein the second set of array tags is different from the first set of array tags. The array tags may include barcodes.
[0374] In some embodiments, the first group includes protected hmC, such as glucosylated hmC.
[0375] In some embodiments, the first group was subjected to one of the conversion procedures discussed herein, such as bisulfite conversion, Ox-BS conversion, TAB conversion, ACE conversion, TAP conversion, TAPSβ conversion, or CAP conversion. In some embodiments, the first group was subjected to deamination of mC and / or C after protection of hmC.
[0376] In some embodiments of the combination, the first population contains or is derived from DNA having cytosine modifications in a larger proportion than the second population, the first population comprises first and second subpopulations, the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity. In some embodiments, the second population does not contain the first nucleic acid base. In some embodiments, the first nucleic acid base is a modified or unmodified cytosine, the second nucleic acid base is a modified or unmodified cytosine, and the modified cytosine is optionally mC or hmC. In some embodiments, the first nucleic acid base is a modified or unmodified adenine, the second nucleic acid base is a modified or unmodified adenine, and the modified adenine is optionally mA.
[0377] In some embodiments, the first nucleic acid base (e.g., modified cytosine) is biotinylated. In some embodiments, the first nucleic acid base (e.g., modified cytosine) is the product of hysgen cycloaddition to β-6-azido-glucosyl-5-hydroxymethylcytosine containing an affinity label (e.g., biotin).
[0378] In any of the combinations described herein, the captured DNA may include cfDNA.
[0379] The captured DNA may have any of the features described herein for the capture set, such as containing a higher concentration of DNA corresponding to a sequence variable target region set (normalized with respect to footprint size as discussed above) than, for example, the concentration of DNA corresponding to an epigenetic target region set. In some embodiments, the DNA of the capture set includes sequence tags that may be added to the DNA described herein. Generally, including sequence tags results in a DNA molecule that differs from its naturally occurring untagged form.
[0380] The combination may further include a probe set or sequencing primer described herein, each of which may differ from naturally occurring nucleic acid molecules. For example, the probe set described herein may include a capture portion, and the sequencing primer may include a label that does not exist in nature. F. Computer System
[0381] The methods of the present disclosure can be implemented using or with the assistance of a computer system. For example, such a method may include the steps of: distributing a sample into a plurality of partial samples, including a first partial sample and a second partial sample, wherein the first partial sample contains DNA having cytosine modifications at a higher rate than the second partial sample; subjecting the first partial sample to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity; and sequencing the DNA in the first partial sample and the DNA in the second partial sample in such a manner that the first nucleic acid base in the DNA of the first partial sample is distinguished from the second nucleic acid base.
[0382] Figure 5 shows a computer system 501 programmed or otherwise configured to implement the method of this disclosure. The computer system 501 can configure various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 501 is configured to perform sample analysis, including sample preparation and nucleic acid sequencing.
[0383] The computer system 501 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 505, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 501 also includes memory or memory locations 510 (e.g., random-access memory, read-only memory, flash memory), an electronic storage unit 515 (e.g., a hard disk), a communication interface 520 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 525, e.g., a cache, other memory, data storage, and / or an electronic display adapter. The memory 510, storage unit 515, interface 520, and peripheral devices 525 communicate with the CPU 505 through a communication network or bus (solid wire), such as a motherboard. The storage unit 515 may be a data storage unit (or data repository) for storing data. The computer system 501 can be operably coupled to a computer network 530 with the help of the communication interface 520. The computer network 530 may be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some examples, the computer network 530 may be a telecommunications and / or data network. The computer network 530 may include one or more computer servers that can enable distributed computing, such as cloud computing. In some examples, the computer network 530 may implement a peer-to-peer network that, with the help of a computer system 501, can couple devices to the computer system 501 and allow them to behave as clients or servers.
[0384] CPU 505 can execute a sequence of machine-readable instructions that can be embodied in a program or software. Instructions may be stored in memory locations such as memory 510. Examples of operations performed by CPU 405 include fetching, decoding, executing, and writing back.
[0385] The storage unit 515 can store files such as drivers, libraries, and saved programs. The storage unit 515 can store user-created programs, recorded sessions, and output associated with programs. The storage unit 515 can store user data, such as user preferences and user programs. In some examples, the computer system 501 may include one or more external data storage units located on remote servers that communicate with the computer system 501, for example, via an intranet or the Internet. Data can be transferred from one location to another, for example, using a communication network or physical data transfer (e.g., using a hard drive, thumb drive, or other data storage mechanism).
[0386] Computer system 501 can communicate with one or more remote computer systems via network 530. In one embodiment, computer system 501 can communicate with a user's (e.g., operator's) remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone®, Android®-enabled devices, Blackberry®), or personal digital assistants. The user can access computer system 501 via network 530.
[0387] The methods described herein can be implemented by machine-executable code (e.g., a computer processor) stored in an electronic storage location of a computer system 501, such as memory 510 or an electronic storage unit 515. The machine-executable or machine-readable code can be provided in the form of software. In use, the code can be executed by the processor 505. In some examples, the code can be read from the storage unit 515 and stored in memory 510 for easy access by the processor 505. In some situations, the electronic storage unit 515 can be omitted, and the machine-executable instructions are stored in memory 510.
[0388] In one embodiment, the present disclosure, when performed by at least one electronic processor, comprises the steps of: distributing a DNA-containing sample into a plurality of subsamples, including a first subsample and a second subsample, wherein the first subsample contains DNA having cytosine modifications at a higher rate than the second subsample; contacting the second subsample with a methylation-dependent nuclease to degrade nonspecifically distributed DNA in the second subsample to produce a treated second subsample; and optionally contacting the first subsample with a methylation-sensitive endonuclease, This provides a non-temporary, computer-readable medium that includes a computer-executable instruction that performs at least part of a method comprising: degrading nonspecifically distributed DNA in a first partial sample to produce a treated first partial sample; capturing a first set of target regions containing epigenetic target regions from the first partial sample and the treated first partial sample; capturing a second set of target regions containing epigenetic target regions from the treated second partial sample; and sequencing the DNA in the first set of target regions and the second set of target regions. In one embodiment, the present disclosure provides a non-temporary computer-readable medium, which includes a computer-executable instruction that, when executed by at least one electronic processor, performs at least part of a method comprising: distributing a sample into a plurality of subsamples, including a first subsample and a second subsample, wherein the first subsample contains DNA having cytosine modifications at a higher rate than the second subsample; contacting the second subsample with a methylation-dependent nuclease to degrade nonspecifically distributed DNA in the second subsample to produce a treated second subsample; optionally contacting the first subsample with a methylation-sensitive endonuclease to degrade nonspecifically distributed DNA in the first subsample to produce a treated first subsample; capturing a first set of target regions containing epigenetic target regions from the first subsample and the treated first subsample; and sequencing the DNA in the first set of target regions and the DNA derived from the second subsample.In some embodiments, the method further includes the steps of: obtaining a plurality of sequence reading data generated from sequencing by a nucleic acid sequencer; mapping the plurality of sequence reading data to one or more reference sequences to generate mapped sequence reading data; and processing the mapped sequence reading data to determine the likelihood that a subject has cancer.
[0389] The code can be configured for use with a machine that has a pre-compiled and adapted processor to run the code, or it can be compiled during execution time. The code can be supplied in a programming language that can be selected so that the code can be pre-compiled or run as compiled.
[0390] Embodiments of the systems and methods provided herein, for example, computer system 501, can be embodied in programming. Various embodiments of the technology can typically be considered “products” or “manufactured goods” in the form of machine (or processor) executable code and / or related data that are executed or embodied in one type of machine-readable medium. Machine executable code can be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. A “storage” type medium may include any or all of the tangible memory, processor, etc., or related modules of a computer, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-temporary storage at any time for software programming.
[0391] All or part of the software may, at times, be communicated through the Internet or various other telecommunications networks. Such communication may, for example, enable loading software from one computer or processor to another computer or processor, for example, from a management server or host computer to an application server computer platform. Thus, other types of media that may have software elements include light waves, radio waves, and electromagnetic waves, for example, wired and optical fixed telephone networks, and media used across physical interfaces between local devices through various air links. Physical elements that transmit such waves, for example, wired or wireless links, optical links, etc., may also be considered media having software. As used herein, unless limited to non-temporary tangible “storage” media, terms such as computer or machine “readable media” refer to any medium involved in providing instructions to a processor for execution.
[0392] Therefore, machine-readable media, such as computer executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, or any storage device in any computer, such as those used to implement databases, as shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wires and optical fibers, including wires including buses in computer systems. Carrier media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper tapes, any other physical storage media having a hole pattern, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier-transmitted data or instructions, cables or links that transport such carriers, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be related to communicating one or more sequences of one or more instructions to a processor for execution.
[0393] The computer system 501 includes, or can communicate with, an electronic display 535, which includes a user interface (UI) 540 for providing, for example, one or more results of sample analysis. Examples of UIs, but not limited to, include graphical user interfaces (GUIs) and web-based user interfaces.
[0394] Further details relating to computer systems and networks, databases, and computer program artifacts are also included herein by reference, for example, in Peterson, Computer Networks: A Systems Approach, each of which is incorporated herein by reference in its entirety. Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd It is also available in Ed. (2006) and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011). G. Applications 1. Cancer and other diseases
[0395] This method can be used to diagnose a condition in a subject, particularly the presence of cancer; to characterize a condition (e.g., to stage cancer or determine cancer heterogeneity); to monitor the response to treatment of a condition; and to achieve prognosis of the risk of developing a condition or the subsequent course of a condition. This disclosure may also be useful in determining the effectiveness of a particular treatment option. In a successful treatment option, if the treatment is successful, more cancer cells may be killed and DNA may be lost, which may increase the amount of copy number variations or rare mutations detected in the subject's blood. In other cases, this may not occur. In another case, perhaps a particular treatment option may be correlated with the genetic profile of cancer over time. This correlation may be useful in selecting a treatment. In some embodiments, hypermethylated variable epigenetic target regions are analyzed to determine whether they exhibit hypermethylation features in tumor cells or cells that do not normally contribute significantly to cfDNA, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit hypomethylation features in tumor cells or cells that do not normally contribute significantly to cfDNA.
[0396] Furthermore, if the cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.
[0397] In some embodiments, the methods and systems disclosed herein may be used to identify customized or targeted therapies for treating a given disease or condition in a patient, based on the classification of nucleic acid variants as being of somatic or germline origin. Typically, the disease under consideration is a type of cancer. Non-specific examples of such cancers include cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia This includes CMML (common myeloma), liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, progenitor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal cancer (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary tumor, pancreatic acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma. The type and / or stage of cancer can be detected from genetic mutations, including mutations, rare mutations, indels, copy number variations, base transpositions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, alterations in chromosomal structure, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplication, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0398] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profiling data may enable the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that particular subtype. This information may also provide subjects or practitioners with clues about the prognosis of specific types of cancer, enabling them to adapt treatment options as the disease progresses. Some cancers may progress, become more invasive, and become genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of this disclosure may be useful in determining disease progression.
[0399] Furthermore, the methods of this disclosure may be used to characterize heterogeneity of an abnormal condition in a subject. Such a method may include, for example, the step of creating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile includes multiple data resulting from the analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition may be a condition resulting in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells from different stages of cancer. In other examples, the heterogeneity may constitute multiple lesions of the disease. Again, in the example of cancer, there may be multiple tumor lesions, and perhaps one or more lesions are the result of metastasis spreading from the primary site.
[0400] This method can be used to create or profile a data fingerprint or set, which is a summary of genetic information derived from different cells in heterogeneous diseases. This data set may include copy number variation, epigenetic variation, and mutation analysis, either individually or in combination.
[0401] This method can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods described herein do not involve diagnosing, prognosing, or monitoring a fetus, and therefore do not apply to non-invasive prenatal testing. In other embodiments, these methodologies may be used on a pregnant subject to diagnose, prognose, monitor, or observe cancer or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.
[0402] Other gene-based diseases, disorders, or conditions that may be evaluated as needed using the methods and systems disclosed herein include, but are not limited to, achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth disease (CMT), Clíduchat disease, Crohn's disease, cystic fibrosis, Darkham disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombosis, familial hypercholesterolemia, familial Mediterranean fever, and brittleness. Examples include weak X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland syndrome, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatocardiafacial syndrome, WAGR syndrome, and Wilson's disease.
[0403] In some embodiments, the method described herein includes the step of detecting the presence or absence of DNA originating from or derived from tumor cells at a pre-selected time point after a previous cancer treatment of a subject previously diagnosed with cancer, using a set of sequence information obtained as described herein. The method may further include the step of determining a cancer recurrence score indicating the presence or absence of DNA originating from or derived from tumor cells with respect to the subject of test.
[0404] If a cancer recurrence score is determined, it may be used to further determine a cancer recurrence status. For example, if the cancer recurrence score is above a predetermined threshold, the cancer recurrence status may be that there is a risk of cancer recurrence. For example, if the cancer recurrence score is above a predetermined threshold, the cancer recurrence status may be that there is a low or less risk of cancer recurrence. In certain embodiments, a cancer recurrence score equal to a predetermined threshold may result in a cancer recurrence status that is either at risk of cancer recurrence or low or less risk of cancer recurrence.
[0405] In some embodiments, the cancer recurrence score is compared to a predetermined cancer recurrence threshold. If the cancer recurrence score is above the cancer recurrence threshold, the subject is classified as a candidate for further cancer treatment; if the cancer recurrence score is below the cancer recurrence threshold, the subject is classified as not a candidate for treatment. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in either a classification as a candidate for further cancer treatment or not a candidate for treatment.
[0406] The methods discussed above may further include any suitability features (one or more) described elsewhere in this Spec, including a section on methods for determining the risk of cancer recurrence in a study subject and / or classifying a study subject as a candidate for subsequent cancer treatment. 2. A method for determining the risk of cancer recurrence in study subjects and / or classifying them as candidates for subsequent cancer treatment.
[0407] In some embodiments, the methods provided herein are for determining the risk of cancer recurrence in a test subject. In some embodiments, the methods provided herein are for classifying a test subject as a candidate for subsequent cancer treatment.
[0408] Any such method may include the step of collecting DNA (e.g., originating from or derived from tumor cells) from a test subject diagnosed with cancer at one or more pre-selected time points after one or more prior cancer treatments to the test subject. The subject may be any of the subjects described herein. The DNA may be cfDNA. The DNA may be obtained from a tissue sample.
[0409] One such method may include a step of capturing a set of target regions from a DNA of interest, wherein the set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby producing a set of captured DNA molecules. The capture step may be carried out according to any embodiment described elsewhere herein.
[0410] In any of these methods, the prior cancer treatment may include surgery, administration of therapeutic compositions, and / or chemotherapy.
[0411] One such method may involve sequencing the captured DNA molecules, thereby producing a set of sequence information. Captured DNA molecules of a sequence-variable target region set may be sequenced to a greater depth than captured DNA molecules of an epigenetic target region set.
[0412] Any such method may include the step of detecting the presence or absence of DNA originating from or derived from tumor cells at a pre-selected point in time using a set of sequence information. The detection of the presence or absence of DNA originating from or derived from tumor cells may be carried out according to any embodiment described elsewhere herein.
[0413] A method for determining the risk of cancer recurrence in a test subject may include the step of determining a cancer recurrence score indicating the presence or absence, or amount, of DNA originating from or derived from the tumor cells of the test subject. The cancer recurrence score may further be used to determine the cancer recurrence status. A cancer recurrence status may indicate a risk of cancer recurrence, for example, if the cancer recurrence score is above a predetermined threshold. A cancer recurrence status may indicate a low or less risk of cancer recurrence, for example, if the cancer recurrence score is above a predetermined threshold. In certain embodiments, a cancer recurrence score equal to a predetermined threshold may result in a cancer recurrence status of either a risk of cancer recurrence or a low or less risk of cancer recurrence.
[0414] A method for classifying a test subject as a candidate for subsequent cancer treatment may include the steps of comparing the test subject's cancer recurrence score to a predetermined cancer recurrence threshold and classifying the test subject as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or classifying it as not a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in either a classification as a candidate for subsequent cancer treatment or not a candidate for treatment. In some embodiments, the subsequent cancer treatment includes chemotherapy or administration of a therapeutic composition.
[0415] One such method may include a step of determining the disease-free survival (DFS) period of the study subject based on a cancer recurrence score, for example, the DFS period may be 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.
[0416] In some embodiments, the set of sequence information includes a sequence variable target region sequence, and the step of determining the cancer recurrence score may include the step of determining at least a first subscore indicating the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence variable target region sequence.
[0417] In some embodiments, the number of mutations in the sequence variable target region, selected from 1, 2, 3, 4, or 5, is sufficient to result in a cancer recurrence score in which the first subscore is classified as positive for cancer recurrence. In some embodiments, the number of mutations is selected from 1, 2, or 3.
[0418] In some embodiments, the set of sequence information includes epigenetic target region sequences, and determining the cancer recurrence score involves determining a second partial score indicating the amount of molecules (derived from the epigenetic target region sequences) that represent an epigenetic state different from the DNA found in a corresponding sample derived from a healthy subject (e.g., cfDNA found in a blood sample derived from a healthy subject, or DNA found in a tissue sample derived from a healthy subject, the tissue sample being of the same type as that obtained from the subject of study). These abnormal molecules (i.e., molecules having an epigenetic state different from the DNA found in a corresponding sample derived from a healthy subject) may correspond to cancer-associated epigenetic changes, such as methylation of hypermethylated variable target regions and / or fragmentation of variable target regions, where “disruption” means different from the DNA found in a corresponding sample derived from a healthy subject.
[0419] In some embodiments, the percentage of molecules corresponding to the hypermethylated variable target region set and / or fragmented variable target region set that exhibit hypermethylation in the hypermethylated variable target region set and / or abnormal fragmentation in the fragmented variable target region set, which is greater than or equal to a value in the range of 0.001% to 10%, is sufficient for the second partial score to be classified as positive for cancer recurrence. The range may be 0.001% to 1%, 0.005% to 1%, 0.01% to 5%, 0.01% to 2%, or 0.01% to 1%.
[0420] In some embodiments, any of such methods may include the step of determining a fraction of tumor DNA from fractions of molecules in a set of sequence information exhibiting one or more features indicating origin from tumor cells. This may be done for molecules corresponding to some or all of epigenetic target regions, including, for example, one or both of hypermethylated variable target regions and fragmentation variable target regions (hypermethylation of hypermethylated variable target regions and / or abnormal fragmentation of fragmentation variable target regions may be considered to indicate origin from tumor cells). This may be done for molecules corresponding to sequence variable target regions, e.g., molecules including modifications consistent with cancer, e.g., SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA may be determined based on a combination of molecules corresponding to epigenetic target regions and molecules corresponding to sequence variable target regions.
[0421] The determination of the cancer recurrence score is obtained at least partially based on the fractionation of tumor DNA, 10 -11 ~1 or 10 -10 Fractions of tumor DNA greater than a threshold in the range of ~1 are sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, 10 -10 ~10 -9 , 10 -9 ~10 -8 , 10 -8 ~10 -7 , 10 -7 ~10 -6 , 10 -6 ~10 -5 , 10 -5 ~10-4 , 10 -4 ~10 -3 , 10 -3 ~10 -2 , or 10 -2 ~10 -1 Fractions of tumor DNA that are greater than or equal to a threshold in the range are sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, at least 10 -7 A tumor DNA fraction greater than a threshold is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. The determination that a tumor DNA fraction is greater than a threshold, e.g., the threshold corresponding to any of the embodiments described above, may be made based on cumulative probability. For example, a sample was considered positive if the cumulative probability that the tumor fraction is greater than the threshold in any of the aforementioned ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999. In some embodiments, the probability threshold is at least 0.95, e.g., 0.99.
[0422] In some embodiments, the set of sequence information includes a sequence variable target region sequence and an epigenetic target region sequence, and the step of determining a cancer recurrence score includes determining a first subscore indicating the amount of SNVs, insertions / deletions, CNVs and / or fusions present in the sequence variable target region sequence, and a second subscore indicating the amount of abnormal molecules in the epigenetic target region sequence, and a step of combining the first and second subscores to provide a cancer recurrence score. When combining the first and second subscores, they may be combined by independently applying thresholds to each subscore (e.g., greater than a predetermined number of mutations (e.g., >1) in the sequence variable target region and greater than a predetermined fraction of abnormal molecules in the epigenetic target region (i.e., molecules having a different epigenetic state than the DNA found in the corresponding sample from a healthy subject, e.g., a tumor)), or by training a machine learning classifier to determine the state based on multiple positive and negative training samples.
[0423] In some embodiments, a combined score value in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.
[0424] In any embodiment in which the cancer recurrence score is classified as positive for cancer recurrence, the subject's cancer recurrence status may be classified as potentially at risk of cancer recurrence, and / or the subject may be a candidate for subsequent cancer treatment.
[0425] In some embodiments, the cancer is one of the types of cancer described elsewhere herein, for example, colorectal cancer. 3. Treatment and related administration
[0426] In certain embodiments, the methods disclosed herein relate to identifying a customized treatment, taking into account whether the nucleic acid variant is of somatic or germline origin, and administering it to a patient. In some embodiments, essentially any cancer treatment (e.g., surgery, radiation therapy, chemotherapy, and / or similar) may be included as part of these methods. Typically, a customized treatment comprises at least one immunotherapy (or immunotherapy agent). Immunotherapy generally refers to a method of enhancing the immune response against a given type of cancer. In certain embodiments, immunotherapy refers to a method of enhancing the T-cell response against a tumor or cancer.
[0427] In certain embodiments, the status of nucleic acid variants from a subject-derived sample, whether somatic or germline origin, may be compared to a database of comparator results from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the subject of study, and / or patients who are receiving or have received the same treatment as the subject of study. Customized or targeted therapies (or multiple therapies) may be identified if the nucleic acid variants and comparator results meet certain classification criteria (e.g., substantial or approximate match).
[0428] In certain embodiments, the customized treatments described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized treatments (e.g., immunotherapeutic agents) may also be administered by means such as oral, sublingual, rectal, vaginal, urethral, topical, intraocular, intranasal, and / or intraaural, and the administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, ointments, etc.
[0429] While preferred embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. It is not intended that the present invention be limited by the specific examples provided herein. Although the present invention is described with reference to the preceding specification, the descriptions and examples of embodiments herein are not intended to be constrained. Those skilled in the art will be able to conceive of numerous variations, modifications, and substitutions without departing from the present invention. Furthermore, it should be understood that not all aspects of the present invention are limited to the specific descriptions, configurations, or relative proportions described herein, and that they depend on various conditions and variables. It should be understood that various alternative forms to the embodiments of the present disclosure described herein can be used in the practice of the present disclosure. Therefore, the present disclosure is also intended to cover any such alternative forms, modifications, variations, or equivalents. The following claims define the scope of the present invention, and the methods and structures within these claims, as well as their equivalents, are intended to be covered thereby.
[0430] While the foregoing disclosure is described in part in detail by illustrations and examples for the purposes of clarity and understanding, it will be apparent to those skilled in the art by reading this disclosure that various variations in form and detail can be made without departing from the true scope of this disclosure and can be implemented within the scope of the appended claims. For example, all methods, systems, computer-readable media, and / or features, steps, elements, or other embodiments thereof can be used in various combinations. H. Kit
[0431] Kits containing the compositions described herein are also provided. The kits may be useful for carrying out the methods described herein. In some embodiments, the kit includes a first reagent for distributing a sample into a plurality of partial samples, e.g., any of the distribution reagents described elsewhere herein, as described herein. In some embodiments, the kit includes a second reagent for subjecting the first partial sample to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA, where the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity (e.g., any of the reagents described elsewhere herein for converting nucleic acid bases, e.g., cytosine or methylated cytosine, to different nucleic acid bases). The kit may include the first and second reagents, as well as additional elements considered below and / or elsewhere herein.
[0432] The kit also includes ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSFIR, CTNNBl, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, and GNAQ. It may contain multiple oligonucleotide probes that selectively hybridize to at least 5, 6, 7, 8, 9, 10, 20, 30, 40, or all genes selected from the group consisting of GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID 1 A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRKl. The number of genes that an oligonucleotide probe can selectively hybridize can vary. For example, the number of genes may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or 54. The kit may include a container containing multiple oligonucleotide probes for carrying out any of the methods described herein, and instructions for use.
[0433] Oligonucleotide probes can selectively hybridize to the exon regions of genes, for example, at least five genes. In some examples, oligonucleotide probes can selectively hybridize to at least 30 exons of genes, for example, at least five genes. In some examples, multiple probes can selectively hybridize to each of at least 30 exons. Each probe hybridizing to an exon may have a sequence that overlaps with at least one other probe. In some embodiments, oligo probes can selectively hybridize to non-coding regions of genes disclosed herein, for example, intron regions of genes. Oligo probes can also selectively hybridize to regions of genes that include both exon and intron regions of genes disclosed herein.
[0434] Any number of exons can be targeted by oligonucleotide probes. For example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 1 55, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 400, 500, 600, 700, 800, 900, 1,000, or more exons can be targeted.
[0435] The kit may include at least four, five, six, seven, or eight different library adapters, each having a distinct molecular barcode and identical sample barcode. The library adapters do not necessarily have to be sequencing adapters. For example, a library adapter may not contain a flow cell sequence or a sequence that allows for the formation of a hairpin loop for sequencing. Different mutations and combinations of molecular and sample barcodes are described throughout this specification and are applicable to the kit. Furthermore, in some examples, the adapters are not sequencing adapters. Additionally, the adapters provided in the kit may also include sequencing adapters. Sequencing adapters may include sequences that hybridize to one or more sequencing primers. Sequencing adapters may further include sequences that hybridize to a solid support, such as a flow cell sequence. For example, a sequencing adapter may be a flow cell adapter. Sequencing adapters can be attached to one or both ends of a polynucleotide fragment. In some examples, the kit may include at least eight different library adapters, each having a distinct molecular barcode and identical sample barcode. The library adapter does not have to be a sequencing adapter. The kit may further include a sequencing adapter having a first sequence that selectively hybridizes to the library adapter and a second sequence that selectively hybridizes to the flow cell sequence. In another example, the sequencing adapter may be hairpin-shaped. For example, a hairpin-shaped adapter may include a complementary double-stranded portion and a loop portion, the double-stranded portion of which can be attached (e.g., ligated) to a double-stranded polynucleotide. A hairpin-shaped sequencing adapter can be attached to both ends of a polynucleotide fragment to create a cyclic molecule which can be sequenced multiple times.The sequencing adapter can be connected in lengths of up to 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or 54 units from end to end. , 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or more bases. A sequencing adapter may contain 20-30, 20-40, 30-50, 30-60, 40-60, 40-70, 50-60, or 50-70 bases from end to end. ...
Claims
[Claim 1] The invention described in the specification.