Methods for quantifying immune cell DNA
By employing DNA methylation-based methods for immune cell type quantification, the method addresses the limitations of current liquid biopsy analysis, enhancing the accuracy of disease detection and diagnosis through improved immune cell type differentiation.
Patent Information
- Application Number
- JP2025517753
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2023-09-27
- Publication Date
- 2025-09-29
AI Technical Summary
Existing methods for analyzing liquid biopsies, such as detecting circulating tumor DNA (ctDNA) from blood samples, are inadequate due to low and variable nucleic acid amounts, and current immune cell type differentiation methods, like complete blood counts and microarray-based profiling, fail to distinguish between certain immune cell types, limiting accurate disease detection and diagnosis.
A method for quantifying immune cell types based on DNA methylation data, using next-generation sequencing and methyl-binding assays to profile differentially methylated regions, allowing for the differentiation of immune cell types like activated and naive lymphocytes.
Enhances the accuracy of disease detection and diagnosis by improving the differentiation of immune cell types, enabling more precise detection of disorders and informing treatment strategies.
Smart Images

Figure 2025532196000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 377,327, filed September 27, 2022, which is incorporated by reference herein in its entirety for all purposes.
[0002] FIELD OF THE INVENTION The present disclosure provides methods and compositions related to analyzing immune cell DNA.In some embodiments, DNA is derived from a subject who has or suspects to have disease or disorder, such as cancer.In some embodiments, the immune cell type that DNA originates from is identified and quantified. [Background technology]
[0003] Introduction and Abstract Invasive diagnostic procedures, including biopsies, are commonly used to detect or diagnose cancer, ulcers, liver disease, infection, transplant rejection, and other diseases and disorders, where analysis of cells or tissues from potential disease sites is analyzed for relevant characteristics. Detection of diseases and disorders based on analysis of bodily fluids such as blood ("liquid biopsy") is an interesting alternative. Liquid biopsies are non-invasive and may only require blood sampling. However, developing accurate and sensitive methods for analyzing liquid biopsies has been difficult, as has recovering nucleic acids from such fluids in an analyzable form, due to the low and variable amounts of nucleic acids released into bodily fluids. For example, detecting the presence of circulating tumor DNA (ctDNA) derived from early-stage cancer is difficult due to its low abundance.
[0004] An alternative or complementary approach is to detect signals related to secondary effects of the presence of diseases such as cancer. One such signal is a signal from the immune response to tumor formation. In the presence of tumor cells, immune cells proliferate, differentiate, and possibly turn over at a faster rate than in healthy subjects. Such phenomena can lead to changes in immune cell populations and the amount of their DNA in the blood. Therefore, secondary immune signals can be useful in detecting diseases or disorders, such as cancer, with high sensitivity, at least in some situations.
[0005] Furthermore, quantifying different blood cell types provides important information about the overall health of a subject in addition to information about disease state. The ability to distinguish between different cell types, including closely related cell types, can be important for distinguishing between different types of disease and disorders. In disease state, the DNA derived from a specific immune cell type can be evaluated and used as an indicator of disease. In some genomic regions, the DNA methylation signatures in different immune cell types can be distinguished from myeloid cells and other immune cell types. Existing methods, such as complete blood counts and microarray-based profiling of differentially methylated regions, can detect many cell types, but cannot distinguish between certain types of immune cells, for example, between naive lymphocytes and activated lymphocytes, and therefore do not provide important information about the state of a subject's immune system. Summary of the Invention [Means for solving the problem]
[0006] The methods herein provide an approach for quantifying the levels of DNA from different immune cell types based on DNA methylation data generated, for example, by partitioning DNA, e.g., DNA obtained from immune cells in a blood sample, e.g., a buffy coat sample or a whole blood sample, based on the degree of methylation and next-generation DNA sequencing. In some embodiments, the methods herein use genomic regions that are differentially methylated between specific immune cell types and other blood cells and a methyl-binding assay to profile the methylation status of DNA fragments derived from these regions, facilitating the quantification of the presence of one or more specific immune cell types. Applications of this approach include cancer detection by detecting tumor-induced immune cell proliferation.
[0007] The methods herein can also provide complex information about other DNA variations and modifications, including but not limited to sequence variations.
[0008] The present disclosure aims to fulfill the need for improved analysis of DNA originating from different immune cell types, e.g., activated and naive lymphocytes, including rare immune cell types. Compared to existing methods, e.g., complete blood count (CBC) or DNA methylation-based methods, which do not distinguish between immune cell types, improved differentiation of immune cell types allows for more accurate detection (diagnosis) of disorders and therefore improved treatment. Accordingly, the following illustrative examples are provided.
[0009] Embodiment 1 is a method for analyzing DNA, comprising: a) sequencing DNA to determine methylation levels for a set of epigenetic target regions comprising a plurality of target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types, wherein the DNA is from a buffy coat sample or the DNA is from cells of the sample; b) determining the quantity of each of a plurality of immune cell types from which the DNA originates based on the methylation level; The method includes:
[0010] Embodiment 2 is a method for analyzing DNA, comprising: a) capturing at least a set of epigenetic target regions from DNA or a subsample thereof, comprising contacting the DNA or subsample thereof with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the DNA is from a buffy coat sample or the DNA is from cells of the sample, and the set of epigenetic target regions includes target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types; b) determining the methylation level for the target region; c) determining the quantity of each of a plurality of immune cell types from which the DNA originates; The method includes:
[0011] Embodiment 3 is a method for analyzing DNA, comprising: a) sequencing DNA to determine methylation levels for a set of epigenetic target regions comprising a plurality of hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in a plurality of immune cell types, wherein the DNA is from a buffy coat sample or the DNA is from cells of the sample; b) determining the quantity of each of a plurality of immune cell types from which the DNA originates based on the methylation level; The method includes:
[0012] Embodiment 3.1 is a method for analyzing DNA, comprising: a) sequencing DNA to determine methylation levels for a set of epigenetic target regions comprising a plurality of hypermethylated variable target regions comprising DNA sequences that are differentially hypermethylated in a plurality of immune cell types, wherein the DNA is from a buffy coat sample or the DNA is from cells of the sample; b) determining the quantity of each of a plurality of immune cell types from which the DNA originates based on the methylation level; The method includes:
[0013] Embodiment 4 is a method for analyzing DNA, comprising: a) capturing at least a set of epigenetic target regions from DNA or a subsample thereof, comprising contacting the DNA or subsample thereof with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the set of epigenetic target regions comprises hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in a plurality of immune cell types, and wherein the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) determining the methylation level for the target region; c) determining the quantity of each of a plurality of immune cell types from which the DNA originates; The method includes:
[0014] Embodiment 5 is a method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) sequencing DNA from one or more of the plurality of aliquots; c) detecting the level of the DNA sequence to determine the quantity of each of the multiple immune cell types from which the DNA originated; The method includes:
[0015] Embodiment 6 is a method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, the step comprising contacting the DNA with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the target regions of the set of epigenetic target regions comprise DNA sequences that are differentially methylated in a plurality of immune cell types; c) sequencing the captured DNA to determine the levels of each of the multiple immune cell types from which the DNA originated; The method includes:
[0016] Embodiment 7 is a method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, wherein at least one of the set of epigenetic target regions comprises a set of hypomethylated variable target regions; c) sequencing the captured DNA; d) detecting the level of the captured DNA sequences and determining the level of each of the multiple immune cell types from which the DNA originated; The method includes:
[0017] Embodiment 8 is a method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, the step comprising contacting the DNA with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the target regions of the set of epigenetic target regions comprise DNA sequences that are hypomethylated in a plurality of immune cell types; c) sequencing the captured DNA; The method includes:
[0018] Embodiment 9 is a method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, the step comprising contacting the DNA with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the target regions of the set of epigenetic target regions comprise DNA sequences that are hypermethylated in a plurality of immune cell types; c) sequencing the captured DNA; The method includes:
[0019] Embodiment 10 comprises the steps of: a) sequencing the CDR3 region of the DNA; b) determining the number of B cells, T cells, or both B and T cells from which the DNA originates based on the CDR3 sequence; 10. The method of any one of the preceding embodiments, further comprising:
[0020] Embodiment 11 is a method for analyzing DNA, comprising: a) sequencing the CDR3 region of DNA, wherein the DNA is from a buffy coat sample or the DNA is from cells of the sample; b) determining the number of B cells, T cells, or both B and T cells from which the DNA originates based on the CDR3 sequence; The method includes:
[0021] Embodiment 12 is the method of any one of embodiments 10 to 14, wherein the step of determining the number of B cells, T cells, or both B cells and T cells comprises determining the number of recombined CDR3 regions and the number of unrecombined CDR3 regions.
[0022] Embodiment 13 is the method of embodiment 10 or embodiment 11, wherein the CDR3 region is a CDR3 region of a B cell receptor chain or a T cell receptor chain.
[0023] Embodiment 14 is the method of embodiment 12, wherein the B cell receptor chain is a heavy chain or a light chain.
[0024] Embodiment 15 is the method of embodiment 12 or embodiment 13, wherein the T cell receptor chain is an alpha chain or a beta chain.
[0025] Embodiment 16 is a method for determining methylation levels for a set of epigenetic target regions, the set including a plurality of target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types, wherein the DNA is from a buffy coat sample or the DNA is from cells of the sample; b) determining the quantity of each of a plurality of immune cell types from which the DNA originates based on the methylation level; 16. The method according to any one of embodiments 11 to 15, further comprising:
[0026] Embodiment 17 is a method for detecting a target region comprising: a) capturing at least a set of epigenetic target regions from DNA or a subsample thereof prior to the sequencing step, the step comprising contacting the DNA or subsample thereof with target-specific probes specific for at least one of the set of epigenetic target regions, the set of epigenetic target regions comprising target regions that include DNA sequences that are differentially methylated in a plurality of immune cell types; b) determining the methylation level for the target region; c) determining the quantity of each of a plurality of immune cell types from which the DNA originates; 17. The method according to any one of embodiments 11 to 16, further comprising:
[0027] Embodiment 18 is a method for sequencing a DNA sequence comprising the steps of: a) prior to sequencing, partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot; b) sequencing DNA from one or more of the plurality of aliquots; c) detecting the level of the DNA sequence to determine the quantity of each of the multiple immune cell types from which the DNA originated; 18. The method according to any one of embodiments 11 to 17, further comprising:
[0028] Embodiment 19 is an embodiment of the present invention, wherein a) the target regions of the set of epigenetic target regions comprise DNA sequences that are hypomethylated in multiple immune cell types; or b) the target regions of the set of epigenetic target regions comprise DNA sequences that are hypermethylated in multiple immune cell types; 19. The method of embodiment 17 or embodiment 18.
[0029] Embodiment 20 is a method according to any one of embodiments 5, 7, 10, 12-15, 18, or 19, comprising capturing at least a set of epigenetic target regions of DNA from at least one of the first and second subsamples, wherein the target regions of the set of epigenetic target regions comprise DNA sequences that are differentially methylated in multiple immune cell types, and wherein the capturing step is performed before the sequencing step.
[0030] Embodiment 21 is the method of any one of the preceding embodiments, wherein the set of epigenetic target regions comprises a set of hypermethylated variable target regions and a set of hypomethylated variable target regions.
[0031] Embodiment 22 is a method for treating a tumour comprising administering to a subject therapies a tumour comprising: a) a plurality of immune cell types, including macrophages (including M1 macrophages and M2 macrophages); B cells, e.g., naive B cells or activated B cells (including regulatory B cells, memory B cells, switched memory B cells, and plasma cells); B cell precursors; T cells, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (cytotoxic T cells, regulatory T cells (Tregs), CD4 T cells (including naive CD4 T cells, activated CD4 T cells, and CD4 effector memory T cells), CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and CD4 effector memory T cells), The method of any one of the preceding embodiments, wherein the cells comprise two or more of: immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), neutrophils, low-density neutrophils, immature neutrophils, and immature granulocytes); dendritic cells; eosinophils; monocytes; erythrocytes; megakaryocytes; and natural killer (NK) cells.
[0032] Embodiment 23 is a method for treating a patient having a multi-cell disease, comprising: i. Naive and activated lymphocytes; ii. monocytes and macrophages; and / or iii. Myelocytes, neutrophils, and eosinophils 10. The method of any one of the preceding embodiments, comprising:
[0033] Embodiment 23.1 is the method of any one of the preceding embodiments, wherein the plurality of immune cell types comprises naive lymphocytes.
[0034] Embodiment 23.2 is the method of any one of the preceding embodiments, wherein the plurality of immune cell types comprises activated lymphocytes.
[0035] Embodiment 23.3 is the method of any one of the preceding embodiments, wherein the multiple immune cell types include monocytes.
[0036] Embodiment 23.4 is the method of any one of the preceding embodiments, wherein the multiple immune cell types include macrophages.
[0037] Embodiment 23.5 is the method of any one of the preceding embodiments, wherein the plurality of immune cell types comprises myeloid cells.
[0038] Embodiment 23.6 is the method of any one of the preceding embodiments, wherein the plurality of immune cell types includes neutrophils.
[0039] Embodiment 23.7 is the method of any one of the preceding embodiments, wherein the multiple immune cell types include eosinophils.
[0040] Embodiment 24 is the method of any one of the preceding embodiments, wherein the multiple immune cell types comprise naive and activated lymphocytes.
[0041] Embodiment 25 is the method of the preceding embodiment, wherein the multiple immune cell types comprise naive T cells, naive B cells, effector CD4 T cells, effector CD8 T cells, Treg cells, plasma cells, and memory cells.
[0042] Embodiment 26 is the method of the preceding embodiment, wherein the effector CD4 T cells comprise effector memory CD4 T cells and central memory CD4 T cells, and the effector CD8 T cells comprise effector memory CD8 T cells and central memory CD8 T cells.
[0043] Embodiment 27 is the method of any one of the preceding embodiments, wherein the multiple immune cell types include monocytes and macrophages.
[0044] Embodiment 28 is the method of embodiment 22, 23, or 27, wherein the macrophages are M1 macrophages or M2 macrophages.
[0045] Embodiment 29 is the method of any one of the preceding embodiments, wherein the multiple immune cell types comprise myelocytes, neutrophils, and eosinophils.
[0046] Embodiment 30 is the method of any one of the preceding embodiments, wherein the plurality of immune cell types comprises metamyelocytes.
[0047] Embodiment 31 is the method of any one of the preceding embodiments, wherein the multiple immune cell types comprise natural killer (NK) cells.
[0048] Embodiment 32 is the method of any one of the preceding embodiments, wherein the level of each of the plurality of immune cell types is determined relative to the level of total blood cells.
[0049] Embodiment 33 is the method of any one of the preceding embodiments, comprising determining a ratio of levels or quantities of immune cell types based on the determined levels or quantities of the plurality of immune cell types.
[0050] Embodiment 34 is the method of the preceding embodiment, wherein the numerator of the ratio comprises the level or quantity of neutrophils, monocytes, or both neutrophils and monocytes.
[0051] Embodiment 35 is the method of embodiment 33 or 34, wherein the denominator of the ratio comprises the level or quantity of T cells, B cells, NK cells, or total lymphocytes.
[0052] Embodiment 36 is the method of any one of embodiments 33 to 35, wherein the numerator of the ratio comprises the level or quantity of neutrophils and the denominator of the ratio comprises the level or quantity of total lymphocytes.
[0053] Embodiment 37 is the method of any one of embodiments 33 to 35, wherein the numerator of the ratio comprises the level or quantity of monocytes and the denominator of the ratio comprises the level or quantity of T cells.
[0054] Embodiment 37.1 is the method of any one of embodiments 33-35, wherein the numerator of the ratio comprises the level or quantity of M1 macrophages and the denominator of the ratio comprises the level or quantity of M2 macrophages.
[0055] Embodiment 37.2 is the method of any one of embodiments 33-35, wherein the numerator of the ratio comprises the level or quantity of NK cells and the denominator of the ratio comprises the level or quantity of total lymphocytes.
[0056] Embodiment 38 is the method of any one of the preceding embodiments, comprising determining the turnover frequency for at least one of the plurality of immune cell types.
[0057] Embodiment 39 is the method of the preceding embodiment, wherein the turnover comprises proliferation.
[0058] Embodiment 40 is the method of embodiment 38, wherein the turnover comprises apoptosis.
[0059] Embodiment 41 is the method of any one of the preceding embodiments, comprising determining the level of at least one cell type other than the immune cell type from which the DNA originated.
[0060] Embodiment 42 is the method of the preceding embodiment, comprising capturing at least one set of epigenetic target regions that contain sequence-independent differences in target regions in DNA originating from a cell type other than an immune cell type compared to the same target regions in DNA originating from all other cell types in the sample or sub-sample.
[0061] Embodiment 43 is the method of embodiment 41 or 42, wherein the cell type other than an immune cell type is not a blood cell type.
[0062] Embodiment 44 is the method of the preceding embodiment, wherein the cell type other than an immune cell type is colorectal, lung, breast, prostate, skin, stomach, bladder, liver, ovary, pancreas, squamous epithelium, salivary gland, larynx, hypopharynx, nasal cavity, paranasal sinuses, nasopharynx, or kidney.
[0063] Embodiment 45 is the method of any one of the preceding embodiments, wherein the sample is obtained from a subject.
[0064] Embodiment 46 is the method of the preceding embodiment, comprising determining the likelihood that the subject has cancer or a precancerous condition.
[0065] Embodiment 47 is the method of the preceding embodiment, further comprising determining the likelihood that the subject has cancer.
[0066] Embodiment 48 is the method of the preceding embodiment, wherein the cancer is an immune cell type cancer.
[0067] Embodiment 49 is the method of the preceding embodiment, wherein the cancer is a lymphocytic cancer.
[0068] Embodiment 50 is the method of the preceding embodiment, wherein the cancer is leukemia, lymphoma, or myeloma.
[0069] Embodiment 51 is the method of any one of embodiments 46 to 50, wherein the cancer is a myeloid cancer.
[0070] Embodiment 52 is the method of embodiment 46 or 47, wherein the cancer is a cancer of a cell or tissue type other than an immune cell type.
[0071] Embodiment 53 is the method of any one of embodiments 46, 47, or 52, wherein the cancer or precancerous condition is a cancer or precancerous condition other than a hematological cancer or precancerous condition, or the cancer or precancerous condition is a solid tumor cancer, optionally wherein the solid tumor cancer is a carcinoma, adenocarcinoma, or sarcoma.
[0072] Embodiment 54 is the method of any one of embodiments 46, 47, 52, or 53, wherein the cancer is colorectal cancer, lung cancer, breast cancer, prostate cancer, skin cancer, stomach cancer, bladder cancer, liver cancer, ovarian cancer, pancreatic cancer, head and neck cancer, or kidney cancer.
[0073] Embodiment 55 is the method of embodiment 46, comprising determining the likelihood that the subject has a precancerous condition.
[0074] Embodiment 56 is the method of the preceding embodiment, wherein the precancerous condition is an adenoma.
[0075] Embodiment 57 is the method of the preceding embodiment, wherein the adenoma is an advanced adenoma.
[0076] Embodiment 58 is the method of any one of embodiments 46, or 55-57, wherein the precancerous condition is colorectal precancerous condition, lung precancerous condition, breast precancerous condition, prostate precancerous condition, skin precancerous condition, stomach precancerous condition, bladder precancerous condition, liver precancerous condition, ovarian precancerous condition, pancreatic precancerous condition, head and neck precancerous condition, or kidney precancerous condition.
[0077] Embodiment 59 is the method of any one of embodiments 45 to 58, comprising determining the likelihood that the subject has an infection.
[0078] Embodiment 60 is the method of any one of embodiments 45 to 59, comprising determining the likelihood that the subject has graft rejection.
[0079] Embodiment 61 is the method of any one of embodiments 35 to 60, comprising predicting response to treatment in a subject.
[0080] Embodiment 62 is the method of the preceding embodiment, wherein the treatment is a chemotherapeutic agent.
[0081] Embodiment 63 is the method of embodiment 61, wherein the treatment is an immunotherapeutic agent.
[0082] Embodiment 64 is the method of any one of embodiments 55 to 63, wherein determining the quantity or sequencing of each of the plurality of immune cell types comprises generating a plurality of sequencing reads, and the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads to determine the likelihood that the subject has cancer, a precancerous condition, an infection, or transplant rejection.
[0083] Embodiment 65 is the method of any one of the preceding embodiments, wherein the sample is obtained from a subject previously diagnosed with cancer and who has undergone one or more previous cancer treatments, and optionally, the sample is obtained at one or more preselected time points after the one or more previous cancer treatments.
[0084] Embodiment 66 is the method of the preceding embodiments, further comprising determining a cancer recurrence score, optionally wherein the subject's cancer recurrence status is determined to be at risk of cancer recurrence if the cancer recurrence score is determined to be at or above a predetermined threshold, or the subject's cancer recurrence status is determined to be at low risk of cancer recurrence if the cancer recurrence score is below the predetermined threshold.
[0085] Embodiment 67 is the method of the preceding embodiment, further comprising the step of comparing the subject's cancer recurrence score to a predetermined cancer recurrence threshold, wherein the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or is not classified as a candidate for subsequent cancer treatment if the cancer recurrence score is below the cancer recurrence threshold.
[0086] Embodiment 68 is the method of any one of the preceding embodiments, further comprising analyzing DNA from an additional sample from the subject, wherein the additional sample was collected at a different time than when the DNA from the buffy coat sample or the DNA from the cells of the sample was collected.
[0087] Embodiment 69 is the method of the preceding embodiment, wherein:
[0088] DNA from a buffy coat sample, or DNA from cells from the sample, was collected before the subject underwent treatment, and an additional sample was collected after the subject underwent treatment; or
[0089] DNA from a buffy coat sample, or DNA from cells from the sample, was collected before the subject was diagnosed with the condition, and an additional sample was collected after the subject was diagnosed with the condition.
[0090] Embodiment 70 is the method of any one of embodiments 2, 4, 6, 10, 12-15, or 17-69, wherein the capturing step comprises capturing a sequence-variable target region.
[0091] Embodiment 71 is the method of the preceding embodiment, wherein the capturing step comprises contacting the DNA with target-specific probes specific for at least one of the set of epigenetic target regions and a target-specific probe specific for a sequence-variable target region.
[0092] Embodiment 72 is the method of any one of embodiments 5 to 10, or 12 to 15, or 18 to 71, wherein the modified cytosine is a methylcytosine.
[0093] Embodiment 73 is the method of any one of embodiments 5 to 10, or 12 to 15, or 18 to 71, wherein the agent that recognizes modified cytosines is a methyl-binding reagent.
[0094] Embodiment 74 is the method of the preceding embodiment, wherein the methyl-binding reagent is an antibody.
[0095] Embodiment 75 is the method of embodiment 73, wherein the methyl-binding reagent is a methyl-binding protein or comprises a methyl-binding domain.
[0096] Embodiment 76 is the method of embodiments 73-75, wherein the methyl-binding reagent specifically recognizes 5-methylcytosine.
[0097] Embodiment 77 is the method of embodiments 73 to 76, wherein the methyl-binding reagent is immobilized on a solid support.
[0098] Embodiment 78 is the method of any one of embodiments 5 to 77, wherein the partitioning step comprises immunoprecipitation of methylated DNA.
[0099] Embodiment 79 is the method of any one of embodiments 5 to 78, wherein the partitioning step comprises partitioning based on binding to a protein, and optionally the protein is a methylated protein, an acetylated protein, an unmethylated protein, or an unacetylated protein; and / or optionally the protein is a histone.
[0100] Embodiment 80 is the method of the preceding embodiment, wherein the distributing step comprises contacting the DNA of the sample with a binding reagent that is specific for the protein and that is immobilized on a solid support.
[0101] Embodiment 81 is the method of any one of the preceding embodiments, wherein determining the methylation level for the target region comprises bisulfite sequencing.
[0102] Embodiment 82 is the method of any preceding embodiment, further comprising performing a procedure that affects a first nucleobase of the DNA or at least one aliquot differently than a second nucleobase of the DNA or at least one aliquot; optionally, the first nucleobase is an unmodified cytosine and the second nucleobase is a modified cytosine, and further optionally, the modified cytosine is 5-methylcytosine or 5-hydroxymethylcytosine.
[0103] Embodiment 83 is the method of embodiment 82, wherein the procedure that affects a first nucleobase of the DNA or at least one aliquot differently from a second nucleobase of the DNA or at least one aliquot chemically converts the first or second nucleobase, thereby altering the base-pairing specificity of the converted nucleobase.
[0104] Embodiment 84 is the method of embodiment 82 or 83, wherein the step of affecting a first nucleobase of the DNA or at least one aliquot differently from a second nucleobase of the DNA or at least one aliquot is carried out as follows:
[0105] before the dispensing step, before the capturing step, or after the capturing step; and
[0106] before the sequencing step.
[0107] Embodiment 85 is the method of any one of embodiments 82 to 84, wherein the procedure that affects a first nucleobase of the DNA or at least one aliquot differently from a second nucleobase of the DNA or at least one aliquot is a methylation-sensitive conversion.
[0108] Embodiment 86 is the method of embodiment 85, wherein the methylation-sensitive conversion is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-coupled epigenetic (ACE) conversion, or enzymatic conversion.
[0109] Embodiment 87 is the method of embodiment 86, wherein the Tet-assisted conversion further comprises a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane.
[0110] Embodiment 87.1 is the method of any one of embodiments 82 to 87, wherein the step of affecting a first nucleobase of the DNA or at least one aliquot differently from a second nucleobase of the DNA or at least one aliquot comprises a TET enzyme.
[0111] Embodiment 87.2 is the method of embodiment 87.1, wherein the TET enzyme comprises TET1, TET2, TETv, or TETcd.
[0112] Embodiment 87.3 is the method of embodiment 87.1 or embodiment 87.2, wherein the TET enzyme comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3-5.
[0113] Embodiment 87.4 is the method of any one of embodiments 87.1 to 87.3, wherein the TET enzyme comprises the amino acid sequence of any one of SEQ ID NOs: 3 to 5.
[0114] Embodiment 87.5 is the method of any one of embodiments 87.1 to 87.3, wherein the TET enzyme comprises the V1900 TET mutant.
[0115] Embodiment 87.6 is the method of embodiment 87.5, wherein the TET enzyme comprises a V1900A, V1900C, V1900G, V1900I, or V1900P TET mutant.
[0116] Embodiment 87.7 is the method of embodiment 87.5 or embodiment 87.6, wherein the TET enzyme comprises a V1900A, V1900C, V1900G, V1900I, or V1900P TET2 mutant.
[0117] Embodiment 87.8 is the method of embodiment 87.7, wherein the TET enzyme comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 6, 7, 8, 9, or 10.
[0118] Embodiment 87.9 is the method of embodiment 87.8, wherein the TET enzyme comprises the amino acid sequence of any one of SEQ ID NOs: 6, 7, 8, 9, or 10.
[0119] Embodiment 87.10 is the method of any one of embodiments 87.1 to 87.9, wherein the TET enzyme comprises a mutation that increases 5-caC formation.
[0120] Embodiment 88 is the method of any one of embodiments 5 to 74, comprising contacting at least one aliquot with a restriction enzyme prior to the capturing or sequencing step, where the contacting is optionally performed (a) after the step of dividing the sample into a plurality of aliquots, (b) before performing a procedure that affects the DNA or a first nucleic acid base of the at least one aliquot differently from the DNA or a second nucleic acid base of the at least one aliquot, or (c) both (a) and (b).
[0121] Embodiment 89 is the method of the preceding embodiment, wherein the restriction enzyme is a methylation-dependent restriction enzyme (MDRE).
[0122] Embodiment 90 is the method of the preceding embodiment, wherein the second aliquot is contacted with an MDRE.
[0123] Embodiment 91 is the method of any one of embodiments 75 to 77, wherein the restriction enzyme is a methylation-sensitive restriction enzyme (MSRE).
[0124] Embodiment 92 is the method of the preceding embodiment, wherein the first aliquot is contacted with an MSRE.
[0125] Embodiment 93 is the method of any one of the preceding embodiments, comprising ligating an adaptor to the DNA, thereby producing an adaptor-ligated DNA.
[0126] Embodiment 94 is the method according to the preceding embodiment, wherein the adaptor-ligated DNA is amplified prior to the sequencing step.
[0127] Embodiment 95 is the method according to any one of embodiments 5 to 94, wherein the aliquots are pooled before the sequencing step.
[0128] Embodiment 96 is the method of any one of embodiments 1-4 or 6-95, wherein the set of epigenetic target regions comprises hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in multiple immune cell types.
[0129] Embodiment 97 is the method of any one of embodiments 3, 4, 7, 8, 10, 12-15, or 19-96, wherein the hypomethylated variable target region comprises a DNA sequence that is differentially hypomethylated in multiple immune cell types.
[0130] Embodiment 98 is the method of any one of embodiments 3, 4, 7, 8, 10, 12-15, 19, 96, or 97, wherein the DNA sequence that is differentially hypomethylated in the plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type in the sample.
[0131] Embodiment 99 is the method of any one of embodiments 3, 4, 7, 8, 10, 12-15, 19, 96, or 97, wherein the DNA sequence that is differentially hypomethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type.
[0132] Embodiment 100 is the method of any one of embodiments 3, 4, 7, 8, 10, 12-15, 19, 96, or 97, wherein the DNA sequence that is differentially hypomethylated in the plurality of immune cell types comprises a degree of methylation in one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type in the sample.
[0133] Embodiment 101 is the method of any one of embodiments 3, 4, 7, 8, 10, 12-15, 19, 96, or 97, wherein the DNA sequence that is differentially hypomethylated in a plurality of immune cell types comprises a degree of methylation in one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type.
[0134] Embodiment 102 is the method of any one of embodiments 3, 4, 7, 8, 10, 12-15, 19, 96, or 97, wherein the DNA sequence that is differentially hypomethylated in the plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in at least one other immune cell type and that is detectably lower than the degree of methylation of the same sequence in at least one non-immune cell type in the sample.
[0135] Embodiment 103 is the method of any one of embodiments 3, 4, 7, 8, 10, 12-15, 19, 96, or 97, wherein the DNA sequence that is differentially hypomethylated in the plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in at least one other immune cell type, and that is detectably lower than the degree of methylation of the same sequence in any non-immune cell type in the sample.
[0136] Embodiment 104 is the method according to any one of embodiments 98 to 103, wherein the detectably lower degree of methylation is one less methylated cytosine than the same sequence in other cell types.
[0137] Embodiment 105 is the method of any one of embodiments 98 to 103, wherein the detectably lower degree of methylation is two fewer methylated cytosines than the same sequence in other cell types.
[0138] Embodiment 106 is the method of any one of embodiments 98 to 103, wherein the detectably lower degree of methylation is 3 fewer methylated cytosines than the same sequence in other cell types.
[0139] Embodiment 107 is the method of any one of embodiments 98 to 103, wherein the detectably lower degree of methylation is four fewer methylated cytosines than the same sequence in other cell types.
[0140] Embodiment 108 is the method of any one of embodiments 98 to 103, wherein the detectably lower degree of methylation is 5 or more fewer methylated cytosines than the same sequence in other cell types.
[0141] Embodiment 109 is the method of any one of embodiments 1-4, 6, 10, 12-15, or 19-108, wherein the set of epigenetic target regions comprises hypermethylated variable target regions comprising DNA sequences that are differentially hypermethylated in multiple immune cell types.
[0142] Embodiment 110 is the method of embodiment 109, wherein the DNA sequence that is differentially hypermethylated in the plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type in the sample.
[0143] Embodiment 111 is the method of embodiment 109, wherein the DNA sequence that is differentially hypermethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type.
[0144] Embodiment 112 is the method of embodiment 109, wherein the DNA sequence that is differentially hypermethylated in the plurality of immune cell types comprises a degree of methylation in one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type in the sample.
[0145] Embodiment 113 is the method of embodiment 109, wherein the DNA sequence that is differentially hypermethylated in a plurality of immune cell types comprises a degree of methylation in one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type.
[0146] Embodiment 114 is the method of embodiment 109, wherein the DNA sequence that is differentially hypermethylated in the plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in at least one other immune cell type and that is detectably lower than the degree of methylation of the same sequence in at least one non-immune cell type in the sample.
[0147] Embodiment 115 is the method of embodiment 109, wherein the DNA sequence that is differentially hypermethylated in the plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in at least one other immune cell type and that is detectably lower than the degree of methylation of the same sequence in any non-immune cell type in the sample.
[0148] Embodiment 116 is the method according to any one of embodiments 110 to 115, wherein the detectably higher degree of methylation is one more cytosine methylation than the same sequence in other cell types.
[0149] Embodiment 117 is the method according to any one of embodiments 110 to 115, wherein the detectably higher degree of methylation is two more cytosine methylations than the same sequence in other cell types.
[0150] Embodiment 118 is the method according to any one of embodiments 110 to 115, wherein the detectably higher degree of methylation is three more cytosine methylations than the same sequence in other cell types.
[0151] Embodiment 119 is the method according to any one of embodiments 110 to 115, wherein the detectably higher degree of methylation is four more cytosine methylations than the same sequence in other cell types.
[0152] Embodiment 120 is the method according to any one of embodiments 110 to 115, wherein the detectably higher degree of methylation is 5 or more cytosine methylations than the same sequence in other cell types.
[0153] Embodiment 121 is the method of any one of the preceding embodiments, comprising determining the quantity or detecting the level of DNA in a sample originating from erythrocytes or erythrocyte precursors, macrophages, B cells, T cells, myeloid cells, or natural killer cells.
[0154] Embodiment 122 is the method of embodiment 121, wherein the macrophages are M1 macrophages or M2 macrophages.
[0155] Embodiment 123 is the method of embodiment 121, wherein the B cells are activated B cells.
[0156] Embodiment 124 is the method of embodiment 123, wherein the activated B cells are regulatory B cells, memory B cells, or plasma cells.
[0157] Embodiment 125 is the method of embodiment 121, wherein the T cells are central memory T cells, naive T cells, or naive-like T cells, and optionally, the central memory T cells are CD4 central memory T cells or CD8 central memory T cells.
[0158] Embodiment 126 is the method of embodiment 121, wherein the T cells are activated T cells.
[0159] Embodiment 127 is the method of embodiment 126, wherein the activated T cells are cytotoxic T cells, regulatory T cells, CD4 effector memory T cells, or CD8 effector memory T cells.
[0160] Embodiment 128 is the method of embodiment 121, wherein the myeloid cells are immature myeloid cells.
[0161] Embodiment 129 is the method of embodiment 128, wherein the immature myeloid cells are myeloid-derived suppressor cells, low-density neutrophils, immature neutrophils, or immature granulocytes.
[0162] Embodiment 130 is the method of any one of embodiments 98 to 129, wherein a detectably lower or higher degree of methylation is present in samples from at least one group of donors compared to another group of donors, and optionally at least one group of donors has a cancer that responds to treatment and another group of donors has a cancer that does not respond to treatment.
[0163] Embodiment 131 is the method of any one of the preceding embodiments, wherein the DNA is derived from a buffy coat sample.
[0164] Embodiment 132 is the method of any one of the preceding embodiments, wherein the DNA is derived from cells of a blood sample.
[0165] Embodiment 133 is the method according to the preceding embodiments, wherein the blood sample is a whole blood sample, a leukoreduced sample, or a PBMC sample.
[0166] Embodiment 134 is the method of embodiment 132 or 133, wherein the blood sample is a PBMC sample. [Brief explanation of the drawings]
[0167] [Figure 1A-1] 1A shows an exemplary workflow according to certain embodiments disclosed herein. DNA useful in the disclosed embodiments can include DNA collected from a sample containing cells (e.g., a buffy coat sample, or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduced sample, or a PBMC sample)). [Figure 1A-2]1A shows an exemplary workflow according to certain embodiments disclosed herein. DNA useful in the disclosed embodiments can include DNA collected from a sample containing cells (e.g., a buffy coat sample, or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduced sample, or a PBMC sample)).
[0168] [Figure 1B] Figure 1B is a heatmap showing the methylation levels of loci identified as having cell type- or cell cluster-specific differential methylation as a function of cell type or cell cluster. DMSs and DMRs indicate differentially methylated sites (e.g., individual CpGs) and differentially methylated regions (containing multiple DMSs).
[0169] [Figure 2] FIG. 2 is a schematic diagram of an example system suitable for use in some embodiments of the present disclosure.
[0170] [Figure 3A] Figures 3A and 3B show cell type populations in healthy and CRC samples from two cohorts (Cohort A and Cohort B). The cell type proportions from the fitted models were renormalized to sum to 1. Figure 3A shows the proportions of B cells, neutrophils, NK cells, T cells, NK cell to total lymphocyte ratio, and neutrophil to lymphocyte ratio (NLR) in healthy and CRC samples. Figure 3B shows the proportions of memory B cells, naive B cells, and plasma cells in healthy and CRC samples. [Figure 3B]Figures 3A and 3B show cell type populations in healthy and CRC samples from two cohorts (Cohort A and Cohort B). The cell type proportions from the fitted models were renormalized to sum to 1. Figure 3A shows the proportions of B cells, neutrophils, NK cells, T cells, NK cell to total lymphocyte ratio, and neutrophil to lymphocyte ratio (NLR) in healthy and CRC samples. Figure 3B shows the proportions of memory B cells, naive B cells, and plasma cells in healthy and CRC samples. DETAILED DESCRIPTION OF THE INVENTION
[0171] Detailed Description of Certain Embodiments Reference will now be made in detail to certain specific embodiments of the invention. While the invention will be described in conjunction with such embodiments, it will be understood that it is not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents which may be included within the scope of the present invention as defined by the appended claims.
[0172] Before describing the present teachings in detail, it should be understood that the present disclosure is not limited to specific compositions or process steps, and therefore may vary.When used in this specification and the appended claims, it should be noted that the singular forms "a", "an" and "the" include plural references unless the context clearly dictates otherwise.Thus, for example, reference to "a nucleic acid" includes plural nucleic acids, reference to "a cell" includes plural cells, etc.
[0173] Numerical ranges are inclusive of the numbers defining the range. It is understood that measured and measurable values are approximate, taking into account significant digits and error associated with measurement. Also, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting. It is to be understood that both the foregoing general and detailed descriptions are exemplary and explanatory only and are not restrictive of the present teachings.
[0174] Unless otherwise stated in the specification above, embodiments herein that recite "comprising" various components are also contemplated as "consisting of" or "consisting essentially of" the recited components, and embodiments herein that recite "consisting of" various components are also contemplated as "comprising" or "consisting essentially of" the recited components, and embodiments herein that recite "consisting essentially of" various components are also contemplated as "consisting of" or "comprising" the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0175] The section headings used herein are for organizational purposes and should not be construed as limiting the disclosed subject matter in any way. In the event that any document or other material incorporated by reference contradicts any express content of this specification, including definitions, the present specification will control.
[0176] definition "Buffy coat" refers to a portion of a blood (e.g., whole blood) or bone marrow sample that contains all or most of the sample's white blood cells and platelets. A buffy coat fraction of a sample can be prepared from the sample using centrifugation, which separates sample components by density. For example, after centrifugation of a whole blood sample, the buffy coat fraction is located between the plasma layer and the erythrocyte (red blood cell) layer. Buffy coats can contain both mononuclear leukocytes (e.g., T cells, B cells, NK cells, dendritic cells, and monocytes) and polymorphonuclear leukocytes (e.g., granulocytes, such as neutrophils and eosinophils).
[0177] As used herein, "fragment" or "fragmenting" refers to the breakdown or separation of a biological component, such as a nucleic acid molecule (e.g., DNA or RNA), into two or more small pieces. Fragmentation, such as DNA fragmentation, can occur naturally or can be intentionally induced, for example, using standard laboratory techniques such as those described herein. DNA fragmentation can be performed, for example, to prepare DNA for sequencing (e.g., genomic DNA and / or DNA isolated from a sample containing cells).
[0178] As used herein, "partitioning" nucleic acids, such as DNA molecules, refers to separating, fractionating, or sorting a sample or population of nucleic acids into multiple subsamples or subpopulations of nucleic acids based on one or more modifications or characteristics that differ in proportion among the multiple subsamples or subpopulations. Partitioning can include physically dividing nucleic acid molecules based on the presence or absence of one or more methylated nucleic acid bases. A sample or population can be divided into one or more divided subsamples or subpopulations based on characteristics that indicate genetic or epigenetic changes or pathologies.
[0179] As used herein, a modification or other feature is present in a "greater proportion" in a first sample or population of nucleic acids than in a second sample or population when the fraction of nucleotides having the modification or other feature is greater in the first sample or population than in the second population. For example, if one in ten nucleotides in a first sample are mC and one in twenty nucleotides in a second sample are mC, then the first sample contains a greater proportion of 5-methylated cytosine modifications than the second sample.
[0180] As used herein, "isolated" refers to a biological component (e.g., a nucleic acid molecule, a protein, or a cell) that has been substantially separated from, produced separately from, or purified separately from other components (e.g., other components in a sample, cell, or organism in which the component naturally occurs). An "isolated" nucleic acid molecule, protein, or cell includes one that has been purified using standard purification methods. The terms "isolated" or "purified" do not require absolute purity; rather, they are intended as relative terms. Thus, for example, an isolated biological component is one in which the biological component is enriched in a preparation relative to the biological component in its natural environment within a cell, organism, sample, or production vessel (e.g., a cell culture system). For example, an isolated biological component can represent at least 50%, e.g., at least 70%, at least 80%, at least 90%, at least 95%, or more of the total biological component content of a preparation.
[0181] As used herein, "leukapheresis" refers to a procedure in which white blood cells (leukocytes) are isolated from a sample of blood collected from a subject. Leukapheresis can be performed, for example, to obtain cells for research, diagnostic, prognostic, or monitoring purposes, such as those described herein. Thus, as used herein, a "leukapheresis sample" refers to a sample containing white blood cells collected from a subject using leukapheresis.
[0182] As used herein, "peripheral blood mononuclear cells" or "PBMCs" refer to immune cells with a single round nucleus that originate from the bone marrow and are found in the peripheral circulation. Such cells include, for example, lymphocytes (T cells, B cells, and NK cells) and monocytes, and are isolated from blood samples (e.g., from whole blood samples collected from a subject) using density gradient centrifugation.
[0183] As used herein, the "originally isolated" form of a sample refers to the composition or chemical structure of the sample when it is isolated and before it is subjected to any procedure that alters the chemical structure of the isolated sample. Similarly, a feature "natively present" in a DNA molecule refers to a feature that is present in the "original DNA molecule" or that is present in the DNA molecule "originally containing" the feature before the DNA molecule is subjected to a procedure that alters the chemical structure of the DNA molecule.
[0184] As used herein, "without substantially changing the base-pairing specificity" of a given nucleobase means that the majority of molecules that contain that nucleobase and can be sequenced have no change in the base-pairing specificity of the given nucleobase compared to the base-pairing specificity of the nucleobase when it was originally isolated in the sample.In some embodiments, 75%, 90%, 95%, or 99% of the molecules that contain that nucleobase and can be sequenced have no change in base-pairing specificity compared to the base-pairing specificity of the nucleobase when it was originally isolated in the sample.As used herein, "altered base-pairing specificity" of a given nucleobase means that the majority of molecules that contain that nucleobase and can be sequenced have base-pairing specificity at that nucleobase compared to the base-pairing specificity of the nucleobase when it was originally isolated in the sample.
[0185] As used herein, "base-pairing specificity" refers to the standard DNA base (A, C, G, or T) with which a given base most preferentially pairs. For example, unmodified cytosine and 5-methylcytosine have the same base-pairing specificity (i.e., specificity for G), whereas uracil and cytosine have different base-pairing specificities, with uracil having base-pairing specificity for A and cytosine having base-pairing specificity for G. Uracil's ability to form a wobble pair with G is not important, since uracil nevertheless most preferentially pairs with A among the four standard DNA bases.
[0186] As used herein, a "combination" containing multiple members refers to either a single composition containing the members or a set of compositions that are in close proximity, e.g., in separate containers or in compartments within a larger container such as a multi-well plate, test tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other form of storage.
[0187] The "capture yield" of a panel of probes for a given target set refers to the amount of nucleic acid corresponding to the target set that the panel of probes captures under typical conditions (e.g., the amount relative to another target set, or the absolute amount). Exemplary capture conditions are incubation of sample nucleic acid and probes in a small reaction volume (approximately 20 μL) containing a stringent hybridization buffer at 65°C for 10-18 hours. Capture yields may be expressed in absolute terms, or relative to multiple panels of probes. When capture yields for multiple sets of target regions are compared, the capture yield is normalized to the footprint size of the target region set (e.g., per kilobase). Thus, for example, if the footprint sizes of the first and second target regions are 50 kb and 500 kb, respectively (normalization factor 0.1), DNA corresponding to the first set of target regions will be captured with a higher yield than DNA corresponding to the second set of target regions when the mass per volume concentration of the captured DNA corresponding to the first set of target regions is greater than 0.1 times the mass per volume concentration of the captured DNA corresponding to the second set of target regions. As a further example, using the same footprint size, if the captured DNA corresponding to the first set of target regions has a mass per volume concentration that is 0.2 times the mass per volume concentration of the captured DNA corresponding to the second set of target regions, the DNA corresponding to the first set of target regions will be captured with a capture yield that is 2 times higher than the DNA corresponding to the second set of target regions.
[0188] "Capturing" one or more target nucleic acids, or one or more nucleic acids containing at least one target region, refers to preferentially isolating or separating one or more target nucleic acids, or one or more nucleic acids containing at least one target region, from non-target nucleic acids or from nucleic acids that do not contain at least one target region.
[0189] A "captured set" of nucleic acids, or a "captured" nucleic acid, refers to the nucleic acids that have undergone capture.
[0190] As used herein, a "capture moiety" is a molecule that allows for affinity separation of a molecule, e.g., a nucleic acid, linked to the capture moiety from a molecule that lacks the capture moiety. Exemplary capture moieties include biotin, which allows affinity separation by binding to streptavidin that is or can be linked to a solid phase, or an oligonucleotide, which allows affinity separation by binding to a complementary oligonucleotide that is or can be linked to a solid phase.
[0191] As used herein, a "cell type" is a set of cells that share common characteristics. For example, an immune cell type can include immune cells of different origins, differentiation types, activation types, or any combination of different origins, differentiation types, and activation types. In fact, the differentiation state and activation state can overlap and often change together in a given immune cell. For example, activation of an immune cell can induce differentiation of the cell. Immune cells of different activation types can include activated cells (e.g., cells activated by inflammatory cytokines or antigens), suppressive cells (e.g., T regulatory cells (Tregs), M2 macrophages, and others, or subsets thereof), or suppressive cells, e.g., cells suppressed by Tregs. Exemplary immune cell types include macrophages (including M1 macrophages and M2 macrophages); activated B cells (including regulatory B cells, memory B cells, and plasma cells); T cell subsets, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (including cytotoxic T cells, regulatory T cells (Tregs), CD4 effector memory T cells, and CD8 effector memory T cells); immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), low-density neutrophils, immature neutrophils, and immature granulocytes); and natural killer (NK) cells. In some embodiments, cell types can be distinguished based on characteristics, e.g., one or more cell surface markers, gene signatures (e.g., expression (or expression levels) of a particular gene or set of genes), and / or epigenetic signatures, e.g., regions of DNA hypermethylation or hypomethylation.
[0192] As used herein, a "cell cluster" or "cluster" refers to a plurality of related cell types, e.g., immune cell types. In some embodiments, the cell types within a cluster have similar DNA methylation profiles, e.g., in multiple hypermethylated and / or hypomethylated variable target regions.
[0193] A "target region" refers to a genomic locus that is targeted for identification and / or capture, e.g., by using a probe (e.g., by sequence complementarity). A "target region set" or "set of target regions" refers to a plurality of genomic loci that are targeted for identification and / or capture, e.g., by using a set of probes (e.g., by sequence complementarity).
[0194] "Specifically bind" in the context of a primer, probe or other oligonucleotide and a target sequence means that under appropriate hybridization conditions, the primer, oligonucleotide or probe hybridizes to its target sequence or a copy thereof to form a stable hybrid, while at the same time minimizing the formation of stable non-target hybrids. In this way, the primer or probe hybridizes to the target sequence or a copy thereof to a sufficiently greater extent than to non-target sequences, ultimately allowing capture or detection of the target sequence. Suitable hybridization conditions are well known in the art, can be predicted based on sequence composition, or can be determined by using routine testing methods (e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2004, incorporated herein by reference). nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) §§ 1.90-1.91, 7.37-7.57, 9.47-9.51, and 11.47-11.57, especially §§ 9.50-9.51, 11.12-11.13, 11.45-11.47, and 11.55-11.57).
[0195] "Sequence variable target region" refers to a target region that may exhibit sequence changes, such as nucleotide substitutions (i.e., single-nucleotide variations), insertions, deletions, or gene fusions or rearrangements, in neoplastic cells (e.g., tumor and cancer cells) compared to normal cells. A sequence variable target region set is a set of sequence variable target regions. In some embodiments, the sequence variable target region is a target region that may exhibit changes affecting 50 or fewer consecutive nucleotides, e.g., 40 or fewer, 30 or fewer, 20 or fewer, 10 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer nucleotides.
[0196] "Epigenetic target region" refers to a target region that may show sequence-independent differences in different cell or tissue types (e.g., different types of immune cells), or in neoplastic cells (e.g., tumor cells and cancer cells) compared to normal cells; or that may show sequence-independent differences (i.e., differences that do not change the nucleotide sequence, such as differences in methylation, nucleosome distribution, or other epigenetic features) in DNA, for example, from different cell types or from subjects with cancer compared to DNA from healthy subjects. Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), nucleosome distribution, fragmentation patterns, CCCTC-binding factor ("CTCF") binding, transcription start site (e.g., with respect to any one or more of RNA polymerase component binding, regulatory protein binding, fragmentation characteristics, and nucleosome distribution), and regulatory protein binding regions. Therefore, epigenetic target region set includes, but is not limited to, hypermethylated variable target region set, hypomethylated variable target region set, and fragmented variable target region set, such as, but not limited to, CTCF binding site and transcription start site.For this purpose, the locus that is prone to the local amplification and / or gene fusion associated with neoplasia, tumor or cancer can also be included in epigenetic target region set, because the detection of copy number change by sequencing or the fusion sequence that maps to more than one locus in reference genome tends to be more similar to the detection of the exemplary epigenetic changes discussed above than the detection of nucleotide substitution, insertion or deletion, for example, in that the detection of local amplification and / or gene fusion does not depend on the accuracy of base call at one or a few individual positions, and therefore can be detected at a relatively low sequencing depth.Epigenetic target region set is a set of epigenetic target regions.
[0197] As used herein, a "differentially methylated region" refers to a region of DNA that has a detectably different degree of methylation in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, or that has a detectably different degree of methylation in at least one cell or tissue type obtained from a subject with a disease or disorder compared to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region) in at least one cell or tissue type, e.g., at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, e.g., another immune cell type, or the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region) in at least one cell or tissue type, e.g., at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, e.g., other immune cell types.
[0198] Nucleic acids are "produced by a tumor" if they originate from tumor cells. Tumor cells are neoplastic cells that originate from a tumor, regardless of whether they remain within the tumor or are separated from the tumor (e.g., as in the case of metastatic cancer cells and circulating tumor cells). As used herein, a "precancerous condition" or "precancerous condition" is an abnormality that has the potential to become cancerous, and the likelihood of becoming cancerous is greater than the likelihood that the abnormality would have if it had not been present, i.e., normal. Examples of precancerous conditions include, but are not limited to, adenoma, hyperplasia, metaplasia, dysplasia, benign neoplasm (benign tumor), premalignant intraepithelial carcinoma, and polyp. It should be noted that certain types of intraepithelial carcinoma are recognized in the art as cancerous rather than premalignant, e.g., stage 0 cancer.
[0199] The term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleobase in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to cytosine at a CpG site (a cytosine-phosphate-guanine site (i.e., cytosine followed by guanine in the 5' to 3' direction of a nucleic acid sequence)). In some embodiments, DNA methylation refers to the addition of a methyl group to adenine, for example, at N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the six-carbon ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine, generating 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the six-carbon ring of cytosine). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of cytosine, generating 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, methylation of DNA within a promoter region can silence gene transcription. DNA methylation is crucial for normal development, and aberrant methylation can disrupt epigenetic regulation. Disruption, e.g., repression, of epigenetic regulation can lead to diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0200] The term "hypermethylated" refers to an increased level or degree of methylation of a nucleic acid molecule(s) within a population of nucleic acid molecules (e.g., a sample) relative to other nucleic acid molecules containing the same genetic information. In some embodiments, hypermethylated DNA can include DNA molecules containing at least one methylated residue, at least two methylated residues, at least three methylated residues, at least five methylated residues, or at least ten methylated residues.
[0201] The term "hypomethylated" refers to a reduced level or degree of methylation of a nucleic acid molecule(s) within a population (e.g., a sample) of nucleic acid molecules relative to other nucleic acid molecules containing the same genetic information. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules containing zero methylated residues, up to one methylated residue, up to two methylated residues, up to three methylated residues, up to four methylated residues, or up to five methylated residues.
[0202] The term "agent that recognizes modified nucleobases in DNA," e.g., "agent that recognizes modified cytosine in DNA," refers to a molecule or reagent that binds to or detects one or more modified nucleobases in DNA, e.g., methylcytosine. A "modified nucleobase" is a nucleobase that contains a difference in chemical structure from an unmodified nucleobase. In the case of DNA, the unmodified nucleobase is adenine, cytosine, guanine, or thymine. In some embodiments, the modified nucleobase is a modified cytosine. In some embodiments, the modified nucleobase is a methylated nucleobase. In some embodiments, the modified cytosine is a methylcytosine, e.g., 5-methylcytosine. In such embodiments, the cytosine modification is methyl. Agents that recognize methylcytosine in DNA include, but are not limited to, "methyl-binding reagents," which, as used herein, refer to reagents that bind to methylcytosine. Methyl-binding reagents include, but are not limited to, methyl-binding domains (MBDs) and methyl-binding proteins (MBPs), and antibodies specific for methylcytosine. In some embodiments, such antibodies bind to 5-methylcytosine in DNA. In some such embodiments, DNA can be single-stranded or double-stranded. Suitable agents include agents that recognize double-stranded DNA, single-stranded DNA, and modified nucleotides in both double-stranded DNA and single-stranded DNA.
[0203] The terms "or a combination thereof" and "or a combination thereof," as used herein, refer to any and all permutations and combinations of the terms listed before the term combination. For example, "A, B, C, or a combination thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and, if order is important in the particular context, also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Continuing this example, combinations containing repeats of one or more items or terms, e.g., BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc., are expressly included. Those skilled in the art will understand that, in general, there is no limit to the number of items or terms in any combination, unless otherwise clear from the context.
[0204] "Or" is used in its inclusive sense, ie, equivalent to "and / or" unless the context requires a different interpretation.
[0205] Exemplary Methods
[0206] A. Identification and Quantification of Immune Cell Types In some embodiments, the methods disclosed herein include sequencing DNA isolated from a buffy coat sample to determine methylation levels for multiple target regions containing DNA sequences that are differentially methylated in multiple immune cell types. In some embodiments, the DNA to be sequenced is isolated from a blood sample, such as a buffy coat sample, a whole blood sample, a leukocyte-reduced sample, or a PBMC sample. In some embodiments, the DNA to be sequenced is isolated from cells of a blood sample, such as a buffy coat sample, a whole blood sample, a leukocyte-reduced sample, or a PBMC sample. In any of the embodiments of the present disclosure, DNA isolated from any type of sample containing cells, including but not limited to a blood sample (e.g., a buffy coat sample, a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), can be DNA isolated from cells of the sample. In some embodiments, the methods disclosed herein include capturing at least a set of epigenetic target regions from DNA or a subsample thereof, comprising contacting the DNA or subsample thereof with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the set of epigenetic target regions comprises target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types, and determining methylation levels for the target regions. In some embodiments, the methods disclosed herein include sequencing the DNA and determining methylation levels for a plurality of hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in the plurality of immune cell types. The target regions that are differentially hypomethylated in the plurality of immune cell types exhibit lower levels of methylation in the plurality of immune cell types than in other cell types.In some embodiments, the methods disclosed herein include capturing at least a set of epigenetic target regions from DNA or a subsample thereof, the capturing step comprising contacting the DNA or a subsample thereof with target-specific probes specific for at least one of the set of epigenetic target regions, the set of epigenetic target regions comprising hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in multiple immune cell types; and determining methylation levels for the target regions. In any of these embodiments, the methylation level can be determined using partitioning, methylation-sensitive conversion, e.g., bisulfite conversion, direct detection during sequencing, or any other suitable approach. Various approaches are described herein. In any of these embodiments, additional DNA, e.g., DNA from a tumor sample, can also be analyzed (e.g., sequenced, captured, converted, and / or partitioned as described above) to obtain additional information, e.g., to quantify the contribution of immune cells to the additional DNA; to identify other cell types that contribute to the additional DNA; to detect mutations in the additional DNA; and / or to detect epigenetic differences, e.g., differential methylation, compared to healthy or normal DNA.
[0207] The methylation level of the DNA isolated from the buffy coat sample and, optionally, additional DNA can be used to determine the quantity of each of multiple immune cell types from which the DNA isolated from the buffy coat sample and, optionally, additional DNA originated. This can be useful, for example, to detect the presence of cancer or precancerous conditions, or other conditions (e.g., infection, graft rejection), because the state of the immune system, as reflected in the distribution of cell types contributing to the DNA isolated from the buffy coat sample and, optionally, additional DNA, can change as a result of such conditions. In some embodiments, the DNA isolated from the buffy coat sample and, optionally, additional DNA originates from tumor cells, and the cancer is a blood cancer. In some embodiments, the DNA isolated from the buffy coat sample and, optionally, additional DNA does not originate from tumor cells. In some such embodiments, the cancer is not a blood cancer. In some such embodiments, the cancer is a solid tumor cancer, e.g., a carcinoma, adenocarcinoma, or sarcoma. Without wishing to be bound by theory, cancers including solid tumor cancers, such as carcinomas, adenocarcinomas, and sarcomas, can cause changes in the distribution of immune cell types, including differentiated immune cell types and immune cell activation states, compared to the distribution of immune cells in healthy or non-cancer subjects. See Nabet et al. Cell. 2020, 183:363-376; Watson et al. Sci. Immunol. 2021, 6: eabj8825; Lozano et al. Nature Medicine. 2022, 28:353-362. Such changes can be detected by the methods herein and can be useful for detecting cancer and determining cancer prognosis and / or treatment options.
[0208] Some embodiments of the present disclosure include isolating DNA from a sample, such as a buffy coat sample. Other sample types containing immune and / or cancer-derived cells, such as blood samples (e.g., whole blood samples, leukoreduced samples, or peripheral blood PBMC samples), can also be used in embodiments of the disclosed methods. DNA can be isolated from cells (e.g., PBMCs) of any such sample. Such DNA isolation methods can include, but are not limited to, organic extraction (e.g., using phenol-chloroform), inorganic methods (e.g., salting out and proteinase K treatment), or adsorption methods (e.g., using silica- or cellulose-based techniques). For example, a method for isolating DNA from a sample using adsorption can include lysing the cells of the sample, separating soluble DNA in the sample from cell debris and other insoluble material, binding the DNA of interest to a purification matrix, washing proteins and other contaminants from the matrix, and eluting the DNA from the matrix.
[0209] In some embodiments, the methods disclosed herein include fragmenting DNA from a sample (e.g., DNA from a buffy coat sample and / or DNA from any other sample containing cells, such as a whole blood sample, a leukoreduced sample, or a PBMC sample). Fragmentation methods can include physical fragmentation, chemical fragmentation, and / or enzymatic fragmentation. Physical fragmentation methods can include, but are not limited to, acoustic shearing, hydrodynamic shearing (e.g., ultrasonic or point-sink shearing), needle shearing, and nebulization. Enzymatic fragmentation methods can include the use of restriction endonucleases (e.g., 4-cutter or 5-cutter restriction endonucleases, such as AluI, DpnI, Eco47I, HaeIII, HpaII, MboI, MseI, MspI, PspGI, RsaI, Sse9I, or TaqI), non-specific nucleases (e.g., micrococcal nuclease), or transposases (e.g., when insertion of adapters into the fragmented double-stranded DNA molecules is desired). DNA can also be fragmented using chemical shearing methods. Chemical fragmentation methods can include, but are not limited to, thermal digestion of DNA in the presence of divalent metal cations (e.g., magnesium or zinc).
[0210] In some embodiments, the methods disclosed herein include the steps of partitioning DNA (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a whole blood sample, a leukoreduction sample, a PBMC sample, and / or additional DNA) by contacting the DNA with an agent that recognizes modified cytosines in the DNA; sequencing the DNA; and determining the level of each of multiple immune cell types from which the DNA originated. The levels of immune cell types can be expressed, for example, as a relative amount or percentage for each quantified cell type. Such determination is exemplified, for example, in Example 3. In some embodiments, the method includes capturing or enriching a set of epigenetic target regions of DNA from one or more partitioned aliquots prior to sequencing. In some embodiments, the modified cytosines are methylcytosines.
[0211] In this way, the methods herein enable the detection and / or identification of immune-specific differentially methylated genomic regions, which can be used to identify and quantify different immune cell types from which DNA in a sample originates. Immune cell types can include immune cells of different origins, different differentiation types, different activation types, or any combination of different origins, different differentiation types, and different activation types. In fact, the differentiation state and activation state significantly overlap and often change together in a given immune cell. For example, immune cell activation can induce cell differentiation. Immune cells of different activation types include activated cells, such as cells activated by inflammatory cytokines or antigens, and suppressor cells, such as cells suppressed by Tregs. Immune cell types include macrophages (including M1 macrophages and M2 macrophages); B cells, e.g., naive or activated B cells (including regulatory B cells, memory B cells, switched memory B cells, and plasma cells); B cell precursors; T cells, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (cytotoxic T cells, regulatory T cells (Tregs)); CD4 T cells (including naive CD4 T cells, activated CD4 T cells, and CD4 effector memory T cells), CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and CD4 effector memory T cells), and CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and activated CD8 T cells). immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), neutrophils, low-density neutrophils, immature neutrophils, and immature granulocytes); dendritic cells; eosinophils; monocytes; erythrocytes; megakaryocytes; and natural killer (NK) cells.In some embodiments of the disclosed methods, the immune-specific differentially methylated genomic regions are determined from DNA of macrophages (including M1 macrophages and M2 macrophages), which are the source of DNA in the sample; B cells, e.g., naive B cells or activated B cells (including regulatory B cells, memory B cells, switched memory B cells, and plasma cells); B cell precursors; T cells, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (cytotoxic T cells, regulatory T cells (Tregs), CD4 T cells (including naive CD4 T cells, activated CD4 T cells, and CD4 effector memory T cells), CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and CD4 effector memory T cells), CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and activated CD8 T cells), and CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and activated CD8 T cells). immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), neutrophils, low-density neutrophils, immature neutrophils, and immature granulocytes); dendritic cells; eosinophils; monocytes; erythrocytes; megakaryocytes; and / or natural killer (NK) cells (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve ... at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, or each).
[0212] In some embodiments, the immunospecific differentially methylated genomic regions are selected from 1 to 40, e.g., 1 to 35, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 5, 2 to 40, 2 to 35, 2 to 30, 2 to 25, 2 to 20, 2 to 15, 2 to 10, 2 to 5, 5 to 40, of the following cell types from which the DNA in the sample originates: Sets of 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-40, 10-35, 10-30, 10-25, 10-20, or 10-15 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 1 6, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 sets) are detected and / or identified as described herein: macrophages (including M1 macrophages and M2 macrophages); B cells, e.g., naive B cells or activated B cells (including regulatory B cells, memory B cells, switched memory B cells, and plasma cells); B cell precursors; T cells, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (cytotoxic T cells, regulatory T cells (Tregs), CD4 T cells (including naive CD4 T cells, activated CD4 T cells, and CD4 effector memory T cells), CD8 T cells (including naive CD8 T cells, activated CD8 T cells, central memory T cells, and CD8 effector memory T cells); immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), neutrophils, low-density neutrophils, immature neutrophils, and immature granulocytes); dendritic cells; eosinophils; monocytes; erythrocytes; megakaryocytes; and / or natural killer (NK) cells.
[0213] In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify B cells, including memory B cells, naive B cells, or both memory and naive B cells. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify macrophages (e.g., M1 macrophages and / or M2 macrophages) and monocytes. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify dendritic cells, macrophages, and monocytes. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify myeloid cells and mature neutrophils. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify naive and activated lymphocytes. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify naive lymphocytes. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify activated lymphocytes. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify myeloid cells, neutrophils, and eosinophils. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify naive T cells, naive B cells, effector CD4 T cells (e.g., effector memory CD4 T cells and / or central memory CD4 T cells), effector CD8 T cells (e.g., effector memory CD8 T cells and / or central memory CD8 T cells), Treg cells, plasma cells, and memory cells.In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify myelocytes, neutrophils, and eosinophils. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify metamyelocytes. In some embodiments, immune-specific differentially methylated genomic regions are detected and / or identified as described herein to quantify NK cells.
[0214] In some embodiments, the ratio of immune cell types is calculated and can be used to determine the presence or absence of cancer in a subject and / or to predict the clinical outcome of a treatment (e.g., a chemotherapeutic or immunotherapeutic agent) in a subject. In some embodiments, the ratio of immune cell types is calculated from the quantities of two of the following cell types from which DNA in a sample originates: macrophages (including M1 macrophages and M2 macrophages); B cells, e.g., naive B cells or activated B cells (including regulatory B cells, memory B cells, switched memory B cells, and plasma cells); B cell precursors; T cells, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (cytotoxic T cells, regulatory T cells (Tregs), CD4 T cells (including naive CD4 T cells, activated CD4 T cells, and CD4 effector memory T cells), CD8 T cells (naive CD8 T cells, activated CD8 T cells, and activated CD8 T cells), and CD4 T cells (including naive CD8 T cells, activated CD8 T cells, and CD4 effector memory T cells). In some embodiments, the ratio is the ratio of NK cells to total lymphocytes. In some embodiments, the ratio is the ratio of neutrophils to lymphocytes. In some embodiments, the ratio is the ratio of M1 macrophages to M2 macrophages. In some embodiments, the numerator of the ratio includes the level or quantity of neutrophils, monocytes, or both neutrophils and monocytes. In some embodiments, the denominator of the ratio includes the level or quantity of T cells, B cells, NK cells, or total lymphocytes. In some embodiments, the numerator of the ratio includes the level or quantity of neutrophils and the denominator of the ratio includes the level or quantity of total lymphocytes. In some embodiments, the numerator of the ratio comprises the level or quantity of NK cells and the denominator of the ratio comprises the level or quantity of total lymphocytes, hi some embodiments, the numerator of the ratio comprises the level or quantity of M1 macrophages and the denominator of the ratio comprises the level or quantity of M2 macrophages.In some embodiments, the numerator of the ratio comprises the level or quantity of monocytes and the denominator of the ratio comprises the level or quantity of T cells.
[0215] In some embodiments, an increase in such a ratio is associated with cancer. In other embodiments, a decrease in such a ratio is associated with cancer. In some embodiments, an increase in such a ratio is associated with a better clinical outcome (e.g., longer survival) of a treatment (e.g., a chemotherapeutic or immunotherapeutic agent) in a subject. In other embodiments, a decrease in such a ratio is associated with a better clinical outcome (e.g., longer survival) of a treatment (e.g., a chemotherapeutic or immunotherapeutic agent) in a subject.
[0216] DNA from such cell types may be rare in samples such as DNA from buffy coat samples, whole blood samples, leukapheresis samples, PBMC samples, and / or additional DNA samples from healthy individuals, but may be abundant in such samples from individuals with a disease or disorder, such as cancer or a precancerous condition. In some embodiments, both hypermethylated and hypomethylated regions may be detected to distinguish DNA from closely related cell types, such as naive and activated B cells, naive and activated T cells, or different stages of myeloid lineage. In some embodiments, at least some of the differentially methylated regions are exclusively hypermethylated or exclusively hypomethylated in only one cell type or in only one cell type within a cluster. In some embodiments, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten differentially methylated regions are exclusively hypermethylated or exclusively hypomethylated in only one cell type to be identified or quantified within a cluster.
[0217] In some embodiments, determining the levels of different immune cell types from which DNA in a sample originates facilitates disease diagnosis or identification of appropriate treatments. In some embodiments, changes in the levels of one or more immune cell types indicate the presence of a disease or disorder in the subject, such as cancer, precancerous conditions, infection, transplant rejection, or other disorders that cause a change in the relative amount of a particular immune cell type compared to the amount present in healthy subjects. In some embodiments, changes in the levels of both one or more immune cell types in combination with sequence-independent changes in epigenetic target regions indicate the presence of a disease or disorder in the subject, such as cancer, precancerous conditions, infection, transplant rejection, or other disorders that cause a change in the relative amount of a particular immune cell type compared to healthy subjects and epigenetic changes. In some embodiments, the method facilitates identification of appropriate treatments based on the likelihood that the subject will respond to the treatment. In some such embodiments, determining the level of DNA from one or more immune cell types in a sample from a subject with a particular type of cancer facilitates prediction of the clinical outcome of immunotherapy in a subject. When determining the levels of different immune cell types facilitates disease diagnosis, identification of appropriate treatment, and association with therapeutic response, the threshold for diagnosing disease and the threshold for identifying appropriate treatment can be the same or different. The levels can be determined based on the count of molecules corresponding to different immune cell types, or the relative frequency of such molecules based on the count of molecules corresponding to one or more different immune cell types, or any value or ratio.
[0218] 1. Determine B and / or T cell numbers using the CDR3 region Complementarity-determining region 3 (CDR3) is the most hypervariable region in B cell receptor (BCR) and T cell receptor (TCR) genes. The CDR3s of the BCR heavy and light chains and the TCR alpha and beta chains undergo recombination during B cell and T cell development, respectively; therefore, recombined CDR3 sequences can be easily distinguished from their germline counterparts by sequencing. In some embodiments of the disclosed methods, sequencing of the CDR3 region of the B cell receptor or T cell receptor can be used as an alternative to or in addition to methylation analysis to quantify the B cell and / or T cell origin of DNA (e.g., DNA from a buffy coat sample or DNA from cells of the sample).
[0219] VDJ recombination (or VJ recombination) is a process in which T cells and B cells randomly assemble different gene segments (i.e., variable (V), diversity (D), and joining (J) genes) to generate unique receptors that can collectively recognize many different types of molecules. The variable region contains three highly diverse loops (CDR1, 2, and 3) that contact antigens. The CDR1 and CDR2 loops are encoded within the V gene segment of the variable region gene, while the CDR3 loop is encoded by the region spanning the V and J or V, D, and J junctions. Therefore, compared with the CDR1 and CDR2 loops, the CDR3 loop is more diverse due to the addition and / or loss of nucleotides during the recombination process. Roth, Microbiol Spectr. 2014, 2(6); Hughes et al., Eur. J. Immunol. 2003, 33:1568-1575; Freeman et al., Genome Res. 2009, 19:1817-1824. Modified B cell-specific CDR3 regions can be identified in B cell receptor-encoding genes, and modified T cell-specific CDR3 regions can be identified in T cell receptor-encoding genes. Because these modified CDR3 regions are specific to B cells or T cells and are not present in other cell types, sequencing of the CDR3 regions in a DNA sample can be used to determine the number of B cells and / or T cells from which the DNA in the sample originated. In some embodiments, the number of B cells from which the DNA in a sample originated can be determined by comparing the number of modified B cell-specific CDR3 regions in the DNA with the number of unmodified CDR3 regions. In some embodiments, the number of T cells from which the DNA in a sample originates can be determined by comparing the number of recombined T-cell-specific CDR3 regions in the DNA with the number of unrecombined CDR3 regions, which indicate cells that are neither B cells nor T cells.
[0220] Thus, some embodiments of the disclosed methods include determining the number of B cells, T cells, or both B and T cells that are the source of DNA (e.g., DNA from a buffy coat sample, or DNA from cells of the sample) based on the CDR3 sequence within the DNA. Such embodiments include sequencing the CDR3 region within the DNA. In some embodiments, determining the number of B cells, T cells, or both B and T cells that are the source of DNA includes determining the number of recombined and unrecombined CDR3 regions within the DNA. In some embodiments, the CDR3 region is a CDR3 region of a B cell receptor chain or a T cell receptor chain. For example, determining the number of B cells from the DNA may include sequencing the CDR3 region within the DNA and determining (1) the number of recombined CDR3 regions and (2) the number of unrecombined CDR3 regions corresponding to the B cells. As another example, determining the number of T cells in the DNA can include sequencing the CDR3 regions in the DNA and determining (1) the number of recombined CDR3 regions and (2) the number of unrecombined CDR3 regions corresponding to the T cells.
[0221] In some embodiments, the B cell receptor chain is a heavy chain. In some embodiments, the B cell receptor is a light chain. Similarly, in some embodiments, the T cell receptor chain is an alpha chain. In some embodiments, the T cell receptor chain is a beta chain.
[0222] Some embodiments that include determining the quantity of B cells, T cells, or both B cells and T cells from which the DNA (e.g., DNA from a buffy coat sample, or DNA from cells of the sample) originates based on the CDR3 sequence further include: a) sequencing the DNA, or a subsample thereof, to determine methylation levels for a set of epigenetic target regions, the set including multiple target regions comprising DNA sequences that are differentially methylated in multiple immune cell types, wherein the DNA is from the buffy coat sample or the DNA is from cells of the sample; and b) determining the quantity of each of the multiple immune cell types from which the DNA originates based on the methylation levels.
[0223] Some embodiments that involve determining the quantity of B cells, T cells, or both B cells and T cells from which DNA (e.g., DNA from a buffy coat sample, or DNA from cells of the sample) originates based on CDR3 sequences further include: a) prior to the sequencing step, capturing at least a set of epigenetic target regions from the DNA or a subsample thereof, comprising contacting the DNA or subsample thereof with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the set of epigenetic target regions includes target regions that comprise DNA sequences that are differentially methylated in a plurality of immune cell types; b) determining the methylation level for the target regions; and c) determining the quantity of each of the plurality of immune cell types from which the DNA originates. In such embodiments, prior to the capturing step, the DNA can be separated into two or more aliquots, with a first aliquot being used to determine the quantity of B cells, T cells, or both B and T cells from which the DNA (e.g., DNA from the buffy coat sample or DNA from the cells of the sample) is derived based on the CDR3 sequences, and a second aliquot being used in the capturing step.
[0224] Some embodiments that include determining the quantity of B cells, T cells, or both B and T cells from which DNA (e.g., DNA from a buffy coat sample, or DNA from cells of the sample) originates based on CDR3 sequences further include: a) prior to the sequencing step, partitioning the DNA into multiple aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot; b) sequencing the DNA from one or more of the multiple aliquots; and c) detecting the level of the DNA sequences to determine the quantity of each of multiple immune cell types from which the DNA originated. In such embodiments, prior to the distributing step, the DNA can be separated into two or more aliquots, with a first aliquot being used to determine the quantity of B cells, T cells, or both B and T cells from which the DNA (e.g., DNA from the buffy coat sample, or DNA from the cells of the sample) originates based on the CDR3 sequences, and a second aliquot being used in the distributing step.
[0225] Certain embodiments that involve determining the quantity of B cells, T cells, or both B and T cells from which DNA (e.g., DNA from a buffy coat sample or DNA from cells of the sample) originates based on CDR3 sequences include: a) partitioning the DNA into multiple aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with modified cytosines than the second aliquot, and wherein the DNA is from the buffy coat sample or the DNA is from cells of the sample; The method further comprises the steps of: b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second aliquots, thereby obtaining captured DNA, the step comprising contacting the DNA with target-specific probes specific for at least one of the set of epigenetic target regions, wherein the target regions of the set of epigenetic target regions comprise DNA sequences that are differentially methylated in a plurality of immune cell types; and c) sequencing the captured DNA to determine the levels of each of the plurality of immune cell types from which the DNA originated. In such embodiments, prior to the partitioning step, the DNA can be separated into two or more aliquots, with a first aliquot being used to determine the quantity of B cells, T cells, or both B and T cells from which the DNA (e.g., DNA from the buffy coat sample, or DNA from cells of the sample) originated based on the CDR3 sequences, and a second aliquot being used in the partitioning step and subsequent capturing step.
[0226] In certain embodiments, the target regions of the set of epigenetic target regions comprise DNA sequences that are hypomethylated in multiple immune cell types. In another specific embodiment, the target regions of the set of epigenetic target regions comprise DNA sequences that are hypermethylated in multiple immune cell types.
[0227] B. Divide the sample into multiple aliquots In some embodiments described herein, different forms of DNA (e.g., hypermethylated DNA and hypomethylated DNA) are physically partitioned based on one or more characteristics of the DNA. This approach can be used, for example, to determine whether a particular site or region is hypermethylated or hypomethylated. Partitioning can be performed, for example, before adapters are attached to DNA molecules in a sample, facilitating the inclusion of partitioning tags on the adapters. The partitioning tag can be used to identify which partition a molecule is found in. After partitioning (and adapter attachment, if applicable), further steps such as amplification, target capture, and sequencing can be performed. Partitioning can be performed on DNA from a buffy coat sample or any other sample containing cells, such as a whole blood sample, a leukoreduction sample, a PBMC sample, and / or additional DNA, such as tumor cells.
[0228] Methylation profiling can include determining the methylation pattern across different regions of genome.For example, after the further steps discussed above, including the distribution and sequencing of molecules based on the degree of methylation (for example, the relative number of methylated nucleic acid bases per molecule), the sequences of molecules in different distributions can be mapped to a reference genome.This can indicate the regions of genome that are more highly methylated or less highly methylated compared to other regions.In this way, genome regions can have different degrees of methylation compared to individual molecules.
[0229] By partitioning nucleic acid molecules in a sample, for example, rare nucleic acid molecules that are predominant in one partition of the sample can be enriched to increase rare signals. For example, genetic variations that are present in hypermethylated DNA but present less (or not at all) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple partitions of a sample, multidimensional analysis of single molecules can be performed, thus achieving greater sensitivity. Partitioning can include physically dividing nucleic acid molecules into partitions or subsamples based on the presence or absence of one or more methylated nucleic acid bases. Samples can be partitioned into partitions or subsamples based on features indicative of differential gene expression or disease state. During the analysis of nucleic acids, such as tumor DNA or circulating tumor DNA (ctDNA), samples can be partitioned based on features or combinations thereof that result in differences in signals between normal and diseased states.
[0230] In some embodiments, hypermethylated and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit differential methylation signatures of particular immune cell types, such as rare immune cell types, tumor cells, or cell types that do not normally contribute to the DNA sample being analyzed.
[0231] In some cases, the heterogeneous DNA in a sample is divided into two or more partitions (for example, at least three, four, five, six or seven partitions).In some embodiments, each partition is differentially tagged.The tagged partitions can then be pooled together for collective sample preparation and / or sequencing.The partition-tagging-pooling step can be carried out more than once, and each partition is tagged using a differential tag that is based on one or more different features (for example, in the examples provided herein) and is distinguished from other partitions and partitioning means.In other cases, the differentially tagged partitions are sequenced separately.
[0232] In some embodiments, sequence reads are obtained from differentially tagged and pooled DNA and analyzed in silico. Tags can be used to distinguish between reads from different distributions. Analysis to detect genetic variants can be performed at the distribution level as well as at the total nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variants, such as CNVs, SNVs, indels, and fusions, in the nucleic acids in each distribution. In some cases, in silico analysis can include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the position of nucleosomes in chromatin. Higher coverage can be correlated with higher nucleosome occupancy in a genomic region, while lower coverage can be correlated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0233] In some embodiments, the partitioning is based on one or more characteristics, such as methylation. Where applicable, as part of the data analysis or partitioning, appropriate techniques can be used to sort molecules according to other characteristics, such as sequence length, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins binding to DNA. The resulting partitioning can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, partitioning based on cytosine modification (e.g., cytosine methylation) or methylation is typically performed, optionally combined with at least one additional partitioning step that can be based on any of the aforementioned DNA characteristics or forms. In some embodiments, a heterogeneous nucleic acid population is partitioned into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (for example, 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation), and the association and level of association with one or more proteins, such as histones.Alternatively or additionally, heterogeneous nucleic acid populations can be divided into nucleic acid molecules associated with nucleosomes and nucleic acid molecules that lack nucleosomes.Alternatively or additionally, heterogeneous nucleic acid populations can be divided into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA).Alternatively or additionally, heterogeneous nucleic acid populations can be divided based on nucleic acid length (for example, molecules that are up to 160 bp and molecules that have a length longer than 160 bp).
[0234] The agent used to partition the population of nucleic acids within a sample can be an affinity agent, such as an antibody with a desired specificity, a natural binding partner or variant thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected to have specificity for a given target, for example, by phage display. In some embodiments, the agent used for partitioning is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine, e.g., a methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is the product of a procedure that affects a first nucleobase in the DNA of the sample differently from a second nucleobase in the DNA. In some embodiments, the modified nucleobase can be a "converted nucleobase," i.e., one whose base-pairing specificity has been altered by a procedure. For example, in certain procedures, unmethylated or unmodified cytosine is converted to dihydrouracil, or more commonly, at least one modified or unmodified form of cytosine undergoes deamination, thereby generating uracil (considered a modified nucleobase in DNA) or a further modified form of uracil. Examples of partitioning agents include antibodies, for example, antibodies that recognize modified nucleobases, such as methylcytosine (e.g., 5-methylcytosine). In some embodiments, the partitioning agent is an antibody that recognizes modified cytosines other than 5-methylcytosine, such as 5-carboxylcytosine (5caC). Alternative partitioning agents include the methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including proteins such as MeCP2.
[0235] Additional non-limiting examples of partitioning agents are histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.
[0236] The binding of the partitioning agent to specific nucleic acids and the partitioning of nucleic acids into subsamples can be carried out to a certain extent, or can be carried out essentially in a binary manner.In some cases, nucleic acids containing a greater proportion of a particular modification bind to the agent to a greater extent than nucleic acids containing a smaller proportion of the modification.Similarly, partitioning can produce subsamples containing a greater proportion and a smaller proportion of nucleic acids containing a particular modification.Alternatively, partitioning can produce subsamples containing essentially all or none of the nucleic acids containing the modification.In all cases, various levels of modification can be sequentially eluted from the partitioning agent.
[0237] In some embodiments, partitioning can include both binary partitioning and partitioning based on the degree / level of modification. For example, methylated fragments can be partitioned by methylated DNA immunoprecipitation (MeDIP), or all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, further partitioning can include eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and bound fragments. As the salt concentration increases, fragments with higher methylation levels are eluted.
[0238] In some cases, the final distribution is enriched for nucleic acids with different degrees of modification (over- or under-representation of the modification). Over- and under-representation can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is two, nucleic acids containing more than two 5-methylcytosine residues are over-represented in this modification, and nucleic acids with one or zero 5-methylcytosine residues are under-represented. The effect of affinity separation is to enrich nucleic acids with over-represented modifications in the binding phase and under-represented modifications in the non-binding phase (i.e., in solution). Nucleic acids in the binding phase can be eluted prior to subsequent processing.
[0239] When using MeDIP or the MethylMiner® Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be separated using sequential elution. For example, the low-methylation (unmethylated) distribution can be separated from the methylated distribution by contacting the nucleic acid population with MBD from the kit attached to magnetic beads. The beads are used to separate the methylated nucleic acids from the unmethylated nucleic acids. Subsequently, one or more sequential elution steps are performed to elute nucleic acids with different levels of methylation. For example, the first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, for example, at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After eluting such methylated nucleic acids, magnetic separation is again used to separate nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps can be repeated to create various partitions, such as a low-methylated partition (enriched for nucleic acids with no methylation), a methylated partition (enriched for nucleic acids with low levels of methylation), and a high-methylated partition (enriched for nucleic acids with high levels of methylation).
[0240] In some methods, nucleic acids bound to the agent used for affinity separation-based partitioning are subjected to a wash step. The wash step washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched for nucleic acids that have modifications at a level close to the average or median (i.e., halfway between the nucleic acids that remain bound to the solid phase when the sample is first contacted with the agent and the nucleic acids that are not bound to the solid phase).
[0241] Affinity separation results in at least two, sometimes three or more, distributions of nucleic acids with different degrees of modification. The distributions are still separate, but at least one distribution of nucleic acid, and usually two or three (or more) distributions, is linked to a nucleic acid tag, usually provided as an adapter component, and the nucleic acids in different distributions receive different tags that distinguish one distribution from another. The tags linked to nucleic acid molecules of the same distribution can be the same or different from each other. However, if different from each other, the tags can have a common code portion so that the molecules to which they are attached are identified as belonging to a specific distribution.
[0242] For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0243] In some embodiments, nucleic acid molecules may be fractionated into different partitions based on which nucleic acid molecules are bound to a particular protein or fragment thereof and which are not bound to that particular protein or fragment thereof.
[0244] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific protein properties. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that can bind to DNA and serve as a basis for fractionation include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate nucleic acid molecules based on protein-bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein-bound regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric flow field separation (AF4).
[0245] In some embodiments, the sample is divided into multiple aliquots by contacting the nucleic acid with an antibody that recognizes a modified nucleic acid base in DNA, and the modified nucleic acid base can be a modified cytosine or the product of a procedure that affects a first nucleic acid base in DNA differently from a second nucleic acid base in the DNA of the sample. In some embodiments, the modified nucleic acid base is 5mC. In some embodiments, the modified nucleic acid base is 5caC. In some embodiments, the modified nucleic acid base is dihydrouracil (DHU). In some embodiments, the single-stranded DNA is divided using an antibody that recognizes a modified nucleic acid base in DNA.
[0246] In some embodiments, partitioning is performed by contacting the nucleic acid with the methyl-binding domain ("MBD") of a methyl-binding protein ("MBP"). In some such embodiments, the nucleic acid is contacted with the entire MBP. In some embodiments, the MBD binds to 5-methylcytosine (5mC), and the MBP comprises an MBD, referred to herein interchangeably as a methyl-binding protein or a methyl-binding domain protein. In some embodiments, the MBD binds to 5mC and 5hmC. In some embodiments, the MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 streptavidin, via a biotin linker. Partitioning into fractions with different degrees of methylation can be performed by eluting the fractions with increasing NaCl concentration.
[0247] In some embodiments, the bound DNA is eluted by contacting the antibody or MBD with a protease, such as proteinase K. This can be performed instead of or in addition to the elution step using NaCl discussed above.
[0248] Examples of agents that recognize modified nucleobases contemplated herein include, but are not limited to, the following:
[0249] (a) MeCP2 is a protein that preferentially binds 5-methyl-cytosine over unmodified cytosine.
[0250] (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 bind preferentially to 5-hydroxymethyl-cytosine over unmodified cytosine.
[0251] (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind 5-formylcytosine over unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).
[0252] (d) An antibody specific for one or more methylated or modified nucleobases or their conversion products, e.g., 5mC, 5caC, or DHU.
[0253] Generally, elution is a function of the number of modifications (e.g., the number of methylation sites) per molecule, with higher salt concentrations resulting in more methylated molecules. A series of elution buffers with increasing NaCl concentrations can be used to separate DNA into distinct populations based on the degree of methylation. Salt concentrations can range from about 100 mM to about 2500 mM NaCl. In one embodiment, the process results in three partitions. The molecules are contacted with a solution containing molecules at a first salt concentration, including an agent that recognizes modified nucleobases, and the molecules may be bound to a capture moiety, e.g., streptavidin. At the first salt concentration, one population of molecules binds to the agent, while another population remains unbound. The unbound population can be separated as a "hypomethylated" population. For example, the first partition enriches for hypomethylated forms of DNA that remain unbound at low salt concentrations, e.g., 100 mM or 160 mM. A second partition enriched for intermediately methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM. This partition is also separated from the sample. A third partition enriched for highly methylated forms of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.
[0254] In some embodiments, methylated DNA is purified using a monoclonal antibody raised against 5-methylcytidine (5mC). To obtain single-stranded DNA fragments, the DNA is denatured, for example, at 95°C. The DNA bound to the antibody is immunoprecipitated using standard or magnetic bead-coupled protein G and washed after incubation with the anti-5mC antibody. The DNA can then be eluted. The partition may include unprecipitated DNA and one or more partitions eluted from the beads.
[0255] In some embodiments, sample DNA (e.g., between 5 and 200 ng) is mixed with a methyl-binding domain (MBD) buffer, conjugated to magnetic beads with MBD protein, and incubated overnight. Methylated DNA (hypermethylated DNA) binds to the MBD protein on the magnetic beads during this incubation. Unmethylated (hypomethylated DNA) or less methylated DNA (intermediate methylation) is washed from the beads with buffers containing increasing concentrations of salt. For example, one, two, or more fractions containing unmethylated, hypomethylated, and / or intermediate methylated DNA can be obtained from such washes. Finally, highly methylated DNA (hypermethylated DNA) is eluted from the MBD protein using a high-salt buffer. In some embodiments, these washes result in three partitions of DNA with increasing methylation levels: a hypomethylated partition, an intermediate methylated fraction, and a hypermethylated partition.
[0256] In some embodiments, the partitioning procedure may result in incomplete sorting of DNA molecules among the subsamples. For example, a small number of molecules in an unmethylated or hypomethylated subsample may be highly modified (e.g., hypermethylated), and / or a small number of molecules in a hypermethylated subsample may be unmodified or mostly unmodified (e.g., unmethylated or mostly unmethylated). Such molecules are considered to be non-specifically partitioned.
[0257] In some embodiments, non-specifically partitioned molecules are removed using a methylation-dependent nuclease, e.g., a methylation-dependent restriction enzyme (MDRE), which digests / cleaves DNA whose restriction enzyme (RE) recognition site contains a methylated nucleotide but does not cleave DNA whose restriction enzyme (RE) recognition site contains an unmethylated nucleotide. In some embodiments, non-specifically partitioned molecules are removed using a methylation-sensitive nuclease, e.g., a methylation-sensitive restriction enzyme (MSRE), which digests / cleaves DNA whose restriction enzyme (RE) recognition site contains an unmethylated nucleotide but does not cleave DNA whose restriction enzyme (RE) recognition site contains a methylated nucleotide. For example, in some embodiments, a low-methylation aliquot is contacted with a methylation-dependent nuclease, e.g., a methylation-dependent restriction enzyme, thereby degrading non-specifically partitioned DNA, e.g., methylated DNA, in the aliquot. Alternatively, or in addition, a high-methylation aliquot is contacted with a methylation-sensitive endonuclease, e.g., a methylation-sensitive restriction enzyme, thereby degrading non-specifically partitioned DNA in the aliquot.
[0258] The degradation of non-specifically distributed DNA in one or more divided aliquots can improve the performance of the method that relies on the accurate distribution of DNA based on cytosine modification.For example, such degradation can improve sensitivity and / or simplify downstream analysis.In some embodiments, DNA is divided based on modification such as methylation, and then non-specifically distributed DNA is removed using MDRE and / or MSRE as described herein, thereby improving efficiency and / or cost compared with the DNA analysis method that comprises a procedure that affects first nucleobase differently from second nucleobase, such as bisulfite sequencing or bisulfite conversion.
[0259] In some embodiments, one or more nucleases are used to degrade non-specifically distributed DNA molecules. In some embodiments, the aliquot is contacted with multiple nucleases. The aliquot can be contacted with the nucleases sequentially or simultaneously. Simultaneous use of nucleases can be advantageous to avoid unnecessary sample manipulation if the nucleases are active under similar conditions (e.g., buffer composition). Contacting the aliquot with more than one methylation-dependent restriction enzyme can more thoroughly degrade non-specifically distributed hypermethylated DNA. Contacting the aliquot with more than one methylation-sensitive restriction enzyme can more thoroughly degrade non-specifically distributed hypomethylated and / or unmethylated DNA.
[0260] In some embodiments, the methylation-dependent nuclease comprises one or more of MspJI, LpnPI, FspEI, or McrBC. In some embodiments, at least two methylation-dependent nucleases are used. In some embodiments, at least three methylation-dependent nucleases are used.
[0261] In some embodiments, the methylation-sensitive nucleases include one or more of AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Pspl406I, PvuI, SacII, SalI, SmaI, and SnaBI. In some embodiments, at least two methylation-sensitive nucleases are used. In some embodiments, at least three methylation-sensitive nucleases are used. In some embodiments, the methylation-sensitive nuclease includes BstUI and HpaII. In some embodiments, the two methylation-sensitive nucleases include HhaI and AccII. In some embodiments, the methylation-sensitive nucleases include BstUI, HpaII, and Hin6I.
[0262] In some embodiments, the DNA fraction is desalted and concentrated in preparation for the enzymatic steps of library preparation.
[0263] C. Adapter Ligation or Addition In some embodiments, adapters are added to DNA. This can be done in conjunction with an amplification procedure, for example, by providing adapters to the 5' portion of primers (when PCR is used, this can be referred to as library prep-PCR or LP-PCR). In some embodiments, adapters are added by other approaches, such as ligation. In some such methods, a first adapter is added to a nucleic acid by ligation to its 3' end before partitioning or capture, which can include ligation to single-stranded DNA. The adapter can be used as a priming site for second-strand synthesis, for example, using a universal primer and DNA polymerase. A second adapter can then be ligated to the 3' end of at least the second strand of the now double-stranded molecule. In some embodiments, the first adapter includes an affinity tag, such as biotin, and the nucleic acid ligated to the first adapter is attached to a solid support (e.g., beads), which can include a binding partner for the affinity tag, such as streptavidin. For further discussion of related procedures, see Gansauge et al., Nature Protocols 8: 737-748 (2013). Commercially available kits for preparing sequencing libraries compatible with single-stranded nucleic acids are available, such as the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adapter ligation, the nucleic acid is amplified.
[0264] Preferably, the adapters contain a sufficient number of different tags such that the number of tag combinations results in a low probability that two nucleic acids with the same start and end points will receive the same tag combination, e.g., less than or equal to 5%, e.g., 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, or less than 0.1%. Adapters, whether with the same or different tags, may contain the same or different primer binding sites, although preferably the adapters contain the same primer binding sites.
[0265] In some embodiments, after binding of the adapters, the nucleic acids are subjected to amplification, which can be, for example, using universal primers that recognize primer binding sites in the adapters.
[0266] In some embodiments, after the adapter is attached, the DNA is partitioned, which includes contacting the DNA with an agent that preferentially binds to nucleic acids with epigenetic modifications. The nucleic acid is then partitioned into at least two subsamples that differ in the degree to which the nucleic acid has the modification after binding with the agent. For example, if the agent has affinity for nucleic acids with the modification, nucleic acids with over-represented modifications (compared to the median representation in the population) will preferentially bind to the agent, while nucleic acids with under-represented modifications will not bind to the agent or will be more easily eluted from the agent. The nucleic acid can then be amplified using primers that bind to the primer binding sites in the adapter. Alternatively, partitioning can be performed before adapter attachment, in which case the adapter can include a differential tag containing a component that identifies which partition the molecule is present in.
[0267] In some embodiments, the nucleic acid is ligated at both ends to a Y-shaped adapter that comprises a primer binding site and a tag. The molecule is amplified.
[0268] For example, in some embodiments, where DNA has been subjected to a denaturing treatment, such as bisulfite conversion, immunoprecipitation using an anti-methylated cytosine antibody that recognizes 5-mC in single-stranded DNA, or any other treatment that renders some or all of the DNA single-stranded, library preparation procedures suitable for samples containing single-stranded DNA can be used. See, e.g., Gansauge & Meyer, Nature Protocols 8, 737-748 (2013). For example, a biotinylated adapter can be ligated to the 3' end of the DNA, followed by immobilization on streptavidin-coated beads. A primer that anneals to the adapter and DNA polymerase can be used to synthesize a complementary strand. A second adapter can then be attached to the now double-stranded molecule, for example, by blunt-end ligation. The biotinylated adapter can include any embodiment of the tags and / or barcodes described elsewhere herein. The molecules can then be amplified, for example, by PCR, and the amplified products can be sequenced.
[0269] D. Tagging "Tagging" a DNA molecule is the procedure of attaching or associating a tag to a DNA molecule. The tag can be a molecule, e.g., a nucleic acid, that contains information that indicates the characteristics of the molecule to which the tag is associated. For example, a molecule can have a sample tag (that distinguishes a molecule in one sample from molecules in a different sample) or a molecular tag / molecular barcode / barcode (that distinguishes different molecules from each other (in both unique and non-unique tagging scenarios)). For methods that involve a partitioning step, a partitioning tag (that distinguishes a molecule in one partition from molecules in a different partition) can be included. In some embodiments, an adapter containing a tag is added to a DNA molecule, e.g., DNA from a buffy coat sample, a whole blood sample, a leukoreduction sample, a PBMC sample, and / or additional DNA. In certain embodiments, a tag can comprise a barcode or a combination of barcodes. As used herein, the term "barcode" can refer to a specific nucleotide sequence, depending on the context. The term "barcode" refers to a nucleic acid molecule having a nucleotide sequence, or the nucleotide sequence itself. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or sequences with a certain Hamming distance, as desired for a specific purpose. Thus, for example, a molecular barcode can be composed of one barcode or a combination of two barcodes, each attached to a different end of the molecule. Additionally or alternatively, different sets of molecular barcodes, or molecular tags, can be used for different distributions and / or samples, such that the barcodes act as molecular tags through their individual sequences and also serve to identify corresponding distributions and / or samples based on the sets they are members of.
[0270] In some embodiments, two or more partitions, e.g., each partition, are differentially tagged. In some embodiments, the partitions contain DNA from a buffy coat sample or any other sample containing cells, e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample. In some embodiments, the partitions contain additional DNA, e.g., DNA from a tumor sample. Tags can be used to label individual polynucleotide population partitions to correlate the tag(s) with a specific partition. Alternatively, tags can be used in embodiments that do not employ a partitioning step. In some embodiments, a single tag can be used to label a specific partition. In some embodiments, multiple different tags can be used to label a specific partition. In embodiments in which multiple tags are used to label a specific partition, the set of tags used to label one partition can be easily distinguished from the set of tags used to label other partitions. In some embodiments, tags may have additional functions, for example, tags may be used to index the origin of a sample or as unique molecular identifiers (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations, e.g., as in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011); Kou et al., PloS ONE, 11: e0146638 (2016)), or as non-unique molecular identifiers, e.g., as described in U.S. Pat. No. 9,598,731. Similarly, in some embodiments, tags may have additional functions, for example, tags may be used to index the origin of a sample or as non-unique molecular identifiers (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations).
[0271] In some embodiments, distribution tagging involves tagging molecules in each distribution with a distribution tag. After recombining the distributions (e.g., to reduce the number of required sequencing runs and avoid unnecessary costs) and sequencing the molecules, the distribution tag identifies the distribution of origin. In some embodiments, the distribution tag can serve as an identifier for the distribution of origin and the molecules; i.e., different distributions are tagged with different sets of molecular tags, e.g., consisting of barcode pairs. In this way, one or more molecular barcodes attached to molecules not only indicate the distribution of origin but also serve to identify molecules within the distribution. For example, a first set of 35 barcodes can be used to tag molecules in a first distribution, while a second set of 35 barcodes can be used to tag molecules in a second distribution.
[0272] In some embodiments, after partitioning and tagging with partition tags, the molecules can be pooled for single sequencing. In some embodiments, sample tags are added to the molecules, for example, in a step after partition tag addition and pooling. Sample tags can facilitate pooling materials generated from multiple samples for single sequencing.
[0273] Alternatively, in some embodiments, the distribution tag may be correlated with the sample and distribution. As a simple example, a first tag may indicate a first distribution of a first sample, a second tag may indicate a second distribution of the first sample, a third tag may indicate a first distribution of a second sample, and a fourth tag may indicate a second distribution of the second sample.
[0274] Tags may be attached to molecules that have already been distributed based on one or more characteristics, but the final tagged molecules in the library may no longer have those characteristics. For example, single-stranded DNA molecules may be distributed and tagged, but the final tagged molecules in the library are likely to be double-stranded. Similarly, DNA may be distributed based on different levels of methylation, but the tagged molecules derived from these molecules in the final library are likely to be unmethylated. Thus, tags attached to molecules in the library typically represent the characteristics of the "parent molecules" from which the final tagged molecules are derived, not necessarily the characteristics of the tagged molecules themselves.
[0275] For example, molecules in a first distribution are tagged and labeled using barcodes 1, 2, 3, 4, etc., molecules in a second distribution are tagged and labeled using barcodes A, B, C, D, etc., molecules in a third distribution are tagged and labeled using barcodes a, b, c, d, etc. Differentially tagged distributions can be pooled before sequencing. Differentially tagged distributions can be sequenced separately or can be sequenced together simultaneously, for example, in the same flow cell of an Illumina sequencer.
[0276] Tags, including barcodes, can be incorporated into or otherwise joined to adapters. Tags can be incorporated by ligation, overlap extension PCR, among other methods.
[0277] 1. Molecular tagging strategies Molecular tagging refers to a tagging method that allows distinguishing between DNA molecules from which sequence reads originate. Tagging strategies can be divided into unique tagging and non-unique tagging. In unique tagging, all or substantially all molecules in a sample have different tags, so that read data can be assigned to the original molecule based on tag information alone. The tags used in such methods may be referred to as "unique tags." In non-unique tagging, different molecules in the same sample may have the same tag, so that sequence read data is assigned to the original molecule using other information in addition to tag information. Such information may include start and end coordinates, coordinates to which the molecule is mapped, start or end coordinates alone, etc. The tags used in such methods may be referred to as "non-unique tags." Therefore, it is not necessary to uniquely tag all molecules in a sample. It is sufficient to uniquely tag molecules contained in an identifiable class within a sample. Therefore, molecules in different identifiable families can have the same tag without losing information about the identity of the tagged molecule.
[0278] In certain embodiments of non-unique tagging, the number of different tags used may be sufficient if there is a very high probability (e.g., at least 99%, at least 99.9%, at least 99.99%, or at least 99.999%) that all DNA molecules in a particular group have different tags. Note that when barcodes are used as tags, and when barcodes are, for example, randomly attached to both ends of molecules, a combination of barcodes together may constitute a tag. This number is a function of the number of molecules included in the call unambiguously. For example, a class may be all molecules that map to the same start-end position on a reference genome. A class may be all molecules that map to a particular locus, for example, a particular base or a particular region (e.g., up to 100 bases or a gene or exon of a gene). In certain embodiments, the number of different tags used to uniquely identify the number z of molecules in a class is 2* z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z, or 100 * Any of z (e.g., the lower bound) and 100,000 * z, 10,000 * z, 1000 * z, or 100 * z (e.g., upper limit).
[0279] For example, in a sample of about 5 ng to 30 ng of DNA, approximately 3,000 molecules are expected to map to a particular nucleotide coordinate, with approximately 3 to 10 molecules with any given start coordinate expected to share the same end coordinate. Therefore, approximately 50 to approximately 50,000 different tags (e.g., approximately 6 to 220 barcode combinations) may be sufficient to uniquely tag all such molecules. To uniquely tag all 3,000 molecules mapped across all nucleotide coordinates, approximately 1 million to approximately 20 million different tags would be required.
[0280] Generally, the assignment of unique or non-unique tag barcodes in the reaction follows the methods and systems described by U.S. Patent Application Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731. Tags can be linked to sample nucleic acids randomly or non-randomly.
[0281] Unique tags can be loaded such that greater than about 1, greater than 2, greater than 3, greater than 4, greater than 5, greater than 6, greater than 7, greater than 8, greater than 9, greater than 10, greater than 20, greater than 50, greater than 100, greater than 500, greater than 1000, greater than 5000, greater than 10000, greater than 50,000, greater than 100,000, greater than 500,000, greater than 1,000,000, greater than 10,000,000, greater than 50,000,000, or greater than 1,000,000,000 unique tags are loaded per genomic sample. In some cases, unique tags may be loaded such that less than about 2, less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 20, less than 50, less than 100, less than 500, less than 1000, less than 5000, less than 10000, less than 50,000, less than 100,000, less than 500,000, less than 1,000,000, less than 10,000,000, less than 50,000,000, or less than 1,000,000,000 unique tags are loaded per genomic sample.In some cases, the average number of unique tags loaded per sample genome is about less than 1 or more, less than 2 or more, less than 3 or more, less than 4 or more, less than 5 or more, less than 6 or more, less than 7 or more, less than 8 or more, less than 9 or more, less than 10 or more, less than 20 or more, less than 50 or more, less than 100 or more per genome sample. , less than 500 or more, less than 1000 or more, less than 5000 or more, less than 10000 or more, less than 50,000 or more, less than 100,000 or more, less than 500,000 or more, less than 1,000,000 or more, less than 10,000,000, less than 50,000,000, or less than 1,000,000,000 unique tags.
[0282] A preferred format uses 20-50 different tags (e.g., barcodes) ligated to both ends of a target nucleic acid. For example, ligating 35 different tags (e.g., barcodes) to both ends of a target molecule creates 35 x 35 permutations, which equals 1225 permutations for 35 tags. Such a number of tags is sufficient to ensure that different molecules with the same start and end points will receive different combinations of tags with a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%). Other barcode combinations include any number between 10 and 500, such as about 15 x 15, about 35 x 35, about 75 x 75, about 100 x 100, about 250 x 250, and about 500 x 500.
[0283] In some cases, the unique tag can be an oligonucleotide of predetermined, random, or semi-random sequence. In other cases, multiple barcodes can be used, where the barcodes are not necessarily unique to each other among the multiple. In this example, the barcode can be ligated to each molecule so that the combination of the barcode and the sequence to which it can be ligated creates a unique sequence that can be tracked individually. As described herein, detection of a non-unique barcode in combination with sequence data at the beginning (start) and end (end) of the sequence read data can allow for the assignment of a unique identity to a specific molecule. The length or number of base pairs of each sequence read data can also be used to assign a unique identity to such a molecule. As described herein, fragments derived from a single strand of nucleic acid that have been assigned a unique identity can thereby allow for the identification of subsequent fragments derived from the parent strand.
[0284] E. Enrichment / Capture Step; Amplification The methods disclosed herein may include capturing DNA, such as a target region of DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample). In some embodiments, a target region of additional DNA, such as DNA from tumor cells, is also captured. Capture of DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or additional DNA can be performed in parallel (e.g., on pooled and / or differentially tagged DNA from the buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample) and / or additional DNA) or separately. Any of the embodiments described herein relating to enrichment or capture can be performed on DNA and / or additional DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample).
[0285] In some embodiments, the capturing step comprises contacting the DNA with a probe (e.g., an oligonucleotide) specific for the target region. Enrichment or capture can be performed on any sample or aliquot described herein using any suitable approach known in the art.
[0286] In some embodiments, enrichment or capture is performed after binding of adapters to sample molecules. In some embodiments, enrichment or capture is performed after a partitioning step. In some embodiments, enrichment or capture is performed after an amplification step. In some embodiments, sample molecules are partitioned, then adapters are bound, then the sample molecules are amplified, and then the amplified molecules are subjected to enrichment or capture. The enriched or captured molecules can then be subjected to another amplification and then sequenced.
[0287] In some embodiments, the probe specific for the target region includes a capture moiety that facilitates enrichment or capture of DNA hybridized to the probe. In some embodiments, the capture moiety is biotin. In some such embodiments, streptavidin bound to a solid support, such as magnetic beads, is used to bind to the biotin. Non-specifically bound DNA that does not contain the target region is washed away from the captured DNA. In some embodiments, the DNA is then dissociated from the probe and eluted from the solid support using a buffer containing a salt wash or another DNA denaturing agent. In some embodiments, the probe is also eluted from the solid support, for example, by disrupting the biotin-streptavidin interaction. In some embodiments, the captured DNA is amplified after elution from the solid support. In some such embodiments, DNA containing an adapter is amplified using PCR primers that anneal to the adapter. In some embodiments, the captured DNA is amplified while bound to the solid support. In some such embodiments, amplification involves the use of PCR primers that anneal to a sequence within the adapter and a sequence within the probe that annealed to the target region of the DNA.
[0288] In some embodiments, the methods herein include enriching or capturing DNA containing epigenetic and / or sequence-variable target regions. Such regions can be captured from an aliquot of a sample (e.g., a sample that has undergone adaptor binding and amplification), while a separate aliquot of the sample is subjected to a step of partitioning the DNA using an agent that recognizes modified cytosines, such as methylcytosines. Enriching or capturing DNA containing epigenetic and / or sequence-variable target regions can include contacting the DNA with a first or second set of target-specific probes. Such target-specific probes can have any of the features of the target-specific probe sets described herein, including, but not limited to, the embodiments above and the probe section below. The capturing step can be performed on one or more aliquots prepared during the methods disclosed herein. In some embodiments, DNA is captured from a first aliquot or a second aliquot, e.g., a first aliquot and a second aliquot. In some embodiments, the aliquots are differentially tagged (e.g., as described herein) and then pooled before being subjected to capture. Exemplary methods for capturing DNA containing epigenetic and / or sequence variable target regions can be found, for example, in WO2020 / 160414, which is hereby incorporated by reference herein.
[0289] The capturing step can be carried out using the conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on the characteristics of the probe, such as length, base composition, etc. Those skilled in the art will be familiar with the appropriate conditions in view of the general knowledge in the art regarding nucleic acid hybridization. In some embodiments, a complex of target-specific probe and DNA is formed.
[0290] In some embodiments, the methods described herein include capturing multiple sets of target regions of DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample), and / or from additional DNA obtained from a subject. The target regions may contain differences depending on whether they originate from a tumor or from healthy cells, or from a particular cell type. The capturing step produces a set of captured DNA molecules. In some embodiments, DNA molecules corresponding to the set of sequence variable target regions are captured with a higher capture yield in the set of captured DNA molecules than DNA molecules corresponding to the set of epigenetic target regions. In some embodiments, the methods described herein include contacting DNA obtained from the subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture DNA corresponding to the set of sequence variable target regions with a higher capture yield than DNA corresponding to the set of epigenetic target regions. For additional discussion of capture steps, capture yields, and related aspects, see WO2020 / 160414, which is incorporated herein by reference for all purposes.
[0291] Because analyzing sequence-variable target regions with sufficient reliability or accuracy may require sequencing at a greater depth than that required for analyzing epigenetic target regions, it may be beneficial to capture DNA corresponding to a set of sequence-variable target regions from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduced sample, or a PBMC sample), and / or from additional DNA, at a higher capture yield than DNA corresponding to a set of epigenetic target regions. The amount of data required to determine fragmentation patterns (e.g., to test for perturbation of transcription start sites or CTCF binding sites) or fragment abundances (e.g., in hypermethylated and hypomethylated distributions) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence mutations. Capturing target region sets at different yields may facilitate sequencing target regions to different depths of sequencing in the same sequencing run (e.g., using pooled mixtures and / or in the same sequencing cell).
[0292] In some embodiments, the DNA is amplified. In some embodiments, the amplification is carried out before the capture step. In some embodiments, the amplification is carried out after the capture step. In some embodiments, the amplification is carried out before or after the capture step. In various embodiments, the method further comprises sequencing the captured DNA to various degrees of sequencing depth, for example, for epigenetic and sequence variable target region sets, consistent with the discussion herein.
[0293] In some embodiments, the capturing step is performed simultaneously in the same container for the probes of the sequence variable target region set and the probes of the epigenetic target region set, for example, the probes of the sequence variable target region set and the epigenetic target region set are in the same composition. This approach results in a relatively streamlined workflow. In some embodiments, the concentration of the probes for the sequence variable target region set is higher than the concentration of the probes for the epigenetic target region set.
[0294] Alternatively, the capturing step can be performed using a sequence-variable target region probe set in a first container and an epigenetic target region probe set in a second container, or the contacting step can be performed using a sequence-variable target region probe set at a first time and in the first container, and an epigenetic target region probe set at a second time before or after the first time. This approach allows for the preparation of separate compositions: a first composition containing captured DNA corresponding to the sequence-variable target region set and a second composition containing captured DNA corresponding to the epigenetic target region set. The compositions can be processed separately as desired (e.g., to partition based on methylation as described herein) and pooled in appropriate proportions to provide material for further processing and analysis, such as sequencing.
[0295] In some embodiments, the adapter is included in the DNA described herein. In some embodiments, a tag, which may be or include a barcode, is included in the DNA. In some embodiments, such a tag is included in the adapter. The tag can facilitate identification of the origin of the nucleic acid. For example, after pooling multiple samples for parallel sequencing, it may be possible to use the barcode to identify the source (e.g., subject) from which the DNA originated. This can be done simultaneously with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer, as described herein. In some embodiments, the adapter and tag / barcode are provided by the same primer or primer set. For example, the barcode can be located 3' of the adapter and 5' of the portion of the primer that hybridizes to the target. Alternatively, the barcode can be added by other approaches, such as ligation, optionally together with the adapter in the same ligation substrate.
[0296] Additional details regarding amplification, tags, and barcodes are discussed herein and can be combined to the extent feasible with any of these embodiments.
[0297] F. A procedure that affects a first nucleobase in DNA differently than a second nucleobase in DNA In some embodiments, the methods disclosed herein comprise subjecting DNA (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample), and / or additional DNA) to a procedure that affects a first nucleobase within the DNA differently than a second nucleobase within the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity. In some embodiments, the procedure chemically converts the first nucleobase or the second nucleobase, such that the base-pairing specificity of the converted nucleobase is altered. In some embodiments, when the first nucleobase is modified or unmodified adenine, the second nucleobase is modified or unmodified adenine; when the first nucleobase is modified or unmodified cytosine, the second nucleobase is modified or unmodified cytosine; when the first nucleobase is modified or unmodified guanine, the second nucleobase is modified or unmodified guanine; and when the first nucleobase is modified or unmodified thymine, the second nucleobase is modified or unmodified thymine (for purposes of this step, modified and unmodified uracil are encompassed by modified thymine).
[0298] In some embodiments, the first nucleobase is modified or unmodified cytosine, and then the second nucleobase is modified or unmodified cytosine.For example, the first nucleobase can comprise unmodified cytosine (C), and the second nucleobase can comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC).Alternatively, the second nucleobase can comprise C, and the first nucleobase can comprise one or more of mC and hmC.For example, as shown in the above summary of the invention and the following discussion, other combinations are also possible, such as when one of the first and second nucleobases comprises mC, and the other comprises hmC.
[0299] In some embodiments, the procedure that affects the first nucleic acid base in DNA differently from the second nucleic acid base in DNA comprises bisulfite conversion.By treating with bisulfite, unmodified cytosine and certain modified cytosines (for example, 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) are converted to uracil, while other modified cytosines (for example, 5-methylcytosine, 5-hydroxymethylcytosine) are not converted.Therefore, when bisulfite conversion is used, the first nucleic acid base comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other cytosine forms that are affected by bisulfite, and the second nucleic acid base can comprise one or more of mC and hmC, for example, mC and optionally hmC.Sequencing the DNA that has been treated with bisulfite identifies the position that is read as cytosine as being mC or hmC position. On the other hand, positions that read as T are identified as T or bisulfite-sensitive forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Performing bisulfite conversion on a DNA sample, for example, as described herein, therefore facilitates identifying positions containing mC or hmC using sequence reads obtained from an exemplary sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.
[0300] In some embodiments, the procedure that affects a first nucleobase in DNA differently from a second nucleobase in DNA includes oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to fC, which is bisulfite-sensitive, followed by bisulfite conversion. Thus, when oxidative bisulfite conversion is used, the first nucleobase includes one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms that are affected by bisulfite, and the second nucleobase includes mC. Sequencing of the DNA converted with Ox-BS identifies positions that are read as cytosine as mC positions. Meanwhile, positions that are read as T are identified as T, hmC, or bisulfite-sensitive forms of C, such as unmodified cytosine, fC, or hmC. Performing Ox-BS conversion on a DNA sample, for example, as described herein, therefore facilitates identifying mC-containing positions using sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, e.g., Booth et al., Science 2012; 336: 934-937.
[0301] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in DNA comprises Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion and mC is oxidized prior to bisulfite treatment, resulting in the conversion of the position originally occupied by mC to U, while the position originally occupied by hmC remains as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, β-glucosyltransferase can be used to protect hmC (forming 5-glucosylhydroxymethylcytosine (ghmC)), and then a TET protein, e.g., mTet1, can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U, while leaving ghmC unaffected. Thus, when TAB conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, mC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises hmC. Sequencing of TAB-converted DNA identifies positions that read as cytosine as being hmC positions. Meanwhile, positions that read as T are identified as being T, mC, or bisulfite-sensitive forms of C, such as unmodified cytosine, fC, or caC. Performing TAB conversion on a DNA sample, for example, as described herein, therefore facilitates identifying positions containing hmC using sequence reads obtained from the sample.
[0302] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in DNA includes Tet-assisted conversion using a substituted borane reducing agent, optionally with 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In Tet-assisted pic-borane conversion using a substituted borane reducing agent, mC and hmC are converted to caC using a TET protein without affecting unmodified C. caC and, if present, fC are then converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, again without affecting unmodified C. See, for example, Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., Supplementary Figure 1 and Supplementary Note 7). DHU is read as T in sequencing. Thus, when this type of conversion is used, the first nucleobase contains one or more of mC, fC, caC, or hmC, and the second nucleobase contains an unmodified cytosine. Sequencing the converted DNA identifies positions that read as cytosine as unmodified C positions. Meanwhile, positions that read as T are identified as T, mC, fC, caC, or hmC. Performing TAP conversion on a DNA sample, for example, as described herein, therefore facilitates identifying positions containing unmodified C using sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), which is described in more detail in Liu et al. 2019 (supra).
[0303] Alternatively, protection of hmC (e.g., using βGT) can be combined with Tet-assisted conversion using a substituted borane reducing agent. hmC can be protected as described above by glucosylation using βGT to form ghmC. Treatment with a TET protein, e.g., mTet1, then converts mC to caC, but not C or ghmC. caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent, e.g., borane pyridine, tert-butylamine borane, or ammonia borane, again without affecting unmodified C or ghmC. Thus, when Tet-assisted conversion using a substituted borane reducing agent is used, the first nucleobase comprises mC, and the second nucleobase comprises one or more of unmodified cytosine or hmC, e.g., unmodified cytosine, and optionally hmC, fC, and / or caC. Sequencing the converted DNA identifies positions that read as cytosine as either hmC or unmodified C positions, while positions that read as T are identified as T, fC, caC, or mC. Performing TAPSβ conversion on a DNA sample, for example, as described herein, therefore facilitates using sequence reads obtained from the sample to distinguish between positions containing unmodified C or hmC and positions containing mC. For an exemplary description of this type of conversion, see, for example, Liu et al., Nature Biotechnology 2019; 37:424-429.
[0304] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in DNA comprises APOBEC-coupled epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme, such as APOBEC3A (A3A), is used to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when ACE conversion is used, the first nucleobase comprises unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase comprises hmC. Sequencing of the ACE-converted DNA identifies positions that read as cytosine as being hmC, fC, or caC positions. Meanwhile, positions that read as T are identified as being T, unmodified C, or mC. Performing ACE conversion on a DNA sample as described herein thus facilitates using sequence reads obtained from the sample to distinguish between positions containing hmC and positions containing mC or unmodified C. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090. In some embodiments, the procedure that affects a first nucleobase in DNA differently from a second nucleobase in DNA comprises enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1.For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), which can then be used to deaminate the unmodified cytosine and convert it to uracil.
[0305] In some embodiments, the procedure that affects a first nucleobase in DNA differently from a second nucleobase in DNA comprises enzymatic conversion of the first nucleobase, e.g., as in SEM-Seq. See, e.g., Vaisvila et al. (2023) Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA. bioRxiv; DOI: 10.1101 / 2023.06.29.547047v1, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047v1. SEM-seq uses a nonspecific, modification-sensitive double-stranded DNA deaminase (MsddA) in a nondestructive, single-enzyme 5-methylcytosine sequencing (SEM-seq) method to deaminate unmodified cytosines. Therefore, SEM-seq does not require the TET2 / T4-βGT protection and denaturation steps that are useful, for example, in APOEC3A-based protocols. Furthermore, MsddA does not deaminate 5-formylated cytosine (5fC) or 5-carboxylated cytosine (5caC). In SEM-seq, unmodified cytosines in DNA are deaminated to uracil and read as "T" during sequencing. Modified cytosines (e.g., 5mC) are not converted and are read as "C" during sequencing. Cytosines that are read as thymine are identified as unmodified (e.g., unmethylated) cytosines or thymines in DNA. Performing SEM-seq conversion therefore facilitates using the resulting sequence reads to identify positions containing 5mC.
[0306] In some embodiments, the procedure for affecting the first nucleobase in the DNA of the first subsample differently from the second nucleobase in the DNA comprises separating the DNA that originally contains the first nucleobase from the DNA that does not originally contain the first nucleobase. In some such embodiments, the first nucleobase is hmC. The DNA that originally contains the first nucleobase can be separated from other DNA using a labeling procedure that includes biotinylating the position that originally contained the first nucleobase. In some embodiments, the first nucleobase is first derivatized with an azide-containing moiety, for example, a glucosyl-azide-containing moiety. The azide-containing moiety can then serve as a reagent for attaching biotin, for example, through Huisgen cycloaddition chemistry. Then, the DNA originally containing the first nucleobase here biotinylated can be separated from the DNA that does not originally contain the first nucleobase using a biotin-binding agent, such as avidin, neutravidin (deglycosylated avidin with an isoelectric point of about 6.3), or streptavidin.An example of a procedure for separating the DNA originally containing the first nucleobase from the DNA that does not originally contain the first nucleobase is hmC-sealing, which involves labeling hmC to form β-6-azido-glucosyl-5-hydroxymethylcytosine, then binding a biotin moiety via Huisgen cycloaddition, and then using a biotin-binding agent to separate the biotinylated DNA from other DNA.For an exemplary description of hmC-sealing, see, for example, Han et al., Mol. Cell 2016; 63: 711-719.This approach is useful for identifying fragments that contain one or more hmC nucleobases.
[0307] In some embodiments, after such separation, the method further comprises the step of differentially tagging each DNA that originally comprises the first nucleobase with the DNA that originally does not comprise the first nucleobase.The method can further comprise the step of pooling the DNA that originally comprises the first nucleobase and the DNA that originally does not comprise the first nucleobase after differential tagging.The DNA that originally comprises the first nucleobase and the DNA that originally does not comprise the first nucleobase can then be used in downstream analysis.For example, the pooled DNA that originally comprises the first nucleobase and the DNA that originally does not comprise the first nucleobase can be sequenced in the same sequencing cell (for example, after being subjected to further processing such as described herein), but still retain the ability to use differential tagging to determine whether a given read originates from a molecule of DNA that originally comprises the first nucleobase or a molecule of DNA that originally does not comprise the first nucleobase.
[0308] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is N6-methyladenine (mA). 6 -Methyladenine (mA), N 6 -hydroxymethyladenine (hmA), or N 6 -formyl adenine (fA).
[0309] Techniques including partitioning based on methylation status or methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases, such as mC, mA, caC (e.g., generated by oxidation of mC or hmC with Tet2 prior to enzymatic conversion of the unmodified C to U, e.g., using a deaminase such as APOBEC3A), or dihydrouracil. See, e.g., Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies against various modified nucleobases, such as mC, caC, and thymine / uracil (including dihydrouracil) forms, or halogenated forms, such as 5-bromouracil, are commercially available. Various modified bases can also be detected based on their altered base pairing specificity. For example, hypoxanthine is a modified form of adenine that can arise by deamination and is read as G in sequencing. See, e.g., U.S. Pat. No. 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, NY, 2002, chapter 14, "Mutation, Repair, and Recombination."
[0310] In some embodiments, the conversion procedure is an enzymatic conversion procedure that alters the base pairing specificity of a modified nucleoside (e.g., a DM-seq conversion that involves adding a protecting group (e.g., a carboxymethyl group) to an unmodified cytosine and deaminating 5mC, e.g., using an APOBEC enzyme), or an enzymatic conversion procedure that alters the base pairing specificity of an unmodified nucleoside (e.g., SEM-seq).
[0311] In some cases, the conversion procedure used in the method of the present disclosure is a conversion procedure that changes the base pairing specificity of modified nucleosides (e.g., methylated cytosine), but does not change the base pairing specificity of corresponding unmodified nucleosides (e.g., cytosine), or does not change the base pairing specificity of any unmodified nucleosides (e.g., cytosine, adenosine, guanosine, and thymidine (or uracil)).The advantages of a method that does not change the base pairing specificity of unmodified nucleosides include reduced loss of sequence complexity, higher sequencing efficiency, and reduced alignment loss.In addition, methods such as DM-seq may be preferable in some cases to methods such as bisulfite sequencing and EM-seq, because these methods are less destructive (especially important for low-yield samples such as cfDNA) and do not require denaturation, which means that non-conversion errors are theoretically more likely to be random. In methods that require denaturation for conversion, failure to denature the DNA molecule results in unconverted bases in the DNA molecule. Because biological changes in methylation are primarily concerted to local regions of interest, these non-random (local) conversions can appear as false negatives (unmethylated regions). Random non-conversion methods can have the greatest impact on a low percentage of bases within a region, and therefore, setting a threshold for the percentage of bases within a methylated / unmethylated region can maximize the specificity of methylation change detection (reducing false positives). Therefore, in some cases, conversion procedures that do not involve denaturation are preferred.
[0312] In other cases, the conversion procedure used in the disclosed methods is one that alters the base pairing specificity of an unmodified nucleoside (e.g., cytosine) but does not alter the base pairing specificity of the corresponding modified nucleoside (e.g., methylated cytosine).
[0313] Those skilled in the art will be able to select a suitable method according to their needs, including which nucleoside modifications are to be detected and / or identified.
[0314] In some embodiments, the conversion procedure converts modified nucleosides. In some embodiments, the conversion procedure for converting modified nucleosides includes enzymatic conversion, such as DM-seq, as described in WO2023 / 288222A1. In DM-seq, unmodified cytosines in DNA are enzymatically protected from a subsequent deamination step, in which 5mC in 5mCpG is converted to T. Enzymatically protected unmodified (e.g., unmethylated) cytosines are not converted and are read as "C" during sequencing. Cytosines that are read as thymines (in the context of CpG) are identified as methylated cytosines in DNA.
[0315] Thus, when this type of conversion is used, the first nucleobase contains an unmodified (e.g., unmethylated) cytosine, and the second nucleobase contains a modified (e.g., methylated) cytosine. Sequencing the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T or 5mC. Performing DM-seq conversion therefore facilitates identifying positions containing 5mC using the resulting sequence reads.
[0316] Exemplary cytosine deaminases for use herein include APOBEC enzymes, such as APOBEC3A.Generally, AID / APOBEC family DNA deaminase enzymes, such as APOBEC3A (A3A), are used to deaminate (unprotected) unmodified cytosine and 5mC.For example, see Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090 for an exemplary description of APOBEC conversion.
[0317] Enzymatic protection of unmodified cytosine in DNA involves adding a protecting group to the unmodified cytosine. Such protecting groups can include alkyl groups, alkyne groups, carboxyl groups, carboxyalkyl groups, amino groups, hydroxymethyl groups, glucosyl groups, glucosylhydroxymethyl groups, isopropyl groups, or dyes. For example, DNA can be treated with a methyltransferase, such as a CpG-specific methyltransferase, which adds a protecting group to the unmodified cytosine. The term methyltransferase is used broadly herein to refer to an enzyme that can transfer methyl or substituted methyl (e.g., carboxymethyl) to a substrate (e.g., cytosine in a nucleic acid). In some embodiments, DNA is contacted with a CpG-specific DNA methyltransferase (MTase), such as a CpG-specific carboxymethyltransferase (CxMTase), and a substituted methyl donor, such as a carboxymethyl donor (e.g., carboxymethyl-S-adenosyl-L-methionine). See, for example, WO2021 / 236778A2. In certain embodiments, CxMTase can facilitate the addition of a protective carboxymethyl group to unmethylated cytosine. In some embodiments, the unmethylated cytosine is unmodified cytosine. The carboxymethyl group can prevent deamination of cytosine during a deamination step (e.g., a deamination step using an APOBEC enzyme, such as A3A). Substituted methyl or carboxymethyl donors useful in the disclosed methods include, but are not limited to, S-adenosyl-L-methionine (SAM) analogs, and optionally, the SAM analogs are carboxy-S-adenosyl-L-methionine (CxSAM). SAM analogs are described, for example, in WO2022 / 197593A1. The MTase can be, for example, CpG methyltransferase from Spiroplasma sp. strain MQ1 (M.SssI), DNA-methyltransferase 1 (DNMT1), DNA-methyltransferase 3 alpha (DNMT3A), DNA-methyltransferase 3 beta (DNMT3B), or DNA adenine methyltransferase (Dam).The CxMTase can be a CpG methyltransferase from Mycoplasma penetrans (M.MpeI). In certain embodiments, the methyltransferase enzyme is SEQ ID NO: 1 or SEQ ID NO: 2, or a variant of M.MpeI having a sequence at least 90%, at least 92%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto, optionally wherein the amino acid corresponding to position 374 is R or K.
[0318] In one embodiment, the methyltransferase enzyme is a variant of M.MpeI having an N374R or N374K substitution. The methyltransferase of SEQ ID NO:1 or SEQ ID NO:2 may further comprise one or more amino acid substitutions selected from: a) substitution of one or both of residues T300 and E305 with S, A, G, Q, D, or N; b) substitution of one or more residues A323, N306, and Y299 with a positively charged amino acid selected from K, R, or H; and / or c) substitution of S323 with A, G, K, R, or H, which may enhance the activity of the enzyme.
[0319] Optionally, the conversion procedure further includes enzymatic protection of 5hmC in DNA, for example, by glucosylation of 5hmC (e.g., using βGT) before deamination of unprotected modified cytosines. In this method, 5hmC can be protected from conversion, for example, through glucosylation using β-glucosyltransferase (βGT), to form 5ghmC (forming 5-glucosylhydroxymethylcytosine). This is described, for example, in Yu et al., Cell 2012; 149: 1368-80. Glucosylation of 5hmC can reduce or eliminate deamination of 5hmC by deaminases such as APOBEC3A. In this case, a protecting group is added to unmodified (unmethylated) cytosines in DNA by treatment with MTase or CxMTase. The 5mC is then deaminated (in the case of 5mC, converted to T; protected unmodified cytosine and 5ghmC are not deaminated) by treatment with a deaminase, e.g., an APOBEC enzyme (e.g., APOBEC3A). Sequencing the converted DNA identifies positions that read as cytosine as being either 5hmC or unmodified C positions, while positions that read as T are identified as being T or 5mC. Performing DM-seq conversion of 5hmC by glycosylation on a sample as described herein thus facilitates using the resulting sequence reads to distinguish between positions containing unmodified C or 5hmC and positions containing 5mC.
[0320] Also provided herein are methods in which alternative base conversion schemes are used, for example, unmethylated cytosines can be left intact, while methylated and hydroxymethyl cytosines are converted to bases that are read as thymine (e.g., uracil, thymine, or dihydrouracil).
[0321] In some embodiments, methylating cytosines in at least one of the first complementary strand or the second complementary strand comprises contacting the cytosines with a methyltransferase, e.g., DNMT1 or DNMT5. In such embodiments, the step of oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., by contacting 5-hydroxymethylcytosines in the first and second strands with KRuO4) can be optional.
[0322] In some embodiments, converting the modified cytosine in at least one of the first or second strands to thymine, or a base that is read as thymine, comprises oxidizing the hydroxymethylcytosine, e.g., oxidizing the hydroxymethylcytosine to formylcytosine. In some embodiments, oxidizing the hydroxymethylcytosine to formylcytosine comprises contacting the hydroxymethylcytosine with a ruthenate, e.g., potassium ruthenate (KRuO).
[0323] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil.
[0324] In some embodiments, the method includes converting formylcytosine and / or methylcytosine to carboxylcytosine as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine. For example, converting formylcytosine and / or methylcytosine to carboxylcytosine can include contacting formylcytosine and / or methylcytosine with a TET enzyme, such as TET1, TET2, or TET3. In some embodiments, the method includes reducing carboxylcytosine and / or reducing carboxylcytosine to dihydrouracil as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine. In some embodiments, reducing carboxylcytosine includes contacting carboxylcytosine with a borane or borohydride reducing agent.
[0325] In some embodiments, the borane or borohydride reducing agent comprises pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, sodium cyanoborohydride (NaBHCN), lithium borohydride (LiBH), ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof. In other embodiments, the reducing agent comprises lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof.
[0326] Various TET enzymes can be used in the disclosed methods as needed. In some embodiments, one or more TET enzymes include TETv. TETv is described in U.S. Patent 10,260,088, the sequence of which is SEQ ID NO: 1 therein (SEQ ID NO: 3 herein). In some embodiments, one or more TET enzymes include TETcd. TETcd is described in U.S. Patent 10,260,088, the sequence of which is SEQ ID NO: 3 therein (SEQ ID NO: 4 herein). In some embodiments, one or more TET enzymes include TET1. In some embodiments, one or more TET enzymes include TET2. TET2 can be represented and used as a fragment comprising TET2 residues 1129-1480 (SEQ ID NO: 5 herein) joined to TET2 residues 1844-1936 by a linker, e.g., as described in U.S. Patent 10,961,525. In some embodiments, one or more TET enzymes include TET1 and TET2. In some embodiments, the one or more TET enzymes comprise a V1900 TET mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET mutant. In some embodiments, the one or more TET enzymes comprise a V1900 TET2 mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET2 mutant. Examples of the V1900A, V1900C, V1900G, V1900I, and V1900P TET2 mutants are provided as SEQ ID NOs: 6-10. In some embodiments, the V1900 TET mutant has at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NOs: 6, 7, 8, 9, or 10. Position 1900 of the wild-type TET2 sequence corresponds to position 438 in each of SEQ ID NOs: 5-10. Using TET enzymes that maximize the formation of 5-carboxylcytosine (5-caC) relative to less oxidized modified cytosines, particularly 5-formylcytosine, can be beneficial because 5-caC is not a substrate for enzymatic deamination by APOBEC enzymes, such as APOBEC3A.Maximizing the formation of 5-caC therefore reduces the risk of false calls, in which a base is identified as unmethylated because it undergoes deamination even if it was methylated (or hydroxymethylated) in the original sample. Thus, in some embodiments, the TET enzyme contains a mutation that increases the formation of 5-caC. Exemplary mutations are described above. A "mutation that increases the formation of 5-caC" means that a TET enzyme with the mutation produces more 5-caC than an otherwise identical TET enzyme lacking the mutation. 5-caC production can be measured, for example, as described in Liu et al., Nat Chem Biol 13:181-187 (2017) (see Online Methods section, TET reactions in vitro subsection, "driving" conditions). Any of the variants and / or mutants described in Liu et al. (2017) can be used in the disclosed methods, as needed.
[0327] G. Captured set; target area In some embodiments, nucleic acids captured or enriched using the methods described herein from DNA and / or additional DNA derived from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample), include captured DNA, e.g., one or more captured sets of DNA. In some embodiments, the captured DNA includes target regions that are differentially methylated in different immune cell types. In some embodiments, the immune cell types include rare or closely related immune cell types, e.g., activated and naive lymphocytes or myeloid cells at different stages of differentiation.
[0328] In some embodiments, the set of captured epigenetic target regions captured from a sample or a first subsample includes a hypermethylated variable target region. In some embodiments, the hypermethylated variable target region is differentially or exclusively hypermethylated in one cell type, one immune cell type, or one immune cell type within a cluster. In some embodiments, the hypermethylated variable target region is hypermethylated to an extent that it is discriminably more or exclusively present in one cell type, one immune cell type, or one immune cell type within a cluster. Such hypermethylated variable target regions may also be hypermethylated in other cell types, but not to the extent observed in one cell type. In some embodiments, the hypermethylated variable target region exhibits lower methylation in healthy DNA (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukapheresis sample, or a PBMC sample)) than in at least one other tissue type. In some embodiments, the hypermethylated variable target region exhibits lower methylation in healthy additional DNA than in at least one other tissue type.
[0329] In some embodiments, the set of captured epigenetic target regions captured from one sample or a second subsample comprises a hypomethylated variable target region. In some embodiments, the hypomethylated variable target region is exclusively hypomethylated in one cell type, one immune cell type, or one immune cell type within a cluster. In some embodiments, the hypomethylated variable target region is hypomethylated to the extent that it is exclusively present in one cell type, one immune cell type, or one immune cell type within a cluster. Such a hypomethylated variable target region may also be hypomethylated in other cell types, but not to the extent observed in one cell type. In some embodiments, the hypomethylated variable target region exhibits lower methylation in healthy DNA (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukapheresis sample, or a PBMC sample)) than in at least one other tissue type. In some embodiments, the hypomethylated variable target region exhibits higher methylation in healthy additional DNA than in at least one other tissue type.
[0330] Without wishing to be bound by any particular theory, individuals with cancer, proliferating, or activated immune cells (and potentially cancer cells) may shed more DNA into the bloodstream than immune cells in healthy individuals (and, respectively, healthy cells of the same tissue type). Thus, the distribution of cell types and / or tissues of origin of DNA and / or additional DNA from a buffy coat sample or any other cell-containing sample, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample), may change upon onset of cancer. For example, the distribution of immune cell types of origin may change in subjects with cancer, precancerous conditions, infection, graft rejection, or other diseases or disorders that directly or indirectly affect the immune system. The status of epigenetic target regions in certain immune cell types may also change in subjects with such diseases compared to healthy subjects or compared to the same subjects prior to developing the disease or disorder. Thus, variations in hypermethylation and / or hypomethylation may be indicative of disease. For example, elevated levels of hypermethylated and / or hypomethylated variable target regions in the aliquots after the partitioning step may be indicative of the presence (or recurrence, depending on the subject's medical history) of cancer.
[0331] Exemplary hypermethylated and hypomethylated variable target regions useful for distinguishing between various cell types, including, but not limited to, immune cell types, have been identified by analyzing DNA obtained from various cell types by whole-genome bisulfite sequencing, as described, for example, in Stunnenberg, HG et al., "The International Human Epigenome Consortium: A Blueprint for Scientific Collaboration and Discovery," Cell 167, 1145 (2016) (doi.org / 10.1186 / sl3059-020-02065-5). Whole-genome bisulfite sequencing data is available from the Blueprint consortium, available online at dcc.blueprint-epigenome.eu.
[0332] In some embodiments, the first captured set of target regions and the second captured set of target regions comprise DNA corresponding to a set of sequence variable target regions and DNA corresponding to a set of epigenetic target regions, respectively, as described, for example, in WO2020 / 160414. The first captured set and the second captured set can be combined to provide a combined captured set.
[0333] If DNA (e.g., a sample or aliquot) has been subjected to a procedure such as bisulfite conversion, treatment with deaminase, or any of the other such procedures mentioned herein that alter the base-pairing specificity of certain bases, enrichment or capture can use oligonucleotides (e.g., primers or probes) specific for the altered or unaltered sequence, as desired.
[0334] In some embodiments, where the captured set includes DNA corresponding to the set of sequence variable target regions and DNA corresponding to the set of epigenetic target regions, including the combined captured sets discussed above, the DNA corresponding to the set of sequence variable target regions is at a higher concentration than the DNA corresponding to the set of epigenetic target regions, e.g., 1.1-fold to 1.2-fold higher, 1.2-fold to 1.4-fold higher, 1.4-fold to 1.6-fold higher, 1.6-fold to 1.8-fold higher, 1.8-fold to 2.0-fold higher, 2.0-fold to 2.2-fold higher, 2.2-fold to 2.4-fold higher, 2.4-fold to 2.6-fold higher, 2.6-fold to 2.8-fold higher, 2.8-fold to 3.0-fold higher, 3.0-fold to 3.5-fold higher, 3.5-fold to 4.0-fold, 4.0-fold to 4.5-fold higher, 4.5-fold to 5.0-fold higher, 5.0-fold to 5.5-fold higher. , 5.5x to 6.0x higher concentration, 6.0x to 6.5x higher concentration, 6.5x to 7.0x higher concentration, 7.0x to 7.5x higher concentration, 7.5x to 8.0x higher concentration, 8.0x to 8.5x higher concentration, 8.5x to 9.0x higher concentration, 9.0x to 9.5x higher concentration, 9.5x to 10.0x higher concentration, 10x to 11x higher concentration, 11x to 12x higher concentration, 12x to 13x higher concentration, 13x to 14x higher concentration, 1 It may be present at a 4-fold to 15-fold higher concentration, a 15-fold to 16-fold higher concentration, a 16-fold to 17-fold higher concentration, a 17-fold to 18-fold higher concentration, a 18-fold to 19-fold higher concentration, a 19-fold to 20-fold higher concentration, a 20-fold to 30-fold higher concentration, a 30-fold to 40-fold higher concentration, a 40-fold to 50-fold higher concentration, a 50-fold to 60-fold higher concentration, a 60-fold to 70-fold higher concentration, a 70-fold to 80-fold higher concentration, a 80-fold to 90-fold higher concentration, or a 90-fold to 100-fold higher concentration. The degree of concentration difference accounts for normalization with respect to the footprint size of the target region, as discussed in the definitions section.
[0335] 1. Epigenetic target region set In some embodiments, epigenetic target region set may comprise one or more types of target region that may distinguish DNA from various immune cell types and DNA from other non-immune cell types, and / or may distinguish DNA from neoplastic (e.g., tumor or cancer) cells and DNA from healthy cells, such as non-neoplastic circulating cells.Exemplary types of such regions are discussed in detail herein.Epigenetic target region set may also comprise one or more control regions, for example, as described herein.
[0336] In some embodiments, the set of epigenetic target regions has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the set of epigenetic target regions has a footprint in the range of 100 to 1,000 kb, e.g., 100 to 200 kb, 200 to 300 kb, 300 to 400 kb, 400 to 500 kb, 500 to 600 kb, 600 to 700 kb, 700 to 800 kb, 800 to 900 kb, and 900 to 1,000 kb. In some embodiments, the set of epigenetic target regions has a footprint of about 100 kb to about 1000 kb, e.g., about 200 kb to about 1000 kb, about 300 kb to about 1000 kb, about 400 kb to about 1000 kb, about 500 kb to about 1000 kb, about 600 kb to about 1000 kb, or about 700 kb to about 1000 kb. Other exemplary footprints of the set of epigenetic target regions are about 100 kb to about 500 kb, about 200 kb to about 500 kb, about 300 kb to about 500 kb, about 400 kb to about 600 kb, or about 250 kb to about 750 kb. In some embodiments, the footprint of the set of epigenetic target regions is about 100 kb, about 150 kb, about 200 kb, about 250 kb, about 300 kb, about 350 kb, about 400 kb, about 450 kb, about 500 kb, about 550 kb, about 600 kb, about 650 kb, about 700 kb, about 750 kb, about 800 kb, about 850 kb, about 900 kb, about 950 kb, or about 1000 kb.
[0337] 1. Hypermethylated variable target regions In some embodiments, the set of epigenetic target regions comprises one or more hypermethylated variable target regions, which are exclusively hypermethylated in one immune cell type, or are hypermethylated in one immune cell type to a greater extent than in any other immune cell type, or in any other immune cell type within the same immune cell cluster. In some such embodiments, the hypermethylated variable target regions indicate the levels of specific immune cell types from which the DNA originated, including macrophages (including M1 macrophages and M2 macrophages); activated B cells (including regulatory B cells, memory B cells, and plasma cells); T cell subsets, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (including cytotoxic T cells, regulatory T cells (Tregs), CD4 effector memory T cells, and CD8 effector memory T cells); immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), low-density neutrophils, immature neutrophils, and immature granulocytes); and rare immune cell types such as natural killer (NK) cells. Methylation patterns of hypermethylated variable target regions useful for deconvolving immune cell types may further alter in certain disease states, e.g., cancer. Thus, in some embodiments, hypermethylated variable target regions useful for deconvolving immune cell types are also useful for determining the likelihood that a subject from whom a sample was obtained has cancer or a precancerous condition. In some such embodiments, hypermethylated variable target regions are useful for determining whether levels of particular immune cell types are abnormal and whether such abnormal levels may be associated with the presence of cancer or a precancerous condition, or whether such abnormal levels are associated with various diseases or conditions other than cancer or a precancerous condition.
[0338] In some embodiments, certain hypermethylated variable target regions indicate elevated levels of methylation, e.g., hypermethylated, observed in DNA produced by neoplastic cells, such as tumor or cancer cells. Detection of such hypermethylated variable target regions, for example, in conjunction with detection of hypermethylated variable target regions indicative of immune cell types, can further increase the specificity and / or sensitivity of the methods described herein. In some embodiments, such increased methylation observed in hypermethylated variable target regions indicates an increased likelihood that a sample (e.g., DNA and / or additional DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukapheresis sample, or a PBMC sample)) was obtained from a subject with cancer. For example, hypermethylation of promoters of tumor suppressor genes has been repeatedly observed. See, e.g., Kang et al., Genome Biol. 18: 53 (2017) and references cited therein. In another example, as discussed above, a hypermethylated variable target region may include a region that is not necessarily differentially methylated in cancerous tissue compared to DNA derived from healthy tissue of the same type, but is differentially methylated (e.g., has more methylation) compared to DNA that is typical in healthy subjects. For example, if the presence of cancer results in cell death, e.g., increased apoptosis of cells of the tissue type corresponding to the cancer, such cancer can be detected, at least in part, using such hypermethylated variable target regions. In some embodiments, hypermethylated variable target regions useful for determining the likelihood that a subject has cancer are different from hypermethylated variable target regions useful for determining the level of a particular immune cell type. In some embodiments, at least some of the hypermethylated variable target regions useful for determining the likelihood that a subject has cancer are the same as the hypermethylated variable target regions useful for determining the level of a particular immune cell type.
[0339] An extensive review of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866: 106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. A set of exemplary hypermethylated variable target regions based on colorectal cancer (CRC) studies is provided in Table 1. Many of these genes may also be relevant to cancers other than colorectal cancer. For example, TP53 is widely recognized as a crucial tumor suppressor, and hypermethylation-based inactivation of this gene may be a common mechanism of carcinogenesis.
[0340] [Table 1]
[0341] In some embodiments, the hypermethylated variable target region comprises multiple loci listed in Table 1, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. For example, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site of the gene and the stop codon (the last stop codon of an alternatively spliced gene) or in the promoter region of the gene. In some embodiments, the one or more probes bind within 300 bp, e.g., within 200 bp or 100 bp, of the transcription start site of a gene in Table 1.
[0342] Methylation variable target regions in various types of lung cancer have been reported, for example, in Ooki et al., Clin. Cancer Res. 23: 7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77: 453-74 (2015); Hulbert et al., Clin. Cancer Res. 23: 1998-2005 (2017); Shi et al., BMC Genomics 18: 901 (2017); Schneider et al., BMC Cancer. 11: 102 (2011); Lissa et al., Transl Lung Cancer Res 5 (5): 492-504 (2016); Skvortsova et al., Br. J. Cancer. 94 (10): 1492-1495 (2006); Kim et al., Cancer Res. 61: 3419-3424 (2001);Furonaka et al., Pathology International 55: 303-309 (2005);Gomes et al., Rev. Port. Pneumol. 20: 20-30 (2014);Kim et al., Oncogene. 20: 1765-70 (2001);Hopkins-Donaldson et al., Cell Death Differ. 10: 356-64 (2003);Kikuchi et al., Clin. Cancer Res. 11: 2954-61 (2005);Heller et al., Oncogene 25: 959-968 (2006);Licchesi et al., Carcinogenesis. 29: 895-904 (2008);Guo et al. al., Clin. Cancer Res. 10: 7917-24 (2004); Palmisano et al., Cancer Res. 63: 4620-4625 (2003); and Toyooka et al., Cancer Res. 61: 4556-4560 (2001).
[0343] A set of exemplary hypermethylated variable target regions based on lung cancer studies is provided in Table 2. Many of these genes may also be relevant to cancers other than lung cancer. For example, Casp8 (caspase 8) is a key enzyme in programmed cell death, and hypermethylation-based inactivation of this gene may be a general mechanism of carcinogenesis not limited to lung cancer. Furthermore, several genes appear in both Tables 1 and 2, indicating their generality.
[0344] [Table 2-1] [Table 2-2]
[0345] Any of the above embodiments relating to target regions identified in Table 2 can be combined with any of the above embodiments relating to target regions identified in Table 1. In some embodiments, the hypermethylated variable target regions include multiple loci listed in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2.
[0346] In some embodiments, the hypermethylated variable target region comprises a region of one or more genes listed in Table 2b, e.g., at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1000, 1100, 1150, or 1200 genes listed in Table 2b. Hypermethylation of these genes can be useful for detecting contributions from immune cells to a DNA sample. In some embodiments, the hypermethylated variable target region comprises regions of multiple genes listed in Table 2b, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the genes listed in Table 2b. In some embodiments, the hypermethylated variable target region comprises regions of all of the genes listed in Table 2b. [Table 2b-1] [Table 2b-2] [Table 2b-3] [Table 2b-4] [Table 2b-5] [Table 2b-6] [Table 2b-7]
[0347] Additional hypermethylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a probabilistic method called CancerLocator using hypermethylated target regions derived from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions can be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0348] In some embodiments, when different epigenetic target regions are captured from the first aliquot and the second aliquot, the epigenetic target region captured from the first aliquot comprises a hypermethylated variable target region.
[0349] a. Hypomethylated variable target region In some embodiments, the set of epigenetic target regions comprises one or more hypomethylated variable target regions, which are exclusively hypomethylated in one immune cell type or are hypomethylated to a greater extent in one immune cell type than in any other immune cell type or in any other immune cell type within the same immune cell cluster. In some such embodiments, the hypomethylated variable target regions indicate the levels of specific immune cell types from which the DNA originated, including macrophages (including M1 macrophages and M2 macrophages); activated B cells (including regulatory B cells, memory B cells, and plasma cells); T cell subsets, such as CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (including cytotoxic T cells, regulatory T cells (Tregs), CD4 effector memory T cells, and CD8 effector memory T cells); immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), low-density neutrophils, immature neutrophils, and immature granulocytes); and rare immune cell types such as natural killer (NK) cells. Methylation patterns of hypomethylated variable target regions useful for deconvolving immune cell types may further alter in certain disease states, such as cancer. Thus, in some embodiments, hypomethylated variable target regions useful for deconvolving immune cell types are also useful for determining the likelihood that a subject from whom a sample was obtained has cancer or a precancerous condition. In some such embodiments, hypomethylated variable target regions are useful for determining whether levels of particular immune cell types are abnormal and whether such abnormal levels may be associated with the presence of cancer or a precancerous condition, or whether such abnormal levels are associated with various diseases or conditions other than cancer or a precancerous condition.
[0350] Furthermore, global hypomethylation is a commonly observed phenomenon in various cancers. See Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (review article mentioning the observation of hypomethylation in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions such as repetitive elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, as well as intergenic regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Thus, in some embodiments, the epigenetic target region set includes hypomethylated variable target regions, where a decreased level of observed methylation indicates an increased likelihood of the presence of cancer. Detection of such hypomethylated variable target regions, for example, in conjunction with detection of hypomethylated variable target regions indicative of immune cell types, can further increase the specificity and / or sensitivity of the methods described herein. In another example, as discussed above, hypomethylated variable target regions may include regions that are differentially methylated (e.g., less methylated) in cancerous tissue compared to DNA typical in healthy subjects, although not necessarily differentially methylated compared to DNA derived from healthy tissue of the same type. For example, if the presence of cancer results in cell death, e.g., increased apoptosis of cells of a tissue type corresponding to the cancer, such cancer can be detected, at least in part, using such hypomethylated variable target regions. In some embodiments, hypomethylated variable target regions useful for determining the likelihood that a subject has cancer are different from hypomethylated variable target regions useful for determining the level of a particular immune cell type. In some embodiments, at least some of the hypomethylated variable target regions useful for determining the likelihood that a subject has cancer are the same as the hypomethylated variable target regions useful for determining the level of a particular immune cell type.
[0351] In some embodiments, the hypomethylated variable target region comprises a repetitive element and / or an intergenic region, hi some embodiments, the repetitive element comprises one, two, three, four, or five of a LINE1 element, an Alu element, a centromeric tandem repeat, a pericentromeric tandem repeat, and / or satellite DNA.
[0352] Exemplary specific genomic regions that exhibit cancer-associated hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 of human chromosome 1. In some embodiments, the hypomethylated variable target region overlaps with or includes one or both of these regions.
[0353] Additionally, hypomethylated variable target regions can be obtained, for example, from Fox-Fisher et al., Elife Nov 29; 10 (2021), the EpiDISH R package, Moss et al., Nat Commun 9:1 (2018), and Loyfer et al. bioRxiv https: / / doi.org / 10.1101 / 2022.01.24.477547 (2022). In some embodiments, the hypomethylated variable target region can be specific to one or more types of immune cells.
[0354] In some embodiments, when different epigenetic target regions are captured from the first and second sub-samples, the epigenetic target region captured from the second sub-sample comprises a hypomethylated variable target region, hi some embodiments, the epigenetic target region captured from the second sub-sample comprises a hypomethylated variable target region and the epigenetic target region captured from the first sub-sample comprises a hypermethylated variable target region.
[0355] b.CTCF binding region CTCF is a DNA-binding protein that contributes to chromatin organization and often co-localizes with cohesin. Disruption of CTCF binding sites has been reported in various different cancers. For example, see Katainen et al., Nature Genetics, doi:10.1038 / ng.3335, published online 8 June 2015; Guo et al., Nat. Commun. 9:1520 (2018). CTCF binding results in a recognizable pattern in DNA that can be detected by sequencing, for example, through fragment length analysis. Details regarding sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57-68 (2016), WO2018 / 009723, and US20170211143A1, each of which is incorporated herein by reference.
[0356] Thus, disruption of CTCF binding leads to variations in DNA fragmentation patterns. Thus, the CTCF binding site is a type of fragmentation variable target region.
[0357] There are many known CTCF binding sites.For example, see CTCFBSDB (CTCF binding site database), available on the Internet at insulatordb.uthsc.edu / , for example, Cuddapah et al., Genome Res.19:24-32 (2009); Martin et al., Nat. Struct. Mol. Biol.18:708-14 (2011); Rhee et al., Cell.147:1408-19 (2011), each of which is incorporated by reference.Exemplary CTCF binding sites are nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13.
[0358] Thus, in some embodiments, the set of epigenetic target regions comprises CTCF binding regions, in some embodiments, the CTCF binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those described in the CTCFBSDB or one or more of the Cuddapah et al., Martin et al., or Rhee et al. articles described above or listed above.
[0359] In some embodiments, at least some of the CTCF sites can be methylated or unmethylated, where the methylation status correlates with whether the cell is a cancer cell or not. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding site.
[0360] C transcription start site Transcription start sites may also show disturbances in neoplastic cells.For example, the nucleosome composition at various transcription start sites in healthy cells of hematopoietic lineage may differ from the nucleosome composition at those transcription start sites in neoplastic cells.This results in different DNA patterns that can be detected by sequencing, as generally discussed in Snyder et al., Cell 164:57-68 (2016), WO2018 / 009723, and US20170211143A1.In another example, transcription start sites in cancerous tissues may not necessarily be epigenetically different from the DNA from the same type of healthy tissue, but they may be epigenetically different (e.g., in terms of nucleosome composition) from the DNA that is typical in healthy subjects.For example, if the presence of cancer leads to cell death, for example, increased apoptosis of cells of the tissue type corresponding to cancer, such cancer can be detected, at least in part, using the difference in such transcription start sites.
[0361] Therefore, perturbations in transcription start sites also result in variations in DNA fragmentation patterns, and thus transcription start sites are also a type of variable fragmentation target region.
[0362] Human transcription start sites are available from DBTSS (DataBase of Human Transcription Start Sites), available on the internet at btss.hgc.jp, and are described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), which is incorporated herein by reference.
[0363] Thus, in some embodiments, the set of epigenetic target regions includes transcription start sites. In some embodiments, the transcription start sites include at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those listed in the DBTSS. In some embodiments, at least some of the transcription start sites may be methylated or unmethylated, where the methylation status correlates with whether a cell is a cancer cell or not. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the transcription start site.
[0364] d. local amplification Although local amplification is somatic mutation, it can be detected by sequencing based on the frequency of read data in a similar manner to the approach of detecting certain epigenetic changes, for example, methylation changes.Therefore, the region that may show local amplification in cancer can be included in epigenetic target region set, and can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA and RAF1.For example, in some embodiments, epigenetic target region set comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 or 18 of the above-mentioned targets.
[0365] e. Methylation control region It may be useful to include control regions to facilitate data validation. In some embodiments, the set of epigenetic target regions includes control regions that are expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is derived from cancer cells or normal cells. In some embodiments, the set of epigenetic target regions includes control hypomethylated regions that are expected to be hypomethylated in essentially all samples. In some embodiments, the set of epigenetic target regions includes control hypermethylated regions that are expected to be hypermethylated in essentially all samples.
[0366] 1. Sequence-variable target region set In some embodiments, the set of sequence variable target regions includes multiple regions known to undergo somatic mutations (e.g., single base variations and / or indels) in cancer. The single base variations and / or indels may be relative to a reference sequence, e.g., a published human genome sequence such as the GRCh38 human genome assembly.
[0367] In some embodiments, the sequence variable target region set targets a plurality of different genes or genomic regions ("panel") selected so that a determined proportion of subjects with cancer exhibit genetic variants or tumor markers in one or more different genes or genomic regions within the panel. The panel can be selected to limit the region for sequencing to a fixed number of base pairs. The panel can be selected to sequence a desired amount of DNA, for example, by adjusting the affinity and / or amount of probes, as described elsewhere herein. The panel can also be selected to achieve a desired sequence read data depth. The panel can be selected to achieve a sequence read data depth or sequence read data coverage desired for the amount of base pairs sequenced. The panel can be selected to achieve a theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for detecting one or more genetic variants in a sample.
[0368] The probe for detecting a panel of regions can include those for detecting genomic regions of interest (hotspot regions).Probe design can take into account information about chromatin structure, and / or probe can be designed to maximize the possibility of capturing specific sites (for example, KRAS codons 12 and 13), and can be designed to optimize capture based on the analysis of DNA coverage and fragment size variation, which are affected by nucleosome binding pattern and GC sequence composition.Region as used herein can also include non-hotspot regions that are optimized based on nucleosome position and GC model.
[0369] Examples of lists of genomic locations of interest can be found in Tables 3 and 4. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least a portion of at least one, at least two, or three of the indels in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least a portion of at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes in Table 4. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4.In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 4. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least a portion of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. Each of these genomic locations of interest can be identified as a scaffold region or a hotspot region for a given panel. An example list of hotspot genomic locations of interest can be found in Table 5. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least a portion of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, at least eighteen, at least nineteen, or at least twenty of the genes in Table 5. Each hotspot genomic region is listed with several characteristics, including the associated gene, the chromosome on which it resides, the genomic start and end positions representing the gene's locus, the length of the gene's locus in base pairs, the exons covered by the gene, and important features (e.g., types of mutations) that a given genomic region of interest may seek to capture. [Table 3] [Table 4] [Table 5-1] [Table 5-2] [Table 5-3]
[0370] In addition, or alternatively, suitable target region set can be obtained from literature.For example, Gale et al., PLoS one 13: e0194630 (2018), which is incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of sequence variable target region set.These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53 and U2AF1.
[0371] In some embodiments, the set of sequence variable target regions includes target regions from at least 10, 20, 30, or 35 cancer-associated genes, such as those listed above.
[0372] H. Eligibility In some embodiments, the DNA (e.g., DNA from a buffy coat sample or any other sample comprising cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or the additional DNA) is obtained from a subject with cancer or a precancerous condition, an infection, a transplant rejection, or other disease that directly or indirectly affects the immune system. In some embodiments, the DNA (e.g., DNA from a buffy coat sample or any other sample comprising cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or the additional DNA) is obtained from a subject suspected of having cancer or a precancerous condition, an infection, a transplant rejection, or other disease that directly or indirectly affects the immune system. In some embodiments, the DNA (e.g., DNA from a buffy coat sample or any other sample comprising cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or the additional DNA) is obtained from a subject with a tumor. In some embodiments, DNA (e.g., DNA from a buffy coat sample or any other sample comprising cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or additional DNA) is obtained from a subject suspected of having a tumor. In some embodiments, DNA (e.g., DNA from a buffy coat sample or any other sample comprising cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or additional DNA) is obtained from a subject suspected of having a neoplasm. In some embodiments, DNA (e.g., DNA from a buffy coat sample or any other sample comprising cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or additional DNA) is obtained from a subject suspected of having a neoplasm.In some embodiments, the DNA (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukapheresis sample, or a PBMC sample), and / or additional DNA) is obtained from a subject in remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the lung. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the breast. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.
[0373] I. Pools of DNA from samples or subsamples or portions thereof In some embodiments, the methods herein include preparing one or more pools containing tagged DNA from multiple distributed aliquots. In some embodiments, the pools contain at least a portion of DNA with a low methylation distribution and at least a portion of DNA with a high methylation distribution. Target regions, including epigenetic target regions and / or sequence-variable target regions, can be captured from the pools. The step of capturing a set of target regions from at least one aliquot or portion of a sample or aliquot described elsewhere herein includes a capture step performed on a pool containing DNA from a first aliquot and a second aliquot. Prior to capturing target regions from the pools, a step of amplifying the DNA in the pools can be performed. The capture step can have any of the features described elsewhere herein for the capture step.
[0374] In some embodiments, the method includes preparing a first pool comprising at least a portion of the DNA with a low methylation distribution. In some embodiments, the method includes preparing a second pool comprising at least a portion of the DNA with a high methylation distribution. In some embodiments, the method includes capturing at least a first set of target regions from the first pool, wherein the first set comprises sequence-variable target regions. This capturing step can be preceded by a step of amplifying the DNA in the first pool. In some embodiments, capturing the first set of target regions from the first pool comprises contacting the DNA of the first pool with a first set of target-specific probes, wherein the first set of target-specific probes comprises target-binding probes specific for the sequence-variable target regions. In some embodiments, the method includes capturing a second plurality of sets of target regions from a second pool, wherein the second plurality comprises sequence-variable target regions and epigenetic target regions. This capturing step can be preceded by a step of amplifying the DNA in the second pool. In some embodiments, capturing a second plurality of sets of target regions from the second pool comprises contacting the DNA of the first pool with a second set of target-specific probes, wherein the second set of target-specific probes comprises target-binding probes specific for sequence-variable target regions and target-binding probes specific for epigenetic target regions.
[0375] In some embodiments, a sequence-variable target region is captured from a second portion of the divided aliquots. The second portion may contain some, a majority, substantially all, or all of the DNA from the aliquots that was not included in the pool. The regions captured from the pool and the regions captured from the aliquots may be combined and analyzed in parallel.
[0376] Epigenetic target regions may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from a particular cell or tissue type, or from a tumor, or from a healthy cell, as discussed elsewhere herein. Sequence-variable target regions may exhibit differences in sequence depending on whether they originate from a tumor or from a healthy cell.
[0377] In some applications, the analysis of epigenetic target regions derived from hypomethylated distribution may be less informative than the analysis of sequence variable target regions derived from hypermethylated distribution and hypomethylated distribution, and epigenetic target regions derived from hypermethylated distribution.Therefore, in the method of capturing sequence variable target regions and epigenetic target regions, the latter can be captured to a lesser extent than one or more of the sequence variable target regions from hypermethylated distribution and hypomethylated distribution, and / or the epigenetic target regions can be captured to a lesser extent than the capture of epigenetic target regions from hypermethylated distribution.For example, sequence variable target regions can be captured from a portion of hypomethylated distribution that is not pooled with hypermethylated distribution, and the pool can be prepared using a portion (e.g., majority, substantially all, or all) of DNA derived from hypermethylated distribution, and no or a portion (e.g., minority) of DNA derived from hypomethylated distribution.This approach can reduce or eliminate the sequencing of epigenetic target regions derived from hypomethylated distribution, thereby reducing the amount of sequencing data that is sufficient for further analysis.
[0378] In some embodiments, including a minority of the hypomethylated distribution of DNA in the pool facilitates, e.g., relatively, quantification of one or more epigenetic traits (e.g., methylation or other epigenetic trait(s) discussed in detail elsewhere herein).
[0379] In some embodiments, the pool contains a minority of DNA with a hypomethylated distribution, e.g., less than about 50% of the DNA with a hypomethylated distribution, e.g., less than or equal to about 45%, less than or equal to 40%, less than or equal to 35%, less than or equal to 30%, less than or equal to 25%, less than or equal to 20%, less than or equal to 15%, less than or equal to 10%, or less than or equal to 5% of the DNA with a hypomethylated distribution. In some embodiments, the pool contains about 5% to 25% of the DNA with a hypomethylated distribution. In some embodiments, the pool contains about 10% to 20% of the DNA with a hypomethylated distribution. In some embodiments, the pool contains about 10% of the DNA with a hypomethylated distribution. In some embodiments, the pool contains about 15% of the DNA with a hypomethylated distribution. In some embodiments, the pool contains about 20% of the DNA with a hypomethylated distribution.
[0380] In some embodiments, the pool comprises a portion of the hypermethylated distribution, which may be at least about 50% of the hypermethylated DNA. For example, the pool may comprise at least about 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the hypermethylated DNA. In some embodiments, the pool comprises 50-55%, 55-60%, 60-65%, 65-70%, 70-75%, 75-80%, 80-85%, 85-90%, 90-95%, or 95-100% of the hypermethylated DNA. In some embodiments, the second pool comprises all or substantially all of the hypermethylated DNA.
[0381] In some embodiments, the first pool contains substantially all or all of the DNA with a hypomethylated distribution (e.g., when the second pool does not contain DNA with a hypomethylated distribution). In some embodiments, the second pool does not contain DNA with a hypomethylated distribution (e.g., when the first pool contains substantially all or all of the DNA with a hypomethylated distribution).
[0382] In some embodiments, the second pool contains a portion of the high methylation distribution, which may be any of the values and ranges described above for the low methylation distribution. In some embodiments, the second pool contains all or substantially all of the DNA of the high methylation distribution.
[0383] In exemplary embodiments, after splitting, each split is separately subjected to end repair and ligation to adapters containing molecular barcodes, and then is separately amplified.After amplification, the amplified molecules are enriched (still keeping each split separate).After enrichment, the enriched DNA is pooled according to any of the embodiments described herein, and then amplified again.After amplification, the molecules are sequenced.
[0384] In various embodiments, consistent with the above discussion, the method further includes sequencing the captured DNA, for example, to different degrees of sequencing depth for the set of epigenetic and sequence-variable target regions.
[0385] J. Sequencing Generally, sample nucleic acid, including the nucleic acid flanked by adapters, whether or not it has been previously amplified, can be subjected to sequencing.Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, digital gene expression (Helicos), next-generation sequencing (NGS), single molecule sequencing by synthesis (SMSS) (Helicos), massively parallel sequencing, clonal single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, and using PacBio, SOLiD, Ion Torrent or Nanopore platform sequencing.
[0386] In some embodiments, sequencing includes detecting and / or distinguishing between unmodified and modified nucleic acid bases. For example, PacBio sequencing (e.g., single-molecule real-time (SMRT) sequencing) provides the ability to directly detect, for example, 5-methylcytosine and 5-hydroxymethylcytosine, as well as unmodified cytosine. See, e.g., Schatz., Nature Methods. 14 (4): 347-348 (2017); and US 9,150,918. Also, an Oxford nanopore sequencing system (e.g., MinION sequencer) capable of directly detecting DNA methylation (e.g., 5-methylcytosine and 5-hydroxymethylcytosine) can be used in this embodiment. Sequencing reactions can be performed in various sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means for processing multiple sample sets substantially simultaneously. The sample processing unit may also include multiple sample chambers, allowing for simultaneous processing of multiple runs. Similarly, Ion Torrent sequencing can be used to directly detect methylation. Thus, in some embodiments, methylation status can be determined during sequencing, for example, without or independently of a conversion procedure such as a partitioning step or bisulfite treatment.
[0387] Sequencing reaction can be carried out on one or more forms of nucleic acid, for example, those known to contain markers of cancer or other diseases.Sequencing reaction can also be carried out on any nucleic acid fragment present in sample.In some embodiments, the sequence coverage of genome can be less than 5%, less than 10%, less than 15%, less than 20%, less than 25%, less than 30%, less than 40%, less than 50%, less than 60%, less than 70%, less than 80%, less than 90%, less than 95%, less than 99%, less than 99.9%, or less than 100%.In some embodiments, sequence reaction can provide at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% sequence coverage of genome. Sequence coverage can be performed for at least 5, 10, 20, 70, 100, 200, or 500 different genes, or for at most 5000, 2500, 1000, 500, or 100 different genes.
[0388] Multiplex sequencing can be used to carry out simultaneous sequencing reactions.In some cases, nucleic acid can be sequenced at least 1000 times, 2000 times, 3000 times, 4000 times, 5000 times, 6000 times, 7000 times, 8000 times, 9000 times, 10000 times, 50000 times, 100,000 times sequencing reactions.In other cases, nucleic acid can be sequenced less than 1000 times, less than 2000 times, less than 3000 times, less than 4000 times, less than 5000 times, less than 6000 times, less than 7000 times, less than 8000 times, less than 9000 times, less than 10000 times, less than 50000 times, less than 100,000 times sequencing reactions.Sequencing reactions can be carried out sequentially or simultaneously. Subsequent data analysis can be carried out on all or part of sequencing reaction.In some cases, data analysis can be carried out on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reaction.In other cases, data analysis can be carried out on less than 1000, less than 2000, less than 3000, less than 4000, less than 5000, less than 6000, less than 7000, less than 8000, less than 9000, less than 10000, less than 50000, less than 100,000 sequencing reaction. An exemplary read depth is 1,000 to 50,000 reads per locus (base).
[0389] 1. Differential depth of sequencing In some embodiments, nucleic acids corresponding to the set of sequence variant target regions are sequenced to a sequencing depth that is greater than nucleic acids corresponding to the set of epigenetic target regions. For example, the sequencing depth of nucleic acids corresponding to the set of sequence variant target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x greater, or between 1.25x and 1. The sequencing depth may be 5x, 1.5x to 1.75x, 1.75x to 2x, 2x to 2.25x, 2.25x to 2.5x, 2.5x to 2.75x, 2.75x to 3x, 3x to 3.5x, 3.5x to 4x, 4x to 4.5x, 4.5x to 5x, 5x to 5.5x, 5.5x to 6x, 6x to 7x, 7x to 8x, 8x to 9x, 9x to 10x, 10x to 11x, 11x to 12x, 13x to 14x, 14x to 15x, or 15x to 100x greater. In some embodiments, the sequencing depth is at least 2x greater. In some embodiments, the sequencing depth is at least 5x greater. In some embodiments, the sequencing depth is at least 10x greater. In some embodiments, the sequencing depth is 4-10 times greater, hi some embodiments, the sequencing depth is 4-100 times greater.
[0390] In some embodiments, DNA corresponding to a set of sequence variable target regions and / or a set of epigenetic target regions are sequenced together, e.g., in the same sequencing cell (e.g., a flow cell of an Illumina sequencer), and / or in the same composition, which may be a combined or pooled composition resulting from recombining separately captured sets, or a composition obtained by capturing DNA corresponding to a set of sequence variable target regions (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample), and / or additional DNA) and / or captured DNA corresponding to a set of epigenetic target regions in the same container.
[0391] K.Analysis In some embodiments, the methods described herein include determining the level of a particular immune cell type from which the DNA originates. The immune cell types may include naive lymphocytes, activated lymphocytes, myeloid cells at different points of differentiation, and / or other types described elsewhere herein. In some methods, determining the level of an immune cell type facilitates determining the likelihood that the subject from which the DNA was obtained has a disease or disorder related to the immune system, such as an infection, transplant rejection, or a cancer or precancerous condition.
[0392] In some embodiments, the methods described herein include identifying the presence of DNA produced by a tumor (or neoplastic or cancer cell) or by a precancerous cell. In some embodiments, the methods described herein include identifying the presence of DNA produced by an immune cell that is not a tumor cell, cancer cell, or precancerous cell. In some such embodiments, determining immune cell distribution facilitates the detection or diagnosis of cancer or precancerous condition, or the prognosis of cancer or determining cancer treatment options (e.g., predicting the clinical outcome of a therapy, e.g., chemotherapy or immunotherapy, in a subject). For example, determining the ratio of different immune cell types can facilitate such detection or determination. In some embodiments, the numerator of the ratio is the number or relative number of neutrophils, monocytes, or both, and the denominator of the ratio is the number or relative number of T cells, B cells, NK cells, or total lymphocytes. In some embodiments, the numerator of the ratio is the number or relative number of neutrophils, and the denominator of the ratio is the number or relative number of T cells, B cells, NK cells, or total lymphocytes. In some embodiments, the numerator of the ratio is the number or relative number of monocytes, and the denominator of the ratio is the number or relative number of T cells, B cells, NK cells, or total lymphocytes. In some embodiments, the numerator of the ratio is the number or relative number of neutrophils and monocytes, and the denominator of the ratio is the number or relative number of T cells, B cells, NK cells, or total lymphocytes. In some embodiments, the ratio is the ratio of neutrophils to lymphocytes. In some embodiments, the ratio is the ratio of NK cells to total lymphocytes. In some embodiments, the numerator of the ratio is the number or relative number of NK cells, and the denominator of the ratio is the number or relative number of total lymphocytes. In some embodiments, the ratio is the ratio of M1 macrophages to M2 macrophages. In some embodiments, the numerator of the ratio is the number or relative number of M1 macrophages, and the denominator of the ratio is the number or relative number of M2 macrophages. In some embodiments, the ratio is the ratio of monocytes to T cells. In some embodiments, an increase in such a ratio is associated with cancer. In other embodiments, a decrease in such ratio is associated with cancer.
[0393] The present method can be used to diagnose the presence of a condition, particularly a cancer or precancerous condition, in a subject, characterize the condition (e.g., stage the cancer or determine the heterogeneity of the cancer), monitor the response to treatment of the condition, or prognose the risk of developing the condition or the subsequent course of the condition. The present disclosure can also be useful for determining the effectiveness of a particular treatment option. In a successful treatment option, if the treatment is successful, more cancer cells may be killed and DNA may be shed, resulting in an increase in the amount of copy number variations or rare mutations detected in the subject's blood. In other examples, this may not occur. In another example, perhaps a particular treatment option can be correlated with the genetic profile of the cancer over time. This correlation can be useful for selecting a treatment.
[0394] Additionally, if the cancer is observed to be in remission after treatment, the method can be used to monitor for residual disease or recurrence of the disease.
[0395] The types and number of cancers that can be detected can include blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc. The type and / or stage of cancer can be detected from genetic variations, including mutations, rare mutations, indels, copy number variations, transversions, translocations, recombinations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, alterations in chromosomal structure, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0396] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous, both in composition and stage classification. Genetic profile data can enable characterization of specific subtypes of cancer, which may be important for diagnosing or treating that specific subtype. This information can also provide the subject or practitioner with clues regarding the prognosis of a particular cancer type, allowing either the subject or practitioner to tailor treatment options as the disease progresses. Some cancers can become more aggressive and genetically unstable as they progress. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure can be useful for determining disease progression.
[0397] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of an abnormal condition in a subject. Such a method can include, for example, generating a genetic profile of extracellular polynucleotides from the subject, where the genetic profile includes multiple data obtained by analyzing copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer or precancer. In some embodiments, the abnormal condition can result in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells at different stages of cancer. In other examples, the heterogeneity can constitute multiple disease foci. Furthermore, in the example of cancer, there can be multiple tumor foci, possibly the result of metastasis, where one or more foci have spread from the primary site.
[0398] The method can be used to generate a profile, fingerprint, or set of data that compiles genetic information from different cells in a heterogeneous disease. Such a data set can include analysis of copy number variations, epigenetic variations, or other mutations, alone or in combination.
[0399] This method can be used to diagnose, prognose, monitor or observe cancer or other diseases.In some embodiments, the method herein does not involve diagnosing, prognosing or monitoring fetus, and therefore is not intended for non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in prenatal subjects, where DNA and other polynucleotides can co-circulate with maternal molecules.
[0400] Generally, after sequencing, analysis of reads can be performed at the level of each distribution and at the level of the whole DNA population. Tags can be used to sort out the reads that come from different distributions. Analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, genome coordinate length, coverage, and / or copy number. In some embodiments, higher coverage can be correlated with higher nucleosome occupancy in genome regions, while lower coverage can be correlated with lower nucleosome occupancy or nucleosome depleted regions (NDRs).
[0401] An exemplary method for identification of immune cell types by NGS includes the following steps: 1. Preparing an extracted DNA sample (e.g., DNA isolated from a human buffy coat sample or any other sample containing cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), and / or additional DNA) by ligating an adapter comprising a molecular tag to the DNA. In some embodiments, e.g., embodiments using DNA isolated from a buffy coat sample or any other sample containing cells, e.g., a blood sample (e.g., a whole blood sample, a leukocyte-reduced sample, or a PBMC sample), the DNA is fragmented, e.g., by sonication or restriction digestion, prior to this ligation. 2. Partitioning the adaptor-ligated DNA into a plurality of differentially methylated subsamples by contacting the DNA with an agent that recognizes modified cytosines, such as methylcytosines, within the DNA. 4. Capturing hypermethylated and hypomethylated variable target regions from the distributed aliquots by contacting the aliquots with target-specific probes. 5. Enriching and / or eluting DNA containing the target region. 6. The DNA is amplified and assayed multiplexed on an NGS instrument. 7. Analyzing the NGS data with molecular tags used to identify unique molecules.
[0402] An exemplary workflow is shown in Figure 1A, in which DNA obtained from a sample (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample), and / or additional DNA) is partitioned based on the methylation level of the DNA molecules. Molecular barcodes are then added to the DNA molecules in each partition, and epigenetic and, optionally, sequence-variable target regions are then captured from the partitioned aliquots to obtain a targeted library. The epigenetic target regions include hypermethylated and / or hypomethylated variable target regions that are differentially methylated in certain types of immune cells. The targeted library can be amplified prior to sequencing, which then provides sequence information that is used to determine the level of immune cell type and / or the likelihood of the presence of a disease or disorder.
[0403] In some embodiments of the methods described herein, the molecular tag consists of nucleotides that are not modified by a procedure that affects a first nucleobase in DNA differently than a second nucleobase in DNA, e.g., any of those described herein (e.g., if the procedure is a bisulfite conversion or any other conversion that does not affect mC, then mC with A, T, and G; if the procedure is a conversion that does not affect hmC, then hmC with A, T, and G, etc.). In some embodiments of the methods described herein, the molecular tag does not include nucleotides that are modified by a procedure that affects a first nucleobase in DNA differently than a second nucleobase in DNA, e.g., any of those described herein (e.g., if the procedure is a bisulfite conversion or any other conversion that affects C, then the tag does not include an unmodified C; if the procedure is a conversion that affects mC, then the tag does not include mC; if the procedure is a conversion that affects hmC, then the tag does not include hmC, etc.).
[0404] II. ADDITIONAL FEATURES OF CERTAIN DISCLOSED METHODS A. Sample The sample can be any biological sample isolated from a subject. The sample can be a bodily sample. Samples can include body tissues, such as known or suspected solid tumors (e.g., carcinoma, adenocarcinoma, or sarcoma), whole blood, buffy coat, PBMC, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid, fluid in the space between cells, including gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a bodily fluid, particularly blood and its fractions and urine. The sample can be in the form originally isolated from the subject, or can be further processed to remove or add components, such as cells, or to enrich one component relative to another. Thus, preferred body fluids for analysis include cell-containing body fluids such as whole blood, buffy coat isolated from whole blood, PBMCs isolated from whole blood, leukocyte-reduced samples, and / or plasma or serum.
[0405] In some embodiments, the population of nucleic acids is obtained from a serum, plasma, or blood sample (e.g., a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukapheresis sample, or a PBMC sample)) from a subject suspected of having or previously diagnosed with a neoplasm, tumor, precancerous condition, or cancer. The population includes nucleic acids with varying levels of sequence variation, epigenetic variation, and / or replication or post-transcriptional modifications. Post-replicative modifications include cytosine modifications, particularly those at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0406] The sample can be isolated or obtained from a subject and transported to the site of sample analysis. The sample can be stored or shipped at a desired temperature, for example, room temperature, 4°C, -20°C, and / or -80°C. The sample can be isolated or obtained from a subject at the site of sample analysis. The subject can be a human, mammal, animal, companion animal, service animal, or pet. The subject can have cancer, a precancerous condition, an infection, a transplant rejection, or other disease or disorder associated with an altered immune system. The subject may not have cancer or detectable symptoms of cancer. The subject may have been treated with one or more cancer therapies, such as any one or more of chemotherapy, antibodies, vaccines, or biologics. The subject may be in remission. The subject may or may not have been diagnosed with cancer or a susceptibility to any cancer-related genetic mutation / disorder.
[0407] In some embodiments, the sample comprises plasma. The volume of plasma obtained can depend on the desired read depth of the region to be sequenced. Exemplary volumes are 0.4-40 mL, 5-20 mL, 10-20 mL, and 3-5 mL. For example, the volume can be 0.5 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 6 mL, 7 mL, 8 mL, 9 mL, 10 mL, 20 mL, 30 mL, or 40 mL. An exemplary volume of plasma sampled can be 5-20 mL. In some embodiments, the sample volume is 3-5 mL of plasma, e.g., 4 mL of plasma, per 10 mL of whole blood.
[0408] In some embodiments, the sample comprises whole blood. Exemplary volumes of whole blood sampled are 0.4 to 40 mL, 5 to 20 mL, 10 to 20 mL, 1 to 6 mL, 1 to 3 mL, and 3 to 5 mL. For example, the volume can be 0.5 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 6 mL, 7 mL, 8 mL, 9 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of whole blood sampled can be 5 to 20 mL. In some embodiments, the sample volume is 1 to 5 mL of whole blood, e.g., 2.5 mL of whole blood.
[0409] In some embodiments, the sample comprises buffy coat separated from whole blood. Exemplary volumes of buffy coat sampled are 0.1 to 20 mL, 1 to 10 mL, 1 to 5 mL, 0.2 to 0.6 mL, and 0.3 to 0.5 mL. For example, the volume can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of buffy coat sampled can be 1 to 10 mL. In some embodiments, the sample volume is 0.1 to 0.5 mL of buffy coat, e.g., 0.3 mL of buffy coat, per 10 mL of whole blood.
[0410] In some embodiments, the sample comprises PBMCs isolated from whole blood. Exemplary volumes of PBMCs sampled are 0.1 to 20 mL, 1 to 10 mL, 1 to 5 mL, 0.2 to 0.6 mL, and 0.3 to 0.5 mL. For example, the volume can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of PBMCs sampled can be 1 to 10 mL. In some embodiments, the sample volume is 0.1 to 0.5 mL of PBMCs, e.g., 0.3 mL of PBMCs, per 10 mL of whole blood.
[0411] In some embodiments, the sample comprises leukocytes separated from the subject's blood using leukocyte reduction. Exemplary volumes of leukocytes sampled from leukocyte reduction are 0.1-20 mL, 1-10 mL, 1-5 mL, 0.2-0.6 mL, and 0.3-0.5 mL. For example, the volume can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of leukocytes sampled from leukocyte reduction can be 1-10 mL. In some embodiments, the sample volume is 0.1-0.6 mL of leukocytes from leukocyte reduction, e.g., 0.4 mL of leukocytes, per 10 mL of whole blood.
[0412] A sample can contain various amounts of nucleic acid, including genome equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genome equivalents. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents.
[0413] The sample may contain nucleic acids from different sources, e.g., cells of the same subject or cells of different subjects. The sample may contain nucleic acids having mutations. For example, the sample may contain DNA having germline mutations and / or somatic mutations. A germline mutation refers to a mutation present in the germline DNA of a subject. A somatic mutation refers to a mutation originating from the somatic cells of a subject, e.g., a precancerous cell or a cancer cell. The sample may contain DNA having a cancer-associated mutation (e.g., a cancer-associated somatic mutation). The sample may contain epigenetic variants (i.e., chemical or protein modifications), where the epigenetic variants are associated with the presence of a genetic variant, e.g., a cancer-associated mutation. In some embodiments, the sample contains epigenetic variants associated with the presence of a genetic variant, where the sample does not contain a genetic variant.
[0414] Exemplary amounts of nucleic acid (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduced sample, or a PBMC sample)) in a sample prior to amplification range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 150 ng, or at least 200 ng of nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of nucleic acid molecule. The method can include obtaining between 1 femtogram (fg) and 200 ng.
[0415] Nucleic acids can be isolated from cells, for example, cells in body fluids. Cells can be lysed and the cellular nucleic acids can be processed. Generally, after adding buffer and washing steps, nucleic acids can be precipitated with alcohol. Additional washing steps, such as silica-based columns, can be used to remove contaminants or salts. Non-specific bulk carrier nucleic acids, such as C1 DNA, DNA, or proteins, for bisulfite sequencing, hybridization, and / or ligation can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.
[0416] After such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA may be converted to double-stranded form so that they can be included in subsequent processing and analysis steps.
[0417] A DNA molecule can be ligated to an adaptor at one or both ends. Typically, double-stranded molecules are blunt-ended by treatment with a polymerase having a 5'-3' polymerase and a 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenow large fragment and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecule can be ligated with an at least partially double-stranded adaptor (e.g., a Y-shaped or bell-shaped adaptor). Alternatively, complementary nucleotides can be added to the blunt ends of the sample nucleic acid and the adaptor to facilitate ligation. Both blunt-end ligation and sticky-end ligation are contemplated herein. In blunt-end ligation, both the nucleic acid molecule and the adaptor tag have blunt ends. In sticky-end ligation, the nucleic acid molecule typically has an "A" overhang, and the adaptor typically has a "T" overhang.
[0418] B. Amplification The sample nucleic acid flanked by adapters can be amplified by PCR and other amplification methods. Amplification is typically primed by primers that anneal or bind to primer-binding sites in the adapters adjacent to the DNA molecule to be amplified. The amplification method can involve cycles of denaturation, annealing, and extension as a result of thermocycling, or can be isothermal, as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication. In some embodiments, the method performs dsDNA ligation using T-tail and C-tail adapters, resulting in amplification of at least 50, 60, 70, or 80% of the double-stranded nucleic acid before ligation to the adapter. Preferably, the method increases the amount or number of amplified molecules by at least 10, 15, or 20% compared to a control method using T-tail adapters alone.
[0419] C. Capture part As discussed above, nucleic acids in a sample can be subjected to a capture step, in which molecules having target regions are captured for subsequent analysis.Target capture can involve the use of a capture moiety, such as a probe (e.g., an oligonucleotide) labeled with biotin, and a second moiety or binding partner, such as streptavidin, that binds to the capture moiety.In some embodiments, the capture moiety and binding partner can have higher and lower capture yields for different sets of target regions, such as a sequence variable target region set and an epigenetic target region set, as discussed elsewhere herein.Methods involving capture moieties are further described, for example, in U.S. Patent No. 9,850,523, issued on December 26, 2017, which is incorporated herein by reference.
[0420] Capture moieties include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically adsorbable particles. Extraction moieties can be members of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, a capture moiety attached to an analyte is captured by its binding pair attached to an isolable moiety, such as a magnetically adsorbable particle or a large particle that can be sedimented by centrifugation. A capture moiety can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties are biotin, which allows affinity separation by binding to streptavidin that is or can be linked to a solid phase, or oligonucleotides, which allow affinity separation through binding to complementary oligonucleotides that are or can be linked to a solid phase.
[0421] D. Collection of target-specific probes In some embodiments, a collection of target-specific probes is used in a method comprising a set of epigenetic target regions and / or a set of sequence-variable target regions described herein. In some embodiments, the collection of target-specific probes comprises target-binding probes specific for the set of sequence-variable target regions and target-binding probes specific for the set of epigenetic target regions. In some embodiments, the capture yield of the target-binding probes specific for the set of sequence-variable target regions is higher (e.g., at least two-fold higher) than the capture yield of the target-binding probes specific for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for the set of sequence-variable target regions that is higher (e.g., at least two-fold higher) than its capture yield specific for the set of epigenetic target regions.
[0422] In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x higher than the capture yield of target binding probes specific for the set of epigenetic target regions. In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is 1.25x to 1.5x, 1.5x to 1.75x, 1.75x to 2x, 2x to 2.25x, 2.25x to 2.5x, 2.5x to 2.75x, 2.75x to 3x, 3x to 3.5x, 3.5x to 4x, 4x to 4.5x, 4.5x to 5x, 5x to 5.5x, 5.5x to 6x, 6x to 7x, 7x to 8x, 8x to 9x, 9x to 10x, 10x to 11x, 11x to 12x, 13x to 14x, or 14x to 15x higher than the capture yield of target binding probes specific for the set of epigenetic target regions.
[0423] In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x higher than its capture yield for a set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is 1.25x to 1.5x, 1.5x to 1.75x, 1.75x to 2x, 2x to 2.25x, 2.25x to 2.5x, 2.5x to 2.75x, 2.75x to 3x, 3x to 3.5x, 3.5x to 4x, 4x to 4.5x, 4.5x to 5x, 5x to 5.5x, 5.5x to 6x, 6x to 7x, 7x to 8x, 8x to 9x, 9x to 10x, 10x to 11x, 11x to 12x, 13x to 14x, or 14x to 15x higher than its capture yield specific for the set of epigenetic target regions.
[0424] A collection of probes can be configured to provide higher capture yields for a set of sequence-variable target regions in a variety of ways, including concentration, varying length, and / or chemistry (e.g., to affect affinity), and combinations thereof. Affinity can be modulated by adjusting the length of the probe and / or including nucleotide modifications, as discussed below.
[0425] In some embodiments, the target-specific probes specific for the set of sequence variable target regions are present at a higher concentration than the target-specific probes specific for the set of epigenetic target regions, hi some embodiments, the concentration of target-binding probes specific for the set of sequence variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the concentration of target-binding probes specific for the set of epigenetic target regions. In some embodiments, the concentration of target-binding probes specific for the set of sequence-variable target regions is 1.25x to 1.5x, 1.5x to 1.75x, 1.75x to 2x, 2x to 2.25x, 2.25x to 2.5x, 2.5x to 2.75x, 2.75x to 3x, 3x to 3.5x, 3.5x to 4x, 4x to 4.5x, 4.5x to 5x, 5x to 5.5x, 5.5x to 6x, 6x to 7x, 7x to 8x, 8x to 9x, 9x to 10x, 10x to 11x, 11x to 12x, 13x to 14x, or 14x to 15x higher than the concentration of target-binding probes specific for the set of epigenetic target regions. In such embodiments, concentration may refer to the average mass per volume concentration of the individual probes in each set.
[0426] In some embodiments, target-specific probes specific to a set of sequence-variable target regions have higher affinity for their targets than target-specific probes specific to a set of epigenetic target regions.Affinity can be modulated by any method known to those skilled in the art, including by using different probe chemistries.Certain nucleotide modifications, such as cytosine 5-methylation (in the context of a certain sequence), modifications that provide a heteroatom in the 2' sugar, and LNA nucleotides, can increase the stability of double-stranded nucleic acids, and oligonucleotides with such modifications have been shown to have relatively high affinity for their complementary sequences.See, for example, Severin et al., Nucleic Acids Res. 39: 8740-8751 (2011); Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997); U.S. Patent No. 9,738,894.Similarly, longer sequence lengths generally provide increased affinity. Other nucleotide modifications, such as substituting the nucleobase hypoxanthine for guanine, reduce affinity by decreasing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, target-specific probes specific for a set of sequence-variable target regions have modifications that increase their affinity for their targets. In some embodiments, alternatively or additionally, target-specific probes specific for a set of epigenetic target regions have modifications that decrease their affinity for their targets. In some embodiments, target-specific probes specific for a set of sequence-variable target regions have a longer average length and / or a higher average melting temperature than target-specific probes specific for a set of epigenetic target regions. These embodiments may be combined with each other and / or with the concentration differences discussed above to achieve a desired fold difference in capture yield, such as any of the fold differences or ranges described above.
[0427] In some embodiments, the target-specific probe comprises a capture moiety. The capture moiety may be any of the capture molecules described herein, such as biotin. In some embodiments, the target-specific probe is covalently or non-covalently linked to a solid support, such as through the interaction of the binding pair of the capture moiety. In some embodiments, the solid support is a bead, such as a magnetic bead.
[0428] In some embodiments, the target-specific probes specific for the set of sequence-variable target regions and / or the target-specific probes specific for the set of epigenetic target regions comprise capture moieties as discussed above, e.g., capture moieties and probes comprising sequences selected for tiling across a panel of regions such as genes.
[0429] In some embodiments, the target-specific probes are provided in a single composition. The single composition can be a solution (liquid or frozen). Alternatively, it can be a lyophilizate.
[0430] Alternatively, target-specific probes can be provided as multiple compositions, including, for example, a first composition containing probes specific to a set of epigenetic target regions and a second composition containing probes specific to a set of sequence-variable target regions. These probes can be mixed in appropriate proportions to provide a combined probe composition with any of the aforementioned fold differences in concentration and / or capture yield. Alternatively, they can be used in separate capture procedures (e.g., using aliquots of a sample or sequentially using the same sample) to provide first and second compositions containing captured epigenetic and sequence-variable target regions, respectively.
[0431] 1. Probes specific to epigenetic target regions The probes for the set of epigenetic target regions may include probes specific for one or more types of target regions that may distinguish DNA originating from different types of immune cells, including rare immune cell types, and / or distinguish DNA from precancerous or neoplastic (e.g., tumor or cancer) cells from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The probes for the set of epigenetic target regions may also include probes for one or more control regions, e.g., as described herein.
[0432] In some embodiments, the probes for the epigenetic target region probe set have a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the probes for the epigenetic target region set have a footprint in the range of 100 to 1,000 kb, e.g., 100 to 200 kb, 200 to 300 kb, 300 to 400 kb, 400 to 500 kb, 500 to 600 kb, 600 to 700 kb, 700 to 800 kb, 800 to 900 kb, and 900 to 1,000 kb. In some embodiments, the probes for the epigenetic target region probe set have a footprint of at least 5 kb, e.g., at least 10 kb, 20 kb, or 50 kb. In some embodiments, probes for an epigenetic target region set have a footprint of about 100 kb to about 1000 kb, e.g., about 200 kb to about 1000 kb, about 300 kb to about 1000 kb, about 400 kb to about 1000 kb, about 500 kb to about 1000 kb, about 600 kb to about 1000 kb, or about 700 kb to about 1000 kb. In other embodiments, probes for an epigenetic target region probe set have a footprint of about 100 kb to about 500 kb, about 200 kb to about 500 kb, about 300 kb to about 500 kb, about 400 kb to about 600 kb, or about 250 kb to about 750 kb. In some embodiments, the probes for an epigenetic target region probe set have a footprint of about 100 kb, about 150 kb, about 200 kb, about 250 kb, about 300 kb, about 350 kb, about 400 kb, about 450 kb, about 500 kb, about 550 kb, about 600 kb, about 650 kb, about 700 kb, about 750 kb, about 800 kb, about 850 kb, about 900 kb, about 950 kb, or about 1000 kb.
[0433] a. Hypermethylated variable target region In some embodiments, the probes for the set of epigenetic target regions include probes specific for one or more hypermethylated variable target regions. The hypermethylated variable target regions may be any of those described above. For example, in some embodiments, the probes specific for the hypermethylated variable target regions include probes specific for multiple loci that are differentially methylated in different immune cell types. In some embodiments, each immune cell type-specific hypermethylated variable target region includes at least one CpG site that is methylated at a frequency greater than or equal to 0.3, 0.4, 0.5, or 0.6 in one immune cell type and less than or equal to 0.1, 0.2, or 0.3 in all other immune cell types. In some embodiments, each immune cell type-specific hypermethylated variable target region comprises at least two CpG sites within 100 base pairs of each other, each of which is methylated at a frequency greater than or equal to 0.3, 0.4, 0.5, or 0.6 in one immune cell type and less than or equal to 0.1, 0.2, or 0.3 in all other immune cell types. In some such embodiments, each immune cell type-specific hypermethylated variable target region comprises a total of at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites within 150 or 200 base pairs, wherein fewer than three of the at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites are methylated at a frequency greater than 0.1, 0.2, or 0.3 in any normal tissue type. In some embodiments, each set of immune cell type-specific epigenetic target regions comprises at least 3, at least 5, at least 10, at least 20, or at least 30 hypermethylated variable target regions that are uniquely hypermethylated in every single immune cell type identified in the method.
[0434] In some embodiments, probes specific for hypermethylated variable target regions comprise probes specific for a plurality of loci listed in Table 1, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, probes specific for hypermethylated variable target regions comprise probes specific for a plurality of loci listed in Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. In some embodiments, probes specific for hypermethylated variable target regions include probes specific for multiple loci listed in Table 1 or Table 2, for example, probes specific for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2.
[0435] In some embodiments, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site of the gene and the stop codon (or the final stop codon in the case of an alternatively spliced gene). In some embodiments, the one or more probes bind within 300 bp, e.g., within 200 or 100 bp, of the listed position. In some embodiments, the probes have hybridization sites that overlap with the listed positions. In some embodiments, the probes specific for the hypermethylated target regions include probes specific for one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0436] b. Hypomethylated variable target region In some embodiments, the probes for the set of epigenetic target regions include probes specific for one or more hypomethylated variable target regions. The hypomethylated variable target regions may be any of those described above. For example, in some embodiments, the probes specific for the hypomethylated variable target regions include probes specific for multiple loci that are differentially methylated in different immune cell types. In some embodiments, each immune cell type-specific hypomethylated variable target region includes at least one CpG site that is methylated at a frequency lower than or equal to 0.1, 0.2, or 0.3 in one immune cell type and at a frequency higher than or equal to 0.3, 0.4, 0.5, or 0.6 in all other immune cell types. In some embodiments, each immune cell type-specific hypomethylated variable target region comprises at least two CpG sites within 100 base pairs of each other, each of which is methylated at a frequency less than or equal to 0.1, 0.2, or 0.3 in one immune cell type and at a frequency greater than or equal to 0.3, 0.4, 0.5, or 0.6 in all other immune cell types. In some such embodiments, each immune cell type-specific hypomethylated variable target region comprises a total of at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites within 150 or 200 base pairs, wherein fewer than 3 of the at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites are methylated at a frequency less than 0.1, 0.2, or 0.3 in any normal tissue type. In some embodiments, each set of immune cell type-specific epigenetic target regions comprises at least 3, at least 5, at least 10, at least 20, or at least 30 hypomethylated variable target regions that are uniquely hypomethylated in every single one of the immune cell types identified in the method.
[0437] In some embodiments, probes specific for one or more hypomethylated variable target regions may include probes for regions such as repetitive elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, and intergenic regions that are normally methylated in healthy cells may show reduced methylation in tumor cells.
[0438] In some embodiments, the probes specific for the hypomethylated variable target region comprise probes specific for repetitive elements and / or intergenic regions, hi some embodiments, the probes specific for repetitive elements comprise probes specific for one, two, three, four, or five of the following: a LINE1 element, an Alu element, a centromeric tandem repeat, a pericentromeric tandem repeat, and / or satellite DNA.
[0439] Exemplary probes specific for genomic regions exhibiting cancer-associated hypomethylation include probes specific for nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1. In some embodiments, probes specific for hypomethylated variable target regions include probes specific for regions overlapping with or including nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1.
[0440] c.CTCF binding region In some embodiments, the probes of the set of epigenetic target regions include probes specific for CTCF binding regions. In some embodiments, the probes specific for CTCF binding regions include probes specific for at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those described above, or in one or more of the CTCFBSDB, or the above-cited Cuddapah et al., Martin et al., or Rhee et al. articles. In some embodiments, the probes of the set of epigenetic target regions include at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding site.
[0441] d. transcription start site In some embodiments, the probes of the set of epigenetic target regions include probes specific for transcription start sites. In some embodiments, the probes specific for transcription start sites include probes specific for at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those listed in the DBTSS. In some embodiments, the probes of the set of epigenetic target regions include probes for sequences at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the transcription start site.
[0442] e. local amplification As mentioned above, local amplification is a somatic mutation, but it can be detected by sequencing based on the frequency of read data in a similar manner to the approach for detecting certain epigenetic changes such as methylation changes.Therefore, the region that may show local amplification in cancer can be included in the epigenetic target region set as discussed above.In some embodiments, the probe specific to the epigenetic target region set comprises the probe specific to local amplification.In some embodiments, the probe specific to local amplification comprises the probe specific to one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA and RAF1. For example, in some embodiments, probes specific for local amplification include probes specific for one or more of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the aforementioned targets.
[0443] f. Control area It may be useful to include control regions to facilitate data validation. In some embodiments, the probes specific for the set of epigenetic target regions include probes specific for control methylated regions that are expected to be methylated in essentially all samples. In some embodiments, the probes specific for the set of epigenetic target regions include probes specific for control hypomethylated regions that are expected to be hypomethylated in essentially all samples.
[0444] 2. Probes specific to sequence-variable target regions The probes of the sequence variable target region set may include probes specific for multiple regions known to undergo somatic mutation in cancer. The probes may be specific for any of the sequence variable target region sets described herein. Exemplary sequence variable target region sets are discussed in detail herein, for example, in the section above regarding captured sets.
[0445] In some embodiments, sequence variable target region probe sets have a footprint of at least 0.5 kb, e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, sequence variable target region probe sets have a footprint in the range of 0.5 to 100 kb, e.g., 0.5 to 2 kb, 2 to 10 kb, 10 to 20 kb, 20 to 30 kb, 30 to 40 kb, 40 to 50 kb, 50 to 60 kb, 60 to 70 kb, 70 to 80 kb, 80 to 90 kb, and 90 to 100 kb. In some embodiments, the sequence variable target region probe set has a footprint of about 10 kb to about 100 kb, e.g., about 20 kb to about 100 kb, about 30 kb to about 100 kb, about 40 kb to about 100 kb, about 50 kb to about 100 kb, about 60 kb to about 100 kb, or about 70 kb to about 100 kb. In other embodiments, the footprint of the sequence variable target region probe set is about 10 kb to about 50 kb, about 20 kb to about 50 kb, about 30 kb to about 50 kb, about 40 kb to about 60 kb, or about 25 kb to about 75 kb. In some embodiments, the footprint of a sequence variable target region probe set is about 10 kb, about 15 kb, about 20 kb, about 25 kb, about 30 kb, about 35 kb, about 40 kb, about 45 kb, about 50 kb, about 55 kb, about 60 kb, about 65 kb, about 70 kb, about 75 kb, about 80 kb, about 85 kb, about 90 kb, about 95 kb, or about 100 kb.
[0446] In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or at least 70 of the genes in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or at least 70 of the SNVs in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least one, at least two, or three indels in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes in Table 4. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4.In some embodiments, probes specific for a set of sequence variable target regions include probes specific for at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 4. In some embodiments, probes specific for a set of sequence variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5.
[0447] In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for target regions from at least 10, 20, 30, or 35 cancer-associated genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0448] E. Computer Systems The methods of the present disclosure can be implemented using or with the aid of a computer system. Figure 2 shows a computer system 201 programmed or otherwise configured to implement the methods of the present disclosure. The computer system 201 may coordinate various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 201 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing, e.g., according to any of the methods disclosed herein.
[0449] The computer system 201 includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") 205, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 201 also includes memory or memory locations 210 (e.g., random access memory, read-only memory, flash memory), electronic storage 215 (e.g., a hard disk), a communication interface 220 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 225, such as cache, other memory, data storage, and / or an electronic display adapter. The memory 210, storage 215, interface 220, and peripheral devices 225 communicate with the CPU 205 via a bus (solid lines), such as a communications network or a motherboard. The storage 215 may be a data storage device (or data repository) for storing data. The computer system 201 may be operably coupled to a computer network 230 utilizing the communication interface 220. Computer network 230 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. Computer network 230, in some cases, is a telecommunications and / or data network. Computer network 230 may include one or more computer servers, which may enable distributed computing such as cloud computing. Computer network 230, in some cases, may utilize computer system 201 to implement a peer-to-peer network, which may enable devices coupled to computer system 201 to act as clients or servers.
[0450] CPU 205 may execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 210. Examples of operations performed by CPU 205 may include fetch, decode, execute, and writeback.
[0451] The storage device 215 may store files, such as drivers, libraries, and saved programs. The storage device 215 may store programs and recorded sessions generated by users, as well as output(s) associated with the programs. The storage device 215 may store user data, such as user preferences and user programs. The computer system 201 may, in some cases, include one or more additional data storage devices located external to the computer system 201, for example, on a remote server with which the computer system 201 communicates over an intranet or the Internet. Data can be transferred from one location to another, for example, by using a communications network or physical data migration (e.g., using a hard drive, thumb drive, or other data storage mechanism).
[0452] Computer system 201 may communicate with one or more remote computer systems over network 230. In some embodiments, computer system 201 may communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android®-enabled device, a Blackberry®), or a personal digital assistant. A user may access computer system 201 over network 230.
[0453] The methods described herein may be implemented by code executable by a machine (e.g., a computer processor) stored in an electronic storage location of the computer system 201, such as memory 210 or electronic storage 215. The machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by the processor 205. In some cases, the code may be retrieved from storage 215 and stored in memory 210 for ready access by the processor 205. In some situations, the electronic storage 215 may be omitted, and the machine-executable instructions may be stored in memory 210.
[0454] In certain aspects, the disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: sequencing DNA obtained from a buffy coat sample or any other sample comprising cells, e.g., a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample) from a subject; determining methylation levels for a set of epigenetic target regions, the set including a plurality of target regions comprising DNA sequences that are differentially methylated in a plur...
Claims
1. 1. A method for analyzing DNA, comprising: a) sequencing the DNA to determine methylation levels for a set of epigenetic target regions comprising a plurality of target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types, wherein the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) determining the quantity of each of a plurality of immune cell types from which the DNA originates based on the methylation level; A method comprising:
2. 1. A method for analyzing DNA, comprising: a) capturing at least a set of epigenetic target regions from the DNA or subsample thereof, comprising contacting the DNA or subsample thereof with target-specific probes specific for the at least one set of epigenetic target regions, wherein the DNA is derived from a buffy coat sample or the DNA is derived from cells of a sample, and the set of epigenetic target regions includes target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types; b) determining the methylation level for said target region; c) determining the quantity of each of the plurality of immune cell types from which the DNA originates; A method comprising:
3. 1. A method for analyzing DNA, comprising: a) sequencing the DNA to determine methylation levels for a set of epigenetic target regions comprising a plurality of hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in a plurality of immune cell types, wherein the DNA is from a buffy coat sample or the DNA is from cells of a sample; b) determining the quantity of each of a plurality of immune cell types from which the DNA originates based on the methylation level; A method comprising:
4. 1. A method for analyzing DNA, comprising: a) capturing at least a set of epigenetic target regions from said DNA or subsample thereof, comprising contacting said DNA or subsample thereof with target-specific probes specific for said at least one set of epigenetic target regions, said set of epigenetic target regions comprising hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in a plurality of immune cell types, said DNA being derived from a buffy coat sample, or said DNA being derived from cells of a sample; b) determining the methylation level for said target region; c) determining the quantity of each of the plurality of immune cell types from which the DNA originates; A method comprising:
5. 1. A method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with the modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of a sample; b) sequencing DNA from one or more of said plurality of aliquots; c) detecting the level of DNA sequences to determine the quantity of each of a plurality of immune cell types from which the DNA originates; A method comprising:
6. 1. A method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with the modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of a sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, comprising contacting the DNA with target-specific probes specific for the at least one set of epigenetic target regions, wherein target regions of the set of epigenetic target regions comprise DNA sequences that are differentially methylated in a plurality of immune cell types; c) sequencing the captured DNA to determine the levels of each of the multiple immune cell types from which the DNA originated; A method comprising:
7. 1. A method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with the modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of a sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, wherein at least one set of epigenetic target regions comprises a set of hypomethylated variable target regions; c) sequencing the captured DNA; d) detecting the level of the captured DNA sequences and determining the level of each of a plurality of immune cell types from which said DNA originated; A method comprising:
8. 1. A method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with the modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of a sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, the step comprising contacting the DNA with target-specific probes specific for the at least one set of epigenetic target regions, wherein target regions of the set of epigenetic target regions comprise DNA sequences that are hypomethylated in a plurality of immune cell types; c) sequencing the captured DNA; A method comprising:
9. 1. A method for analyzing DNA, comprising: a) partitioning the DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting the DNA with an agent that recognizes modified cytosines in the DNA, wherein the first aliquot contains a greater proportion of DNA with the modified cytosines than the second aliquot, and the DNA is derived from a buffy coat sample or the DNA is derived from cells of a sample; b) capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, thereby obtaining captured DNA, the step comprising contacting the DNA with target-specific probes specific for the at least one set of epigenetic target regions, wherein target regions of the set of epigenetic target regions comprise DNA sequences that are hypermethylated in a plurality of immune cell types; c) sequencing the captured DNA; A method comprising:
10. a) sequencing the CDR3 region of said DNA; b) determining the number of B cells, T cells, or both B and T cells from which the DNA originates based on the CDR3 sequence; 10. The method of any one of the preceding claims, further comprising:
11. 1. A method for analyzing DNA, comprising: a) sequencing the CDR3 region of said DNA, wherein said DNA is derived from a buffy coat sample or said DNA is derived from cells of a sample; b) determining the number of B cells, T cells, or both B and T cells from which the DNA originates based on the CDR3 sequence; A method comprising:
12. 15. The method of any one of claims 10 to 14, wherein the step of determining the number of B cells, T cells, or both B and T cells comprises determining the number of recombined CDR3 regions and the number of unrecombined CDR3 regions.
13. 12. The method of claim 10 or claim 11, wherein the CDR3 region is a CDR3 region of a B cell receptor chain or a T cell receptor chain.
14. The method of claim 12, wherein the B cell receptor chain is a heavy chain or a light chain.
15. 14. The method of claim 12 or claim 13, wherein the T cell receptor chain is an alpha chain or a beta chain.
16. a) sequencing the DNA to determine methylation levels for a set of epigenetic target regions comprising a plurality of target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types, wherein the DNA is derived from a buffy coat sample or the DNA is derived from cells of the sample; b) determining the quantity of each of a plurality of immune cell types from which the DNA originates based on the methylation level; The method of any one of claims 11 to 15, further comprising:
17. a) prior to said sequencing step, capturing at least a set of epigenetic target regions from said DNA or subsample thereof, said capturing step comprising contacting said DNA or subsample thereof with target-specific probes specific for said at least one set of epigenetic target regions, said set of epigenetic target regions comprising target regions comprising DNA sequences that are differentially methylated in a plurality of immune cell types; b) determining the methylation level for said target region; c) determining the quantity of each of the plurality of immune cell types from which the DNA originates; The method of any one of claims 11 to 16, further comprising:
18. a) prior to said sequencing step, partitioning said DNA into a plurality of aliquots, including a first aliquot and a second aliquot, by contacting said DNA with an agent that recognizes modified cytosines in said DNA, wherein said first aliquot contains a greater proportion of DNA with modified cytosines than said second aliquot; b) sequencing DNA from one or more of said plurality of aliquots; c) detecting the level of DNA sequences to determine the quantity of each of a plurality of immune cell types from which the DNA originates; The method of any one of claims 11 to 17, further comprising:
19. a) the target regions of the set of epigenetic target regions comprise DNA sequences that are hypomethylated in multiple immune cell types; or b) the target regions of the set of epigenetic target regions comprise DNA sequences that are hypermethylated in multiple immune cell types; 19. The method of claim 17 or claim 18.
20. 20. The method of any one of claims 5, 7, 10, 12-15, 18, or 19, comprising capturing at least a set of epigenetic target regions of DNA from at least one of the first and second sub-samples, wherein the target regions of the set of epigenetic target regions comprise DNA sequences that are differentially methylated in a plurality of immune cell types, and wherein the capturing step is performed before the sequencing step.
21. 10. The method of any one of the preceding claims, wherein the set of epigenetic target regions comprises a set of hypermethylated variable target regions and a set of hypomethylated variable target regions.
22. The plurality of immune cell types may be selected from the group consisting of macrophages (including M1 macrophages and M2 macrophages); B cells, e.g., naive B cells or activated B cells (including regulatory B cells, memory B cells, switched memory B cells, and plasma cells); B cell precursors; T cells, e.g., CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (cytotoxic T cells, regulatory T cells (Tregs), CD4 T cells (including naive CD4 T cells, activated CD4 T cells, and CD4 effector memory T cells), CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and CD4 effector memory T cells), and CD8 T cells (including naive CD8 T cells, activated CD8 T cells, and activated CD8 T cells).
6. The method of any one of the preceding claims, wherein the target cells comprise two or more of: immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), neutrophils, low-density neutrophils, immature neutrophils, and immature granulocytes); dendritic cells; eosinophils; monocytes; erythrocytes; megakaryocytes; and natural killer (NK) cells.
23. the plurality of immune cell types i. Naive and activated lymphocytes; ii. monocytes and macrophages; and / or iii. Myelocytes, neutrophils, and eosinophils 10. The method of any one of the preceding claims, comprising:
24. 10. The method of any one of the preceding claims, wherein the multiple immune cell types comprise naive and activated lymphocytes.
25. 10. The method of claim 1, wherein the plurality of immune cell types comprises naive T cells, naive B cells, effector CD4 T cells, effector CD8 T cells, Treg cells, plasma cells, and memory cells.
26. 10. The method of claim 9, wherein the effector CD4 T cells comprise effector memory CD4 T cells and central memory CD4 T cells, and the effector CD8 T cells comprise effector memory CD8 T cells and central memory CD8 T cells.
27. 10. The method of any one of the preceding claims, wherein the multiple immune cell types comprise monocytes and macrophages.
28. 28. The method of claim 22, 23, or 27, wherein the macrophage is an M1 macrophage or an M2 macrophage.
29. 10. The method of any one of the preceding claims, wherein the multiple immune cell types comprise myeloid cells, neutrophils, and eosinophils.
30. 10. The method of any one of the preceding claims, wherein the multiple immune cell types comprise metamyelocytes.
31. 10. The method of any one of the preceding claims, wherein the multiple immune cell types comprise natural killer (NK) cells.
32. 10. The method of any one of the preceding claims, wherein the level of each of the plurality of immune cell types is determined relative to the level of total blood cells.
33. 10. The method of any one of the preceding claims, comprising determining a ratio of levels or quantities of immune cell types based on the determined levels or quantities of said plurality of immune cell types.
34. 10. The method of the preceding claim, wherein the numerator of the ratio comprises the level or quantity of neutrophils, monocytes, or both neutrophils and monocytes.
35. 35. The method of claim 33 or 34, wherein the denominator of the ratio comprises the level or quantity of T cells, B cells, NK cells or total lymphocytes.
36. 36. The method of any one of claims 33-35, wherein (a) the numerator of the ratio comprises the level or quantity of neutrophils and the denominator of the ratio comprises the level or quantity of total lymphocytes; or (b) the numerator of the ratio comprises the level or quantity of NK cells and the denominator of the ratio comprises the level or quantity of total lymphocytes.
37. 36. The method of any one of claims 33 to 35, wherein the numerator of the ratio comprises the level or number of monocytes and the denominator of the ratio comprises the level or number of T cells.
38. 10. The method of any one of the preceding claims, comprising determining a turnover frequency for at least one of said plurality of immune cell types.
39. 10. The method of claim 1, wherein the turnover comprises proliferation.
40. 39. The method of claim 38, wherein the turnover comprises apoptosis.
41. 10. The method of any one of the preceding claims, comprising determining the level of at least one cell type other than the immune cell type from which the DNA originated.
42. 10. The method of the preceding claim, comprising capturing at least one set of epigenetic target regions that contain sequence independent differences in target regions in DNA originating from a cell type other than an immune cell type compared to the same target regions in DNA originating from all other cell types in the sample or sub-sample.
43. 43. The method of claim 41 or 42, wherein the cell type other than an immune cell type is not a blood cell type.
44. 10. The method of the preceding claim, wherein said cell type other than an immune cell type is colorectal, lung, breast, prostate, skin, stomach, bladder, liver, ovary, pancreas, squamous epithelium, salivary gland, larynx, hypopharynx, nasal cavity, paranasal sinuses, nasopharynx or kidney.
45. 10. The method of any one of the preceding claims, wherein the sample is obtained from a subject.
46. 10. The method of the preceding claim, comprising determining the likelihood that the subject has cancer or a precancerous condition.
47. 10. The method of the preceding claim, comprising determining the likelihood that the subject has cancer.
48. 10. The method of the preceding claim, wherein the cancer is an immune cell type cancer.
49. 10. The method of the preceding claim, wherein the cancer is a lymphocytic cancer.
50. 10. The method of the preceding claim, wherein the cancer is leukemia, lymphoma, or myeloma.
51. The method of any one of claims 46 to 50, wherein the cancer is a myeloid cancer.
52. 48. The method of claim 46 or 47, wherein the cancer is a cancer of a cell or tissue type other than an immune cell type.
53. 53. The method of any one of claims 46, 47, or 52, wherein the cancer or precancerous condition is a cancer or precancerous condition other than a hematological cancer or precancerous condition, or wherein the cancer or precancerous condition is a solid tumor cancer, optionally wherein the solid tumor cancer is a carcinoma, adenocarcinoma, or sarcoma.
54. 54. The method of any one of claims 46, 47, 52 or 53, wherein the cancer is colorectal cancer, lung cancer, breast cancer, prostate cancer, skin cancer, stomach cancer, bladder cancer, liver cancer, ovarian cancer, pancreatic cancer, head and neck cancer, or kidney cancer.
55. 47. The method of claim 46, comprising determining the likelihood that the subject has a precancerous condition.
56. 10. The method of the preceding claim, wherein the precancerous condition is an adenoma.
57. 10. The method of the preceding claim, wherein the adenoma is an advanced adenoma.
58. 58. The method of any one of claims 46 or 55-57, wherein the precancerous condition is colorectal precancerous condition, lung precancerous condition, breast precancerous condition, prostate precancerous condition, skin precancerous condition, stomach precancerous condition, bladder precancerous condition, liver precancerous condition, ovarian precancerous condition, pancreatic precancerous condition, head and neck precancerous condition, or kidney precancerous condition.
59. 59. The method of any one of claims 45 to 58, comprising determining the likelihood that the subject has an infection.
60. 60. The method of any one of claims 45 to 59, comprising determining the likelihood that the subject has graft rejection.
61. 61. The method of any one of claims 35 to 60, comprising predicting response to treatment in said subject.
62. 10. The method of the preceding claim, wherein the treatment is a chemotherapeutic agent.
63. 62. The method of claim 61, wherein the treatment is an immunotherapeutic agent.
64. 64. The method of any one of Claims 55-63, wherein said determining the quantity or sequencing of each of the plurality of immune cell types comprises generating a plurality of sequencing reads, said method further comprising mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads to determine the likelihood that the subject has cancer, a precancerous condition, an infection, or a transplant rejection.
65. 10. The method of any one of the preceding claims, wherein the sample is obtained from a subject previously diagnosed with cancer and who has undergone one or more previous cancer treatments, optionally wherein the sample is obtained at one or more pre-selected time points after said one or more previous cancer treatments.
66. 10. The method of any preceding claim, further comprising determining a cancer recurrence score, optionally wherein the subject's cancer recurrence status is determined to be at risk of cancer recurrence if the cancer recurrence score is determined to be at or above a predetermined threshold, or wherein the subject's cancer recurrence status is determined to be at low risk of cancer recurrence if the cancer recurrence score is below the predetermined threshold.
67. 10. The method of any preceding claim, further comprising comparing said cancer recurrence score of said subject to a predetermined cancer recurrence threshold, wherein said subject is classified as a candidate for subsequent cancer treatment if said cancer recurrence score is above said cancer recurrence threshold, or is not classified as a candidate for subsequent cancer treatment if said cancer recurrence score is below said cancer recurrence threshold.
68. 10. The method of any one of the preceding claims, further comprising analyzing DNA of an additional sample from the subject, wherein the additional sample was collected at a different time than when the DNA from the buffy coat sample or the DNA from the cells of the sample was collected.
69. the DNA from the buffy coat sample or the DNA from cells of the sample was collected before the subject underwent a treatment, and the additional sample was collected after the subject underwent the treatment; or the DNA from the buffy coat sample or the DNA from cells of the sample was collected before the subject was diagnosed with the condition, and the additional sample was collected after the subject was diagnosed with the condition.
10. The method of the preceding claims.
70. 70. The method of any one of claims 2, 4, 6, 10, 12-15, or 17-69, wherein the capturing step comprises capturing a sequence-variable target region.
71. 10. The method of claim 9, wherein the capturing step comprises contacting the DNA with target-specific probes specific for the set of at least one epigenetic target region and with target-specific probes specific for the sequence-variable target region.
72. 72. The method of any one of claims 5 to 10, or 12 to 15, or 18 to 71, wherein the modified cytosine is a methylcytosine.
73. 72. The method of any one of claims 5 to 10, or 12 to 15, or 18 to 71, wherein the agent that recognizes modified cytosines is a methyl-binding reagent.
74. 10. The method of the preceding claim, wherein the methyl-binding reagent is an antibody.
75. 74. The method of claim 73, wherein the methyl-binding reagent is a methyl-binding protein or comprises a methyl-binding domain.
76. 76. The method of claims 73-75, wherein the methyl-binding reagent specifically recognizes 5-methylcytosine.
77. 77. The method of claims 73-76, wherein the methyl-binding reagent is immobilized on a solid support.
78. 78. The method of any one of claims 5 to 77, wherein the partitioning step comprises immunoprecipitation of methylated DNA.
79. 79. The method of any one of claims 5 to 78, wherein the partitioning step comprises partitioning based on binding to a protein, optionally wherein the protein is a methylated protein, an acetylated protein, an unmethylated protein, an unacetylated protein; and / or optionally wherein the protein is a histone.
80. 10. The method of claim 9, wherein the distributing step comprises contacting the DNA of the sample with a binding reagent specific for the protein and immobilized on a solid support.
81. 10. The method of any one of the preceding claims, wherein determining the methylation level for the target region comprises bisulfite sequencing.
82. 10. The method of any preceding claim, further comprising performing a procedure that affects a first nucleobase of said DNA or at least one sub-sample differently than a second nucleobase of said DNA or said at least one sub-sample, optionally wherein said first nucleobase is an unmodified cytosine and said second nucleobase is a modified cytosine, and further optionally wherein said modified cytosine is 5-methylcytosine or 5-hydroxymethylcytosine.
83. 83. The method of Claim 82, wherein said procedure affecting a first nucleobase of said DNA or said at least one aliquot differently than a second nucleobase of said DNA or said at least one aliquot chemically converts said first or second nucleobase such that the base-pairing specificity of the converted nucleobase is altered.
84. said procedure affecting a first nucleic acid base of said DNA or said at least one aliquot differently from a second nucleic acid base of said DNA or said at least one aliquot, before the dispensing step, before the capturing step, or after the capturing step; and Prior to the sequencing step 84. The method of claim 82 or 83, wherein
85. 85. The method of any one of claims 82 to 84, wherein the procedure that affects a first nucleobase of the DNA or the at least one aliquot differently from a second nucleobase of the DNA or the at least one aliquot is a methylation-sensitive conversion.
86. 86. The method of claim 85, wherein the methylation-sensitive conversion is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-coupled epigenetic (ACE) conversion, or enzymatic conversion.
87. 87. The method of claim 86, wherein the Tet-assisted conversion further comprises a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane.
88. 88. The method of any one of claims 5 to 87, wherein the aliquots are pooled prior to the sequencing step.
89. 89. The method of any one of claims 1-4 or 6-88, wherein the set of epigenetic target regions comprises hypomethylated variable target regions comprising DNA sequences that are differentially hypomethylated in multiple immune cell types.
90. 90. The method of any one of claims 3, 4, 7, 8, 10, 12-15, or 19-89, wherein the hypomethylated variable target region comprises a DNA sequence that is differentially hypomethylated in multiple immune cell types.
91. 91. The method of any one of claims 3, 4, 7, 8, 10, 12-15, 19, 89, or 90, wherein the DNA sequence that is differentially hypomethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type in the sample.
92. 91. The method of any one of claims 3, 4, 7, 8, 10, 12-15, 19, 89, or 90, wherein the DNA sequence that is differentially hypomethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type.
93. 91. The method of any one of claims 3, 4, 7, 8, 10, 12-15, 19, 89, or 90, wherein the DNA sequence that is differentially hypomethylated in a plurality of immune cell types comprises a degree of methylation in one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type in the sample.
94. 91. The method of any one of claims 3, 4, 7, 8, 10, 12-15, 19, 89, or 90, wherein the DNA sequence that is differentially hypomethylated in multiple immune cell types comprises a degree of methylation in one immune cell type that is detectably lower than the degree of methylation of the same sequence in any other immune cell type.
95. 91. The method of any one of claims 3, 4, 7, 8, 10, 12-15, 19, 89, or 90, wherein the DNA sequence that is differentially hypomethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in at least one other immune cell type and that is detectably lower than the degree of methylation of the same sequence in at least one non-immune cell type in the sample.
96. 91. The method of any one of claims 3, 4, 7, 8, 10, 12-15, 19, 89, or 90, wherein the DNA sequence that is differentially hypomethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably lower than the degree of methylation of the same sequence in at least one other immune cell type and that is detectably lower than the degree of methylation of the same sequence in any non-immune cell type in the sample.
97. 97. The method of any one of claims 91 to 96, wherein said detectably lower degree of methylation is one less methylated cytosine than the same sequence in said other cell type.
98. 97. The method of any one of claims 91 to 96, wherein said detectably lower degree of methylation is two fewer methylated cytosines than the same sequence in said other cell type.
99. 97. The method of any one of claims 91 to 96, wherein said detectably lower degree of methylation is 3 fewer methylated cytosines than the same sequence in said other cell type.
100. 97. The method of any one of claims 91 to 96, wherein said detectably lower degree of methylation is four fewer methylated cytosines than the same sequence in said other cell type.
101. 97. The method of any one of claims 91 to 96, wherein said detectably lower degree of methylation is 5 or more fewer methylated cytosines than the same sequence in said other cell type.
102. 102. The method of any one of claims 1-4, 6, 10, 12-15, or 19-101, wherein the set of epigenetic target regions comprises hypermethylated variable target regions comprising DNA sequences that are differentially hypermethylated in multiple immune cell types.
103. 103. The method of Claim 102, wherein said DNA sequence that is differentially hypermethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type in said sample.
104. 103. The method of Claim 102, wherein said DNA sequence that is differentially hypermethylated in a plurality of immune cell types comprises a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type.
105. 103. The method of Claim 102, wherein said DNA sequence that is differentially hypermethylated in a plurality of immune cell types comprises a degree of methylation in one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type in said sample.
106. 103. The method of claim 102, wherein said DNA sequence that is differentially hypermethylated in a plurality of immune cell types comprises a degree of methylation in one immune cell type that is detectably higher than the degree of methylation of the same sequence in any other immune cell type.
107. 103. The method of Claim 102, wherein said DNA sequences that are differentially hypermethylated in a plurality of immune cell types comprise a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in at least one other immune cell type and that is detectably lower than the degree of methylation of the same sequence in at least one non-immune cell type in said sample.
108. 103. The method of Claim 102, wherein said DNA sequences that are differentially hypermethylated in a plurality of immune cell types comprise a degree of methylation in at least one immune cell type that is detectably higher than the degree of methylation of the same sequence in at least one other immune cell type and that is detectably lower than the degree of methylation of the same sequence in any non-immune cell type in the sample.
109. 109. The method of any one of claims 103 to 108, wherein said detectably higher degree of methylation is one more cytosine methylation than the same sequence in said other cell type.
110. 109. The method of any one of claims 103 to 108, wherein said detectably higher degree of methylation is two more cytosine methylations than the same sequence in said other cell type.
111. 109. The method of any one of claims 103 to 108, wherein said detectably higher degree of methylation is three more cytosine methylations than the same sequence in said other cell type.
112. 109. The method of any one of claims 103 to 108, wherein said detectably higher degree of methylation is four more cytosine methylations than the same sequence in said other cell types.
113. 109. The method of any one of claims 103 to 108, wherein said detectably higher degree of methylation is 5 or 6 or more cytosine methylations than the same sequence in said other cell type.
114. 10. The method of any one of the preceding claims, comprising determining the quantity or detecting the level of DNA in the sample originating from erythrocytes or erythrocyte precursors, macrophages, B cells, T cells, myeloid cells, or natural killer cells.
115. 115. The method of claim 114, wherein the macrophage is an M1 macrophage or an M2 macrophage.
116. 115. The method of claim 114, wherein the B cells are activated B cells.
117. 117. The method of claim 116, wherein the activated B cells are regulatory B cells, memory B cells, or plasma cells.
118. 115. The method of claim 114, wherein the T cells are central memory T cells, naive T cells, or naive-like T cells, and optionally the central memory T cells are CD4 or CD8 central memory T cells.
119. 115. The method of claim 114, wherein the T cells are activated T cells.
120. 120. The method of claim 119, wherein the activated T cells are cytotoxic T cells, regulatory T cells, CD4 effector memory T cells, or CD8 effector memory T cells.
121. 115. The method of claim 114, wherein the myeloid cells are immature myeloid cells.
122. 122. The method of claim 121, wherein the immature myeloid cells are myeloid-derived suppressor cells, low-density neutrophils, immature neutrophils, or immature granulocytes.
123. 123. The method of any one of claims 91-122, wherein the detectably lower or higher degree of methylation is present in samples from at least one group of donors compared to another group of donors, optionally wherein said at least one group of donors has a cancer that responds to treatment and said another group of donors has a cancer that does not respond to treatment.
124. 10. The method of any one of the preceding claims, wherein the DNA is derived from a buffy coat sample.
125. 10. The method of any one of the preceding claims, wherein the DNA is derived from cells of a blood sample.
126. 10. The method of the preceding claim, wherein the blood sample is a whole blood sample, a leukoreduced sample, or a PBMC sample.
Citation Information
Cited By
Immune condition prediction system, immune condition prediction method, and program
JP2026027793A