Methods and systems for tissue informed differentially methylated region analysis
A method using a universal panel of genomic regions and advanced sequencing techniques for ctDNA analysis addresses the challenge of low ctDNA fraction in blood, enabling accurate cancer detection and monitoring with high sensitivity and specificity.
Patent Information
- Application Number
- PCT/US2025/016672
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-24
- Filing Date
- 2025-02-20
- Publication Date
- 2025-08-28
AI Technical Summary
Existing methods for analyzing circulating tumor DNA (ctDNA) as a non-invasive biomarker for cancer detection face challenges due to the low fraction of ctDNA in blood-derived samples and the need for improved methods to accurately identify differentially methylated regions (DMRs) for early cancer detection and monitoring.
A method involving the use of a universal panel of genomic regions to enrich and analyze nucleic acid molecules from samples, including sequencing at a depth of up to 50 million single reads, and comparing these regions to reference sets to generate methylation scores for cancer detection and monitoring, utilizing techniques such as Support Vector Machine and logistic regression.
This approach enables accurate detection and monitoring of cancer with high specificity, allowing for the identification of minimal residual disease and guiding therapy decisions, with sensitivity and specificity up to 90%.
Smart Images

Figure US2025016672_28082025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR TISSUE INFORMED DIFFERENTIALLYMETHYLATED REGION ANALYSISCROSS REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 556,347, filed February 21, 2024, and U.S. Provisional Application No. 63 / 711,616, filed October 24, 2024, each of which is incorporated herein by reference in its entirety.BACKGROUND
[0002] Circulating tumor deoxyribonucleic acid (ctDNA) may be used as a non-invasive, tumor-specific biomarker for clinical use. ctDNA may be derived from tumor cells undergoing cell-death and released into circulation of various bodily fluids including blood. In certain subjects with cancer, the majority of blood-derived cell-free DNA may originate from healthy (e.g., non-cancerous) tissues. In addition, the fraction of ctDNA observed may range from <0.1% to 90% of the total cell-free DNA depending on factors including the primary site of the tumor and disease burden. ctDNA provides non-invasive access to the tumor’s molecular landscape and disease burden.SUMMARY
[0003] In an aspect, the present disclosure provides a method for analyzing a sample derived from a subject, comprising: (a) obtaining a sample comprising nucleic acid molecules obtained or derived from the subject; (b) assaying the nucleic acid molecules to generate a data set comprising methylation states of one or more genomic regions comprising differentially methylation regions (DMRs); and (c) processing at least a portion of the data set to generate an output indicative of presence or absence cancer in the subject, wherein the at least the portion of the data set pertains to a set of DMRs specific to the subject.
[0004] In some embodiments, the method further comprises providing a universal panel of genomic regions. In some embodiments, the method further comprises using the universal panel of genomic regions to enrich for the one or more genomic regions comprising the DMRs.
[0005] In some embodiments, the assaying further comprises sequencing the nucleic acid molecules at a depth of at most 50 Million (M) single reads. In some embodiments, the assaying further comprises sequencing the nucleic acid molecules at a depth of at most 10 M single reads.
[0006] In some embodiments, the processing of (c) further comprises comparing the one or more genomic regions to a set of DMRs specific to one or more reference subjects to generate the set of DMRs specific to the subject. In some embodiments, the processing of (c) further comprises comparing the one or more genomic regions to a set of anti-DMRs to generate a set of anti-DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of anti-DMRs specific to the subject. In some embodiments, the method further comprises normalizing the one or more counts of the set of DMRs specific to the subject to the one or more counts of the set of anti-DMRs specific to the subject. In some embodiments, the normalizing further comprises generating a methylation score. In some embodiments, the method further comprises comparing the methylation score to a threshold score, thereby generating the output indicative of the presence or absence of the cancer in the subject.
[0007] In some embodiments, the method further comprises obtaining one or more control samples comprising control nucleic acid molecules. In some embodiments, the one or more control samples comprise one or more non-tissue samples and / or one or more tissue samples. In some embodiments, the one or more control nucleic acid molecules are derived from one or more tissue samples. In some embodiments, the one or more control nucleic acid molecules comprise one or more cell-free nucleic acid molecules. In some embodiments, the one or more control samples are derived from one or more control subjects without cancer.
[0008] In some embodiments, the method further comprises assaying the control nucleic acid molecules to generate a control data set comprising methylation states of one or more control genomic regions.
[0009] In some embodiments, the assaying the control nucleic acid molecules further comprises conducting one or more methylation reactions. In some embodiments, the assaying the control nucleic acid molecules further comprises sequencing the control nucleic acids, or derivatives thereof. In some embodiments, the methylation states of the one or more control genomic regions comprise hypermethylated states, methylated states, non-methylated states, or hypomethylated states, or any combinations thereof.
[0010] In some embodiments, the method further comprises processing at least a portion of the control data set to identify one or more hypomethylated regions, regions of nonmethylation, or regions that are amenable to methylation enrichment, or any combinations thereof, thereby generating the universal panel of regions. In some embodiments, theuniversal panel of genomic regions comprises the regions of non-methylation and the regions that are amenable to methylation enrichment.
[0011] In some embodiments, the method further comprises obtaining one or more reference samples from one or more reference subjects. In some embodiments, the one or more reference subjects have cancer. In some embodiments, the one or more reference samples comprise one or more reference nucleic acid molecules. In some embodiments, the one or more reference nucleic acid molecules comprise cell-free nucleic acid molecules. In some embodiments, the one or more reference nucleic acid molecules are derived from one or more non-tissue samples and / or one or more tissue samples. In some embodiments, the method further comprises obtaining another one or more control samples comprising another one or more control nucleic acid molecules.
[0012] In some embodiments, the method further comprises assaying the one or more reference nucleic acid molecules and the another one or more control nucleic acid molecule to generate a reference data set comprising methylation states of one or more regions. In some embodiments, the method further comprises processing the reference data set with the universal panel of genomic regions to identify the set of DMRs specific to one or more reference subjects and the set of anti -DMRs.
[0013] In some embodiments, the processing of (c) further comprises comparing the one or more genomic regions to a set of DMRs specific to a reference subject to generate the set of DMRs specific to the subject. In some embodiments, the processing of (c) further comprises comparing the one or more genomic regions to a set of anti-DMRs to generate a set of anti- DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of anti-DMRs specific to the subject. In some embodiments, the method further comprises normalizing the one or more counts of the set of DMRs specific to the subject to the one or more counts of the set of anti-DMRs specific to the subject. In some embodiments, the normalizing generates another methylation score. In some embodiments, the method further comprises comparing the another methylation score to a threshold score, thereby generating the output indicative of the cancer in the subject.
[0014] In some embodiments, the method further comprises obtaining from the reference subject, a reference sample comprising one or more another reference nucleic acid molecules. In some embodiments, the one or more another reference nucleic acid molecules are derived from one or more non-tissue samples and / or one or more tissue samples. In someembodiments, the one or more another reference nucleic acid molecules comprise cell-free nucleic acid molecules. In some embodiments, the reference subject is same subject as the subject. In some embodiments, the reference sample is obtained prior to obtaining the sample. In some embodiments, the reference sample is obtained or derived from the subject subsequent to diagnosis with the cancer. In some embodiments, the reference subject is obtained or derived from the subject prior to treatment with a therapy.
[0015] In some embodiments, the method further comprises assaying the one or more another reference nucleic acid molecules to generate a reference data set comprising methylation states of one or more reference genomic regions. In some embodiments, the assaying comprises enriching for the one or more reference genomic regions using the universal panel of genomic regions. In some embodiments, the assaying comprises enriching for the one or more reference genomic regions without using the universal panel of genomic regions. In some embodiments, the method further comprises processing the reference data set to identify the set of DMRs specific to the reference subject.
[0016] In some embodiments, the method further comprises integrating the methylation score and the additional methylation score to generate a single score. In some embodiments, the single score is indicative of the cancer. In some embodiments, the integrating comprises Support Vector Machine, logistic regression, Bayesian Interference Model, weighted average, decision trees, and / or random forests.
[0017] In some embodiments, the nucleic acid molecules are derived from one or more nontissue samples. In some embodiments, the nucleic acid molecules are derived from one or more tissue samples. In some embodiments, the nucleic acid molecules comprise cell-free deoxyribonucleic acid (DNA) molecules. In some embodiments, the sample comprises a tissue sample, a blood sample, and / or a plasma sample. In some embodiments, the cancer is a late-stage cancer. In some embodiments, the cancer is an early-stage cancer.
[0018] In some embodiments, the method further comprises generating an output indicative of presence or absence of minimal residual disease in the subject. In some embodiments, the method further comprises, based at least on the processing, treating the subject with a therapy capable of treating the cancer. In some embodiments, the therapy comprises a chemotherapy, a radiation therapy, an immunotherapy, a targeted therapy, a surgical resection, or a combination thereof. In some embodiments, the method further comprises, based at least on the processing, recommending a therapy regimen for the subject or changing a therapy regimen for the subject.
[0019] In some embodiments, the assaying further comprises mixing the nucleic acid molecules with filler nucleic acid molecules. In some embodiments, the assaying does not comprise mixing the nucleic acid molecules with filler nucleic acid molecules. In some embodiments, the assaying further comprises enriching methylated nucleic acids. In some embodiments, the enriching further comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG-binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody is an anti 5-mC antibody. In some embodiments, the antibody is an anti 5 -hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule.
[0020] In some embodiments, the assaying further comprises sequencing the nucleic acids, or derivatives thereof. In some embodiments, the sequencing does not comprise bisulfite sequencing. In some embodiments, the sequencing further comprises bisulfite sequencing with methylation specific PCR. In some embodiments, the sequencing further comprises targeted sequencing.
[0021] In some embodiments, the sequencing further comprises using a plurality of capture probes. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to the one or more genomic regions. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation. In some embodiments, the sequencing generates sequencing reads corresponding to the one or more genomic regions.
[0022] In some embodiments, the processing of (c) comprises counting a number of sequencing reads corresponding to a region of the one or more genomic regions.
[0023] In another aspect, the present disclosure provides a method for monitoring a subject for regression or progression of a disease or condition comprising: assaying a biological sample of the subject for one or more markers specific to the subject, using a universal panel of genomic regions.
[0024] In some embodiments, the assaying is without use of a primer or bait set specific to the one or more markers.
[0025] In some embodiments, the assaying further comprises sequencing nucleic acid molecules from the biological sample at a depth of at most 50 Million (M) single reads. In some embodiments, the assaying further comprises sequencing the nucleic acid molecules at a depth of at most 10 M single reads.
[0026] In some embodiments, the biological sample is a blood sample or a plasma sample. In some embodiments, the biological sample is a tissue sample and / or a non-tissue sample.
[0027] In some embodiments, the universal panel of genomic regions are derived from one or more control samples are derived from one or more control subjects without cancer. In some embodiments, the universal panel of genomic regions comprises non-methylated regions and regions that are amenable to methylation enrichment.
[0028] In some embodiments, the one or more markers comprise differentially methylated regions (DMRs) specific to the subject.
[0029] In some embodiments, the method further comprises comparing one or more genomic regions of the biological sample to a set of anti-DMRs specific to the one or more reference samples to generate anti-DMRs specific to the subject.
[0030] genomic regions to one or more reference genomic regions to one or more reference samples to generate a set of DMRs specific to the one or more reference samples and / or a set of anti-DMRs specific to the one or more reference samples.
[0031] In some embodiments, the one or more reference samples are derived from one or more reference subjects with cancer. In some embodiments, the one or more reference samples comprise non-tissue samples or tissue samples.
[0032] In some embodiments, the method further comprises comparing the DMRs specific to the subject to the set of DMRs specific to the one or more reference samples to generate one or more counts of the DMRs specific to the subject. In some embodiments, the method further comprises comparing the anti-DMRs specific to the subject to the set of anti-DMRs specific to the one or more reference samples to generate one or more counts of the anti- DMRs specific to the subject. In some embodiments, the method further comprises normalizing the one or more counts of the DMRs specific to the one or more counts of anti- DMRs specific to the subject, thereby generating a first methylation score.
[0033] In some embodiments, the method further comprises obtaining a reference sample. In some embodiments, the reference sample is derived from the subject prior to obtaining the biological sample. In some embodiments, the reference sample is a non-tissue sample and / or a tissue sample.
[0034] In some embodiments, the method further comprises comparing the universal panel of genomic regions to one or more genomic regions of the reference sample to generate a set of DMRs specific to the reference sample and / or a set of anti-DMRs specific to the reference sample. In some embodiments, the method further comprises comparing the DMRs specific to the subject to the set of DMRs specific to the reference sample to generate one or more counts of the DMRs specific to the subject. In some embodiments, the method further comprises comparing the anti-DMRs specific to the subject to the set of anti-DMRs specific to the one or more reference samples to generate one or more counts of the anti-DMRs specific to the subject. In some embodiments, the method further comprises normalizing the one or more counts of the DMRs specific to the subject to the one or more counts of the anti- DMRs specific to the subject, thereby generating a second methylation score.
[0035] In some embodiments, another biological sample is obtained at a time before or after obtaining the biological sample. In some embodiments, the another biological sample comprises another one or more markers specific to the subject. In some embodiments, the another one or more markers comprise another DMRs specific to the subject.
[0036] In some embodiments, the method further comprises comparing the another biological sample to the set of anti-DMRs specific to the one or more reference samples to generate another anti-DMRs specific to the subject. In some embodiments, the method further comprises comparing the another DMRs specific to the subject to the set of DMRs specific to the reference sample to generate one or more counts of the another DMRs specific to the subject. In some embodiments, the method further comprises comparing the another anti- DMRs specific to the subject to the set of anti-DMRs specific to the one or more reference samples to generate one or more counts of the another anti-DMRs specific to the subject. In some embodiments, the method further comprises normalizing the one or more DMR counts of the another DMRs specific to the subject to the one or more counts of the another anti- DMRs specific to the subject, thereby generating a third methylation score. In some embodiments, the method further comprises comparing the second methylation score and the third methylation score, thereby generating an output indicative of the regression or the progression of the disease or the condition.
[0037] In some embodiments, the method further comprises integrating the first methylation score with the second methylation score to generate a single methylation score. In some embodiments, the single methylation score is an output indicative of the regression or the progression of the disease or the condition.
[0038] In some embodiments, the disease or condition comprises a cancer. In some embodiments, the cancer is a late-stage cancer. In some embodiments, the cancer is an early- stage cancer. In some embodiments, the disease or condition is a pre-cancer.
[0039] In some embodiments, the biological sample comprises nucleic acid molecules. In some embodiments, the nucleic acid molecules are derived from one or more non-tissue samples. In some embodiments, the nucleic acid molecules are derived from one or more tissue samples. In some embodiments, the nucleic acid molecules comprise cell-free nucleic acid molecules. In some embodiments, the biological sample comprises a blood sample and / or a plasma sample.
[0040] In some embodiments, the assaying further comprises sequencing the nucleic acid molecules. In some embodiments, the sequencing does not comprise bisulfite sequencing. In some embodiments, the sequencing further comprises bisulfite sequencing with methylation specific PCR. In some embodiments, the sequencing further comprises generating sequencing reads corresponding to the one or more markers. In some embodiments, the sequencing further comprises targeted sequencing.
[0041] In some embodiments, the sequencing further comprises using a plurality of capture probes. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to the plurality of regions. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation. In some embodiments, the assaying further comprises counting a number of sequencing reads corresponding to a marker of the one or more markers.
[0042] In some embodiments, the assaying further comprises enriching methylated nucleic acids. In some embodiments, the enriching further comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG-binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody is an anti 5-mC antibody. In some embodiments, the antibody is an anti 5- hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule.
[0043] In some embodiments, the method further comprises, based at least on the assaying, treating the subject with a therapy capable of treating the cancer. In some embodiments, thetherapy comprises a chemotherapy, a radiation therapy, an immunotherapy, a targeted therapy, a surgical resection, or a combination thereof. In some embodiments, the method further comprises, based at least on the assaying, recommending a therapy regimen for the subject or changing a therapy regimen for the subject.
[0044] In another aspect, provided herein is a method for analyzing a sample derived from a subject, comprising: assaying the sample for at least a portion of a set of differentially methylation regions (DMRs) specific to the subject to generate an output indicative of presence or absence of cancer, wherein the assaying comprises sequencing, wherein the sequencing has a depth of at most 50 Million (M) single reads.
[0045] In some embodiments, the sequencing has a depth of at most 10 M single reads.
[0046] In some embodiments, the method further comprises assaying the sample to generate a data set comprising methylation states of one or more genomic regions.
[0047] In some embodiments, the method further comprises providing a universal panel of genomic regions. In some embodiments, the method further comprises comparing the universal panel of genomic regions to one or more reference genomic regions to generate a set of DMRs specific to a reference sample and / or a set of anti -DMRs specific to the reference sample. In some embodiments, the method further comprises comparing the universal panel of genomic regions to another one or more reference genomic regions to generate a set of DMRs specific to a one or more reference samples and / or a set of anti- DMRs specific to the one or more reference samples.
[0048] In some embodiments, the sample comprises nucleic acid molecules. In some embodiments, the nucleic acid molecules are derived from one or more tissue samples and / or one or more non-tissue samples. In some embodiments, the nucleic acid molecules comprise cell-free nucleic acid molecules. In some embodiments, the sample comprises a tissue sample. In some embodiments, the sample does not comprise a tissue sample. In some embodiments, the sample comprises a blood sample or a plasma sample.
[0049] In some embodiments, the method further comprises comparing the one or more genomic regions to the set of DMRs specific to the one or more reference samples to generate the set of DMRs specific to the subject. In some embodiments, the method further comprises comparing the one or more genomic regions to the set of anti-DMRs specific to the one or more reference samples to generate the set of anti-DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of anti-DMRs specific to the subject. In someembodiments, the method further comprises normalizing the one or more counts of the set of DMRs specific to the subject to the one or more counts of the set of anti-DMRs specific to the subject. In some embodiments, the normalizing further comprises generating a methylation score.
[0050] In some embodiments, the method further comprises comparing the one or more genomic regions to the set of DMRs specific to the reference sample to generate the set of DMRs specific to the subject. In some embodiments, the method further comprises comparing the one or more genomic regions to the set of anti-DMRs specific to the one or more reference samples to generate the set of anti-DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of DMRs specific to the subject. In some embodiments, the comparing further comprises generating one or more counts of the set of anti-DMRs specific to the subject. In some embodiments, the method further comprises normalizing the one or more counts of the set of DMRs specific to the subject to the one or more counts of the set of anti-DMRs specific to the subject. In some embodiments, the normalizing further comprises generating another methylation score.
[0051] In some embodiments, the method further comprises using the output indicative of the presence or absence of cancer to determine progression of the cancer. In some embodiments, the method further comprises using the output indicative of the presence or absence of cancer to determine regression of the cancer. In some embodiments, the method further comprises using the output indicative of the presence or absence of cancer to determine therapy of the cancer.
[0052] In another aspect, the present disclosure provides a method for detecting Minimal Residual Disease (MRD), comprising: assaying a biological sample from a subject, wherein the assaying does not comprise analyzing a solid tumor sample of the subject, wherein the assaying comprises sequencing one or more genomic regions in the biological sample, wherein the sequencing has a depth of at most 50 million single reads.
[0053] In another aspect, the present disclosure provides a method comprising: assaying the sample for at least a portion of a set of differentially methylation regions (DMRs) specific to the subject to generate an output indicative of presence or absence of cancer at a specificity of at least 90%.
[0054] In another aspect, the present disclosure provides a method comprising: assaying a sample for at least a portion of a set of differentially methylation regions (DMRs) specific toa subject to generate an output indicative of presence or absence of cancer at a specificity of at least 90%.
[0055] In another aspect, the present disclosure provides a method for classifying a sample derived from subject, the method comprising: (a) obtaining a sample comprising cell-free nucleic acids from a subject; (b) assaying the cell-free nucleic acids to generate a data set comprising methylation states of one or more genomic regions comprising differentially methylation regions (DMRs); and (c) processing at least a portion of the data set to generate an output indicative of cancer in the subject, wherein the portion of the data set pertains to a set of DMRs specific to the subject.
[0056] In some embodiments, the cell-free nucleic acid molecule is a cell-free deoxyribonucleic acid (DNA) molecule. In some embodiments, the sample comprises a tissue sample, a blood sample, or a plasma sample. In some embodiments, the assaying comprises sequencing the cell-free nucleic acids, or derivatives thereof. In some embodiments, the sequencing does not comprise bisulfite sequencing. In some embodiments, the sequencing comprises targeted sequencing. In some embodiments, the sequencing comprises using a plurality of capture probes. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to the one or more genomic regions. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation. In some embodiments, the sequencing generates sequencing reads corresponding to one or more genomic regions. In some embodiments, the assaying comprises counting a number of sequencing reads corresponding to a region of a plurality of regions. In some embodiments, the assaying comprises enriching methylated nucleic acids. In some embodiments, the enriching comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG-binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody is an anti 5-mC antibody. In some embodiments, the antibody is an anti 5- hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule. In some embodiments, the sample obtained in (a) is a test sample. In some embodiments, the method further comprises obtaining a control sample. In some embodiments, the one or more genomic regions is identified by differentialmethylation analysis of a test sample and control sample. In some embodiments, the control sample is derived from a subject without cancer. In some embodiments, the one or more regions exhibits hypermethylation in the test sample compared to the control sample. In some embodiments, the test sample is obtained from the subject at a time prior to obtaining the sample. In some embodiments, the test sample is a tissue sample. In some embodiments, the test sample is a blood sample. In some embodiments, the further comprising obtaining a second test sample at a time subsequent to obtain the sample. In some embodiments, the second test sample is a tissue sample. In some embodiments, the second test sample is a blood sample. In some embodiments, the method further comprising assaying cell-free nucleic acids of the second sample to generate a second data set comprising methylation states of one or more genomic regions comprising differentially methylation regions. In some embodiments, the method further comprises based at least on the processing, treating the subject with a therapy. In some embodiments, the method further comprises based at least on the processing, recommending a therapy regimen for the subject. In some embodiments, the method further comprises based at least on the processing, changing a therapy regimen for the subject.
[0057] In some embodiments, the cancer is breast cancer, bladder cancer, colorectal cancer, endometrial cancer, prostate cancer, renal cancer, pancreatic cancer, or lung cancer. In some embodiments, the cancer is a late-stage cancer. In some embodiments, the cancer is an early- stage cancer. In some embodiments, the one or more genomic regions comprise sites that are amenable to methylation enrichment. In some cases, sites that are amenable to methylation enrichment can comprise sites that can be enriched after in vitro methylation of the sites.
[0058] In another aspect, the present disclosure provides a method for monitoring a subject for regression or progression of a disease or condition comprising: assaying a biological sample of the subject for one or more markers specific to the subject, without use of a primer or bait set specific to the one or more markers. In some embodiments, the disease comprises a cancer. In some embodiments, the cancer is breast cancer, bladder cancer, colorectal cancer, endometrial cancer, prostate cancer, renal cancer, pancreatic cancer, or lung cancer. In some embodiments, the cancer is a late-stage cancer. In some embodiments, the cancer is an early- stage cancer. In some embodiments, the disease or condition is a pre-cancer. In some embodiments, the one or more markers comprises DMRs. In some embodiments, the biological sample comprises cell-free nucleic acid molecules. In some embodiments, the cell- free nucleic acid molecule is a cell-free deoxyribonucleic acid (DNA) molecule. In some embodiments, the biological sample comprises a blood sample and / or a plasma sample. Insome embodiments, the assaying comprises sequencing the cell-free nucleic acids, or derivatives thereof. In some embodiments, the sequencing does not comprise bisulfite sequencing. In some embodiments, the sequencing generates sequencing reads corresponding to the one or more markers. In some embodiments, the sequencing comprises targeted sequencing. In some embodiments, the sequencing comprises using a plurality of capture probes. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to the one or more markers. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation. In some embodiments, the assaying comprises counting a number of sequencing reads corresponding to a marker of the one or more markers. In some embodiments, the assaying comprises enriching methylated nucleic acids. In some embodiments, the enriching comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG-binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody is an anti 5-mC antibody. In some embodiments, the antibody is an anti 5 -hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule. In some embodiments, the one or more markers is identified by differential methylation analysis of a test sample and control sample. In some embodiments, the control sample is derived from a subject without cancer. In some embodiments, the one or more markers exhibits hypermethylation in the test sample compared to the control sample. In some embodiments, the test sample is obtained from the subject at a time prior to obtaining the sample. In some embodiments, the test sample is a tissue sample. In some embodiments, the test sample is a blood sample. In some embodiments, the further comprising obtaining a second sample at a time subsequent to obtain the sample. In some embodiments, the sample is a tissue sample. In some embodiments, the second sample is a blood sample. In some embodiments, the method further comprising assaying the second sample of the subject for one or more markers specific to the subject, without use of a primer or bait set specific to the one or more markers. In some embodiments, the assaying comprises (i) generating a data set comprising data pertaining to additional markers and the one or more marker, and (ii) processing a portion of the data set corresponding to the one or more markers. In someembodiments, the processing does not comprise processing the additional markers. In some embodiments, the method further comprises based at least on the assaying, treating the subject with a therapy. In some embodiments, the method further comprises based at least on the assaying, recommending a therapy regimen for the subject. In some embodiments, the method further comprises based at least on the assaying, changing a therapy regimen for the subject.
[0059] In one aspect, the present disclosure provides a method of nucleic acids processing comprising: (a) assaying methylation levels of a plurality of regions in a biological sample comprising cell-free nucleic acids, wherein the plurality of regions have been identified as regions of a genome that (i) comprises one or more sites that are amenable to methylation enrichment, and (ii) comprises methylation at below a threshold in a non-diseased control; (b) processing methylation levels of the plurality of regions to identify a methylation background; (c) processing the methylation levels for a subset of the plurality of regions, wherein the subset comprises differentially methylated regions (DMRs), to identify a DMR specific methylation level; (d) generating a normalized DMR methylation level by normalizing the DMR specific level against the methylation background.
[0060] In some embodiments, the cell-free nucleic acid molecule is a cell-free deoxyribonucleic acid (DNA) molecule. In some embodiments, the biological sample comprises a blood sample and / or a plasma sample. In some embodiments, the region of the plurality of regions has a length of 300 base pairs (bp).
[0061] In some embodiments, (a) comprises sequencing the cell-free nucleic acids, and / or derivatives thereof. In some embodiments, the sequencing does not comprise bisulfite sequencing. In some embodiments, the sequencing comprises targeted sequencing. In some embodiments, the sequencing comprises using a plurality of capture probes. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to the plurality of regions. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation. In some embodiments, the sequencing generates sequencing reads corresponding to the plurality of regions.
[0062] In some embodiments, the assaying comprises counting a number of sequencing reads corresponding to a region of a plurality of regions. In some embodiments, the (b) comprises counting a number of sequencing reads for each region of the plurality of regions. In someembodiments, the method further comprises generating an average of sequencing read counts for all regions of the plurality of regions thereby generating the methylation background. In some embodiments, the (c) comprises counting a number of sequencing reads for each region of the subset of plurality of regions. In some embodiments, the method further comprises generating an average of sequencing read counts for all regions of the subset of the plurality of regions thereby generating the DMR specific methylation level. In some embodiments, the generating the normalized DMR methylation level comprises dividing an average number of sequencing reads associated with the subset by an average number of sequencing reads associated with the plurality of regions.
[0063] In some embodiments, the (a) comprises enriching methylated nucleic acids. In some embodiments, the enriching comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG- binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody is an anti 5-mC antibody. In some embodiments, the antibody is an anti 5 -hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule.
[0064] In some embodiments, the DMR subset is identified by differential methylation analysis of a test sample and control sample. In some embodiments, the test sample is derived from a subject with cancer. In some embodiments, the control sample is derived from a subject with cancer. In some embodiments, the DMR subset comprises one or more regions that exhibits hypermethylation in the test sample compared to the control sample. In some embodiments, the DMR subset comprises one or more regions that comprises DMRs specific to a particular cancer type. In some embodiments, the DMR subset comprises one or more regions that comprises DMRs that are not specific to a particular cancer type.
[0065] In some embodiments, the method further comprises identifying the subject as having a disease, based at least on the normalized DMR methylation level. In some embodiments, the identifying comprises comparing the normalized DMR methylation level against a control level. In some embodiments, the control level corresponds to an expected value for a non- cancerous sample. In some embodiments, the control level corresponds to an expected value for a cancerous sample. In some embodiments, the disease or condition is a cancer or a tumor. In some embodiments, the disease or condition is a pre-cancer. In some embodiments, the cancer is breast cancer, bladder cancer, colorectal cancer, endometrial cancer, prostatecancer, renal cancer, pancreatic cancer, or lung cancer. In some embodiments, the cancer is a late-stage cancer. In some embodiments, the cancer is an early-stage cancer.
[0066] In some embodiments, the one or more sites that are amenable to methylation enrichment are determined by determining methylation levels in a fully methylated control sample. In some embodiments, the fully methylated control sample comprises nucleic acids subjected to in vitro methylation (e.g., in vitro enzymatic methylation). In some embodiments, the determining methylation levels in the fully methylated control sample comprises enriching methylated nucleic acids. In some embodiments, the enriching comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG-binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody is an anti 5-mC antibody. In some embodiments, the antibody is an anti 5 -hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule. In some embodiments, the methylation control sample further comprises filler deoxyribonucleic acid (DNA) molecule. In some embodiments, the filler DNA has a length of about 50 bp to about 800 bp. In some embodiments, the methylation control sample further comprises genomic DNA. In some embodiments, the genomic DNA is subjected to shearing.
[0067] In some embodiments, the methylation control sample further comprises cell-free DNA.
[0068] In some embodiments, the method comprises a reduction in a noise level compared to a noise level of a corresponding sample that has a normalized DMR methylation level generated by normalizing the DMR specific level against a background derived from a whole genome. In some embodiments, the method comprises a reduction in a noise level compared to a noise level of a corresponding sample that has a normalized DMR methylation level generated by normalizing the DMR specific level against a background derived from all genomic regions that are amenable to methylation enrichment.
[0069] In one aspect, the present disclosure provides a method of identifying a subject as having cancer, the method comprising: (a) obtaining a sample comprising cell-free nucleic acid from a subject; (b) assaying the cell-free nucleic acids to identifying a methylation level of a subset of the cell-free nucleic acids, wherein the subset of cell-free nucleic acids correspond to (e.g., complementary to) regions of a genome that (i) comprise one or more sites that are amenable to methylation enrichment and (ii) comprise substantially nomethylation in a healthy control; and (c) based at least on the methylation state of the subset of the cell-free nucleic acids, identifying the subject as having cancer. In some cases, the subset of cell-free nucleic acids can correspond to regions of a genome by having one or more genomic sequences that are complementary to the sequences of the genome.
[0070] In some embodiments, the assaying comprises sequencing. In some embodiments, based at least on (b), identifying a plurality of differentially methylated regions (DMRs). In some embodiments, the method further comprises processing the plurality of DMRs using a classifier to identify the subject as having cancer.
[0071] In one aspect, the present disclosure provides a method of generating a trained classifier, the method comprising (a) determining the presence of differentially methylated regions (DMRs) in a plurality of regions in a set of biological samples comprising cell-free nucleic acids to generate a training data set, wherein the plurality of regions have been identified as regions of a genome that (i) comprises one or more sites that are amenable to methylation enrichment, and (ii) comprises methylation at below a threshold in a nondiseased control; (b) computer processing the training data set using machine learning to train an untrained classifier, thereby generating the trained classifier.
[0072] In some embodiments, the training data set comprises sample parameter data corresponding to characteristics of the subject from which a sample is derived from. In some embodiments, the characteristics comprises a cancer stage of the subject. In some embodiments, the characteristics comprises a cancer organ origin. In some embodiments, the characteristics comprises an age or sex of a subject. In some embodiments, the characteristics comprises one or more co-morbidities. In some embodiments, a subset of the set of biological samples are derived from subjects having cancer. In some embodiments, a subset of the set of biological samples are derived from subjects that do not have cancer.
[0073] In one aspect, the present disclosure provides a method of nucleic acids processing comprising: (a) assaying methylation levels of a plurality of regions in a biological sample comprising cell-free nucleic acids, wherein the plurality of regions have been identified as regions of a genome that (i) comprises one or more sites that are amenable to methylation enrichment, and (ii) comprises methylation at below a threshold in a healthy control, wherein the assaying comprises sequencing the cell-free nucleic acids to generate sequencing reads;(b) processing sequencing reads corresponding to a subset of the plurality of regions, wherein the subset comprises differentially methylated regions (DMRs), to identify (i) a total read count of sequencing reads corresponding to the subset of the plurality of regions and (ii) a DMR read count of sequencing reads corresponding to nucleic acids of less than 150 bp andcomprising at least 5 CpGs; (c) generating a normalized DMR methylation level by normalizing the DMR read count against the total read count. In some embodiments, the cell- free nucleic acid molecule is a cell-free deoxyribonucleic acid (DNA) molecule. In some embodiments, the biological sample comprises a blood sample and / or a plasma sample. In some embodiments, the region of the plurality of regions has a length of 300 base pairs (bp).
[0074] In some embodiments, the sequencing does not comprise bisulfite sequencing. In some embodiments, the sequencing comprises targeted sequencing. In some embodiments, the sequencing comprises using a plurality of capture probes. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to the plurality of regions. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation. In some embodiments, the sequencing generates sequencing reads corresponding and / or complementary to the plurality of regions.
[0075] In some embodiments, (a) comprises enriching methylated nucleic acids. In some embodiments, enriching comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG- binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody can be an anti 5-mC antibody. In some embodiments, the antibody can be an anti 5 -hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a nonspecific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule.
[0076] In some embodiments, the subset of the plurality of regions is identified by differential methylation analysis of a test sample and control sample. In some embodiments, the test sample is derived from a subject with cancer. In some embodiments, the control sample is derived from a subject without cancer. In some embodiments, the subset of the plurality of regions comprises one or more regions that exhibits hypermethylation in the test sample compared to the control sample. In some embodiments, the subset of the plurality of regions comprises one or more regions that comprises DMRs specific to a particular cancer type. In some embodiments, the subset of the plurality of regions comprises one or more regions that comprises DMRs that are not specific to a particular cancer type.
[0077] In some embodiments, the method further comprises identifying the subject as having a disease, based at least on the normalized DMR methylation level. In some embodiments,identifying comprises comparing the normalized DMR methylation level against a control level. In some embodiments, the control level corresponds to an expected value for a non- cancerous sample. In some embodiments, the control level corresponds to an expected value for a cancerous sample. In some embodiments, the disease or condition is a cancer or a tumor. In some embodiments, the cancer is breast cancer, bladder cancer, colorectal cancer, endometrial cancer, prostate cancer, renal cancer, pancreatic cancer, or lung cancer. In some embodiments, the cancer is a late-stage cancer. In some embodiments, the cancer is an early- stage cancer.
[0078] In some embodiments, the one or more sites that are amenable to methylation enrichment are determined by determining methylation levels in a fully methylated control sample. In some embodiments, the fully methylated control sample comprises nucleic acids subjected to in vitro methylation (e.g., in vitro enzymatic methylation). In some embodiments, determining methylation levels in the fully methylated control sample comprises enriching methylated nucleic acids. In some embodiments, the enriching comprises using a binder that binds to one or more methylated nucleotides. In some embodiments, the binder comprises a protein comprising a methyl-CpG-binding domain. In some embodiments, the protein is a MBD2 protein. In some embodiments, the binder comprises an antibody. In some embodiments, the antibody is an anti 5-mC antibody. In some embodiments, the antibody is an anti 5 -hydroxymethyl cytosine antibody. In some embodiments, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule.
[0079] In some embodiments, the fully methylated control sample further comprises filler deoxyribonucleic acid (DNA) molecule. In some embodiments, the filler DNA has a length of about 50 bp to about 800 bp.
[0080] In some embodiments, the fully methylated control sample further comprises genomic DNA. In some embodiments, the genomic DNA is subjected to shearing.
[0081] In some embodiments, the fully methylated control sample further comprises cell-free DNA.
[0082] In some embodiments, the method comprises a reduction in a noise level compared to a noise level of a corresponding sample that has a normalized DMR methylation level generated by normalizing the DMR specific level against a background derived from all regions of a whole genome. In some embodiments, the method comprises a reduction in a noise level compared to a noise level of a corresponding sample that has a normalized DMRmethylation level generated by normalizing the DMR specific level against a background derived from all genomic regions that are amenable to methylation enrichment.
[0083] In one aspect, the present disclosure provides a method, comprising: (a) obtaining a sample comprising cell-free nucleic acids from a first set of one or more subjects; (b) obtaining a control sample comprising cell-free nucleic acids from a second set of one or more subjects; (c) assaying the cell-free nucleic acids from the first set of one or more subjects and the second set of one or more subjects to generate a training data set comprising methylation states of one or more genomic regions comprising differentially methylation regions (DMRs); and (d) generating a classifier using the training data set to generate an output indicative of cancer.
[0084] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0085] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0086] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure.
[0087] Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0088] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0090] FIG. 1 shows a diagram illustrating a process for collecting flow-through of unmethylated / hypomethylated deoxyribonucleic acid (DNA) fragments.
[0091] FIG. 2 shows a computer system that is programmed or otherwise configured to implement methods provided herein.
[0092] FIG. 3 shows a diagram illustrating how blood quiet regions are selected.
[0093] FIGs. 4A-4B illustrate using blood quiet regions to assess blood quiet circulating tumor DNA (ctDNA) specific differentially methylated regions (DMRs) and ctDNA specific methylation level. FIG. 4A shows a diagram illustrating how to identify blood quiet ctDNA specific DMRs. FIG. 4B shows a diagram illustrating how to utilize the signal -to-noise (SNR) normalization method to calculate ctDNA specific methylation level.
[0094] FIG. 5A-5B illustrates a schematic of an example assay.
[0095] FIG. 6 shows data relating to limits of detection for a targeted panel workflow versus a non-targeted panel workflow.
[0096] FIG. 7 shows a diagram illustrating a process for developing a head and neck cancer (HNC) signature using an algorithm.
[0097] FIG. 8 shows methylation levels of signature DMRs in peripheral blood leukocytes(PBL), normal solid tissue, and primary solid tumor.
[0098] FIG. 9 shows correlation between tumor purity and mean signature signal within The Cancer Genome Atlas (TCGA) HNC tumor tissue.
[0099] FIG. 10 shows enrichment of signature DMRs in differential HNC CpG sites inTCGA, CpG islands, shores, and shelves.
[0100] FIG. 11 shows identified signature DMRs mapped to top 15 genes.
[0101] FIG. 12 shows pathway enrichment in signature DMRs.
[0102] FIG. 13 shows methylation levels of signature DMRs in peripheral blood leukocytes(PBL), normal tissue, head and neck squamous cell carcinoma (HNSC), lung squamous cell carcinoma (LUSC), and cervical squamous cell carcinoma (CESC).
[0103] FIG. 14 illustrates a workflow for selecting proto-DMRs.
[0104] FIG. 15 illustrates a workflow for the baseline / tissue-informed approach.
[0105] FIG. 16 illustrates a workflow for selecting cancer specific anti-DMRs.
[0106] FIG. 17 illustrates a workflow for selecting cancer specific DMRs.
[0107] FIG. 18 illustrates a workflow for the tissue / baseline agnostic approach.
[0108] FIG. 19 illustrates a workflow for the joint model approach.
[0109] FIG. 20 shows the baseline / tissue agnostic scores generated upon subjecting various dilutions of FaDu cell line to the baseline / tissue agnostic (whole methylome) approach. The FaDu cell line were diluted in non-cancer control cell-free DNA (cfDNA).
[0110] FIG. 21 shows the baseline-informed scores generated upon subjecting various dilutions of FaDu cell line to the baseline-informed (whole methylome) approach. The FaDu cell line were diluted in non-cancer control cfDNA.
[0111] FIG. 22 shows the baseline / tissue agnostic scores generated upon subjecting various dilutions of FaDu cell line to the baseline / tissue agnostic (proto-DMR panel) approach. The FaDu cell line were diluted in non-cancer control cfDNA.
[0112] FIG. 23 shows the baseline-informed scores generated upon subjecting various dilutions of FaDu cell line to the baseline-informed (proto-DMR panel) approach. The FaDu cell line were diluted in non-cancer control cfDNA.
[0113] FIG. 24 shows the baseline-informed scores generated upon subjecting various dilutions of cfDNA from colorectal, stage III subjects to the baseline-informed (whole methylome) approach. The cfDNA from colorectal, stage III subjects were diluted in non- cancer control cfDNA.
[0114] FIG. 25 shows the baseline / tissue agnostic scores generated upon subjecting various dilutions of cfDNA from colorectal, stage III subjects to the baseline / tissue agnostic (whole methylome) approach. The cfDNA from colorectal, stage III subjects were diluted in non- cancer control cfDNA.
[0115] FIG. 26 shows the baseline-informed scores generated upon subjecting various dilutions of cfDNA from colorectal, stage IV subjects to the baseline-informed (proto-DMR panel) approach. The cfDNA from colorectal, stage IV subjects were diluted in non-cancer control cfDNA.
[0116] FIG. 27 shows the baseline / tissue agnostic scores generated upon subjecting various dilutions of cfDNA from colorectal, stage IV subjects to the baseline / tissue agnostic (proto- DMR panel) approach. The cfDNA from colorectal, stage IV subjects were diluted in non- cancer control cfDNA.
[0117] FIG. 28A shows the sensitivity and the specificity for detecting minimal residual disease using the whole methylome in the baseline-informed approach, the baseline / tissue agnostic approach, and the joint model approach.
[0118] FIG. 28B shows the sensitivity and the specificity for detecting minimal residual disease using the proto-DMR panel in the baseline-informed approach, the baseline / tissue agnostic approach, and the joint model approach.
[0119] FIG. 29 shows an example schematic of evaluating and thresholding joint predictive models.DETAILED DESCRIPTION
[0120] The present disclosure provides methods and / or systems for the processing and analysis of nucleic acids present in biological samples through the generation of libraries of methylated genomic regions, which can be useful in determining a risk or likelihood of a subject having cancer or a tumor with high sensitivity and / or high specificity. The methods and / or systems disclosed herein can process and / or analyze nucleic acids present in biological samples through different approaches. The one or more approaches can comprise baseline- informed approach, baseline-agnostic approach, and / or joint approach. The different approaches disclosed herein (e.g., baseline-informed approach, baseline-agnostic approach, and / or joint approach) can utilize the same panel comprising genomic regions (e.g., proto- DMRs, control genomic regions) with little to no methylation signals in non-cancer controls (e.g., a pool of non-cancer subjects) to identify and / or select one or more differentially methylated regions (DMRs) that can be used for monitoring a cancer or disease. Methylation patterns of nucleic acid molecules derived from a tissue sample or a non-tissue sample (e.g., cell free nucleic acid molecules, circulating tumor DNA) of a subject can be useful for predicting, screening, diagnosing and / or monitoring for a cancer. Utilizing the panel from non-cancer controls (e.g., a pool of non-cancer subjects) can offer various advantages compared to panels that require customization for different cancer indications. For example, the methods and / or systems disclosed herein may not need large cancer cohorts to determine one or more DMRs for monitoring a subject, and / or may not need prior knowledge of cancerspecific DMRs. The panel can be applicable across any cancer types without modification, and / or expensive panel redesigns. Furthermore, since the panel comprises genomic regions (e.g., proto-DMRs, control genomic regions) with little to no methylation signals in non- cancer controls, low sequencing depth may be needed to identify methylated in a subject, reducing sequencing cost. This can be in contrast to bisulfite sequencing that requiressequencing of the non-methylated regions, meaning that as the size of the panel expands there can be more required sequencing depth.
[0121] Methods and systems provided herein can comprise assaying the cell-free nucleic acids to identifying a methylation level of a subset of the cell-free nucleic acids, which can be processed to monitor a subject. The methylation states of various nucleic acids can allow for identification of differential methylation regions (DMRs) in a subject as compared to another sample (e.g., control). The presence of specific DMRs may be used to differentiate between, for example, cancerous and non-cancerous tissue. Specific DMR may be specific to types of tissues (e.g., a tumor or cancer cell) and may be used to monitor the presence or absence of methylated states of the tissue sample.
[0122] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0123] The term “subject,” as used herein, generally refers to any member of the animal kingdom. The subject may be a human. The subject may be an individual exhibiting a disease (e.g., cancer) or an individual not exhibiting the disease. The subject may be considered to have a risk of developing the disease, such as cancer. The subject may be symptomatic or asymptomatic for a disease. The subject may be a patient. The subject may be a patient receiving medical care for a disease or condition (e.g., cancer).
[0124] The term “genome,” as used herein, generally refers to genomic information from a subject, which may be, for example, at least a portion or an entirety of a subject’s hereditary information. A genome can be encoded either in DNA or in RNA. A genome can comprise coding regions (e.g., that code for proteins) as well as non-coding regions. A genome can include the sequence of all chromosomes together in an organism. For example, the human genome ordinarily has a total of 46 chromosomes. The sequence of all of these together may constitute a human genome.
[0125] The term “methylome,” used herein, generally refers to measure of an amount of DNA methylation and / or DNA methylation level at a plurality of sites or loci in a genome. DNA methylation is a process by which methyl groups are added to a DNA molecule. DNA methylation can act to modulate (e.g., repress) gene transcription. The methylome may correspond to (e.g., complementary to) all of the genome (whole genome methylation), a substantial part of the genome, or relatively small portion(s) of the genome. The term“methylome” as used herein can also refer to the set of methylation modifications (e.g., on a nucleic acid) in an organism, in a cell, or in a sample. A methylome can depend on the method of methylation measurement. For example, when using the antibody 5mC, a methylome can represent all the information of DNA methylation on the cytosines of a genome.
[0126] The term “nucleic acid,” used herein, generally refers to a polynucleotide comprising two or more nucleotides, i.e., a polymeric form of nucleotides of various lengths (e.g., at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 1000, 10000, or more nucleotides in length), either deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or analogs thereof. Nonlimiting examples of nucleic acids include deoxyribonucleic (DNA), ribonucleic acid (RNA), coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A nucleic acid may comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be made before or after assembly of the nucleic acid. The sequence of nucleotides of a nucleic acid may be interrupted by non-nucleotide components. A nucleic acid may be further modified after polymerization, such as by conjugation or binding with a reporter agent. A “variant” nucleic acid is a polynucleotide having a nucleotide sequence identical to that of its original nucleic acid except having at least one nucleotide modified, for example, deleted, inserted, or replaced, respectively. The variant may have a nucleotide sequence at least about 80%, 90%, 95%, or 99%, identity to the nucleotide sequence of the original nucleic acid.
[0127] Cell-free methylated DNA generally includes DNA that can be one or more nucleic acid molecules circulating freely in the blood stream. In some cases, cell-free methylated DNA can be methylated at various regions of the DNA. Samples, for example, plasma samples may be taken to analyze cell-free methylated DNA. Studies reveal that much of the circulating nucleic acids in blood arise from necrotic or apoptotic cells and greatly elevated levels of nucleic acids from apoptosis is observed in diseases such as cancer. Particularly for cancer, where the circulating DNA bears hallmark signs of the disease including mutations in oncogenes, microsatellite alterations, and, for certain cancers, viral genomic sequences, DNA or RNA in plasma has become increasingly studied as a potential biomarker for disease. For example, a quantitative assay for low levels of circulating tumor DNA in total circulatingDNA may serve as a better marker for detecting the relapse of colorectal cancer compared with carcinoembryonic antigen, the biomarker used clinically. Cell-free DNA (e.g., circulating cfDNA) may comprise circulating tumor DNA (ctDNA).
[0128] As used herein, “sequencing,” also referred to as “genomic sequencing,” is a process for determining the order of the chemical building blocks (e.g., adenine, cytosine, guanine, thymine, uracil) that make up a nucleic acid molecule (e.g., DNA, RNA, cDNA).
[0129] As used herein, “library preparation” generally includes list end-repair, A-tailing, adapter ligation, or any other preparation performed on the cell free DNA to permit subsequent sequencing of DNA. Library preparation can allow for a nucleic acid sample (e.g., DNA, cDNA) to adhere to the sequencing apparatus (e.g., a flow cell, a bead). Nonlimiting examples of library preparation include ligation-based library preparation and tagmentation-based library preparation. Library preparation can result in the creation of a sequencing library, a pool of nucleic acid (e.g., DNA) fragments with adapters attached. The type of adapter attached during library preparation can depend on the sequencing platform / apparatus used.
[0130] The output of sequencing can be a “sequencing read.” As used herein, a “sequencing read” is an inferred sequence of base pairs or base pair probabilities corresponding to all or part of a nucleic acid fragment (e.g., a DNA fragment). The length of a sequencing read can depend on the sequencing platform / apparatus used. The length of a sequencing read can be about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, or more base pairs in length.
[0131] As used herein, “sequencing depth” refers to the ratio of the total number of bases obtained by sequencing to the size of the genome. Sequencing depth can also refer to the average number of times each base is measured in a genome during sequencing. Factors that can determine sequencing depth can include the error rate of the sequencing methods, the assembly algorithm used during sequencing, the repeat complexity of the nucleic acid molecule, region, or genome that is being sequenced, and the length of the sequencing read.
[0132] As used herein, “supplemental processed DNA” (e.g., “filler DNA”) generally may be noncoding DNA or it may consist of amplicons.
[0133] The term “proto-DMR” or “proto-differentially methylated region,” or “quiet regions” can refer a genomic region that is able to be methylated by a methylation enrichment assay (e.g., in vitro methylation assay) but does not necessarily show methylation in a given sample type. The proto-DMRs refer to one or more genomic regions that can be 1) pulled down upon in vitro methylation (e.g., enzymatic methylation) of the genomic regions, and 2) confirmed to be non-methylated in non-cancer subjects. These proto-DMRs can be areas of interest in differential methylation analysis as they are shown to be capable of having methylation, however in a particular sample type (e.g., a non-diseased or healthy control), no methylation is observed. As such, a cancer sample, or other sample of interest may comprise a methylation state that is different from a healthy or non-diseased control, thus giving rise to a differentially methylated region (DMR), which be analyzed. The one or more proto-DMRs can be captured (e.g., via hybrid capture and / or multiplex PCR) and / or analyzed in a subject to identify DMRs and anti-DMRs.
[0134] The term “anti-DMR” or “anti-differentially methylated region” can refer a genomic region that is able to be methylated by a methylation enrichment assay (e.g., in vitro methylation assay) but does not show a change in methylation state in when comparing a sample without a particular condition to a sample with a particular condition. The anti-DMR can identified from within a proto-DMR. These regions may be used for normalization or as a reference, for example, in the methods disclosed herein. As these regions may be methylated, these regions, under certain conditions, can be analyzed using methylation specific methods, while representing a background level for samples relating to a particular condition.
[0135] The term “DMR” or “differentially methylated region” can refer a genomic region that is methylated and / or hypermethylated in a sample with a particular condition, compared to a sample without a particular condition. The DMR can be identified from within a proto- DMR.
[0136] The term “non-diseased control” or “non-diseased sample” can refer to a sample that is derived from a sample that does not have a particular disease. The non-diseased control or sample may be substantially free of a particular disease (e.g., cancer), and can be used to as a control or reference for use in detection of the particular disease. The non-diseased control or non-diseased sample may comprise aberrations, genetic variants, or be infected with other diseases that are not the particular disease of interest. Various non-diseased samples or controls may be used in conjunction with other non-diseased samples or controls such to create a more unbiased control, which may more accurately represent the range of samples that are free of a particular condition (while potentially having other conditions)
[0137] In some embodiments, the fragment length metric can be fragment length. In some embodiments, the subject cell-free methylated DNA can be limited to fragments having a length of < 170 bp, < 165 bp, < 160 bp, < 155 bp, < 150 bp, < 145 bp, < 140 bp, < 135 bp, < 130 bp, < 125 bp, < 120 bp, < 115 bp, < 110 bp, < 105 bp, or < 100 bp. In other embodiments, the subject cell-free methylated DNA can be limited to fragments having a length of between about 100 - about 150 bp, 110 - 140 bp, or 120 - 130 bp.
[0138] In some embodiments, the fragment length metric can be the fragment length distribution of the subject cell-free methylated DNA. In some embodiments, the subject cell- free methylated DNA can be limited to fragments within the bottom 50th, 45th, 40th, 35th, 30th, 25th, 20th, 15th, or 10th percentile based on length.
[0139] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 can be equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0140] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 can be equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0141] As used herein, “sheared genomic nucleic acid molecule” or “sheared DNA,” also referred to as “cfDNA mimic” can comprise a subset of whole-genome DNA. In some cases, sheared DNA comprises randomly cleaved DNA. Alternatively, in some cases, sheared DNA comprises DNA cleaved at specified locations along the genome. In some cases, sheared DNA comprises DNA that can be fragmented to a predetermined fragment range. Physical shearing can be performed using, for example, probe sonication or nebulization. Enzymatic shearing or fragmentation can also be performed to generate sheared DNA. Sheared genomic nucleic acid molecules (e.g., sheared DNA) can be about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, or more base pairs in length.
[0142] As used herein, the “control” may comprise both positive and negative control, or at least a positive control.Methods for Comparative DNA Methylation
[0143] Cell-free nucleic acids, such as cell-free DNA (cfDNA), which can be present in biological samples that can be collected non-invasively can be a heterogeneous population comprising both cfDNA derived from healthy tissues and cfDNA derived from tumor or cancer cells (e.g., ctDNA). For example, samples that can be collected noninvasively can be blood, urine, saliva, or CSF. Cancer development can be associated with focal gain of 5’ methylcytosines (5mC), for instance, at cytosine-phosphate-guanine (CpG) islands and CpG island shores. Cancer development can also be associated with global cytosine demethylation. Global cytosine demethylation can be a genome- wide loss of 5mC. In some cases, ctDNA can be distinguished from cfDNA molecules derived from healthy tissue (e.g., non-tumor and / or non-cancer tissue) by the methylation level (e.g., the percentage of nucleotide residues that are methylated) of the nucleic acid molecules. In some cases, nucleic acid molecules of or derived from tumor tissue and / or cancer tissue can be hypomethylated (e.g., can comprise a lower level of methylation, for instance, wherein there are fewer methylated nucleotide residues and / or a lower percentage of methylated nucleotide residues) compared to nucleic acid molecules of or derived from healthy tissue, or non-diseased (e.g., nucleic acid molecules of or derived from healthy tissue that consist of or comprise nucleotide sequences corresponding to the same region(s) of the genome of the subject). For example, tumor- derived nucleic acid molecules (e.g., ctDNA molecules) can comprise one or more regions having fewer methylated nucleotide residues than nucleic acid molecules (e.g., cfDNA molecules) derived from healthy tissues (e.g., non-tumor and / or non-cancer tissues) in the same biological sample. In some cases, nucleic acid molecules of or derived from tumor tissue and / or cancer tissue can be hypermethylated (e.g., can comprise a higher level of methylation, for instance, wherein there are greater methylated nucleotide residues and / or a greater percentage of methylated nucleotide residues) compared to nucleic acid molecules of or derived from healthy tissue (e.g., nucleic acid molecules of or derived from healthy tissue that consist of or comprise nucleotide sequences corresponding to the same region(s) of the genome of the subject). For example, tumor-derived nucleic acid molecules (e.g., ctDNA molecules) can comprise one or more regions having greater methylated nucleotide residues than nucleic acid molecules (e.g., cfDNA molecules) derived from healthy tissues (e.g., non- tumor and / or non-cancer tissues) in the same biological sample. In some cases, all or a portion of a tumor-derived fraction of a plurality of cell-free DNA molecules (e.g., ctDNA) can be distinguished from cfDNA molecules derived from healthy tissue by one or more biophysical properties (e.g., the length of the cfDNA molecules or the presence ofstereotypical 5’ and 3’ end sequence motifs) and / or one or more fragmentomics patterns. For instance, ctDNA molecules can have shorter nucleic acid lengths than cfDNA molecules derived from healthy tissues. In some cases, ctDNA molecules may comprise stereotypical 5’ and 3’ end motifs. In some cases, one or more of these distinguishing features may be used to deplete a population of nucleic acid molecules of cfDNA derived from healthy tissue and / or to enrich a population of nucleic acid molecules for ctDNA. In some cases, ctDNA can have shorter fragment length compared to cfDNA derived from a healthy tissue.
[0144] Nucleic acid molecules derived from tumor or cancer cells or tissue (e.g., ctDNA) may be present in a biological sample (and / or a population of nucleic acids derived from the biological sample) in substantially lower quantities than nucleic acid molecules (e.g., cfDNA) derived from healthy tissue. It can be difficult to detect or sequence (e.g., determine a sequence identity of) ctDNA present in a plurality of nucleic acid molecules (e.g., cfDNA) in or derived from a biological sample, for instance, because they are present in the sample in lower quantities relative to cfDNA derived from healthy tissue (e.g., which may require using a greater amount of potentially scarce biological sample and / or which may require significantly higher sequencing depth). Various methods and systems disclosed herein may alleviate potential issues relating to the low quantities of ctDNA (e.g., via enrichment of methylated nucleic acids).
[0145] In some cases, global or whole-genome methylation techniques can provide information about methylation across the genome, which can differentiate between healthy and cancerous tissue. In some cases, a plurality of nucleic acids (e.g., cfDNA molecules or amplicons thereof derived from a biological sample) may be subjected to genome-wide depletion of nucleic acid molecules methylated in one or more specific regions of the genomic sequence of the nucleic acid molecules (e.g., CpG islands, CpG island shores, or repetitive sequences of the genome, such as long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), or LTRs (long terminal repeats)) to achieve increased sensitivity and / or increased specificity in assays for determining the presence or absence or the sequence identity of ctDNA molecules in the plurality. In some cases, a whole genome comprises all genomic regions. Alternatively, a whole genome can comprise all hypomethylated and / or all hypermethylated genomic regions. In some cases, a whole genome file comprises all the genomic regions of all the chromosomes of a sample (e.g., all human chromosomes). Alternatively, a whole genome file can comprise all the genomic regions of the autosomes (e.g., human chromosomes 1-22).
[0146] Alternatively, a subset of the global or whole genome can be used to provide specific information about methylation at specified regions or a plurality of regions of a genome. In some cases, specified regions or a plurality of regions can comprise one or more sites that are amenable to methylation enrichment. In some cases, one or more sites that are amenable to methylation enrichment can comprise one or more sites that can be enriched for one or more methylated site after in vitro methylation. In some cases, specified regions or a plurality of regions can comprise substantially no methylation at below a threshold in a healthy control. In some cases, specified regions or a plurality of regions can comprise (i) one or more sites that are amenable to methylation enrichment and (ii) substantially no methylation at below a threshold in a healthy control. Such specified regions or a plurality of regions can be experimentally validated and can provide additional information for using or distinguishing one or more biomarkers. The one or more biomarkers can be differentially methylated regions (DMRs) for the purposes of distinguishing cancer from control samples. Using a subset of the whole genome, as opposed to the whole genome may provide advantages such as reducing the overall noise of the method. For example, by selecting specific regions as a background, noise reduction may be more specifically tuned to eliminate noise specific to those regions. For example, the background may comprise regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment, or (ii) methylation at below a threshold in a healthy control. For example, the background may comprise a selection of specific DMRs. The use of a smaller region (e.g., as opposed to whole genome) may decrease the sequencing footprint or may allow for increased sequencing depth for areas that may be relevant for a given subject (e.g., subject-specific DMR).
[0147] Specific DMRs in a sample may be used to monitor a sample for the presence of a tumor or cancer. For example, the change in methylation (e.g., hypermethylation) in a DMR may be detected in a sample derived from a subject having cancer. This increase or decrease in methylation may be used as a marker for the subject’s cancer. The sample may comprise a plurality of DMRs that are indicative of the subject’s cancer. The plurality of DMRs may be specific to a subject’s cancer. In this way, the presence of one or more of the pluralities of DMRs may indicate that the cancer in present in the sample. By observing one or more of the pluralities of DMRs, the cancer may be monitored. For example, a cancer may be monitored subsequent to a therapy to determine an efficacy of a therapy. The cancer can be monitored between two time points. The cancer can be monitored for regression and / or progression. In some cases, by observing one or more of the pluralities of DMRs, recurrence can be determined.
[0148] Different subject may comprise a different set of DMRs. Monitoring subject may comprise monitoring DMRs that are specific to a given subject. In some cases, the DMRs that are specific to a given subject can be used to detect cancer, detect cancer progression and / or regression, or detect minimal residual disease, or any combinations thereof. For example, a first subject may comprise a DMR A, DMR B, and DMR C. A second subject may comprise a DMR X, DMR Y, and DMR Z. To monitor, the first subject, data relating to DMR A, DMR B, and DMR C may be used to detect the presence or absence of a tumor in the first subject. To monitor the second subject, data relating to DMR X, DMR Y, and DMR Z may be used to detect the presence or absence of a tumor in the second subject. Although the first and second subject comprise different DMRs, data pertaining regions of other possible DMRs can still be collected. For example, data pertaining to regions corresponding to DMR A, DMR B, and DMR C can still be collected for the second subject. Downstream data analysis may selectively analyze specific DMRs. For example, for the second subject, data pertaining to regions corresponding to DMR A, DMR B, and DMR C may be collect and then be omitted or ignored when monitoring the second subject. Collecting data for multiple DMR while analyzing a subset of the DMRs may allow for less sample specific reactions and eliminate the need to generate custom panels for each individuals and / or detecting cancer, while still providing data relevant for a given subject. For example, the sequencing reactions may be the same for libraries derived from different subject, with the down stream analysis customized or personalized (e.g., algorithmically) for a given subject. This may reduce variability in the data, or reduce or eliminate the need for custom built probes, primers or other nucleic acid tools, while still allowing for personalization for a given subject.
[0149] The subject-specific DMRs may comprise DMRs that are specific to a cancer subtype or cancer tissue of origin. The subject-specific DMRs may comprise a set of DMRs specific to a patient’s tumor. The subject-specific DMRs may comprise DMRs specific to a non- cancerous cell of a subject (e.g., patient).
[0150] In various embodiments, the methods of the disclosure allow for monitoring of a subject using one or more markers (e.g., DMRs) specific to the subject. In some cases, the one or more markers can be DMRs. The methods may be performed without the use of a primer set or bait set specific to the one or more markers specific to the subject. The primer set can be nucleic acid sequences designed to anneal one or more regions corresponding to one or more markers. The bait set can be labeled probes that can capture one or more regions corresponding to one or more markers. For example, the method may comprise the use of a primer set, or bait set that generated prior to determination of the one or more markers asbeing indicative of a tissue in a subject. The primer set or bait set may comprise primers or bait set that can anneal to the regions corresponding to the one or more markers, however the primer or bait sets may also anneal to regions that do not correspond to the one or more markers. For example, the primer or bait sets may anneal to regions that do not have complementary regions to that of one or more markers. The methods described in this disclosure may filter (e.g., computationally filter) or reduce the data set to comprise data to the one or more markers. For example, the methods may comprise obtaining data for a targeted sequencing reaction and then may be filtered to analyze data that can be deemed relevant for a given subject.
[0151] The method disclosed herein may specifically analyze regions that are amenable to methylation, as opposed to a whole genome. Whole genome sequencing may be agnostic to regions that are able to or are otherwise amenable to methylation in a biologically suitable manner (e.g., via enzymatic methylation). Generating panels specific to methylatable regions (e.g., proto-DMRs) may improve the sequencing data by enriching for areas that are able to be methylated (and therefore relevant for detection of DMRs). Generating panels for targeting one or more regions that have little to no methylation signals in non-cancer controls (e.g., proto-DMRs) can improve sequencing by lowering the sequencing depth requirement. For example, methylation signals for cancer can be detected with low sequencing depth by using the panel since the panel was generated to target regions that can have little to no methylation signals and / or amenable to methylation enrichment in non-cancer controls.
[0152] In an aspect, disclosed herein is a method of nucleic acids processing. The method can comprise assaying methylation levels of a plurality of regions in a biological sample comprising cell-free nucleic acids. The plurality of regions may have been identified as regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment, and / or (ii) methylation at below a threshold in a healthy control. The method can further comprise processing methylation levels of the plurality of regions to identify a methylation background. The methylation levels can be processed for a subset of the plurality of regions. The subset can comprise differentially methylated regions (DMRs), to identify a DMR specific methylation level. A normalized DMR methylation level can be generated by normalizing the DMR specific level against the methylation background (e.g., anti-DMRs).
[0153] In another aspect, disclosed herein is a method for classifying a sample derived from a subject. The sample can be obtained from the subject. The sample can comprise nucleic acid molecules (e.g., cell-free nucleic acid molecules) from the subject. The nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed to generate a data set. Thedata set can comprise methylation states of one or more genomic regions. The one or more genomic regions can comprise differentially methylation regions (DMRs). At least a portion of the data set can be processed to generate an output indicative of cancer in the subject. In some cases, the portion of the data set can pertain to a set of DMRs specific to the subject. For example, DMRs specific to the subject can comprise DMRs that can be unique to the subject. As another example, DMRs specific to the subject can comprise DMRs that may be present in the subject but not in one or more other subjects. As another example, DMRs specific to the subject can comprise DMRs that can be identified from the subject. As another example, DMRs specific to the subject can comprise DMRs that can be personal to the subject. As another example, DMRs specific to the subject can comprise DMRs that have different magnitude of methylation states that can be unique to the subject.
[0154] In some cases, the portion of the data set can be processed by different approaches. For example, the portion of the data set can be processed by a baseline-informed approach and / or a baseline-agnostic approach. The baseline-informed approach can be used when a baseline sample or a reference sample can be available. “Baseline sample” and “reference sample” can be used interchangeably. The reference sample can be obtained at a time prior to obtaining the sample from the subject. For example, the reference sample can be obtained from the subject at a time subsequent to diagnosis and / or prior to treatment of a therapy. The reference sample can be a tissue sample (e.g., cancer tissue sample) and / or a non-tissue sample (e.g., plasma sample). In some cases, when both a tissue sample and a non-tissue sample that were obtained at a time subsequent to obtaining the sample are available, both or either of the samples can be used in the baseline-informed approach. In some cases, when a reference sample can be available, the baseline-informed approach can be performed alone. In some cases, when a reference sample can be available, the baseline-informed approach and the baseline-agnostic approach can both be performed. In some cases, when the reference sample can not be available, the baseline-agnostic approach can be performed alone.
[0155] The baseline-informed approach and / or the baseline-agnostic approach disclosed herein can utilize a universal panel of genomic regions. A “universal panel of genomic regions” can be used interchangeably with “a panel of control genomic regions,” and “proto- DMR panel” herein. The universal panel of genomic regions can comprise regions of nonmethylation and / or regions that are amenable to methylation enrichment in non-cancer controls (e.g., a pool of non-cancer subjects). The universal panel of genomic regions can comprise both regions of non-methylation and regions that are amenable to methylation enrichment in non-cancer controls (e.g., a pool of non-cancer subjects). In some cases, theuniversal panel of genomic regions can comprise regions of little to no methylation signals in non-cancer controls (e.g., a pool of non-cancer subjects).Universal Panel of Genomic Regions (Proto- D MR panel)
[0156] The universal panel of genomic regions (e.g., proto-DMRs panel) can be universally adaptable across any cancer type without modification to the panel. By utilizing one or more control subjects without cancer to generate the universal panel of genomic regions, the need for expensive, hard-to-source cancer subjects may not be needed to generate a panel that can capture DMRs that may be associated with cancer. For example, the universal panel of genomic regions (e.g., proto-DMRs panel) can be used (e.g., may be sensitive) to identifying cancer-associated hypermethylation events (e.g., DMRs) across one or more types of cancer, since the universal panel of genomic can comprise regions that may have no to little methylation signal in non-cancer controls. The universal panel of genomic regions can be used to identify cancer-associated DMRs across various cancer types, which may eliminate the need to generate custom panels for each cancer types, and / or reduce the turn around time to identify cancer-associated DMRs across various cancer types.
[0157] In some cases, the universal panel of genomic regions can be optimized for use following a methylation enrichment assay to target specific genomic regions in the subject and / or one or more reference subjects (e.g., a pool of cancer subjects). For example, after subjecting cell-free nucleic acid molecules to methylation enrichment assay (e.g., cfMeDIP), the universal panel of genomic regions can be used to capture specific genomic regions of the methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) from the subject and / or one or more reference subjects (e.g., a pool of cancer subjects) for sequencing. For example, cfMeDIP-seq + proto-DMR panel targeted enrichment can yield minimal sequencing signal in non-cancer samples, dramatically reducing unnecessary sequencing costs. By using the universal panel of genomic regions, high sequencing depth may not be required to identify genomic regions comprising DMRs from the subject and / or one or more reference subjects (e.g., a pool of cancer subjects). With low sequencing depth, the turn around time and / or the cost to identify DMRs from the subjects and / or one or more reference subjects (e.g., a pool of cancer subjects) can be reduced compared to not using the universal panel of genomic regions. For example, using the universal panel of genomic regions can ensure that non-cancer samples generate almost no sequencing signal, while in cancer samples, nearly all detected reads can come from ctDNA. The universal panel of genomic regions sequencing depth requirements can be determined by the number of molecules present rather than the number of regions sequenced. Expansion of the universalpanel of genomic regions to a larger size (e.g., to 1MB) may not increase sequencing depth requirements.
[0158] With the universal panel of genomic regions, instead of designing a new assay for each patient or cancer type, informative DMRs and Anti-DMRs from a predefined single panel based on non-cancer subjects. The same static Proto-DMR panel can be applied across different cancer types and patient-specific models without modification. The DMRs and Anti- DMRs disclosed herein can be selected (via a custom algorithm) based on the cancer type or individual patient.
[0159] The universal panel of genomic regions can be generated from one or more control samples derived from one or more control subjects without cancer. In some cases, the one or more control samples are one or more non-tissue samples and / or one or more tissue samples. In some cases, the one or more control samples (e.g., from a pool of non-cancer subjects) can comprise control nucleic acid molecules. In some cases, the one or more control samples can comprise control cell-free nucleic acid molecules. The control nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed to generate one or more control data sets comprising methylation states of one or more control genomic regions. In some cases, the methylation states of one or more control genomic regions can comprise hypermethylated states, methylated states, non-methylated states, or hypomethylated states, or any combinations thereof. The control nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed by conducting one or more methylation reactions. In some cases, the one or methylation reactions can comprise in vitro methylation reactions s. In some cases, the one or more methylation reactions can result in methylation and / or hypermethylation of the one or more control genomic regions. In some cases, the one or more methylation reactions can result in one or more fully methylated control genomic regions. One or more control genomic regions that are fully methylated can have methylation and / or hypermethylation at every nucleobases within the one or more control genomic regions. For example, all of the nucleobases within the control genomic regions can be methylated. In some cases, the one or more methylation reactions can result in one or more partially methylated control genomic regions. One or more control genomic regions that are partially methylated can have methylation and / or hypermethylation at a portion of nucleobases within the one or more control genomic regions. For example, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 99% of nucleobases can be methylated and / or hypermethylated within the control genomic regions. For example, at most 50%, at most 60%, at most 70%, at most 80%, at most 90%, or at most 99% of nucleobases can bemethylated and / or hypermethylated within the control genomic regions. In some cases, the control nucleic acid molecules can be assayed by sequencing. In some cases, the control nucleic acid molecules can be sequenced after conducting the one or more methylation reactions (e.g., in vitro enzymatic methylation) and / or after conducting the one or more methylation enrichment reactions. In some cases, the one or more methylation reactions can be conducted, followed by the one or more methylation enrichment reactions. For example, the control nucleic acid molecules can be subjected to in vitro methylation, and then can be enriched for one or more methylated regions by performing a methylation enrichment reaction. Sequencing the control nucleic acid molecules (e.g., cell-free nucleic acid molecules) after conducting the one or more methylation reactions and / or one or more methylation enrichment reactions can generate a control data set comprising one or more control regions that can be amenable to methylation enrichment (e.g., enzymatic methylation enrichment). In some cases, the control nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be sequenced without conducting the one or more methylation reactions. Sequencing the control nucleic acid molecules (e.g., cell-free nucleic acid molecules) without conducting the one or more methylation regions (e.g., in vitro methylation) can generate another control data set comprising one or more controls regions that can be hypomethylated and / or non-methylated. In some cases, the another control data set can comprise one or more controls regions that can have little to no methylation signals in non-cancer subjects. In some cases, the set of the control nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be sequenced after conducting the one or more methylation reactions, and another set of the control nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be sequenced without conducting the one or more methylation reactions. In some cases, the data set comprising one or more control regions that can be amenable to methylation enrichment and the another data set comprising one or more controls regions that are hypomethylated and / or non-methylated can be processed. In some cases, one or more control regions that are amenable to methylation enrichment can comprise one or more control regions that can be enriched for one or more methylated control regions after in vitro methylation. Processing the control data set and the another control data set can comprise intersecting the control data set and the another control data set to identify one or more regions that can be common between the control data set and the another control data set, thereby generating the universal panel of genomic regions (e.g., proto-DMRs). The one or more regions that are common between the two data set (e.g., control data set and another control data set) can comprise one or more control regions that can be amenable to methylation enrichment and regions that can behypomethylated and / or non-methylated in non-cancer controls (e.g., a pool of non-cancer subjects). The universal panel of genomic regions can comprise hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation enrichment, or any combinations thereof. For example, the universal panel of genomic regions can comprise regions of non-methylation and regions that are amenable to methylation enrichment in non- cancer controls (e.g., a pool of non-cancer subjects).
[0160] An another aspect disclosed herein is a method of processing a nucleic acid sample of a subject. Nucleic acid molecules can be obtained from the nucleic acid sample of the subject. One or more methylated regions of the nucleic acid molecules can be enriched. A universal panel of genomic regions can be used to target one or more genomic regions of the one or more methylated regions.
[0161] The universal panel of genomic regions can be used following methylation enrichment to target specific genomic regions of the subject and / or one or more reference subjects (e.g., a pool of cancer subjects). For example, after subjecting nucleic acid molecules to methylation enrichment assay (e.g., cfMeDIP), the universal panel of genomic regions can be used to capture specific genomic regions of the methylated nucleic acid molecules from the subject for sequencing. Since the universal panel of genomic regions can be generated by identifying one or more control genomic regions that can be hypomethylated, nonmethylated, and / or amendable to methylation enrichment, or any combinations thereof, high sequencing depth may not be required to identify genomic regions comprising DMRs in the subject and / or one or more reference subjects (e.g., a pool of cancer subjects). In some cases, the sequencing depth can be a depth of at most 10 million (M) single reads, at most 20 M single reads, at most 30 M single reads, at most 40 M single reads, at most 50 M single reads, at most 60 M single reads, at most 70 M single reads, at most 80 M single reads, at most 90M single reads, or at most 100 M reads. In some cases, the sequencing depth can be a depth from 1 M single reads to 10 M single reads, from 10 M single reads to 20 M single reads, from 20 M single reads to 30 M single reads, from 30 M single reads to 40 M single reads, from 40 M single reads to 50 M single reads, from 50 M single reads to 60 M single reads, from 60 M single reads to 70M single reads, from 70 M single reads to 80 M single reads, from 80 M single reads to 90 M single reads, or from 90 M single reads to 100 M single reads. In some cases, the sequencing depth can be a depth of at least 1 M single reads, at least 10 M single reads, at least 20 M single reads, at least 30 M single reads, at least 40 M single reads, at least 50 M single reads, at least 60 M single reads, at least 70 M single reads, at least80 M single reads, at least 90 M single reads, at least 100 M single reads, or at least 200 M single reads.
[0162] An example of workflow of the method 1400 for generating the universal panel of genomic regions (e.g., proto-DMRs) is shown in FIG. 14. One or more control samples can be obtained. The one or more control samples can be derived from one or more non-cancer controls (e.g., a pool of non-cancer subjects) 1402. The one or more control samples can be derived from one or more blood sample or plasma sample. The one or more control samples can be derived from one or more tissue sample. The one or more control samples can comprise nucleic acid molecules derived from one or more non-tissue samples or one or more tissue samples (e.g., cell free nucleic acid molecules or nucleic acid molecules derived from a tissue sample). The nucleic acid molecules (e.g., cell free methylated nucleic acid molecules) from the one or more control samples can be further assayed for methylation enrichment. For example, methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP) 1403. cfMeDIP can pulldown cell free methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl-CpG binding domain (MBD), methylationdependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET-assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Next, the methylated nucleic acids (e.g., cell free methylated nucleic acids) can be sequenced. Sequencing can be performed with next generation sequencing (NGS) 1404 or with any sequencing methods disclosed herein. Sequencing can generate one or more data sets comprising sequencing reads corresponding to methylated regions, hypermethylated regions, hypomethylated regions, or non-methylated regions, or combination thereof that can be mapped along a genome (e.g., human genome). The one or more data sets can be analyzed to identify one or more regions with low number of sequencing reads 1405. In some cases, when two or more data sets are generated, the sequencing reads from the two or more data sets can be averaged to determine one or more regions with low number of average sequencing reads. The average can be a weighted average. Low number of sequencing reads, or low average sequencing reads can be at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1 sequencing reads. The one or more regions with low number of sequencing reads canrepresent one or more regions that have low methylation signals 1408. In some cases, the one or more regions with low methylation signals can be hypomethylated regions and / or nonmethylated regions.
[0163] The one or more control samples derived from non-cancer controls (e.g., a pool of non-cancer subjects) can also be subjected to one or more methylation reactions. The one or more methylation reactions can be an in vitro methylation. The one or more methylation reactions can generate one or more in-vitro fully methylated samples 1401. For example, the one or more in vitro fully methylated samples can have genomic regions with nucleobases that are all methylated. The one or more in vitro fully methylated samples can comprise methylated control nucleic acid molecules (e.g., cell free nucleic acid molecules). The methylated control nucleic acid molecules (e.g., cell free nucleic acid molecules) can be further assayed for methylation enrichment. For example, methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP) 1403. cfMEDIP can pulldown methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl-CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylationsensitive restriction enzyme (MSRE), TET-assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Next, the methylated nucleic acids (e.g., cell free methylated nucleic acids) can be sequenced. Sequencing can be performed with next generation sequencing (NGS) 1404 or with any sequencing methods disclosed herein. Sequencing can generate one or more data set comprising sequencing reads corresponding to methylated regions, hypermethylated regions, hypomethylated regions, or non-methylated regions, or combination thereof that can be mapped along a genome (e.g., human genome). The one or more data set can be analyzed to identify one or more regions with high number of sequencing reads 1406. One or more regions with high number of sequencing reads can be regions with high binding affinity, where the regions can be pulled down for sequencing upon methylation enrichment. In some cases, when two or more data sets are generated, the sequencing reads from the two or more data sets can be averaged to determine one or more regions with high number of average sequencing reads. The average can be a weighted average. In some cases, a computational method can be applied to identify one or more regions with high number of sequencing reads. For example, a computationalmodel-based analysis of ChlP-Seq (MACS) peak calling method can be used to identify peaks associated with high sequencing reads. The one or more regions with high number of sequencing reads can represent one or more regions with high methylation enrichment 1407. In some cases, the one or more regions with high methylation enrichment can be regions that are amenable to methylation enrichment (e.g., enzymatic methylation enrichment).
[0164] The identified one or more high methylation enrichment regions and the one or more low signal regions in non-cancer controls can be intersected 1409 to generate the proto- DMRs 1410. For example, the intersecting can comprise finding one or more regions that can be both regions of high methylation enrichment and low signals in non-cancer controls (e.g., a pool of non-cancer subjects).Baseline-Informed Approach
[0165] The baseline-informed approach can identify one or more subject- specific DMRs derived from using a reference sample obtained from the subject. The reference sample can be obtained prior to obtaining the sample. By using a reference sample obtained from the subject rather than from one or more samples obtained from a different subject or a plurality of different subjects, the one or more subject-specific DMRs can be personalized to the subject. For example, the one or more subject-specific DMRs can be unique to the subject. The baseline-informed approach can generate personalized ctDNA quantification and / or classification while maintaining high specificity. By leveraging individual non-tissue and / or tissue samples, a customized DMR signature (e.g., DMRs / anti-DMRs specific to a reference subject) can be generated with the baseline-informed approach, which can help avoid reliance on generic indication specific signature, and can make more precise in identifying DMRs relevant to the subject. The baseline-informed approach comprises DMR selection that can be restricted to regions that can also be present in a known positive control cohort via prevalence filtering, and / or can ensure statistical significance by applying strict outlier detection techniques.
[0166] In some cases, the one or more subject-specific DMRs can be used for tracking a subject’s progression and / or regression of a cancer. For example, the reference sample (e.g., plasma sample) obtained subsequent to diagnosis and / or prior to treatment to a therapy can be used in the baseline-informed approach to generate a personalized set of one or more subjectspecific DMRs that may be used to track in the subject over time. In some cases, the reference sample can be a tissue sample. For example, the tissue sample can be a resected tumor tissue. The baseline-informed approach can be a tissue informed and / or a tissue naive approach. In some cases, the one or more subject-specific DMRs generated from using thetissue sample and / or non-tissue sample in the baseline-informed approach can have enhanced specificity. The baseline-informed approach can use a single universal panel of genomic regions (e.g., proto-DMRs), eliminating the need for patient-specific panels, reducing cost and complexity. The baseline-informed approach can generate unique patient specific DMRs and anti-DMRs based on the reference sample (e.g., tissue sample and / or non-tissue sample).
[0167] In some cases, the sample used in the baseline-informed approach can be a non-tissue sample (e.g., a blood sample and / or a plasma sample) and / or a tissue sample. In some cases, the nucleic acid molecules from a non-tissue sample and / or a tissue sample. In some cases, the sample can comprise nucleic acid molecules can be derived from a tissue sample or a non-tissue sample. In some cases, the nucleic acid molecules (e.g., cell-free nucleic acid molecules) obtained from the sample can be assayed for methylation enrichment. For example, the methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP). cfMeDIP can pulldown cell free methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl -CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET- assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Methylation enrichment can generate methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules). In some cases, the methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to proto-DMR enrichment by using a universal panel of genomic regions disclosed herein. For example, performing proto-DMR enrichment can target and isolate genomic regions that have complementarity (e.g., sequence complementarity) to the one or more proto-DMRs. The universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more genomic regions that corresponds to one or more control regions that are hypomethylated regions, regions of non-methylation, or regions that are amenable to methylation enrichment, or any combinations thereof in controls subject without cancer. In another example, the universal panel of genomic regions can be used to enrich for one or more genomic regions that correspond to one or more control regions that are regions of non-methylation and regions that are amenable to methylation enrichment in control subjects without cancer. For example, corresponding can mean that one or more genomic regions can have genomicsequences that are complementary to one or more sequences in the proto-DMRs. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by hybrid capture. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by multiplex PCR. For example, with the proto-DMR panel, enriching can be performed by targeting the one or more genomic regions that can have complementarity (e.g., sequence complementarity) to one or more control regions that are regions of non-methylation and regions that are amenable to methylation enrichment in control subjects without cancer. In some cases, the methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) after proto-DMR enrichment can be subjected to sequencing to generate a data set comprising methylation states of one or more genomic regions. In some cases, the one or more methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to whole methylome sequencing to generate the data set comprising methylation states of the one or more genomic regions. Whole methylome sequencing can be sequencing performed on one or more methylated nucleic acid molecules without further enrichment with the universal panel of genomic regions. In some cases, the methylation states of the one or more genomic regions can comprise hypermethylated states, methylated states, non-methylated states, or hypomethylated states, or any combinations thereof.I. Identifying a set of DMRs specific to the reference subject and a set of anti-DMRs specific to the reference subject
[0168] In some cases, the baseline-informed approach can comprise identifying a set of DMRs specific to the reference subject and / or a set of anti-DMRs specific to the reference subject. For example, a set of anti-DMRs and / or a set of DMRs specific to one or more reference subjects can refer to anti-DMRs and / or a set of DMRs that can be identified by assaying the one or more reference samples. The set of DMRs specific to the reference subject and / or the set of anti-DMRs specific to the reference subject can be identified with a reference sample by using the universal panel of genomic regions (e.g., proto-DMRs). For example, by using the panel of genomic regions, the set of DMRs specific to the reference subject and / or the set of anti-DMRs specific to the reference subject can be selected from the universal panel of genomic regions. For example, the set of DMRs specific to the reference subject and / or the set of anti-DMRs specific to the reference subject can be a subset of the universal panel of genomic regions. In some cases, the reference subject can be the subject. In some cases, the method can comprise obtaining the reference sample from the subject. In some cases, the reference sample can be obtained at a time prior to obtaining the sample. Insome cases, the reference sample can be obtained from the subject subsequent to diagnosis with the cancer. In some cases, the reference sample can be obtained from the subject prior to treatment with a therapy. In some cases, the reference sample can be obtained from the subject subsequent to diagnosis with the cancer and prior to treatment with a therapy. In some cases, the reference sample can be a blood sample and / or a plasma sample. In some cases, the reference sample can be a tissue sample. In some cases, the tissue sample can be a cancer tissue sample from the subject. For example, the tissue sample can be a resected tumor tissue. In some cases, the reference sample can comprise reference nucleic acid molecules derived from a tissue sample. In some cases, the reference sample can comprise reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) derived from a non-tissue sample. In some cases, the reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed for methylation enrichment. For example, the methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP). cfMeDIP can pulldown cell free methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl-CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET-assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Methylation enrichment can generate methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules). In some cases, the methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to proto-DMR enrichment by using the universal panel of genomic regions disclosed herein. For example, proto-DMR enrichment can mean to target and isolate genomic regions that have complementarity (e.g., sequence complementarity) to the one or more proto-DMRs. The universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomic regions that corresponds to one or more control regions that are hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation enrichment, or any combinations thereof in control subjects without cancer. In another example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomic regions that corresponds to one or more control regions that are regions of non-methylation and regions that are amenable to methylationenrichment in control subjects without cancer. For example, corresponding can mean that one or more genomic regions can have genomic sequences that are complementary to one or more sequences in the proto-DMRs. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by hybrid capture. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto- DMRs by multiplex PCR. For example, with the proto-DMR panel, enriching can be performed by targeting the one or more genomic regions that can have complementarity (e.g., sequence complementarity) to one or more control regions that are regions of nonmethylation and regions that are amenable to methylation enrichment in control subjects without cancer. In some cases, the methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) after proto-DMR enrichment can be subjected to sequencing to generate a reference data set comprising methylation states of one or more reference genomic regions. In some cases, the methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to whole methylome sequencing to generate a reference data set comprising methylation states of one or more reference genomic regions. Whole methylome sequencing can be sequencing performed on one or more methylated nucleic acid molecules without further enrichment with the universal panel of genomic regions. In some cases, the methylation states of the one or more reference genomic regions can comprise hypermethylated states, methylated states, non-methylated states, or hypomethylated states, or any combinations thereof. In some cases, the reference data set can be processed to identify a set of DMRs specific to the reference subject. For example, the reference data set can be analyzed to determine regions with high sequencing reads. The set of DMRs specific to the reference subject can comprise reference genomic regions that are hypermethylated and / or methylated. In some cases, the reference data set can be processed to identify a set of anti-DMRs specific to the reference subject. For example, the reference data set can be analyzed to determine regions with little to no sequencing reads. The set of anti-DMRs specific to the reference subject can comprise reference genomic regions that are non-methylated and / or hypomethylated. The set of anti-DMRs reference specific to the subject can comprise genomic regions that are non-methylated and / or hypomethylated in non-cancer subjects.IL Identifying a set of anti-DMRs specific to one or more reference subjects
[0169] In some cases, the baseline-informed approach can comprise identifying a set of anti- DMRs specific to one or more reference subjects. For example, a set of anti-DMRs specific to one or more reference subjects can refer to the anti-DMRs that can be identified byassaying the one or more reference samples. A set of anti-DMRs specific to one or more reference subjects and / or a set of DMRs specific to one or more reference subjects can be identified with one or more reference samples by using the universal panel of genomic regions (e.g., proto-DMRs). For example, by using the panel of genomic regions, the set of anti-DMRs specific to the one or more reference subjects and / or a set of DMRs specific to one or more reference subjects can be selected and / or targeted from the universal panel of genomic regions. For example, the set of anti-DMRs specific to the one or more reference subjects can be a subset of the universal panel of genomic regions. In some cases, the method can comprise obtaining one or more reference samples from one or more reference subjects. The one or more reference subjects can have cancer. In some cases, the one or more reference samples can be one or more blood sample or one or more plasma sample. In some cases, one or more reference samples can be one or more tissue samples. In some cases, one or more reference samples can comprise reference nucleic acid molecules derived from one or more tissue sample. In some cases, one or more reference samples can comprise reference nucleic acid molecules derived from one or more non-tissue samples (e.g., cell-free nucleic acid molecules). In some cases, the one or more reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed for methylation enrichment. For example, the methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP). cfMeDIP can pulldown cell free methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl -CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET- assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Methylation enrichment can generate one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules). In some cases, the one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to proto-DMR enrichment by using the universal panel of genomic regions disclosed herein. For example, performing proto-DMR enrichment can target and isolate genomic regions that have complementarity (e.g., sequence complementarity) to the one or more proto-DMRs. The example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomicregions that corresponds to one or more control regions that are hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation enrichment, or any combinations thereof in controls subject without cancer. In another example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomic regions that corresponds to one or more control regions that are regions of non-methylation and regions that are amenable to methylation enrichment in control subjects without cancer. For example, corresponding can mean that one or more genomic regions can have genomic sequences that are complementary to one or more sequences in the proto- DMRs. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by hybrid capture. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by multiplex PCR. For example, with the proto-DMR panel, enriching can be performed by targeting the one or more genomic regions that can have complementarity(e.g., sequence complementarity) to one or more control regions that are regions of non-methylation and regions that are amenable to methylation enrichment in control subjects without cancer. In some cases, the one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) after proto-DMR enrichment can be subjected to sequencing to generate a reference data set comprising methylation states of one or more reference genomic regions. In some cases, the one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to whole methylome sequencing to generate a reference data set comprising methylation states of one or more reference genomic regions. Whole methylome sequencing can be sequencing performed on one or more methylated nucleic acid molecules without further enrichment with the universal panel of genomic regions. In some cases, the methylation states of the one or more reference genomic regions can comprise hypermethylated states, methylated states, non-methylated states, or hypomethylated states, or any combinations thereof. In some cases, the reference data set can be processed to identify a set of DMRs specific to one or more reference subjects. For example, the reference data set can be analyzed to determine regions with high sequencing reads. The set of DMRs specific to the one or more reference subjects can comprise reference genomic regions that are hypermethylated and / or methylated. In some cases, the reference data set can be processed to identify a set of anti-DMRs specific to the one or more reference subjects. For example, the reference data set can be analyzed to determine regions with little to no sequencing reads. The set of anti-DMRs specific to the one or more reference subjects can comprise reference genomic regions that are non-methylatedand / or hypomethylated. The set of anti-DMRs specific to the one or more reference subjects can comprise genomic regions that are non-methylated and / or hypomethylated in non-cancer subjects.
[0170] Alternatively, the set of anti-DMRs specific to one or more reference subjects can be identified using the one or more reference samples and one or more control samples. The one or more control samples can be obtained from control subjects without cancer. In some cases, the one or more control samples can be one or more tissue samples and / or one or more nontissue samples. In some cases, the one or more control samples can comprise one or more control nucleic acid from one or more tissue samples and / or one or more non-tissue samples. In some cases, the one or more control samples can comprise one or more control nucleic acid molecules (e.g., cell-free nucleic acid molecules). In some cases, the one or more control nucleic acid molecules (e.g., cell-free nucleic acid molecules) and the one or more reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed for methylation enrichment. For example, the methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP). cfMeDIP can pulldown methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl-CpG binding domain (MBD), methylationdependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET-assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Methylation enrichment can generate one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) and / or one or more methylated control nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules). In some cases, the one or more methylated reference nucleic acid molecules and / or the one or more methylated control nucleic acid molecules can be subjected to sequencing to generate another reference data set comprising i) methylation states of one or more reference genomic regions and / or ii) methylation states of one or more control genomic regions. In some cases, the methylation states of one or more reference genomic regions and one or more control genomic regions can be compared to identify one or more nonmethylated regions and / or hypomethylated regions in both reference genomic regions and control genomic regions. In some cases, the identified one or more non-methylated regions and / or hypomethylated regions can be compared with the universal panel of genomic regions(e.g., proto-DMRs) to identify anti-DMRs specific with the one or more reference subjects. For example, the identified one or more non-methylated regions and / or hypomethylated regions can be intersected with the universal panel of genomic regions (e.g., proto-DMRs) to identify regions of that overlap. For example, one or more genomic regions that overlap (e.g., have same genomic coordinate) to non-methylated regions, hypomethylated regions, and / or regions that are amenable to methylation enrichment that can be comprised in the panel can be identified. The set of anti-DMRs specific to the one or more reference subjects can comprise reference genomic regions that are non-methylated and / or hypomethylated.III. Methylation Score
[0171] In some cases, the one or more genomic regions of the subject can be compared to the set of DMRs specific to the reference subject. In some cases, comparing can generate one or more counts of the set of DMRs specific to the subject. In some cases, counts can be methylation signal obtained after analyzing the sequencing data set. In some cases, counts can be sequencing reads that map to the region of interest and / or average of the sequencing reads that map to the region of interest. For example, counts can be one or more sequencing reads in one or more genomic regions that map to the set of DMRs specific to the reference subject to generate one or more counts of the set of DMRs specific to the subject. For example, mapping can mean finding one or more genomic regions that share the same genomic coordinates as the set of DMRs specific to the reference subject. In some cases, the one or more genomic regions of the subject can be compared to the set of anti-DMRs specific to the reference subject and / or anti-DMRs specific to the one or more reference subjects. In some cases, comparing can generate one or more counts of the set of anti-DMRs specific to the subject. In some cases, counts can be methylation signal obtained after analyzing the sequencing data set. In some cases, counts can be sequencing reads that map to the region of interest and / or average of the sequencing reads that map to the region of interest. For example, comparing can comprise counting for one or more sequencing reads in one or more genomic regions that map to the set of anti-DMRs specific to the reference subject and / or anti-DMRs specific the one or more reference subjects to generate one or more counts of the set of anti-DMRs specific to the subject. For example, mapping can mean finding one or more genomic regions that share the same genomic coordinates as the set of DMRs and / or the set of anti-DMRs specific to the reference subject. The one or more counts of the set of DMRs specific to the subject can be normalized to the one or more counts of the set of anti- DMRs specific to the subject. Normalizing can be performed to ensure data consistency across different data sets and reduce noise that may exist within the data sets. In some cases,normalizing can generate a methylation score. In some cases, normalizing can be performed by computing the ratio of the set of the one or more counts of the set of DMRs specific to the subject to the one or more counts of the set of anti -DMRs specific to the subject. In some cases, the one or more counts of the set of DMRs specific to the subject can be averaged to generate an average count of the set of DMRs specific to the subject. In some cases, the one or more counts of the set of anti -DMRs specific to the subject can be averaged to generate an average count of the set of anti -DMRs specific to the subject. In some cases, ratio of the average count of the set of DMRs specific to the subject to the average count of the set of anti -DMRs specific to the subject can be computed. In some cases, the ratio computed can be a methylation score. The methylation score can be a signal-to-noise ratio. In some cases, the methylation score can be generated by using one or more Bayesian-based statistical inference methods. The Bayesian-based statistical inference methods can be for estimating parameters and / or performing regression to generate the methylation score. In some cases, the methylation score can be a circulating tumor DNA quantification. The methylation score can be compared to a threshold score, thereby generating the output indicative of the cancer in the subject. The methylation score above the threshold score can be indicative of the presence of cancer. The methylation score above the threshold score can be indicative of the presence of circulating tumor DNAs. The methylation score below the threshold score can be indicative of the absence of circulating tumor DNAs.
[0172] The threshold score can be generated with a plurality of control samples obtained from a plurality of control subjects. For example, a plurality of control samples can be subjected to the baseline-informed approach disclosed herein to compute a plurality of methylation scores. The threshold score can be set where at least 5% of the control subjects had a score defined to be false positives. In some cases, the threshold score can be set where at least 1%, at least 2%, at least 3%, 4 at least %, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, or more of the control subjects had a score defined to be false positives. In some cases, the threshold score can be set where at most 1%, at most 2%, at most 3%, at most 4%, at most 5%, at most 6%, at most 7%, at most 8%, at most 9%, at most 10%, or more of the control subjects had a score defined to be false positives.
[0173] An example of the workflow for the baseline-informed approach 1500 is shown in FIG. 15. A baseline sample can be obtained from a subject. The term “baseline sample” can be used interchangeably with the “reference sample” herein. The term “subject” can be used interchangeably with “patient” herein. The baseline sample can be obtained from the subject at a time prior to obtaining the sample of interest. For example, the baseline sample can beobtained from the subject subsequent to diagnosis and / or prior to treatment with a therapy. In some cases, the baseline sample can be a tissue sample 1501. In some cases, the baseline sample can be a blood sample and / or a plasma sample. In some cases, nucleic acid molecules can be derived from a tissue sample and / or a non-tissue sample (e.g., blood sample, plasma sample). In some cases, the baseline sample can comprise nucleic acid molecules (e.g., cell- free nucleic acid molecules) 1501. The baseline sample can be subjected to a methylation enrichment assay 1502 to enrich for one or more methylation regions. For example, the nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be subjected to methylated DNA immunoprecipitation (cfMeDIP) or similar methylation enrichment thereof disclosed herein to enrich for one or more methylated nucleic acids (e.g., cell-free methylated nucleic acid molecules) for subsequent sequencing. In some cases, the baseline sample can be subjected to further proto-DMR enrichment by using the universal panel of genomic regions disclosed herein (e.g., Proto-DMRs) 1503. For example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomic regions of interest. For example, performing proto-DMR enrichment can target and isolate genomic regions that have complementarity (e.g., sequence complementarity) to one or more proto- DMRs. The one or more reference genomic regions of interest can correspond to one or more control regions that are hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation enrichment, or any combinations thereof in controls subject without cancer (e.g., correspond to proto-DMRs). For example, corresponding can mean that one or more genomic regions (e.g., reference genomic regions) can have genomic sequences that are complementary to one or more sequences in the proto-DMRs. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by hybrid capture. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by multiplex PCR. For example, with the proto- DMR panel, enriching can be performed by targeting the one or more genomic regions that correspond to one or more control regions that are regions of non-methylation and regions that are amenable to methylation enrichment in control subjects without cancer. The enriched methylated nucleic acids (e.g., cell free methylated nucleic acids) can next be sequenced. Sequencing can be performed with next generation sequencing (NGS) 1504 or with any sequencing methods disclosed herein. Sequencing can generate a dataset comprising methylated regions, hypermethylated regions, hypomethylated regions, or non-methylated regions, or combination thereof that can be mapped along a genome (e.g., human genome). The dataset can be processed to identify one or more patient specific hypermethylated DMRs1505. For example, the sequencing reads can be analyzed to identify one or more regions with high number of reads to identify methylated and / or hypermethylated regions (e.g., DMRs). In some cases, the identified patient specific hypermethylated DMRs can be analyzed for prevalence. For example, the patient specific hypermethylated DMRs can be compared to DMRs of non-cancer subject cohort to determine whether the identified patient specific hypermethylated DMRs are statistical outliers. For example, DMRs can be statistical outliers when compared to one or more non-cancer cohorts. In some cases, the identified patient specific methylated DMRs and / or hypermethylated DMRs can be compared to DMRs of cancer subject cohort to ensure the patient specific methylated DMRs and / or hypermethylated DMRs are present in both the subject’s sample and cancer subject cohort. By analyzing for prevalence, false positives in identifying the patient specific methylated DMRs and / or hypermethylated DMRs can be reduced.
[0174] Another sample can be obtained from the subject at a time point subsequent to obtaining the baseline sample. The another sample can be a tissue sample and / or a non-tissue sample (e.g., a blood sample and / or a plasma sample) from the subject. In some cases, the another sample can be nucleic acid molecules derived from a tissue sample and / or a nontissue sample from the subject. The another sample can comprise cell-free nucleic acid molecules 1506. The another sample can be subjected to a methylation enrichment assay to enrich for one or more methylation regions 1507. For example, the nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be subjected to methylated DNA immunoprecipitation (cfMeDIP) or similar methylation enrichment thereof disclosed herein to enrich for one or more methylated nucleic acids (e.g., cell-free methylated nucleic acid molecules) for subsequent sequencing. In some cases, the sample can be subjected to further proto-DMR enrichment by using the universal panel of genomic regions disclosed herein (e.g., Proto-DMRs) 1508. For example, proto-DMR enrichment can mean to target and isolate genomic regions that have complementarity (e.g., sequence complementarity) to the one or more proto-DMRs. The example, the universal panel of genomic regions (e.g., proto- DMRs) can be used to enrich for one or more reference genomic regions of interest. The one or more reference genomic regions of interest can correspond to one or more control regions that are hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation, or any combinations thereof in controls subject without cancer (e.g., correspond to proto-DMRs). For example, corresponding can mean that one or more genomic regions can have genomic sequences that are complementary to one or more sequences in the proto-DMRs. The universal panel of genomic regions can be used to enrichfor one or more genomic regions of the proto-DMRs by hybrid capture. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto- DMRs by multiplex PCR. For example, with the proto-DMR panel, enriching can be performed by targeting the one or more genomic regions that complementarity (e.g., sequence complementarity) to one or more control regions that are regions of nonmethylation and regions that are amenable to methylation enrichment in control subjects without cancer. The enriched methylated nucleic acids (e.g., cell free methylated nucleic acids) can next be sequenced. Sequencing can be performed with next generation sequencing (NGS) 1509 or with any sequencing methods disclosed herein. Sequencing can generate a dataset comprising methylated regions, hypermethylated regions, hypomethylated regions, or non-methylated regions, or combination thereof that can be mapped along a genome (e.g., human genome) and / or the genomic coordinates of the patient specific anti-DMRs and / or DMRs. The dataset can be processed to identify DMR counts and anti-DMR 1510. For example, the sequencing reads can be counted in genomic regions that correspond or map to the identified patient specific hypermethylated DMRs 1512 to generate DMR counts. In some cases, the DMR counts can be averaged. In some cases, the sequencing reads can be counted in genomic regions that correspond or map to the identified cancer specific anti-DMRs 1511 to generate anti-DMR counts. In some cases, the anti-DMR counts can be averaged. In some cases, the average of the DMR counts can be normalized to the average of the anti-DMR counts 1510. In some cases, normalizing the average of the DMR counts to the average of the anti-DMR counts can generate a patient specific methylation score 1513.Baseline-Agnostic Approach
[0175] The baseline-agnostic approach can be used when a reference sample from the subject may not be available. For example, the baseline-agnostic approach can be used when a sample obtained from the subject subsequent to diagnosis and prior to treatment may not be available. The baseline-agnostic approach can be used to identify DMRs and / or anti-DMRs across multiple cancer types using a single universal panel of genomic regions (e.g., proto- DMRs), eliminating the need for patient-specific panels, reducing cost and complexity. The baseline-agnostic approach can be used with a predefined subset of DMRs and / or anti-DMRs derived from a cohort of cancer subjects using the universal panel of genomic regions. The predefined subject of DMRs and / or anti-DMRs can be identified through a refined selection process as described herein for improved accuracy. In some cases, the baseline-agnostic approach can be used in addition to the baseline-informed approach disclosed herein. In some cases, the baseline-agnostic approach can be a tumor naive and / or tumor agnostic. Forexample, tumor naive and / or tumor agnostic approaches can be approaches that use nontissue samples. This can be in contrast to tumor-informed approach where tumor samples can be used for assaying and / or analyzing. The baseline-agnostic approach can provide robust classification of cancer. The baseline-agnostic approach can provide robust classification of cancer.
[0176] In some cases, the sample used in the baseline-agnostic approach can be a non-tissue sample (e.g., a blood sample and / or a plasma sample). In some cases, the sample can be a tissue sample. In some cases, the sample can comprise nucleic acid molecules can be derived from a tissue sample and / or a non-tissue sample. In some cases, the nucleic acid molecules the nucleic acid molecules (e.g., cell-free nucleic acid molecules) obtained from the sample can be assayed for methylation enrichment. For example, the methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP). cfMeDIP can pulldown cell free methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl-CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET-assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Methylation enrichment can generate methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules). In some cases, the methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to proto-DMR enrichment by using a panel of one and / or more control regions disclosed herein. For example, performing proto- DMR enrichment can target and isolate genomic regions that have complementarity (e.g., sequence complementarity) to the one or more proto-DMRs. The universal panel of genomic regions can be used to enrich for one or more genomic regions that corresponds to one or more control regions that are hypomethylated regions, regions of non-methylation, or regions that are amenable to methylation, or any combinations thereof in controls subject without cancer (e.g., proto-DMRs). In another example, universal panel of genomic regions can be used to enrich for one or more genomic regions that correspond to one or more control regions that are regions of non-methylation and regions that are amenable to methylation enrichment in control subjects without cancer. For example, the one or more genomic regions can have one or more genomic sequences that are complementary to one or more sequencesin the proto-DMRs. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by hybrid capture. The universal panel of genomic regions can be used to enrich for one or more genomic regions of the proto-DMRs by multiplex PCR. For example, with the proto-DMR panel, enriching can result in targeting the one or more genomic regions that can have complementarity (e.g., sequence complementarity) to one or more control regions that are regions of non-m ethylation and regions that are amenable to methylation enrichment in control subjects without cancer. In some cases, the methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) after proto-DMR enrichment can be subjected to sequencing to generate a data set comprising methylation states of one or more genomic regions. In some cases, the one or more methylated nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to whole methylome sequencing to generate the data set comprising methylation states of the one or more genomic regions. Whole methylome sequencing can be sequencing performed on one or more methylated nucleic acid molecules without further enrichment with the universal panel of genomic regions. In some cases, the methylation states of the one or more genomic regions can comprise hypermethylated states, methylated states, non-methylated states, or hypomethylated states, or any combinations thereof.I. Identifying a set of DMRs specific to one or more reference subjects and a set of anti- DMRs specific to one or more reference subjects
[0177] In some cases, the baseline-agnostic approach can comprise identifying a set of anti- DMRs specific to one or more reference subjects. For example, a set of anti-DMRs specific to one or more reference subjects can refer to anti-DMRs were identified by assaying the one or more reference samples. A set of anti-DMRs specific to one or more reference subjects and / or a set of DMRs specific to one or more reference subjects can be identified with one or more reference samples by using the universal panel of genomic regions (e.g., proto-DMRs). For example, by using the panel of genomic regions, the set of anti-DMRs specific to the one or more reference subjects and / or the set of anti-DMRs specific to the one or more reference subjects can be selected from the universal panel of genomic regions. For example, the set of anti-DMRs specific to the one or more reference subjects can be a subset of the universal panel of genomic regions. In some cases, the method can comprise obtaining one or more reference samples from one or more reference subjects. The one or more reference subjects can have cancer. In some cases, the one or more reference samples can be one or more blood sample or one or more plasma sample. In some cases, one or more reference samples can be one or more tissue samples. In some cases, one or more reference samples can comprisereference nucleic acid molecules derived from one or more tissue sample. In some cases, one or more reference samples can comprise reference nucleic acid molecules derived from one or more non-tissue samples (e.g., cell-free nucleic acid molecules). In some cases, the one or more reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed for methylation enrichment. For example, the methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP). cfMeDIP can pulldown cell free methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl-CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET-assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. In some cases, the one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to proto-DMR enrichment by using a panel of one or more control regions disclosed herein. For example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomic regions that corresponds to one or more control regions that are hypomethylated regions, regions of nonmethylation, and / or regions that are amenable to methylation enrichment, or any combinations thereof in controls subject without cancer. In another example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomic regions that corresponds to one or more control regions that are regions of nonmethylation and regions that are amenable to enrichment methylation enrichment in control subjects without cancer. In some cases, the one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) after proto-DMR enrichment can be subjected to sequencing to generate a reference data set comprising methylation states of the one or more reference genomic regions. In some cases, the one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) can be subjected to whole methylome sequencing to generate a reference data set comprising methylation states of one or more reference genomic regions. Whole methylome sequencing can be sequencing performed on one or more methylated nucleic acid molecules without further enrichment with the universal panel of genomic regions. In some cases, the methylation states of the one or more reference genomic regions can comprisehypermethylated states, methylated states, non-methylated states, or hypomethylated states, or any combinations thereof. In some cases, the reference data set can be processed to identify a set of DMRs specific to one or more reference subjects. For example, the reference data set can be analyzed to determine regions with high sequencing reads. The set of DMRs specific to the one or more reference subjects can comprise reference genomic regions that are hypermethylated and / or methylated. In some cases, the reference data set can be processed to identify a set of anti-DMRs specific to the one or more reference subjects. For example, the reference data set can be analyzed to determine regions with little to no sequencing reads. The set of anti-DMRs specific to the one or more reference subjects can comprise reference genomic regions that are non-methylated and / or hypomethylated. The set of anti-DMRs specific to the one or more reference subjects can comprise genomic regions that are non-methylated and / or hypomethylated in non-cancer subjects. For example, the set of anti-DMRs specific to the one or more reference subjects can comprise genomic regions that are non-methylated in non-cancer subjects.
[0178] Alternatively, the set of DMRs specific to one or more reference subjects and / or the set of anti-DMRs specific to one or more reference subjects can be identified using the one or more reference samples and one or more control samples. The one or more control samples can be obtained from control subjects without cancer. In some cases, the one or more control samples can be one or more tissue samples and / or one or more non-tissue samples. In some cases, the one or more control samples can comprise one or more control nucleic acid from one or more tissue samples and / or one or more non-tissue samples. In some cases, the one or more control samples can comprise one or more control nucleic acid molecules (e.g., cell-free nucleic acid molecules). In some cases, the one or more control nucleic acid molecules (e.g., cell-free nucleic acid molecules) and the one or more reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be assayed for methylation enrichment. For example, the methylation enrichment can be performed with cell free methylated DNA immunoprecipitation (cfMeDIP). cfMeDIP can pulldown cell free methylated nucleic acids (e.g., cell free nucleic acid molecules) for subsequent sequencing. Alternatively, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl -CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET- assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfite conversion with methylation specific PCR, and / or methylation-specific hybrid capture, orother derivatives thereof. Methylation enrichment can generate one or more methylated reference nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules) and / or one or more methylated control nucleic acid molecules (e.g., cell-free methylated nucleic acid molecules). In some cases, the one or more methylated reference nucleic acid molecules and / or the one or more methylated control nucleic acid molecules can be subjected to sequencing to generate another reference data set comprising i) methylation states of one or more reference genomic regions and / or ii) one or more control genomic regions. In some cases, the methylation states of one or more reference genomic regions and one or more control genomic regions can be compared to identify one or more non-methylated regions and / or hypomethylated regions in both reference genomic regions and control genomic regions. In some cases, the identified one or more non-methylated regions and / or hypomethylated regions can be compared with the universal panel of genomic regions (e.g., proto-DMRs) to identify anti-DMRs specific with the one or more reference subjects. For example, the identified one or more non-methylated regions and / or hypomethylated regions can be intersected with the panel of one or more control genomic regions (e.g., proto-DMRs) to identify regions that overlap. For example, one or more regions reference genomic regions that overlap (e.g., have same genomic coordinate) to non-methylated regions, hypomethylated regions, and / or regions that are amenable to methylation enrichment in the panel can be identified. The set of anti-DMRs specific to the one or more reference subjects can comprise reference genomic regions that are non-methylated states and / or hypomethylated states.
[0179] In some cases, the methylation states of one or more reference genomic regions and one or more control genomic regions can be compared. Comparing can identify one or more regions that are hypermethylated and / or methylated in the one or more reference genomic regions as opposed to in the one or more control genomic regions. In some cases, the identified one or more methylated regions and / or hypermethylated regions can be compared with the panel of one or more control genomic regions (e.g., proto-DMRs) to identify DMRs specific with the one or more reference subjects. For example, the identified one or more methylated regions and / or hypermethylated regions can be intersected with the panel of one or more control genomic regions (e.g., proto-DMRs) to identify regions that overlap. For example, one or more reference genomic regions that overlap (e.g., have same genomic coordinate) to non-methylated regions, hypomethylated regions, and / or regions that are amenable to methylation enrichment in the panel can be identified. The set of DMRs specificto the one or more reference subjects can comprise reference genomic regions that are methylated states and / or hypermethylated states.IL Methylation Score
[0180] In some cases, the one or more genomic regions of the subject can be compared to the set of DMRs specific to the reference subject. In some cases, comparing can generate one or more counts of the set of DMRs specific to the subject. In some cases, counts can be methylation signal obtained after analyzing the sequencing data set. In some cases, counts can be sequencing reads that map to the region of interest and / or average of the sequencing reads that map to the region of interest. For example, comparing can comprise counting for one or more sequencing reads in one or more genomic regions that map to the set of DMRs specific to the one or more reference subjects to generate one or more counts of the set of DMRs specific to the subject. For example, mapping can mean finding one or more genomic regions that share the same genomic coordinates as the set of DMRs specific to the one or more reference subjects. In some cases, the one or more genomic regions of the subject can be compared to the set of anti -DMRs specific to the one or more reference subjects. In some cases, comparing can generate one or more counts of the set of anti-DMRs specific to the subject. In some cases, counts can be methylation signal obtained after analyzing the sequencing data set. In some cases, counts can be sequencing reads that map to the region of interest and / or average of the sequencing reads that map to the region of interest. For example, comparing can comprise counting for one or more sequencing reads in one or more genomic regions that map to the set of anti-DMRs specific the one or more reference subjects to generate one or more counts of the set of anti-DMRs specific to the subject. For example, mapping can mean finding one or more genomic regions that share the same genomic coordinates as the set of DMRs and / or a set of anti-DMRs specific to the one or more reference subjects. The one or more counts of the set of DMRs specific to the subject can be normalized to the one or more counts of the set of anti-DMRs specific to the subject. Normalizing can be performed to ensure data consistency across different data sets and reduce noise that may exist within the data sets. In some cases, normalizing can generate a methylation score. In some cases, normalizing can be performed by computing the ratio of the set of the one or more counts of the set of DMRs specific to the subject to the one or more counts of the set of anti-DMRs specific to the subject. In some cases, the one or more counts of the set of DMRs specific to the subject can be averaged to generate an average count of the set of DMRs specific to the subject. In some cases, the one or more counts of the set of anti- DMRs specific to the subject can be averaged to generate an average count of the set of anti-DMRs specific to the subject. In some cases, ratio of the average count of the set of DMRs specific to the subject to the average count of the set of anti -DMRs specific to the subject can be computed. In some cases, the ratio computed can be a methylation score. The methylation score can be referred to as a signal-to-noise ratio. In some cases, the methylation score can be a circulating tumor DNA quantification. The methylation score can be compared to a threshold score, thereby generating the output indicative of the cancer in the subject. The methylation score above the threshold score can be indicative of the presence of cancer. The methylation score above the threshold score can be indicative of the presence of circulating tumor DNAs. The methylation score below the threshold score can be indicative of the absence of circulating tumor DNAs.
[0181] The threshold score can be generated with a plurality of control samples obtained from a plurality of control subjects. For example, the plurality of control samples can be subjected to the baseline-informed approach disclosed herein to compute a plurality of methylation scores. The threshold score can be set where at least 5% of the control subjects had a score defined to be false positives. In some cases, the threshold score can be set where at least at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, or more of the control subjects had a score defined to be false positives. In some cases, the threshold score can be set where at most 1%, at most 2%, at most 3%, at most 4%, at most 5%, at most 6%, at most 7%, at most 8%, at most 9%, at most 10%, or more of the control subjects had a score defined to be false positives.
[0182] As shown in FIG. 16 and FIG. 17, cancer specific DMRs and / or cancer specific anti- DMRs can be first identified in the baseline-agnostic approach. The term “cancer specific DMRs” can be used interchangeably with “DMRs specific to one or more reference subjects” herein. The term “cancer specific anti-DMRs” can be used interchangeably with “anti-DMRs specific to one or more reference subjects” herein. The identified cancer specific anti-DMRs can also be used in the baseline-informed approach.
[0183] As shown in the example workflow 1600 in FIG. 16, a set of cancer samples 1601 (e.g., one or more reference samples) and a set of non-cancer samples 1602 (e.g., control samples) can be obtained. The cancer samples can be tissue samples and / or non-tissue samples from cancer subjects. The non-cancer samples can be tissue samples and / or nontissue samples (e.g., blood sample and / or a plasma sample) from non-cancer subjects. In some cases, the cancer samples and non-cancer samples can comprise nucleic acid molecules derived from tissue samples and / or non-tissue samples. In some cases, the cancer samples and non-cancer samples can comprise nucleic acid molecules (e.g., cell-free nucleic acidmolecules). The cancer samples and the non-cancer samples can be assayed to generate data set comprising genomic regions of the cancer samples and genomic regions of the non-cancer samples. The one or more genomic regions of the cancer samples and of the non-cancer samples can be compared to identify one or more regions with stable signal in both cancer and non-cancer samples 1603, thereby generating a panel of stable regions 1604. Stable signal can be hypomethylated regions and / or non-methylated regions in both non-cancer samples and cancer samples. By identifying hypomethylated regions and / or non-methylated regions that are present in both the one or more genomic regions of the cancer samples and one or more genomic regions of the non-cancer samples, the panel of stable regions can be generated. The panel of one or more stable regions can be intersected with a panel of one or more control regions (e.g., Proto-DMRs) 1605, 1606 to generate cancer specific anti-DMRs 1607. For example, the intersecting can comprise finding one or more regions that are common in the stable region and in the panel of one or more genomic regions, thereby generating the cancer specific anti-DMRs.
[0184] Cancer specific DMRs can be generated in a similar workflow. As shown in the example workflow 1700 in FIG. 17, a set of cancer samples 1701 (e.g., one or more reference samples) and a set of non-cancer samples 1702 (e.g., control samples) can be obtained. The cancer samples can be tissue samples and / or non-tissue samples from cancer subjects. The non-cancer samples can be tissue samples and / or non-tissue samples (e.g., blood sample and / or a plasma sample) from non-cancer subjects. In some cases, the cancer samples and non-cancer samples can comprise nucleic acid molecules derived from tissue samples and / or non-tissue samples. In some cases, the cancer samples and non-cancer samples can comprise nucleic acid molecules (e.g., cell-free nucleic acid molecules). The cancer samples and the non-cancer samples can be assayed to generate data set comprising genomic regions of the cancer samples and of the non-cancer samples. The one or more genomic regions of the cancer samples and of the non-cancer samples can be compared to identify one or more regions that are hypermethylated DMRs 1703, thereby generating a set of one or more DMRs 1704. For example, the identified one or more regions that are hypermethylated DMRs are one or more genomic regions of the cancer samples are hypermethylated and / or methylated as opposed to the one or more genomic regions of the non-cancer samples. The set of one or more DMRs can be intersected with a universal panel of genomic regions (e.g., Proto-DMRs) 1705, 1706 to generate cancer specific DMRs 1707. For example, the intersecting can comprise finding one or more regions that are common in the set of DMRs and in the universal panel of one or more genomic regions. In some cases,the identified cancer specific hypermethylated DMRs can be analyzed for prevalence. For example, the cancer specific hypermethylated DMRs can be compared to DMRs of noncancer cohort to determine whether the identified cancer specific hypermethylated DMRs are statistical outliers. For example, DMRs can be statistical outliers when compared to one or more non-cancer cohorts. In some cases, the identified cancer specific methylated and / or hypermethylated DMRs can be compared to DMRs of cancer cohort to ensure the cancer specific methylated and / or hypermethylated DMRs are present in both the reference subjects and another cancer cohort. By analyzing for prevalence, false positives in identifying the cancer specific methylated DMRs and / or hypermethylated DMRs can be reduced.
[0185] The identified one or more cancer specific DMRs and / or one or more cancer specific anti-DMRs can be applied to a baseline-agnostic approach. An example of the workflow for the baseline-agnostic approach 1800 is shown in FIG. 18. A sample can be first obtained from a subject. The sample can be a tissue sample and / or a non-tissue sample. In some cases, the same can comprise cell-free nucleic acid molecules derived from a tissue sample and / or a non-tissue sample. The sample can comprise a nucleic acid molecules (e.g., cell-free nucleic acid molecules) 1801. The sample can be subjected to a methylation enrichment assay to enrich for one or more methylation regions 1802 to enrich for one or more methylation regions. For example, the nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be subjected to methylated DNA immunoprecipitation (cfMeDIP) or similar methylation enrichment thereof disclosed herein to enrich for one or more methylated nucleic acids (e.g., cell-free methylated nucleic acid molecules) for subsequent sequencing. In some cases, the sample can be subjected to further proto-DMR enrichment by using a universal panel of genomic regions (e.g., Proto-DMRs) 1803. For example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more genomic regions of interest. The one or more reference genomic regions of interest can correspond to one or more control regions that are hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation enrichment (e.g., enzymatic methylation enrichment), or any combinations thereof in controls subject without cancer (e.g., correspond to proto-DMRs). For example, the one or more reference genomic regions can have one or more genomic sequences that are complementary to one or more sequences in the proto-DMRs. The enriched methylated nucleic acids (e.g., cell free methylated nucleic acids) can next be sequenced. Sequencing can be performed with next generation sequencing (NGS) 1804 or with any sequencing methods disclosed herein. Sequencing can generate a dataset comprising methylated regions, hypermethylated regions, hypomethylated regions, or non-methylatedregions, or combination thereof that can be mapped along a genome (e.g., human genome). The dataset can be processed to identify DMR counts and anti-DMR 1807. For example, the sequencing reads can be counted in genomic regions that correspond or map to the identified cancer specific hypermethylated DMRs 1805 to generate DMR counts. In some cases, the DMR counts can be averaged. In some cases, the sequencing reads can be counted in genomic regions that correspond or map to the identified cancer specific anti-DMRs 1806 to generate anti-DMR counts. In some cases, the anti-DMR counts can be averaged. In some cases, the average of the DMR counts can be normalized to the average of the anti-DMR counts 1807. In some cases, normalizing the average of the DMR counts to the average of the anti-DMR counts can generate a cancer specific methylation score 1808.Joint Approach
[0186] The “joint approach” or “joint model” can be used interchangeably. The join approach can use both the baseline-informed approach disclosed herein and the baseline-agnostic approach disclosed herein. The joint approach can leverage both the baseline-informed approach and the baseline-agnostic approach to maximize sensitivity and specificity in detection cancer. Detection of cancer can be measured by detection of circulating tumor DNA (ctDNA). The joint approach can integrate both scores from the baseline-informed approach and the baseline-agnostic approach for improved detection of cancer. In some cases, the joint model can comprise subject monitoring of a cancer without a personalized signature. For examples, the joint model can comprise subject monitoring of a cancer without a personalized signature derived from the baseline-informed approach. In some cases, the joint approach can comprise using a personalized signature derived from the baseline-informed approach, enhance sensitivity and specificity
[0187] In some cases, the one or more methylation scores obtained for the baseline-informed approach and the baseline-agnostic approach can be integrated to generate a single score. For example, the methylation score generated from the baseline-informed approach and the additional methylation score generated from the baseline-agnostic approach can be integrated to generate a single score. In some cases, the single score can be indicative of the cancer. In some cases, integrating comprise Support Vector Machine, logistic regression, Bayesian Interference Model, weighted average, decision trees, and / or random forests.
[0188] An example of the workflow for a joint model approach 1900 is shown in FIG. 19. Following the example workflow shown in FIG. 15 and FIG. 18, cancer specific DMRs and cancer specific anti-DMRs can be generated for use in the joint model approach. Patient specific DMRs can be identified following the baseline-informed approach, as shown in FIG.19, and as also shown in FIG. 15. As shown in FIG. 19, patient specific DMRs can be generated by using a baseline sample. The baseline sample can be obtained from the subject at a time prior to obtaining the sample of interest. For example, the baseline sample can be obtained from the subject subsequent to diagnosis and / or prior to treatment with a therapy. In some cases, the baseline sample can be a tissue sample 1901. In some cases, the baseline sample can be a blood sample and / or a plasma sample. In some cases, the baseline sample can comprise nucleic acid molecules derived from a tissue sample and or a non-tissue sample. In some cases, the baseline sample can comprise nucleic acid molecules (e.g., cell-free nucleic acid molecules) 1901. The baseline sample can be subjected to a methylation enrichment assay 1902 to enrich for one or more methylation regions. For example, the nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be subjected to methylated DNA immunoprecipitation (cfMeDIP) or similar methylation enrichment thereof disclosed herein to enrich for one or more methylated nucleic acids (e.g., cell-free methylated nucleic acid molecules) for subsequent sequencing. In some cases, the baseline sample can be subjected to further proto-DMR enrichment by using a panel of one or more control regions disclosed herein (e.g., Proto-DMRs) 1903. For example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more reference genomic regions of interest. The one or more reference genomic regions of interest can correspond to one or more control regions that are hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation enrichment, or any combinations thereof in controls subject without cancer (e.g., correspond to proto-DMRs). For example, the one or more reference genomic regions can have one or more genomic sequences that are complementary to one or more sequences in the proto-DMRs. The enriched methylated nucleic acids (e.g., cell free methylated nucleic acids) can next be sequenced. Sequencing can be performed with next generation sequencing (NGS) 1904 or with any sequencing methods disclosed herein. Sequencing can generate a dataset comprising methylated regions, hypermethylated regions, hypomethylated regions, or non-methylated regions, or combination thereof that can be mapped along a genome (e.g., human genome) and / or the genomic coordinates of the cancerspecific DMRs. The dataset can be processed to identify one or more patient specific hypermethylated DMRs 1905. For example, the sequencing reads can be analzyed to identify one or more regions with high number of reads to identify methylated and / or hypermethylated regions (e.g., DMRs). In some cases, the identified patient specific hypermethylated DMRs can be analyzed for prevalence. For example, the patient specific hypermethylated DMRs can be compared to DMRs of non-cancer cohort to determinewhether the identified patient specific hypermethylated DMRs are statistical outliers. In some cases, the identified patient specific hypermethylated DMRs can be compared to DMRs of cancer cohort to ensure the patient specific hypermethylated DMRs are present in both the subject’s sample and cancer cohort. By analyzing for prevalence, false positives can be reduced.
[0189] Another sample can be obtained from the subject at a time point subsequent to obtaining the baseline sample to compute the cancer specific methylation score generated from the baseline-agnostic approach and the patient specific methylation score generated from the baseline-informed approach. In some cases, a 3rd sample can be obtained at a 3rd time point after obtaining baseline sample. In some cases, a 4th sample can be obtained at a 4th time point after obtaining baseline sample. In some cases, a 5th sample can be obtained at a 5th time point after obtaining baseline sample. In some cases, a 6th sample can be obtained at a 6th time point after obtaining baseline sample. In some cases, a 7th sample can be obtained at a 7th time point after obtaining baseline sample. In some cases, an 8th sample can be obtained at an 8th time point after obtaining baseline sample. In some cases, one or more addition samples can be obtained at one or more additional time points (e.g., subsequent to one another sample).
[0190] The another sample or any subsequent one or more samples can comprise cell-free nucleic acid molecules 1906 derived from a tissue sample and / or a non-tissue sample from the subject. The another sample can be subjected to a methylation enrichment assay to enrich for one or more methylation regions 1907. For example, the nucleic acid molecules (e.g., cell-free nucleic acid molecules) can be subjected to methylated DNA immunoprecipitation (cfMeDIP) or similar methylation enrichment thereof disclosed herein to enrich for one or more methylated nucleic acids (e.g., cell-free methylated nucleic acid molecules) for subsequent sequencing. In some cases, the sample can be subjected to further proto-DMR enrichment by using a panel of one or more control regions disclosed herein (e.g., Proto- DMRs) 1908. For example, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more genomic regions of interest. The one or more genomic regions of interest can correspond to one or more control regions that are hypomethylated regions, regions of non-methylation, and / or regions that are amenable to methylation enrichment, or any combinations thereof in controls subject without cancer (e.g., correspond to proto- DMRs). For example, the one or more reference genomic regions can have one or more genomic sequences that are complementary to one or more sequences in the proto-DMRs. The enriched methylated nucleic acids (e.g., cell free methylated nucleic acids) can next besequenced. Sequencing can be performed with next generation sequencing (NGS) 1909 or with any sequencing methods disclosed herein. Sequencing can generate a dataset comprising methylated regions, hypermethylated regions, hypomethylated regions, or non-methylated regions, or combination thereof that can be mapped along a genome (e.g., human genome). The dataset can be processed to identify DMR counts and anti-DMR 1914. For example, the sequencing reads can be counted in genomic regions that correspond or map to the identified patient specific hypermethylated DMRs from the baseline sample 1912 to generate DMR counts. In some cases, the DMR counts can be averaged. In some cases, the sequencing reads can be counted in genomic regions that correspond or map to the identified cancer specific anti-DMRs from one or more cancer subjects 1910 to generate anti-DMR counts. In some cases, the anti-DMR counts can be averaged. In some cases, the average of the DMR counts can be normalized to the average of the anti-DMR counts 1914. In some cases, normalizing the average of the DMR counts to the average of the anti-DMR counts can generate a patient specific methylation score. In some cases, a machine learning classifier can be used to generate the patient specific score. In some cases, a Bayesian-based statistical inference method for estimating parameters ad / or performing regression can be used to generate the patient specific score.
[0191] In addition, the sequencing reads can be counted in genomic regions that correspond or map to the identified cancer specific hypermethylated DMRs from one or more cancer subjects 1911 to generate DMR counts. In some cases, the DMR counts can be averaged. In some cases, the average of the DMR counts can be normalized to the average of the anti- DMR counts 1915. In some cases, normalizing the average of the DMR counts to the average of the anti-DMR counts can generate a cancer specific methylation score.
[0192] The score from the baseline-agnostic approach and the score form the baseline- informed approach can be integrated. For example, the scores can be integrated by Support Vector Machine (SVM) 1915. Integration of the score can generate a single joint methylation score 1916 that can be the indicative of the cancer.
[0193] In some cases, a 3rd sample can be obtained at a 3rd time point after obtaining baseline sample. In some cases, a 4th sample can be obtained at a 4th time point after obtaining baseline sample. In some cases, a 5th sample can be obtained at a 5th time point after obtaining baseline sample. In some cases, a 6th sample can be obtained at a 6th time point after obtaining baseline sample. In some cases, a 7th sample can be obtained at a 7th time point after obtaining baseline sample. In some cases, an 8th sample can be obtained at an 8th time point after obtaining baseline sample.
[0194] In another aspect, disclosed herein is a method for monitoring a subject. The subject can be monitored for regression or progression of a disease or condition. The method can comprise assaying a biological sample of the subject. The biological sample can be assayed for one or more markers specific to the subject, using a universal panel of genomic regions.
[0195] In some cases, assaying comprises sequencing the nucleic acid molecules from the biological sample at a depth of at most 50 M single reads. In some cases, assaying comprises sequencing the nucleic acid molecules at a depth of at most 10 M single reads. In some cases, the sequencing depth can be a depth of at most 10 million (M) single reads, at most 20 M single reads, at most 30 M single reads, at most 40 M single reads, at most 50 M single reads, at most 60 M single reads, at most 70 M single reads, at most 80 M single reads, at most 90M single reads, or at most 100 M reads. In some cases, the sequencing depth can be a depth from 1 M single reads to 10 M single reads, from 10 M single reads to 20 M single reads, from 20 M single reads to 30 M single reads, from 30 M single reads to 40 M single reads, from 40 M single reads to 50 M single reads, from 50 M single reads to 60 M single reads, from 60 M single reads to 70M single reads, from 70 M single reads to 80 M single reads, from 80 M single reads to 90 M single reads, or from 90 M single reads to 100 M single reads. In some cases, the sequencing depth can be a depth of at least 1 M single reads, at least 10 M single reads, at least 20 M single reads, at least 30 M single reads, at least 40 M single reads, at least 50 M single reads, at least 60 M single reads, at least 70 M single reads, at least 80 M single reads, at least 90 M single reads, at least 100 M single reads, or at least 200 M single reads.
[0196] In some cases, the method for monitoring the subject comprise using a baseline- informed approach disclosed herein. Baseline-informed approach can involve using a reference sample from the subject, one or more control samples from non-cancer subjects, and / or one or more reference samples from cancer subjects as disclosed herein. The reference sample can be from the same subject. In some cases, the reference sample can be obtained at a time prior to obtaining the sample. In some cases, the reference sample can be obtained from the subject subsequent to diagnosis with the cancer. In some cases, the reference sample can be obtained from the subject prior to treatment with a therapy. In some cases, the reference sample can be obtained from the subject subsequent to diagnosis with the cancer and prior to treatment with a therapy. In some cases, the reference sample can be a blood sample and / or a plasma sample. In some cases, the reference sample can be a non-tissue sample. In some cases, the one or more control samples from non-cancer subjects, and / or one or more reference samples can be one or more blood samples or one or more plasma samples.In some cases, the one or more control samples from non-cancer subjects and / or one or more reference samples can be a tissue sample. In some cases, the reference sample can comprise reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) derived from a nontissue sample and / or a tissue sample. In some cases, the one or more control samples from non-cancer subjects comprise control nucleic acid molecules (e.g., cell-free nucleic acid molecules) derived from one or more tissue samples and / or one or more non-tissue samples. In some cases, the one or more reference samples can comprise one or more reference nucleic acid molecules (e.g., cell-free nucleic acid molecules) derived from one or more tissue samples and / or one or more non-tissue samples.
[0197] In some cases, the method for monitoring comprises comparing the universal panel of genomic regions to one or more reference genomic regions of the reference sample to generate a set of DMRs specific to the reference sample and / or a set of anti-DMRs specific to the reference sample. In some cases, the method for monitoring comprises comparing the universal panel of genomic regions to one or more reference genomic regions of the one or more reference samples to generate a set of DMRs specific to the one or more reference samples and / or a set of anti-DMRs specific to the one or more reference samples. In some cases, the set of DMRs specific to the one or more reference samples can further comprise comparing one or more reference genomic regions to one or more control regions of the one or more control samples. Comparing can identify one or more regions that hypermethylated and / or methylated in the reference genomic regions as opposed to the one or more control regions. The one or more identified hypermethylated and / or methylated regions can be compared to the universal panel of genomic regions to further identify the set of DMRs specific to the one or more reference samples. In some cases, the set of anti-DMRs specific to the one or more reference samples can further comprise comparing one or more reference genomic regions to one or more control regions of the one or more control samples.Comparing can identify one or more regions that non-methylated and / or hypomethylated in both the reference genomic regions and the one or more control regions. The one or more identified hypomethylated and / or non-methylated regions can be compared to the universal panel of genomic regions to further identify the set of anti-DMRs specific to the one or more reference samples.
[0198] In some cases, one or more markers comprise DMRs specific to the subject. In some cases, the anti-DMRs specific to the subject can be generated by comparing one or more genomic regions of the biological sample to a set of anti-DMRs specific to the one or more reference samples. In some cases, DMRs specific to the subject can be compared to the set ofDMRs specific to the reference sample to generate one or more counts of the DMRs specific to the subject. The one or more counts of the DMRs specific to the subject can be sequencing reads of the one or more genomic region of the biological sample that map to one or more regions corresponding to the DMRs specific to the reference samples. In some cases, the anti- DMRs specific to the subject can be compared to the set of anti -DMRs specific to the reference sample to generate one or more counts of the anti -DMRs specific to the subject. The one or more counts of the anti -DMRs specific to the subject can be sequencing reads of the one or more genomic region of the biological sample that map to one or more regions corresponding to the anti-DMRs specific to the reference sample. In some cases, the anti- DMRs specific to the subject can be compared to the set of anti-DMRs specific to the one or more reference samples to generate one or more counts of the anti-DMRs specific to the subject. The one or more counts of the anti-DMRs specific to the subject can be sequencing reads of the one or more genomic region of the biological sample that map to one or more regions corresponding to the anti-DMRs specific to the one or more reference samples. In some cases, the one or more counts of the DMRs specific to the subject can be normalized to one or more counts of the anti-DMRs specific to the subject. In some cases, the one or more counts of the DMRs specific to the subject can be averaged to generate an average count of the DMRs specific to the subject. In some cases, the one or more counts of the anti-DMRs specific to the subject can be averaged to generate an average count of the anti-DMRs specific to the subject. In some cases, the average count of the DMRs specific to the subject and be normalized to the average count of the anti-DMRs specific to the subject. For example, the ratio of the average count of the DMRs specific to the subject to the average count of the anti-DMRs specific to the subject can be computed. The computed ratio can represent a methylation score. A first methylation score can be an output indicative of cancer. The first methylation score can be compared to a threshold score disclosed herein. For example, a methylation score higher than a threshold score can be indicative of the presence of cancer. In some cases, a methylation score lower than a threshold score can be indicative of the absence of cancer.
[0199] In some cases, another biological sample can be obtained at a time before or after obtaining the biological sample. The another biological sample can be used to monitor regression or progression of the disease or condition. In some cases, the another biological sample can comprise another one or more markers specific to the subject. In some cases, the another one or more markers can comprise another DMRs specific to the subject. In some cases, the another DMRs specific to the subject can be generated by comparing the genomicregions of the another biological sample to the set of DMRs specific to the reference sample. In some cases, the method further comprises comparing the genomic regions of another biological sample to the set of anti-DMRs to generate another anti-DMRs specific to the subject. In some cases, the another DMRs specific to the subject can be compared to the set of DMRs specific to the reference sample to generate one or more counts of the another DMRs specific to the subject. The one or more counts of the another DMRs specific to the subject can be sequencing reads of the one or more genomic region of the another biological sample that map to one or more regions corresponding to the DMRs specific to the reference sample. In some cases, the another anti-DMRs specific to the subject can be compared to the set of anti-DMRs specific to the reference sample to generate one or more counts of the another anti-DMRs specific to the subject. The one or more counts of the another anti-DMRs specific to the subject can be sequencing reads of the one or more genomic region of the another biological sample that map to one or more regions corresponding to the anti-DMRs specific to the reference sample. In some cases, the another anti-DMRs specific to the subject can be compared to the set of anti-DMRs specific to the one or more reference samples to generate one or more counts of the another anti-DMRs specific to the subject. The one or more counts of the another anti-DMRs specific to the subject can be sequencing reads of the one or more genomic region of the another biological sample that map to one or more regions corresponding to the anti-DMRs specific to the one or more reference samples. In some cases, the one or more counts of the another DMRs specific to the subject can be normalized to one or more counts of the another anti-DMRs specific to the subject. In some cases, the one or more counts of the another DMRs specific to the subject can be averaged to generate an average count of the another DMRs specific to the subject. In some cases, the one or more counts of the another anti-DMRs specific to the subject can be averaged to generate an average count of the another anti-DMRs specific to the subject. In some cases, the average count of the another DMRs specific to the subject and be normalized to the average count of the another anti-DMRs specific to the subject. For example, the ratio of the average count of the another DMRs specific to the subject to the average count of the another anti-DMRs specific to the subject can be computed. The computed ratio can represent a second methylation score.
[0200] In some cases, the first methylation score and the second methylation score can be compared, thereby generating an output indicative of the regression or the progression of the disease or the condition. For example, a first methylation score greater than the second methylation score can be indicative of progression of the disease or condition. A firstmethylation score smaller than the second methylation score can be indicative of regression of the disease or condition. Similarly, an additional biological sample can be obtained to after the another biological sample to generate another methylation score for comparison.
[0201] In some cases, the method for monitoring the subject comprise using a baselineagnostic approach disclosed herein. Baseline-agnostic approach can involve using one or more control samples from non-cancer subjects, and / or one or more reference samples from cancer subjects. In some cases, the one or more control samples from non-cancer subjects, and / or one or more reference samples can be one or more blood samples or one or more plasma samples. In some cases, the one or more control samples from non-cancer subjects, and / or one or more reference samples can be one or more tissue samples. In some cases, the one or more control samples from non-cancer subjects comprise control nucleic acid molecules derived from one or more tissue samples and / or non-tissue samples. In some cases, the one or more reference samples can comprise one or more reference nucleic acid molecules derived from one or more tissue samples and / or non-tissue samples.
[0202] In some cases, the method for monitoring comprises comparing the universal panel of genomic regions to one or more reference genomic regions of the one or more reference samples to generate a set of DMRs specific to the one or more reference samples and / or a set of anti -DMRs specific to the one or more reference samples. In some cases, the set of DMRs specific to the one or more reference samples can further comprise comparing one or more reference genomic regions to one or more control regions of the one or more control samples. Comparing can identify one or more regions that hypermethylated and / or methylated in the reference genomic regions as opposed to the one or more control regions. The one or more identified hypermethylated and / or methylated regions can be compared to the universal panel of genomic regions to further identify the set of DMRs specific to the one or more reference samples. In some cases, the set of anti-DMRs specific to the one or more reference samples can further comprise comparing one or more reference genomic regions to one or more control regions of the one or more control samples. Comparing can identify one or more regions that non-methylated and / or hypomethylated in both the reference genomic regions and the one or more control regions. The one or more identified hypomethylated and / or nonmethylated regions can be compared to the universal panel of genomic regions to further identify the set of anti-DMRs specific to the one or more reference samples.
[0203] In some cases, one or more markers comprise DMRs specific to the subject. In some cases, the anti-DMRs specific to the subject can be generated by comparing one or more genomic regions of the biological sample to a set of anti-DMRs specific to the one or morereference samples. In some cases, the anti-DMRs specific to the subject can be compared to the set of anti-DMRs specific to the one or more reference samples to generate one or more counts of the anti-DMRs specific to the subject. The one or more counts of the anti-DMRs specific to the subject can be sequencing reads of the one or more genomic region of the biological sample that map to one or more regions corresponding to the anti-DMRs specific to the one or more reference samples. In some cases, the one or more counts of the DMRs specific to the subject can be normalized to one or more counts of the anti-DMRs specific to the subject. In some cases, the one or more counts of the DMRs specific to the subject can be averaged to generate an average count of the DMRs specific to the subject. In some cases, the one or more counts of the anti-DMRs specific to the subject can be averaged to generate an average count of the anti-DMRs specific to the subject. In some cases, the average count of the DMRs specific to the subject and be normalized to the average count of the anti-DMRs specific to the subject. For example, the ratio of the average count of the DMRs specific to the subject to the average count of the anti-DMRs specific to the subject can be computed. The computed ratio can represent a methylation score. A first methylation score can be an output indicative of cancer. The first methylation score can be compared to a threshold score disclosed herein. For example, a methylation score higher than a threshold score can be indicative of the presence of cancer. In some cases, a methylation score lower than a threshold score can be indicative of the absence of cancer. In another aspect, disclosed herein is a method for classifying a sample derived from a subject. The sample can be analyzed for at least a portion of a set of DMRs specific to the subject. The analysis for at least a portion of a set of DMRs specific to the subject can generate an output indicative of cancer.Analyzing can comprise sequencing, wherein the sequencing can have a depth of at most 50 M single reads.
[0204] In some cases, one or more genomic regions of the subject can be sequenced. In some cases, the sequencing depth can be a depth of at most 10 million (M) single reads. In some cases, the sequencing depth can be a depth of at most 10 million (M) single reads, at most 20 M single reads, at most 30 M single reads, at most 40 M single reads, at most 50 M single reads, at most 60 M single reads, at most 70 M single reads, at most 80 M single reads, at most 90M single reads, or at most 100 M reads. In some cases, the sequencing depth can be a depth from 1 M single reads to 10 M single reads, from 10 M single reads to 20 M single reads, from 20 M single reads to 30 M single reads, from 30 M single reads to 40 M single reads, from 40 M single reads to 50 M single reads, from 50 M single reads to 60 M single reads, from 60 M single reads to 70M single reads, from 70 M single reads to 80 M singlereads, from 80 M single reads to 90 M single reads, or from 90 M single reads to 100 M single reads. In some cases, the sequencing depth can be a depth of at least 1 M single reads, at least 10 M single reads, at least 20 M single reads, at least 30 M single reads, at least 40 M single reads, at least 50 M single reads, at least 60 M single reads, at least 70 M single reads, at least 80 M single reads, at least 90 M single reads, at least 100 M single reads, or at least 200 M single reads.
[0205] In some cases, the method for classifying further comprises providing a universal panel of genomic regions (e.g., proto-DMRs). In some cases, the universal panel of genomic regions (e.g., proto-DMRs) can be used to enrich for one or more genomic regions of the subject or one or more subjects (e.g., cancer subjects, non-cancer subjects). In some cases, the universal panel of genomic regions can be compared to one or more reference genomic regions to generate a set of DMRs specific to a reference sample disclosed herein and / or a set of anti-DMRs specific to the reference sample disclosed herein. In some cases, the universal panel of genomic regions can be compared to one or more reference genomic regions to generate a set of DMRs specific to one or more reference samples disclosed herein and / or a set of anti-DMRs specific to the one or more reference samples disclosed herein. The anti- DMRs specific to the one or more reference samples, the DMRs specific to the one or more reference samples, the DMRs specific to the reference sample, and / or the DMRs specific to the reference sample can be used as disclosed herein to generate a set of DMRs specific to the subject and a set of anti-DMRs specific to the subject. The set of DMRs specific to the subject and the set of anti-DMRs specific to the subject can be used as disclosed herein to compute a methylation score indicative of the cancer.
[0206] In some cases, the output indicative of the cancer can be used to determine progression of the cancer. In some cases, the output indicative of the cancer can be used to determine regression of the cancer. In some cases, the output indicative of the cancer can be used to determine therapy of the cancer. In some cases, the output indicative of the cancer can be used to determine progression of the cancer. In some cases, the output indicative of the cancer can be used to determine regression of the cancer. In some cases, the output indicative of the cancer can be used to determine therapy response for the cancer.
[0207] In another aspect, provided herein is a method for detecting Minimal Residual Disease (MRD). A biological sample from a subject can be assayed. Assaying may not comprise assaying a solid tumor sample of the subject. The biological sample can be a nontissue sample (e.g., a blood sample and / or a plasma sample). The biological sample can be a blood sample and / or a plasma sample. In some cases, the method can be a tumor naiveapproach. In some cases, the method disclosed herein can detect MRD can be detected at a specificity of at least 90%. In some cases, the method disclosed herein can detect MRD at a specificity of at least about 60%, at least about 70%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect MRD at a specificity of at most about 80%, at most about 81%, at most about 82%, at most about 83%, at most about 84%, at most about 85%, at most about 86%, at most about 87%, at most about 88%, at most about 89%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.5%, at most about 99.6%, at most about 99.7%, at most about 99.8%, at most about 99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect MRD at a specificity of 100%.
[0208] In some cases, the method disclosed herein can detect MRD at a sensitivity of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least 50%, at least about 60%, at least about 70%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect MRD at a sensitivity of at most about 10%, at most about 20%, at most about 30%, at most about 40%, at most 50%, at most about 60%, at most about 70%, at most about 80%, at most about 81%, at most about 82%, at most about 83%, at most about 84%, at most about 85%, at most about 86%, at most about 87%, at most about 88%, at most about 89%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.5%, at most about 99.6%, at most about 99.7%, at most about 99.8%, at most about99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect MRD at a sensitivity of 100%.
[0209] In another aspect, provided herein is a method of detecting Minimal Residual Disease (MRD). A biological sample from a subject can be assayed. In some cases, assaying may not comprise assaying a solid tumor sample of the subject. In some cases, assaying can comprise assaying a tumor sample of the subject. In some cases, assaying can comprise assaying one or more non-tissue sample of the subject (e.g., a blood sample and / or plasma sample). In some cases, assaying may not comprise assaying one or more non-tissue sample of the subject (e.g., a blood sample and / or plasma sample). In some cases, assaying comprise sequencing one or more genomic regions in the biological sample. Sequencing can have a depth of at most 50 M single reads. In some cases, the MRD can be head and neck cancer.
[0210] In another aspect, provided herein is a method for detecting an output indicative of cancer. In some cases, the method can comprise assaying at least a portion of a set of methylated genomic regions of a sample to generate an output indicative of cancer. In some cases, the assaying can comprise sequencing the at least portion of the set of methylated genomic regions. In some cases, the sequencing depth of one or more methylated genomic regions can be at most 50 M single reads.
[0211] In some cases, the sequencing depth of one or more genomic regions can be a depth of at most 10 million (M) single reads, at most 20 M single reads, at most 30 M single reads, at most 40 M single reads, at most 50 M single reads, at most 60 M single reads, at most 70 M single reads, at most 80 M single reads, at most 90M single reads, or at most 100 M reads. In some cases, the sequencing depth can be a depth from 1 M single reads to 10 M single reads, from 10 M single reads to 20 M single reads, from 20 M single reads to 30 M single reads, from 30 M single reads to 40 M single reads, from 40 M single reads to 50 M single reads, from 50 M single reads to 60 M single reads, from 60 M single reads to 70M single reads, from 70 M single reads to 80 M single reads, from 80 M single reads to 90 M single reads, or from 90 M single reads to 100 M single reads. In some cases, the sequencing depth can be a depth of at least 1 M single reads, at least 10 M single reads, at least 20 M single reads, at least 30 M single reads, at least 40 M single reads, at least 50 M single reads, at least 60 M single reads, at least 70 M single reads, at least 80 M single reads, at least 90 M single reads, at least 100 M single reads, or at least 200 M single reads.
[0212] In some cases, the sequencing depth of one or more methylated genomic regions can be a depth of at most 10 million (M) single reads, at most 20 M single reads, at most 30 M single reads, at most 40 M single reads, at most 50 M single reads, at most 60 M single reads,at most 70 M single reads, at most 80 M single reads, at most 90M single reads, or at most 100 M reads. In some cases, the sequencing depth can be a depth from 1 M single reads to 10 M single reads, from 10 M single reads to 20 M single reads, from 20 M single reads to 30 M single reads, from 30 M single reads to 40 M single reads, from 40 M single reads to 50 M single reads, from 50 M single reads to 60 M single reads, from 60 M single reads to 70M single reads, from 70 M single reads to 80 M single reads, from 80 M single reads to 90 M single reads, or from 90 M single reads to 100 M single reads. In some cases, the sequencing depth can be a depth of at least 1 M single reads, at least 10 M single reads, at least 20 M single reads, at least 30 M single reads, at least 40 M single reads, at least 50 M single reads, at least 60 M single reads, at least 70 M single reads, at least 80 M single reads, at least 90 M single reads, at least 100 M single reads, or at least 200 M single reads.
[0213] In some cases, the sequencing depth of one or more methylated genomic regions after proto-DMR enrichment (e.g., using Universal Panel of one or more genomic regions) can be a depth of at most 10 million (M) single reads, at most 20 M single reads, at most 30 M single reads, at most 40 M single reads, at most 50 M single reads, at most 60 M single reads, at most 70 M single reads, at most 80 M single reads, at most 90M single reads, or at most 100 M reads. In some cases, the sequencing depth can be a depth from 1 M single reads to 10 M single reads, from 10 M single reads to 20 M single reads, from 20 M single reads to 30 M single reads, from 30 M single reads to 40 M single reads, from 40 M single reads to 50 M single reads, from 50 M single reads to 60 M single reads, from 60 M single reads to 70M single reads, from 70 M single reads to 80 M single reads, from 80 M single reads to 90 M single reads, or from 90 M single reads to 100 M single reads. In some cases, the sequencing depth can be a depth of at least 1 M single reads, at least 10 M single reads, at least 20 M single reads, at least 30 M single reads, at least 40 M single reads, at least 50 M single reads, at least 60 M single reads, at least 70 M single reads, at least 80 M single reads, at least 90 M single reads, at least 100 M single reads, or at least 200 M single reads.
[0214] In another aspect, disclosed herein is a method for an output indicative of cancer comprising assaying a sample for at least a portion of a set of differentially methylated regions (DMRs) specific to a subject to generate an output indicative of cancer. The output indicative of cancer can be the presence or absence of cancer. The output indicative of cancer can be generated at a specificity of at least 90%.
[0215] In some cases, the output indicative of cancer can be generated at a specificity of at least about 60%, at least about 70%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at leastabout 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage in between the numbers. In some cases, the output indicative of cancer can be generated at at a specificity of most about 60%, at most about 70%, at most about 80%, at most about 81%, at most about 82%, at most about 83%, at most about 84%, at most about 85%, at most about 86%, at most about 87%, at most about 88%, at most about 89%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.5%, at most about 99.6%, at most about 99.7%, at most about 99.8%, at most about 99.9%, or any percentage in between the numbers. In some cases, the output indicative of cancer can be generated at a specificity of 100%.
[0216] In some cases, the output indicative of cancer can be generated at a sensitivity of at least at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least 50%, about 60%, at least about 70%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage in between the numbers. In some cases, the output indicative of cancer can be generated at a sensitivity of at most about 10%, at most about 20%, at most about 30%, at most about 40%, at most 50%, at most about 60%, at most about 70%, at most about 80%, at most about 81%, at most about 82%, at most about 83%, at most about 84%, at most about 85%, at most about 86%, at most about 87%, at most about 88%, at most about 89%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.5%, at most about 99.6%, at most about 99.7%, at most about 99.8%, at most about 99.9%, or any percentage in between the numbers. In some cases, the output indicative of cancer can be generated at a sensitivity of 100%.
[0217] In some cases, the method disclosed herein can detect one or more cancers. In some cases, the method disclosed herein can detect one or more cancers early on (e.g., multi-cancer early detection). In some cases, the method disclosed herein can detect one or more cancers ata specificity of at least about 60%, at least about 70%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect one or more cancers at a specificity of at most about 60%, at most about 70%, at most about 80%, at most about 81%, at most about 82%, at most about 83%, at most about 84%, at most about 85%, at most about 86%, at most about 87%, at most about 88%, at most about 89%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.5%, at most about 99.6%, at most about 99.7%, at most about 99.8%, at most about 99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect one or more cancers at a specificity of 100%.
[0218] In some cases, the method disclosed herein can detect one or more cancers at a sensitivity of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least 50%, at least about 60%, at least about 70%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect one or more cancers at a sensitivity of at most about 10%, at most about 20%, at most about 30%, at most about 40%, at most 50%, at most about 60%, at most about 70%, at most about 80%, at most about 81%, at most about 82%, at most about 83%, at most about 84%, at most about 85%, at most about 86%, at most about 87%, at most about 88%, at most about 89%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.5%, at most about 99.6%, at most about 99.7%, at most about 99.8%, at most about 99.9%, or any percentage in between the numbers. In some cases, the method disclosed herein can detect one or more cancers at a sensitivity of 100%.
[0219] In some cases, a sample from a biological sample (e.g., a nucleic acid sample, a cell- free nucleic acid sample, a tissue sample) can be taken from a subject. A nucleic acid sample can be derived from a sample obtained from a subject. In some cases, the biological sample comprises a blood sample and / or a plasma sample. In some cases, the biological sample comprises a tissue or cell sample. In some cases, the subject can be healthy. In some cases, the subject can be non-diseased. In some cases, the subject can be a non-cancer subject. In some cases, the subject can have or be suspected of having a disease or condition. In some cases, the disease or condition can be a cancer or a tumor. Non-limiting examples of cancer include breast cancer, bladder cancer, colorectal cancer, endometrial cancer, prostate cancer, renal cancer, pancreatic cancer, or lung cancer.
[0220] In some cases, the cell-free nucleic acids (e.g., cell-free DNA (cfDNA)) in a biological sample can be further treated to increase methylation level (e.g., in vitro enzymatic methylation). In some cases, cell-free nucleic acids (e.g., cfDNA) can be partially methylated. In some cases, cell-free nucleic acids (e.g., cfDNA) can be fully methylated. In some cases, the one or more sites that are amenable to methylation enrichment are determined by determining methylation levels in a fully methylated control sample. In some cases, the methylation control sample comprises cell-free DNA. In some cases, the methylation control sample comprises genomic DNA. In some cases, the genomic DNA can be subjected to shearing. In some cases, the fully methylated control sample comprises nucleic acids subjected to in vitro methylation (e.g., in vitro enzymatic methylation).
[0221] In some cases, increasing methylation level can be performed using in vitro methylation by an enzyme directed to alter methylation levels of nucleic acid fragments. In some embodiments, the enzyme can be CpG methyltransferase. In some cases, in vitro methylation can result in partial methylation. In some cases, in vitro methylation can result in full methylation. In some cases, in vitro methylation can result in hypermethylation. In some cases, a methylated DNA binder (e.g., a 5mC antibody) can be used to obtain or enrich for methylated or hypermethylated nucleic acids (e.g., fully methylated control sample). In some cases, cfMeDIP can be used to obtain and / or enrich for methylated or hypermethylated nucleic acids (e.g., fully methylated control sample). In some cases, methylation enrichment can be performed by subjecting the nucleic acid molecules (e.g., cell free nucleic acid molecules) to methylated DNA immunoprecipitation (MeDIP), cell-free methyl-CpG binding domain (cfMBD), methyl-CpG binding domain (MBD), methylation-dependent immunoprecipitation (MDIP), methylation-sensitive restriction enzyme (MSRE), TET- assisted pyridine borane sequencing (TAPS) with methylation specific PCR, bisulfiteconversion with methylation specific PCR, and / or methylation-specific hybrid capture, or other derivatives thereof. Hypermethylated nucleic acids can be further incubated and amplified. Hypermethylated nucleic acids can create a validated set of genomic regions which can more accurately distinguish between cancer and healthy samples in cfDNA. A hypermethylated section of DNA may be methylated at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or more across all bases. A hypermethylated section of DNA may be methylated at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, or more across all bases. A hypermethylated section of DNA may be tightly packed, resulting in a silenced gene. In some cases, assaying methylation levels of a plurality of regions comprises sequencing the cell-free nucleic acids, or derivatives thereof. In some cases, sequencing can be bisulfite sequencing. In some cases, sequencing does not comprise bisulfite sequencing. In some cases, sequencing comprises targeted sequencing. In some cases, sequencing comprises using a plurality of capture probes. In some cases, the plurality of capture probes comprises probes that are homologous or complementary to the plurality of regions. In some cases, the plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation. In some embodiments, the plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation. In some cases, sequencing generates sequencing reads corresponding to the plurality of regions. In some cases, assaying comprises counting a number of sequencing reads corresponding to a region of a plurality of regions.
[0222] In some cases, assaying methylation levels of a plurality of regions comprises preparing cell-free nucleic acids (e.g., cfDNA) to undergo Cell-free Methylated DNA ImmunoPrecipitation sequencing (cfMeDIP-seq), as illustrated in FIG. 1. For example, cell- free nucleic acids (e.g., cfDNA) can be supplemented with spike-in control DNA, undergo end-pairing, A-tailing, adapter ligation, and / or other preparation thereof to permit sequencing. In some cases, a spike-in control DNA can not be supplemented. In some cases, a first plurality of nucleic acid molecules (e.g., comprising nucleic acid molecules, such as cfDNA, from a biological sample of a subject) may be combined (e.g., mixed) with a second plurality of nucleic acid molecules (e.g., wherein the second plurality of nucleic acid molecules can not be from the subject from whom the biological sample was taken), for instance, as shown in FIG. 1. In some cases, the second plurality of nucleic acid molecules comprises supplemental processed nucleic acid (e.g., supplemental processed DNA), (e.g., comprisinglambda DNA). In some cases, each of the second plurality of nucleic acid molecules does not align to a human genome. In some cases, cell-free nucleic acids (e.g., cfDNA) can further undergo enrichment of methylated nucleic acids. For example, cell-free nucleic acids can undergo immunoprecipitation by pulling down the hypermethylated regions with a binder disclosed herein (e.g., protein comprises a methyl-CpG domain). In some cases, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule. In some cases, a binder can be a binder selective for a methylated region of nucleic acid molecules (e.g., a methylcytosine binder (MBD), such as an MBD-Fc fusion protein). In some cases, a binder may be specific to one or more methylated nucleotide species (e.g., 5-methylcytosine (5mC)), for instance, as shown in FIG. 1. Filler DNAs as disclosed herein can also be added to the cell-free nucleic acids (e.g., cfDNA). In some cases, filler DNAs may not be added. In some cases, enriching of methylated nucleic acids can comprise enriching for one or more methylated regions that correspond to one or more control genomic regions. For example, the one or more reference genomic regions can have one or more genomic sequences that are complementary to one or more sequences in the proto-DMRs. The one or more control genomic regions can comprise one or more hypomethylated regions, regions of nonmethylation, or regions that are amenable to methylation enrichment, or any combinations thereof. In some cases, methylated regions of nucleic acid molecules may be purified (e.g., after library creation) to yield a plurality of purified nucleic acid molecules, for example, prior to or as part of a process of determining or identifying a sequence of all or a portion of the methylated nucleic acid molecule population. In some cases, all, or a portion of the plurality of purified nucleic acid molecules may be amplified (e.g., via polymerase chain reaction), for instance, prior to or as part of a process of determining or identifying a sequence of all or a portion of the methylated nucleic acid molecule population. In some cases, a population of amplified nucleic acid molecules or a derivative thereof (e.g., comprising amplicons of all or a portion of the plurality of purified nucleic acid molecules) may be subjected to sequencing (e.g., for the determination and / or identification of a sequence of the nucleic acid molecules). In some cases, cell-free nucleic acids (e.g., cfDNA) with hypermethylated regions enriched can undergo sequencing that does not comprise bisulfite sequencing.
[0223] In some cases, determining methylation levels in the fully methylated control sample (e.g., nucleic acids subjected to in vitro methylation enrichment or in vitro enzymatic methylation), or derivatives thereof comprises sequencing. In some cases, sequencing can bebisulfite sequencing. In some cases, sequencing can be bisulfite sequencing paired with methylation specific PCR. In some cases, sequencing does not comprise bisulfite sequencing. In some cases, sequencing comprises targeted sequencing.
[0224] In some cases, determining methylation levels in the fully methylated control sample (e.g., nucleic acids subjected to in vitro methylation or in vitro enzymatic methylation), or derivatives thereof comprises preparing methylation control sample (e.g., nucleic acids subjected to in vitro methylation enrichment) to undergo cfMeDIP-seq, as illustrated in FIG. 1. For example, methylation control sample (e.g., nucleic acids subjected to in vitro methylation) can be supplemented with spike-in control DNA, undergo end-pairing, A- tailing, adapter ligation, and / or other preparation thereof to permit sequencing. In some cases, the methylation control sample may not be supplemented with a spike-in control DNA. In some cases, a first plurality of nucleic acid molecules (e.g., comprising nucleic acid molecules, such as cfDNA, from a biological sample of a subject) may be combined (e.g., mixed) with a second plurality of nucleic acid molecules (e.g., wherein the second plurality of nucleic acid molecules may not be from the subject from whom the biological sample was taken), for instance, as shown in FIG. 1. In some cases, the second plurality of nucleic acid molecules comprises supplemental processed nucleic acid (e.g., comprising lambda DNA). In some cases, each of the second plurality of nucleic acid molecules does not align to a human genome. In some cases, methylation control sample (e.g., nucleic acids subjected to in vitro methylation) can further undergo enrichment of methylated nucleic acids. For example, methylation control sample (e.g., nucleic acids subjected to in vitro methylation) can undergo immunoprecipitation by pulling down the hypermethylated regions with a binder disclosed herein (e.g., protein comprises a methyl-CpG domain). In some cases, the binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of the cell free nucleic acid molecule or a sheared genomic nucleic acid molecule. In some cases, a binder can be a binder selective for a methylated region of nucleic acid molecules (e.g., a methylcytosine binder (MBD), such as an MBD-Fc fusion protein). In some cases, a binder may be specific to one or more methylated nucleotide species (e.g., 5-methylcytosine (5mC)), for instance, as shown in FIG. 1. Filler DNAs as disclosed herein can also be added to the methylation control sample. In some cases, filler DNAs may not be added. In some cases, methylated regions of the methylation control sample (e.g., nucleic acids subjected to in vitro methylation) may be purified (e.g., after library creation) to yield a plurality of purified nucleic acid molecules, for example, prior to or as part of a process of determining or identifying a sequence of all or a portion of the methylated nucleic acid molecule population.In some cases, all, or a portion of the plurality of purified nucleic acid molecules may be amplified (e.g., via polymerase chain reaction), for instance, prior to or as part of a process of determining or identifying a sequence of all or a portion of the methylated nucleic acid molecule population. In some cases, a population of amplified nucleic acid molecules or a derivative thereof (e.g., comprising amplicons of all or a portion of the plurality of purified nucleic acid molecules) may be subjected to sequencing (e.g., for the determination and / or identification of a sequence of the nucleic acid molecules). In some cases, cell-free nucleic acids (e.g., cfDNA) with hypermethylated regions enriched can undergo sequencing that does not comprise bisulfite sequencing.
[0225] In some cases, a methylation level of a particular nucleic acid fragments (e.g., DNA fragments) may be considered to reach the threshold methylation level when a binder with a sufficient specificity for methylated cytosines can be able to bind to the particular nucleic acid fragments either with or without using filler DNA as described here. In some cases, a methylation level of particular nucleic acid fragments (e.g., DNA fragments, plurality of regions described herein) may be considered to be below the threshold methylation level when a binder with a sufficient specificity for methylated cytosines may not be able to bind to the particular nucleic acid fragments either with or without using filler DNA, as described here. In some cases, depletion of a plurality of nucleic acid molecules (e.g., in the creation of a depleted sequencing library and / or the determination of a presence or sequence of a nucleic acid molecule) results in a remainder population of the plurality of nucleic acid molecules, wherein the remainder of the plurality of nucleic acid molecules comprises (or, in some cases, consists of) nucleic acid molecules having a methylation level below the threshold methylation level (e.g., wherein the remainder population can be hypomethylated / unmethylated relative to one or more nucleic acid molecules removed from the plurality of nucleic acid molecules during depletion). A methylation level may be calculated as a percentage of hypermethylated nucleic acid fragments compared to all the nucleic acid fragments contained in a sample. A methylation level may be calculated as a percentage of hypermethylated nucleic acid fragments compared to nucleic acid fragments that are in regions that are amenable to methylation enrichment that are contained in a sample. In some cases, a threshold methylation level can be from 0.1% to 1%, 1% to 5%, 5% to 10%, 10% to 15%, 15% to 20%, 20% to 25%, 25% to 30%, 30% to 35%, 35% to 40%, 40% to 45%, 45% to 50%, 50% to 55%, 55% to 60%, 65% to 70%, 70% to 75%, 75% to 80%, 80% to 85%, 85% to 90%, 95% to 100%, at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at most 1%, at most 5%, at most 10%, at most 15%, at most 20%, at most 25%, at most 30%, at most 35%, at most 40%, at most 45%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, or at most 100%.
[0226] In some cases, processing methylated levels of the plurality of regions to identify a methylation of background comprises counting a number of sequencing reads for each region of the plurality of regions. In some cases, processing further comprises generating an average of sequencing read counts for all regions of the plurality of regions thereby generating the methylation background. In some cases, the methylation background corresponds to the total number of reads in a region. For example, the methylation background may correspond to the total number of reads in a region comprising an anti-DMR.
[0227] In some cases, processing methylation levels for a subset of the plurality of regions, wherein the subset comprises DMRs (e.g., DMR subsets), comprises counting a number of sequencing reads for each region of the subset of plurality of regions. In some cases, processing further comprises generating an average of sequencing read counts for all regions of the subset of the plurality of regions thereby generating the DMR specific methylation level. Sequences reads may be filtered for particular characteristics, (e.g., length, or number of CpGs), and reads that are filter out may be ignored or omitted for the read counts. The average of sequencing read counts may be a weighted average. In some cases, the DMRs (e.g., DMR subsets) are identified by differential methylation analysis of test sample and control sample. In some cases, the test sample can be derived from a subject with cancer. In some cases, the test sample can be derived from a subject without cancer. In some cases, the control sample can be derived from a subject with cancer. In some cases, the control sample can be derived from a subject without cancer. In some cases, the DMRs (e.g., DMR subsets) comprises one or more regions that exhibits hypermethylation in the test sample compared to the control sample. In some cases, the DMRs (e.g., DMR subsets) comprise one or more regions that comprises DMRs specific to a particular cancer type. In some cases, the DMRs (e.g., DMR subset) comprises one or more regions that comprises DMRs that are not specific to a particular cancer type. In some cases, the DMRs (e.g., DMR subset) comprises one or more regions that comprises DMRs that are specific to a subject. In some cases, the DMRs (e.g., DMR subset) comprises one or more regions that comprises DMRs that are not specific to a subject.
[0228] In some cases, processing methylation levels for a subset of the plurality of regions, wherein the subset comprises DMRs (e.g., DMR subsets), comprises counting a number of sequencing reads of nucleic acid regions with certain characteristics. For example, a count for sequencing reads corresponding to length (e.g., fragment length). For example, sequencing reads for fragments or nucleic acids that are <150 bp can be identified and can be counted. For example, the lengths may be < 170 bp, < 165 bp, < 160 bp, < 155 bp, < 150 bp, < 145 bp, < 140 bp, < 135 bp, < 130 bp, < 125 bp, < 120 bp, < 115 bp, < 110 bp, < 105 bp, or < 100 bp. A count for sequencing reads corresponding amounts to methylation, such as the number of CpG motifs in a region may be performed. For example, sequence read for regions of at least 5 CpGs may be counted. For example, sequence read for regions of at least 1 CpG, 2 CpG, 3 CpG, 4 CpG, 5 CpG, 6 CpG, 7 CpG, 8 CpG, 9 CpG, 10 CpG, 11 CpG, 12 CpG, 13 CpG, 14 CpG, 15 CpG, 16 CpG, 17 CpG, 18 CpG, 19 CpG, 20 CpG, or more, may be counted. The sequences read may be counted for a methylation amount or state and a length. For example, the sequencing reads with a methylation amount of 5 CpG and of less that 150 bp may be counted.
[0229] As described throughout this disclosure, levels of methylations may be normalized against a background. This normalization may reduce or eliminate noise, or otherwise improve the sensitivity. The normalization may also remove or reduce non relevant signals. For example, the levels of methylations may be determined for a set of DMRs. This set of DMRs may be normalized against a background of the methylation level of the plurality of regions comprising (i) one or more sites that are amenable to methylation enrichment and (ii) methylation at below a threshold in a non-diseased control (e.g., anti-DMRs). In another example, the levels of methylations may be determined for a set of DMRs by counting the number of reads that have certain characteristics (e.g., 5 CpG or more and / or a length of < 150 bp) in those DMRs. For example, the levels of methylations may be determined for a set of DMRs by counting the number of reads that have 2 CpG or more and / or a length of < 150 bp in those DMRs. This set of sequencing reads may be normalized against a background of the total number of sequencing read counts for a set of regions that comprise anti-DMRs .
[0230] In some cases, generating the normalized DMR methylation level comprises dividing an average number of sequencing reads associated with the subset by an average number of sequencing reads associated with the plurality of regions. The average may be a weighted average. In some cases, generating the normalized DMR methylation level can comprise one or more Baysian inference approaches and / or machine learning classifiers. In some cases, the method of nucleic acids processing further comprises identifying a subject as having adisease, based on the normalized DMR methylation level. In some cases, identifying a subject as having a disease comprises comparing the normalized DMR methylation level against a control level. In some cases, the control level corresponds to an expected value for a non- cancerous sample. In some cases, the control level corresponds to an expected value for a cancerous sample. In some cases, the disease or condition can be a cancer or a tumor. Nonlimiting examples of cancer include breast cancer, bladder cancer, colorectal cancer, endometrial cancer, prostate cancer, renal cancer, pancreatic cancer, or lung cancer. In some cases, the cancer can be a late-stage cancer. In some cases, the cancer can be an early-stage cancer.
[0231] In some cases, the methylation levels of a plurality of regions disclosed herein (e.g., regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment and (ii) methylation at below a threshold in a healthy control) can be used to identify a background region that can be used for normalization of DMR methylation level to distinguish between healthy and cancerous cell-free nucleic acid samples. A background region can comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, or more genomic windows or window regions. A genomic window or window region can have a length of at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 550, at least about 600, at least about 650, at least about 700, at least about 750, at least about 800, at least about 850, at least about 900, at least about 950, at least about 1,000, at least about2,000, at least about 3,000, at least about 4,000, at least about 5,000, at least about 6,000, at least about 7,000, at least about 8,000, at least about 9,000, at least about 10,000, at least about 20,000, at least about 30,000, at least about 40,000, at least about 50,000, or more base pairs (bp). In some cases, a genomic window can be about 300 bp in length. In some cases, the genomic windows are adjacent to one another in the genome. Alternatively, or in addition to, the genomic windows can be non-adjacent regions on the genome. In some cases, the genomic windows can be of dynamic length.
[0232] In some cases, identified DMR specific methylation level can be normalized by dividing the DMRs to the methylation background (e.g., methylation levels of a plurality of regions disclosed herein, e.g., regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment and (ii) methylation at below a threshold in a healthy control). In some cases, a background region can be used to identify DMRs from healthy and / or cancerous cell-free nucleic acid samples. In some cases, the DMRs identified by using background region comprises a subset of the DMRs identified by using genome-wide background. In some cases, the DMRs identified by using a background region identify more DMRs as compared to using genome-wide background.
[0233] In some cases, the use of a background region (e.g., methylation levels of a plurality of regions disclosed herein, e.g., regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment and (ii) methylation at below a threshold in a healthy control) can improve signal-to-noise (SNR) ratio. In some cases, any one of the methods disclosed herein comprises a reduction in a noise level compared to a noise level of a corresponding sample that has a normalized DMR methylation level generated by normalizing a DMR specific level against a background derived from a whole genome. In some cases, any one of the methods disclosed herein comprises a reduction in noise level compared to a noise level of a corresponding sample that has a normalized DMR methylation level generated by normalizing a DMR specific level against a background derived from all genomic regions that are amenable to methylation enrichment. In some cases, the use of a background region can identify more than at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 550, at least about 600, at least about 650, at least about 700, at least about 750, at least about 800, at least about 850, at least about 900, at least about 1000, or more regions (e.g., DMRs, such as ctDNA specific DMRs, anti-DMRs, or proto- DMRs) compared to using genome-wide background or background derived from all genomic regions that are amenable to methylation enrichment. In some cases, the use of a background region can identify more than at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7- fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least 45-fold, at least about 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500-fold,or more regions (e.g., DMRs, such as ctDNA specific DMRs, anti-DMRs, or proto-DMRs) compared to using genome-wide background or background derived from all genomic regions that are amenable to methylation enrichment. In some cases, the use of a background region can identify more cell-free nucleic acids (e.g., cfDNAs) derived from tumor or cancer cells (e.g., ctDNAs) specific methylation level more than at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about400, at least about 450, at least about 500, at least about 550, at least about 600, at least about650, at least about 700, at least about 750, at least about 800, at least about 850, at least about900, at least about 1000, or more regions (e.g., DMRs, such as ctDNA specific DMRs, anti- DMRs, or proto-DMRs) compared to using genome-wide background or background derived from all genomic regions that are amenable to methylation enrichment. In some cases, the use of a background region can identify more than at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35- fold, at least about 40-fold, at least 45-fold, at least about 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500- fold, or more DMRs compared to using genome-wide background or background derived from all genomic regions that are amenable to methylation enrichment.
[0234] In some cases, the use of a background region reduce the run time to identify DMRs (e.g., ctDNA specific DMRs) by at least about 5 minutes, at least about 10 minutes, at least about 15 minutes, at least about 20 minutes, at least about 25 minutes, at least about 30 minutes, at least about 35 minutes, at least about 40 minutes, at least about 45 minutes, at least about 50 minutes, at least about 55 minutes, at least about 1 hour, at least about 1.5 hours, at least about 2 hours, at least about 2.5 hours, at least about 3 hours, at least about 3.5 hours, at least about 4 hours, at least about 4.5 hours, at least about 5 hours, at least about 5.5 hours, at least about 6 hours, at least about 6.5 hours, at least about 7 hours, at least about 7.5 hours, at least about 8 hours, at least about 8.5 hours, at least about 9 hours, at least about 9.5 hours, at least about 10 hours, or more as compared to using genome-wide background or background derived from all genomic regions that are amenable to methylation enrichment to identify DMRs (e.g., ctDNA specific DMRs).
[0235] In some cases, the use of a background region reduce the run time to identify DMRs (e.g., ctDNA specific DMRs) by about 1-fold, at least about 2-fold, at least about 3-fold, atleast about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40- fold, at least 45-fold, at least about 50-fold, or more as compared to using genome-wide background or background derived from all genomic regions that are amenable to methylation enrichment to identify DMRs (e.g., ctDNA specific DMRs) to identify DMRs (e.g., ctDNA specific DMRs).
[0236] In some cases, the background region (e.g., methylation levels of a plurality of regions disclosed herein, e.g., regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment and (ii) methylation at below a threshold in a healthy control) can be used to further generate a panel of DMRs for assessing hypermethylated regions in cancer. In some cases, the use of the background region can be not limited to any specific type of cancer (e.g., pan cancer). For example, since the panel of the background region (e.g., methylation levels of a plurality of regions disclosed herein, e.g., regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment and (ii) methylation at below a threshold in a healthy control) can be not specific to any cancer, but rather the absence of cancer signal (e.g., methylation at below a threshold in a healthy control), the same panel can be used to assess samples from any cancer types to identify DMRs within the plurality of regions described herein.Supplemental Processed DNA (filler DNA)
[0237] In some cases, supplemental processed nucleic acid (e.g., filler DNA) may be added to a first plurality of nucleic acids (e.g., a plurality of nucleic acids from a biological sample, which may comprise cfDNA from healthy tissue and / or cfDNA from tumor tissue, such as ctDNA), for instance as shown in FIG. 1. In some cases, the supplemental processed nucleic acid can be DNA and / or RNA. In some cases, addition of supplemental processed nucleic acid (e.g., a second plurality of nucleic acid molecules) to a first plurality of nucleic acid molecules can increase the specificity and / or sensitivity of a method, or system described herein, for instance, with respect to the detection and / or identification of a nucleic acid sequence of the first plurality of nucleic acid molecules. In some cases, addition of supplemental processed nucleic acid (e.g., a second plurality of nucleic acid molecules) to a first plurality of nucleic acid molecules may increase the rate of depletion of a methylated region of a nucleic acid sequence, e.g., during the practice of some embodiments of methods and systems described herein. In some cases, addition of supplemental processed nucleic acid (e.g., a second plurality of nucleic acid molecules) to a first plurality of nucleic acidmolecules (e.g., comprising cfDNA of a biological sample) may increase a binder’s selectivity for one or more (e.g., a plurality of) methylated regions of the first plurality of nucleic acid molecules. In some cases, supplemental processed nucleic acid (e.g., the second plurality of nucleic acid molecules) may be added to the first plurality of nucleic acid molecules in an amount sufficient to bring the combined mixture of nucleic acid molecules to a predetermined total mass. In some cases, a predetermined total mass for use in a method or system described herein can be from 20 ng to 30 ng, from 30 ng to 40 ng, from 40 ng to 50 ng, from 50 ng to 60 ng, from 60 ng to 70 ng, from 70 ng to 80 ng, from 80 ng to 90 ng, from 90 ng to 100 ng, from 100 ng to 110 ng, from 110 ng to 120 ng, from 120 ng to 130 ng, from 130 ng to 140 ng, from 140 ng to 150 ng, from 150 ng to 160 ng, from 160 ng to 170 ng, from 170 ng to 180 ng, from 180 ng to 190 ng, from 190 ng to 200 ng, greater than 200 ng, or less than 20 ng. In some cases, an amount of supplemental processed nucleic acid from 1 ng to 5 ng, from 5 ng to 10 ng, from 10 ng to 20 ng, from 20 ng to 30 ng, from 30 ng to 40 ng, from 40 ng to 50 ng, from 50 ng to 60 ng, from 60 ng to 70 ng, from 70 ng to 80 ng, from 80 ng to 90 ng, from 90 ng to 100 ng, from 100 ng to 110 ng, from 110 ng to 120 ng, from 120 ng to 130 ng, from 130 ng to 140 ng, from 140 ng to 150 ng, from 150 ng to 160 ng, from 160 ng to 170 ng, from 170 ng to 180 ng, from 180 ng to 190 ng, from 190 ng to 200 ng, greater than 200 ng, less than 20 ng, less than 10 ng, or less than 5 ng can be added to a first plurality of nucleic acid molecules (e.g., to bring the total mixture of the supplemental processed nucleic acid and the first plurality of nucleic acid molecules to the predetermined total mass). In some embodiments, the present disclosure comprises methods and systems for filling in the sample with an amount of supplemental processed nucleic acid (e.g., filler DNA) to generate a mixture sample, wherein the mixture sample comprises at least about 50ng, 55ng, 60ng, 65ng, 70ng, 75ng, 80ng, 85ng, 90ng, 95ng, lOOng, 120ng, 140ng, 160ng, 180ng, 200ng, or any amount in between the numbers of the total amount of the nucleic acid mixture. In some embodiments, the supplemental processed DNA comprises at least about 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% methylated supplemental processed nucleic acid with remainder being unmethylated supplemental processed nucleic acid, in some cases between 5% and 50%, or between 10%-40%, or between 15%-30% methylated supplemental processed DNA. In some embodiments, the mixture sample comprise an amount of supplemental processed DNA from about 20 ng to about 100 ng, about 30 ng to about 100 ng, or about 50 ng to about 100 ng. In some embodiments, the cell-free DNA from the sample and the first amount of supplemental processed DNA together comprises at least 50 ng of total nucleic acid. In some embodiments,the cfDNA from the sample and the first amount of supplemental processed DNA together comprises at least 100 ng of total nucleic acid. In some cases, the supplemental processed DNA may not be needed and / or required for any one of the methods disclosed herein.
[0238] In some cases, supplemental processed nucleic acid may be produced by fragmentation (e.g., via sonication). In some embodiments, the supplemental processed nucleic acid may be about 50 bp to about 800 bp long, about 100 bp to about 600 bp long, or about 200 bp to about 600 bp long. In some embodiments, the supplemental processed nucleic acid can be double stranded. The supplemental processed nucleic acid may be double stranded DNA. For example, the supplemental processed nucleic acid may be junk nucleic acid. The supplemental processed nucleic acid may also be endogenous or exogenous nucleic acid. For example, the supplemental processed DNA can be non-human nucleic acid, such as DNA. As used herein, “ DNA” refers to Enterobacteria phage DNA. In some embodiments, the supplemental processed nucleic acid has no alignment to human nucleic acid. In some embodiments, supplemental nucleic acid can be hypermethylated.Samples
[0239] A sample can be any biological sample isolated from a subject. For example, a sample may comprise, without limitation, bodily fluid, whole blood, platelets, serum, plasma, stool, white blood cells or leukocytes, endothelial cells, tissue biopsies, synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, urine, fluid from nasal brushings, fluid from a pap smear, or any other bodily fluids. A bodily fluid may include saliva, blood, or serum. A sample may also be a tumor sample, which may be obtained from a subject by various approaches, including, but not limited to, venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage, scraping, surgical incision, or intervention or other approaches. A sample may be a cell-free sample (e.g., substantially free of cells). DNA samples may be denatured, for example, using sufficient heat.
[0240] The sample may be taken from a subject with a disease or disorder. The sample may be taken from a subject suspected of having a disease or a disorder. The sample can be taken from a subject suspected of having minimal residual disease. The sample can be taken from a subject suspected of with minimal residual disease. In some cases, the sample may be obtained before and / or after treatment of a subject with a disease or disorder. Samples may be obtained from a subject during a treatment or a treatment regime. Multiple samples may be obtained from a subject to monitor the effects of the treatment over time. The disease ordisorder may be a cancer. A cancer can be a late-stage or an early-stage cancer. Specific examples of cancer types include suitable for detection with the methods according to the disclosure include acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancers, AIDS-related lymphoma, anal cancer, appendix cancer, astrocytomas, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancers, brain tumors, such as cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumors, visual pathway and hypothalamic glioma, breast cancer, bronchial adenomas, Burkitt lymphoma, carcinoma of unknown primary origin, central nervous system lymphoma, cerebellar astrocytoma, cervical cancer, childhood cancers, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative disorders, colon cancer, cutaneous T-cell lymphoma, desmoplastic small round cell tumor, endometrial cancer, ependymoma, esophageal cancer, Ewing's sarcoma, germ cell tumors, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, gliomas, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular (liver) cancer, Hodgkin lymphoma, Hypopharyngeal cancer, intraocular melanoma, islet cell carcinoma, Kaposi sarcoma, kidney cancer, laryngeal cancer, lip and oral cavity cancer, liposarcoma, liver cancer, lung cancers, such as non-small cell and small cell lung cancer, lymphomas, leukemias, macroglobulinemia, malignant fibrous histiocytoma of bone / osteosarcoma, medulloblastoma, melanomas, mesothelioma, metastatic squamous neck cancer with occult primary, mouth cancer, multiple endocrine neoplasia syndrome, myelodysplastic syndromes, myeloid leukemia, nasal cavity and paranasal sinus cancer, nasopharyngeal carcinoma, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, oral cancer, oropharyngeal cancer, osteosarcoma / malignant fibrous histiocytoma of bone, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, pancreatic cancer, pancreatic cancer islet cell, paranasal sinus and nasal cavity cancer, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal astrocytoma, pineal germinoma, pituitary adenoma, pleuropulmonary blastoma, plasma cell neoplasia, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell carcinoma, renal pelvis and ureter transitional cell cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcomas, skin cancers, skin carcinoma merkel cell, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, stomach cancer, T-cell lymphoma, throat cancer, thymoma, thymic carcinoma, thyroid cancer, trophoblastic tumor (gestational), cancers of unknown primary site, urethral cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrommacroglobulinemia, and Wilm’s tumor. In some cases, the cancer can be head and neck squamous cell carcinoma. In some cases, the sample can be taken from a subject that may have or suspect to have minimal residual disease.
[0241] The sample may be taken from a healthy individual. The sample may be taken from an individual that may not be suffering from a disease. For example, the sample may be taken from an individual that may not be suffering from cancer. In some cases, samples may be taken longitudinally from the same individual. In some cases, samples acquired longitudinally may be analyzed with the goal of monitoring individual health and early detection of health issues. For example, the sample may be taken at a first time to analyze for a set of markers (e.g., DMRs). The sample may be used to initial determine a set of markers that are specific to the subject. A second sample may be taken from the sample at a later time to monitor the markers. For example, a subject with a cancer may initially be analyzed. A second sample may be analyzed after the initiation of a therapy regimen. The monitoring may allow for the efficacy of a therapy regimen or treatment to be analyzed. For example, a DMR associated with a subject’s cancer may be analyzed and observed to be absent from a subject. In some embodiments, the sample may be collected at a home setting or at a point-of-care setting and subsequently transported by a mail delivery, courier delivery, or other transport method prior to analysis. For example, a home user may collect a blood sample through a finger prick, which blood sample may be subsequently transported by mail delivery prior to analysis. In some cases, samples acquired longitudinally may be used to monitor response to stimuli expected to impact healthy, athletic performance, or cognitive performance. Nonlimiting examples include response to medication, dieting, or an exercise regimen.
[0242] In some embodiments, the present disclosure provides a system, method, or kit that includes or uses one or more biological samples. The one or more samples used herein may comprise any substance containing or presumed to contain nucleic acids. A sample may include a biological sample obtained from a subject. In some embodiments, a biological sample can be a liquid sample.
[0243] In some embodiments, the sample comprises less than about 100 ng, 90 ng, 80 ng, 75 ng, 70ng, 60 ng, 50 ng, 40 ng, 30 ng, 20 ng, 10 ng, 5 ng, or any amount in between the numbers of cell-free nucleic acid molecules. Further, in some embodiments, the sample comprises less than about 1 pg, less than about 5 pg, less than about 10 pg, less than about 20 pg, less than about 30 pg, less than about 40 pg, less than about 50 pg, less than about 100 pg, less than about 200 pg, less than about 500 pg, less than about 1 ng, less than about 5 ng, less than about 10 ng, less than about 20 ng, less than about 30 ng, less than about 40 ng, less thanabout 50 ng, less than about 100 ng, less than about 200 ng, less than about 500 ng, less than about 1000 ng, or any amount in between the numbers of cell-free nucleic acid molecules.
[0244] In some cases, creation or provision of a plurality of nucleic acid molecules from a biological sample can comprise performing one or more of end-repair, A-tailing, and / or adapter ligation on the plurality of nucleic acid molecules (e.g., after purification from the biological sample).
[0245] In some embodiments, a sample may be taken at a first time point and sequenced, and then another sample may be taken at a subsequent time point and sequenced. Such methods may be used, for example, for longitudinal monitoring purposes to track the development or progression of a disease. In some embodiments, the progression of a disease may be tracked before treatment, after treatment, or during the course of treatment, to determine the treatment’s effectiveness. For example, a method as described herein may be performed on a subject prior to, and after, a medical treatment to measure the disease’s progression or regression in response to the medical treatment.
[0246] After obtaining a sample from the subject, the sample may be processed to generate datasets indicative of a disease or disorder of the subject. For example, a presence, absence, or quantitative assessment of cell-free nucleic acid molecules (e.g., ctDNA molecules) of the sample at a panel of cancer-associated genomic loci or microbiome-associated loci may be indicative of a cancer of the subject. Processing the sample obtained from the subject may comprise (i) subjecting the sample to conditions that are sufficient to isolate, enrich, or extract a plurality of cell-free nucleic acid molecules, and (ii) assaying the plurality of cell- free nucleic acid molecules to generate the dataset (e.g., nucleic acid sequences). In some embodiments, a plurality of cell-free nucleic acid molecules can be extracted from the sample and subjected to sequencing to generate a plurality of sequencing reads.
[0247] In some embodiments, the cell- free nucleic acid molecules may comprise cell-free ribonucleic acid (cfRNA) or cell-free deoxyribonucleic acid (cfDNA). The cell-free nucleic acid molecules (e.g., cfRNA or cfDNA) may be extracted from the sample by a variety of methods. The cell-free nucleic acid molecule may be enriched by a plurality of probes configured to enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to a panel of cancer-associated genomic loci. The probes may have sequence complementarity with nucleic acid sequences from one or more of the genomic loci of the panel of cancer- associated genomic loci. The panel of cancer-associated genomic loci may comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100, at least 200, at least 300, at least 400, at least 500, at least 600. at least 700. at least 800. at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000. at least 7000. at least 8000. at least 9000, at least 10000, at least 20000, 30000, at least 40000, at least 50000, at least 60000. at least 70000. at least 80000. at least 90000, at least 100000, or more distinct cancer-associated genomic loci. The probes may be nucleic acid molecules (e.g., RNA or DNA) having sequence complementarity with nucleic acid sequences (e.g., RNA or DNA) of the one or more genomic loci (e.g., cancer-associated genomic loci). These nucleic acid molecules may be primers or enrichment sequences. The assaying of the sample using probes that are selective for the one or more genomic loci (e.g., cancer-associated genomic loci or microbiome- associated loci) may comprise use of array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing).
[0248] Certain methods of capturing cell-free methylated DNA are described in WO 2017 / 190215 and WO 2019 / 010564, both of which are incorporated by reference in their entireties and for all purposes.Nucleic Acid Molecule Sequencing
[0249] Various assays may be used in methods of the present disclosure, such as library preparation (which may include polymerase chain reaction (PCR)) followed by sequencing (e.g., next-generation sequencing, Sanger sequencing, etc.). Next-generation sequencing (NGS) techniques, also referred to as high-throughput sequencing, may include various sequencing technologies including: Illumina (Solexa) sequencing, Roche 454 sequencing, Ion torrent: Proton / PGM sequencing, SOLiD sequencing, long reads sequencing (Oxford Nanopore and Pactbio). NGS allow for the sequencing of DNA and RNA much more quickly and cheaply than the Sanger sequencing. In some embodiments, the sequencing can be optimized for short read sequencing. In some embodiments, the sequencing assays do not require amplification.
[0250] Sequencing libraries that are hypermethylated may improve the specificity, the sensitivity, and / or the efficiency of methods and systems for processing nucleic acids. For example, hypermethylated sequencing libraries may improve the specificity, the sensitivity, and / or the efficiency of assays for determining the presence and / or sequence identity of a nucleic acid sequence. A hypermethylated sequencing library may comprise a plurality ofnucleic acids and / or fragments thereof. In some cases, a hypermethylated sequencing library may comprise a plurality of nucleic acid molecules (e.g., a population of nucleic acids and / or fragments thereof). The plurality of nucleic acid molecules may comprise all or a portion of a first plurality of nucleic acid molecules, e.g., wherein the first plurality of nucleic acid molecules comprises one or more nucleic acid molecules that comprise a methylated nucleic acid residue and one or more nucleic acid molecules that does not comprise a methylated nucleic acid residue. In some cases, a methylated nucleic acid may comprise one or more methylated nucleic acid residues. For instance, a methylated nucleic acid may comprise one or more methylated cytosines (e.g., one or more 5-methylcytosines (5mC) and / or one or more 5-hydroxymethylcytosines (5hmC)).
[0251] A plurality of nucleic acid molecules (e.g., a plurality of nucleic acid molecules derived from a biological sample) may be hypermethylated and enriched by using a binder, e.g., as described herein, to form a hypermethylated sequencing library which can be used as a background as opposed to a whole-genome background for use in analysis of cfDNA. In some cases, DNA may be hypermethylated before use of a binder to create a sequencing library with a background. The background sequencing library may comprise a set of background genomic regions that are enriched by the binder.
[0252] The present disclosure provides methods and technologies for determining the sequence of nucleotide bases in one or more polynucleotides. The polynucleotides may be, for example, nucleic acid molecules such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single stranded DNA). Sequencing may be performed by various systems currently available, such as, without limitation, a sequencing system by Illumina®, Pacific Biosciences (PacBio®), Oxford Nanopore®, or Life Technologies (Ion Torrent®). Further, any sequencing methods that provide fragment length such as paired-end sequencing may be utilized. Alternatively or additionally, sequencing may be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real time PCR), or isothermal amplification. Such systems may provide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., human), as generated by the systems from a sample provided by the subject. In some examples, such systems provide sequencing reads (also “reads” herein). A read may include a string of nucleic acid bases corresponding to a sequence of a nucleic acid molecule that has been sequenced. In some situations, systems and methods provided herein may be used with proteomic information.
[0253] In some embodiments, the sequencing reads are obtained via a next-generation sequencing method or a next-next-generation sequencing method. In some embodiments, the sequencing methods comprise cfMeDIP sequencing, e.g., comprising operations or systems as described by Shen et al., (“Sensitive tumor detection and classification using plasma cell- free DNA methyl omes,” (2018) Nature), which is incorporated herein in its entirety. In some embodiments, sequencing can be performed using methyl-CpG-binding domain sequencing (MBD-seq). In some cases, MBD-seq can comprise capture (e.g., via a binder, such as an antibody specific to a species of methylated nucleotide) of double-stranded, methylated DNA fragments for sequencing of methylation-enriched DNA fragment libraries. In some embodiments, the sequencing methods comprise CAncer Personalized Profiling by deep Sequencing (CAPP-Seq), which can be a next-generation sequencing based method used to quantify circulating DNA in cancer (ctDNA). In some embodiments, the sequencing methods comprise chromatin immunoprecipitation sequencing (ChlP-Seq). This method may be generalized for any cancer type that may recurrent mutations and may detect one molecule of mutant DNA in 10,000 molecules of healthy DNA. In some embodiments, the sequencing comprises bisulfite sequencing. Alternatively, in some embodiments, the sequencing does not comprise bisulfite sequencing.
[0254] Sequencing may comprise targeted sequencing. For example, the sequencing reactions may comprise capture probes that are specific to regions of interest. The use of targeted sequencing may increase the amount of reads that are specific to regions that are informative (e.g., related to a DMR, or usable for distinguishing healthy subject vs a subject suffering from a disease). The capture probes may comprise one or more probes that are complementary or homologous to regions that comprises (i) one or more sites that are amenable to methylation enrichment and (ii) substantially no methylation in a healthy control. The capture probes may comprise one or more probes that are complementary or homologous to regions that have a known methylation state. The use of a target panel that enriches for regions comprising (i) one or more sites that are amenable to methylation enrichment and (ii) substantially no methylation in a healthy control may allow for increasing the sequencing depth and signal for areas that can be used with differential methylation analysis without restricting the sequences to a specific disease or a specific type of cancer. For example, the sequences generated from the targeted sequence may be pan-cancer or cancer type agnostic. The target sequencing may not comprise generating a custom panel for a specific subject or cancer. The use of panels that are not custom to a subject allows for an improved consistency in the sequencing and wet-lab protocols while still allowing fordownstream in silico analysis to be personalized or customized to a subject. In some cases the use of panels that are not customized to a subject (e.g., universal panel of genomic region) can reduce the need for high sequencing depth regardless of the size of the panels.
[0255] Sequencing can comprise analysis of the results of sequencing methods. In some embodiments, sequencing analysis can comprise using Model-based Analysis for ChlP-Seq (MACS) software. The MACS algorithm captures the influence of genome complexity to evaluate the significance of enriched regions. Sequencing analysis can comprise identifying broad peaks (broad peak calling) or identifying narrow peaks (narrow peak calling). In some cases, hypermethylated regions can be identified using narrow peak calling. Peak annotations that note regions of interest can be produced by the MACS algorithm by determining signals that differ significantly between two samples (e.g., between a sample and the background region). In some cases, hypermethylated regions can be identified using broad peak calling. In some cases, both narrow peak calling and broad peak calling may identify the same hypermethylated regions. In some cases, narrow peak calling and broad peak calling may identify different hypermethylated regions. Additional processing of peak annotations can merge regions of interest across multiple samples to result in higher resolution and more accurate results. More accurate results can comprise the inclusion of regions that have very few reads in samples, but which can be leveraged to differentiate between healthy and disease samples. In some cases, sequencing analysis can be illustrated using a gene heatmap. Alternatively or in addition to, sequencing analysis can be illustrated using a uniform manifold approximation and projection (UMAP) plot.
[0256] In some cases, a sample or portion thereof (e.g., a plurality of nucleic acids of a sample) may be subjected to library preparation before sequencing. In short, after end-repair and A-tailing, the samples are ligated to nucleic acid adapters and digested using enzymes.
[0257] In some embodiments, sequencing comprises modification of a nucleic acid molecule or fragment thereof, for example, by ligating a barcode, a unique molecular identifier (UMI), or another tag to the nucleic acid molecule or fragment thereof. Ligating a barcode, UMI, or tag to one end of a nucleic acid molecule or fragment thereof may facilitate analysis of the nucleic acid molecule or fragment thereof following sequencing. In some embodiments, a barcode is a unique barcode (e.g., a UMI). In some embodiments, a barcode can be nonunique, and barcode sequences may be used in connection with endogenous sequence information such as the start and stop sequences of a target nucleic acid (e.g., the target nucleic acid can be flanked by the barcode and the barcode sequences, in connection with the sequences at the beginning and end of the target nucleic acid, creates a uniquely taggedmolecule). A barcode, UMI, or tag may be a known sequence used to associate a polynucleotide or fragment thereof with an input or target nucleic acid molecule or fragment thereof. A barcode, UMI, or tag may comprise natural nucleotides or non-natural (e.g., modified) nucleotides (e.g., as described herein). A barcode sequence may be contained within an adapter sequence such that the barcode sequence may be contained within a sequencing read. A barcode sequence may comprise at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more nucleotides in length. In some cases, a barcode sequence may be of sufficient length and may be sufficiently different from another barcode sequence to allow the identification of a sample based on a barcode sequence with which it can be associated. A barcode sequence, or a combination of barcode sequences, may be used to tag and subsequently identify an “original” nucleic acid molecule or fragment thereof (e.g., a nucleic acid molecule or fragment thereof present in a sample from a subject). In some cases, a barcode sequence, or a combination of barcode sequences, can be used in conjunction with endogenous sequence information to identify an original nucleic acid molecule or fragment thereof. For example, a barcode sequence, or a combination of barcode sequences, may be used with endogenous sequences adjacent to a barcode, UMI, or tag (e.g., the beginning and end of the endogenous sequences).
[0258] As described herein, the prepared libraries may be combined with filler nucleic acids (e.g., filler DNAs) to minimize the effect of low abundance ctDNA in the prepared libraries and generate mixed samples. In some embodiments, when the disease and / or condition can be a locoregionally (non-metastatic) cancer, the amount of ctDNA can be low and may not be easily and accurately measured and quantified. The mixed samples are brought to at least about 50ng, 80ng, lOOng, 120ng, 150ng, or 200ng and are subjected to further enrichment.
[0259] Processing a nucleic acid molecule or fragment thereof may comprise performing nucleic acid amplification. For example, any type of nucleic acid amplification reaction may be used to amplify a target nucleic acid molecule or fragment thereof and generate an amplified product. Non-limiting examples of nucleic acid amplification methods include reverse transcription, primer extension, polymerase chain reaction (PCR), ligase chain reaction, asymmetric amplification, rolling circle amplification, and multiple displacement amplification (MDA). Examples of PCR include, but are not limited to, quantitative PCR, real-time PCR, digital PCR, emulsion PCR, hot start PCR, multiplex PCR, asymmetric PCR, nested PCR, and assembly PCR. Nucleic acid amplification may involve one or more reagents such as one or more primers, probes, polymerases, buffers, enzymes, anddeoxyribonucleotides. Nucleic acid amplification may be isothermal or may comprise thermal cycling, and / or with the length of the endogenous sequence.Binders
[0260] A binder may be used to deplete a population of nucleic acid molecules (e.g., a plurality of nucleic acid molecules derived from a biological sample). In some cases, a binder can be used to deplete a plurality of nucleic acid molecules of one or more nucleic acid molecules having a methylation level at or above a threshold methylation level (e.g., by binding to one or more methylated nucleotides of the one or more nucleic acid molecules). A binder may be used to enrich a population of nucleic acid molecules (e.g., a plurality of nucleic acids derived from a biological sample). In some cases, a binder can exhibit a reduced level of non-specific binding to non-methylated nucleotides or non-m ethylated genomic regions or sheared genomic nucleic acid molecule.
[0261] In some cases, a binder can be specific to one or more methylated nucleotide species (e.g., 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 4-methylcytosine (4mC), or 6-methyladenine (6mA)). In some cases, a binder can be selected from the group consisting of an anti-5-methylcytosine antibody or a derivative thereof, an anti-5- carboxylcytosine antibody or a derivative thereof, an anti-5-formylcytosine antibody or a derivative thereof, an anti-5-hydroxymethylcytosine antibody or a derivative thereof, an anti- 3 -methylcytosine antibody or a derivative thereof, and any combinations thereof. In some cases, the binder can be an anti-5-methylcytosine antibody or a derivative thereof. In some embodiments, the binder can be a protein comprising a Methyl-CpG-binding domain. One such protein may be MBD2 protein. As used herein, “Methyl-CpG-binding domain (MBD)” refers to certain domains of proteins and enzymes that can be approximately 70 residues long and binds to DNA that contains one or more symmetrically methylated CpGs. The MBD of MeCP2, MBD1, MBD2, MBD4 and BAZ2 mediates binding to DNA, and in cases of MeCP2, MBD1 and MBD2, preferentially to methylated CpG. Human proteins MECP2, MBD1, MBD2, MBD3, and MBD4 comprise a family of nuclear proteins related by the presence in each of a methyl-CpG-binding domain (MBD). Each of these proteins, with the exception of MBD3, can be capable of binding specifically to methylated DNA.
[0262] In other cases, the binder can be an antibody and capturing cell-free methylated DNA comprises immunoprecipitating the cell-free methylated DNA using the antibody. As used herein, “immunoprecipitation” refers a technique of precipitating an antigen (such as polypeptides and nucleotides) out of solution using an antibody that specifically binds to that particular antigen. This process may be used to isolate and concentrate a particular protein orDNA from a sample and requires that the antibody be coupled to a solid substrate at some point in the procedure. The solid substrate includes for example beads, such as magnetic beads. Other types of beads and solid substrates may be used.
[0263] For example, a 5-mC antibody (e.g., wherein the 5-mC antibody specifically binds to 5-methylcytosine) may be used as a binder. For the immunoprecipitation procedure, in some embodiments at least 0.05 pg of the antibody can be added to the sample; Alternatively, at least 0.16 pg of the antibody can be added to the sample. In some cases, 0.05 pg to 0.80 pg, 0.16 pg to 0.80 pg, 0.40 pg to 0.80 pg, 0.16 pg to 0.40 pg, 0.10 pg to 0.80 pg, 0.20 pg to 0.60 pg, 0.30 pg to 0.50 pg, or 0.40 pg to 0.50 pg of the antibody can be used. To confirm the immunoprecipitation reaction, in some cases, the method described herein further comprises adding a second amount of control DNA to the sample.Methylation Profile
[0264] The present disclosure provides methods and systems for producing a methylation profile of a subject that has a disease and / or condition or can be suspected of having such disease and / or condition, wherein the methylation profile may be used to determine whether the subject has the disease and / or condition or can be at risk of having the disease and / or condition. In some cases, a methylation profile can comprise analysis (e.g., comprising sequencing) of a plurality of nucleic acids (e.g., a plurality of nucleic acid molecules of a depleted sequencing library, as described herein). In some cases, a methylation profile can comprise detection of methylated nucleotides and / or quantification of methylated nucleotide counts. In some cases, a methylation profile can comprise determination of a methylated signal, e.g., in a population of nucleic acids of a depleted sequencing library, as described herein. In some cases, a methylation profile can be compared to a genome-wide background profile. In some cases, a methylation profile can be compared to a background profile created using hypermethylated cfDNA. In some cases, a methylation profile can be compared to a background profile comprising a plurality of regions, wherein the plurality of regions have been identified as regions of a genome that comprises (i) one or more sites that are amenable to methylation enrichment, (ii) methylation at below a threshold in a healthy control, or combination of (i) and (ii).
[0265] The methylation profile may comprise a subset of possible DMRs or markers that are specific to subject’s cancer or tumor. For example, the methylation profile may include a subset of regions that are determined to be specific to subject cancer. The methylation profile may comprise a smaller subset to allow for efficient data analysis. For example, by restricting the methylation profile to a smaller subset, the profile may maintain accuracy of monitoring asubject’s cancer, while reducing the amount of data to be processed. In some cases, by restricting the methylation profile to a universal panel of genomic regions disclosed herein,
[0266] In various embodiments, generation of a methylation profile does not comprise individual analysis of the methylation state of each DMR or interest. Instead, a profile may comprise comparing an aggregate signal of a subset of DMRs or markers in a test sample compared to a control sample, or set of control samples. For example, methylation profile may comprise DMRs that are specific to subject’s cancer or tumor. The generation may comprises identifying the aggregate signal from the subject-specific DMRs and comparing this aggregate signal against a reference or control signal.Genomic Mutation Profile
[0267] The present disclosure provides methods and systems for producing a mutation profile of a subject that has a disease and / or condition or can be suspected of having such disease and / or condition, wherein the methylation profile may be used to determine whether the subject has the disease and / or condition or can be at risk of having the disease and / or condition. The samples disclosed herein can be subjected to library preparation and next generation deep sequencing, for example to a depth of 1 million (M) to 60 M single reads, 10 M to 60 M single reads, 10 M to 100 M single reads, 40 M to 60 M single reads, 40 M to 100 M single reads, 60 M to 100 M single reads, 60 M to 200 M single reads, 1 M to 10 M single reads, 1 M to 40 M single reads, 1 M single reads to 100 M single reads, 1 M single reads to 200 M single reads, at least 1 M single reads, at least 10 M single reads, at least 40 M single reads, at least 60 M single reads, at least 100 M single reads, or at least 200 M single reads. In some cases, sequencing can be performed at low sequencing depth (e.g., 10 M single reads, 20 M single reads, 30 M single reads, 40 M single reads, from 1 M single reads to 10 M single reads, from 10 M single reads to 20 M single reads, from 20 M single reads to 30 M single reads, from 30 M single reads to 40 M single reads, at most 10 M single reads, at most 20 M single reads, at most 30 M single reads, or at most 40 M single reads). In some cases, a sample disclosed herein can be subjected to 1 sequencing at a depth of 0.1X to 100X, 0.1X to 60X, O.IX to 40X, 0.1X to 30X, O. IX to 20X, O.IX to 10X, O.IX to 5.0X, 0.5X to 100X, 0.5X to 60X, 0.5X to 40X, 0.5X to 30X, 0.5X to 20X, 0.5X to 10X, 0.5X to 5.0X, l.OX to 100X, l.OX to 60X, 1.0X to 40X, 1.0X to 30X, 1.0X to 20X, l.OX to lOX, 1. OX to 5. OX, at least 0.1X, at least 0.5X, at least 1.0X, at least 2. OX, at least 3. OX, at least 4. OX, at least 5. OX, at least 10. OX, at least 20. OX, at least 30. OX, at least 40. OX, at least 50. OX, at least 60. OX, at least 100X, at least 200X, at most 0.1X, at most 0.5X, at most 1.0X, at most 2. OX, at most 3. OX, at most 4. OX, at most 5. OX, at most 10. OX, at most 20. OX, at most 30. OX, at most40. OX, at most 50. OX, at most 60. OX, at most 100X, or at most 200X. A plurality of sequencing reads can be generated and analyzed. In some embodiments, deep sequencing may be configured to maximize identifying genomic mutations associated with the disease and / or condition.
[0268] In some embodiments, the relative measure of ctDNA abundance can be calculated from the mean mutant allele fractions (MAFs). In some embodiments, the mean MAF of mutations identified a subject and comprised in his / her mutation profile ranges from at least about 0.01% to at least about 10%. In some cases, the MAF of a ctDNA fraction of a sample can be about at least 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.15%, 0.2%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, or any percentage in between. In some cases, the MAF of a ctDNA fraction of a sample can be about at most 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.15%, 0.2%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, or any percentage in between.
[0269] In some embodiments, a generated mutation profile of a subject can be generated from sequencing results. In some embodiments, the mutation profile comprises genetic polymorphisms, such as missense variant, a nonsense variant, a deletion variant, an insertion variant, a duplication variant, an inversion variant, a frameshift variant, or a repeat expansion variant. In some embodiments, the mutation profile may comprise mutation variant derived from a fraction of cell-free nucleic acid molecules of a specific size range. The present disclosure provides methods, systems, and kits for producing a mutation profile of a subject that has a disease and / or condition or can be suspected of having such disease and / or condition, wherein the mutation profile may be used to determine whether the subject has the disease and / or condition or can be at risk of having the disease and / or condition. Producing a genomic mutation profile can comprise subjecting a plurality of nucleic acid molecules to library preparation and next generation deep sequencing (e.g., cfMeDIP-seq). A plurality of sequencing reads can be generated and analyzed, and, in some cases, deep sequencing may be configured to maximize identifying genomic mutations associated with the disease and / or condition. For example, a panel of canonical cancer driver genes may be included in a selector for sequencing results analysis. In some embodiments, including genes without known driver effects in a particular cancer type in the analysis of sequencing data may increase the sensitivity of ctDNA detection.
[0270] In some embodiments, the relative measure of ctDNA abundance can be calculated from the mean mutant allele fractions (MAFs). In some embodiments, the mean MAF of mutations identified a subject and comprised in his / her mutation profile ranges from at least about 0.01% to at least about 10%. The ctDNA fraction of a sample disclosed herein can be about at least 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.15%, 0.2%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, or any percentage in between.
[0271] In some embodiments, the generated mutation profile of a subject does not include mutation variants derived from cell-free nucleic acid molecules derived from a biological sample. In some embodiments, the mutation profile comprises genetic polymorphisms, such as missense variant, a nonsense variant, a deletion variant, an inserti...
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method for analyzing a sample derived from a subject, comprising:(a) obtaining a sample comprising nucleic acid molecules obtained or derived from said subject;(b) assaying said nucleic acid molecules to generate a data set comprising methylation states of one or more genomic regions comprising differentially methylation regions (DMRs); and(c) processing at least a portion of said data set to generate an output indicative of presence or absence cancer in said subject, wherein said at least said portion of said data set pertains to a set of DMRs specific to said subject.
2. The method of claim 1, further comprising providing a universal panel of genomic regions.
3. The method of claim 2, further comprising using said universal panel of genomic regions to enrich for said one or more genomic regions comprising said DMRs.
4. The method of any one of claims 1-3, wherein said assaying further comprises sequencing said nucleic acid molecules at a depth of at most 50 Million (M) single reads.
5. The method of any one of claims 1-3, wherein said assaying further comprises sequencing said nucleic acid molecules at a depth of at most 10 M single reads.
6. The method of any one of claims 1-5, wherein said processing of (c) further comprises comparing said one or more genomic regions to a set of DMRs specific to one or more reference subjects to generate said set of DMRs specific to said subject.
7. The method of any one of claims 1-6, wherein said processing of (c) further comprises comparing said one or more genomic regions to a set of anti-DMRs to generate a set of anti-DMRs specific to said subject.
8. The method of claim 6, wherein said comparing further comprises generating one or more counts of said set of DMRs specific to said subject.
9. The method of claim 7, wherein said comparing further comprises generating one or more counts of said set of anti-DMRs specific to said subject.
10. The method of claim 8 or 9, further comprising normalizing said one or more counts of said set of DMRs specific to said subject to said one or more counts of said set of anti- DMRs specific to said subject.
11. The method of claim 10, wherein said normalizing further comprises generating a methylation score.
12. The method of claim 11, further comprising comparing said methylation score to a threshold score, thereby generating said output indicative of said presence or absence of said cancer in said subject.
13. The method of any one of claims 1-12, further comprising obtaining one or more control samples comprising control nucleic acid molecules.
14. The method of claim 13, wherein said one or more control samples comprise one or more non-tissue samples and / or one or more tissue samples.
15. The method of claim 13 or 14, wherein said one or more control nucleic acid molecules are derived from one or more tissue samples.
16. The method of any one of claims 13-15, wherein said one or more control nucleic acid molecules comprise one or more cell-free nucleic acid molecules.
17. The method of any one of claims 13-16, wherein said one or more control samples are derived from one or more control subj ects without cancer.
18. The method of any one of claims 13-17, further comprising assaying said control nucleic acid molecules to generate a control data set comprising methylation states of one or more control genomic regions.
19. The method of claim 18, wherein said assaying said control nucleic acid molecules further comprises conducting one or more methylation reactions.
20. The method of claim 18 or 19, wherein said assaying said control nucleic acid molecules further comprises sequencing said control nucleic acids, or derivatives thereof.
21. The method of any one of claims 18-20, wherein said methylation states of said one or more control genomic regions comprise hypermethylated states, methylated states, nonmethylated states, or hypomethylated states, or any combinations thereof.
22. The method of any one of claims 18-21, further comprising processing at least a portion of said control data set to identify one or more hypomethylated regions, regions of non-methylation, or regions that are amenable to methylation enrichment, or any combinations thereof, thereby generating said universal panel of regions.
23. The method of claim 22, wherein said universal panel of genomic regions comprises said regions of non-methylation and said regions that are amenable to methylation enrichment.
24. The method of any one of claims 1-23, further comprising obtaining one or more reference samples from one or more reference subjects.
25. The method of claim 24, wherein said one or more reference subjects have cancer.
26. The method of claim 24 or 25, wherein said one or more reference samples comprise one or more reference nucleic acid molecules.
27. The method of claim 26, wherein said one or more reference nucleic acid molecules comprise cell-free nucleic acid molecules.
28. The method of claim 26 or 27, wherein said one or more reference nucleic acid molecules are derived from one or more non-tissue samples and / or one or more tissue samples.
29. The method of any one of claims 1-28, further comprising obtaining another one or more control samples comprising another one or more control nucleic acid molecules.
30. The method of claim 29, further comprising assaying said one or more reference nucleic acid molecules and said another one or more control nucleic acid molecule to generate a reference data set comprising methylation states of one or more regions.
31. The method of claim 30, further comprising processing said reference data set with said universal panel of genomic regions to identify said set of DMRs specific to one or more reference subjects and said set of anti -DMRs.
32. The method of any one of claims 1-31, wherein said processing of (c) further comprises comparing said one or more genomic regions to a set of DMRs specific to a reference subject to generate said set of DMRs specific to said subject.
33. The method of any one of claims 1-32, wherein said processing of (c) further comprises comparing said one or more genomic regions to a set of anti-DMRs to generate a set of anti-DMRs specific to said subject.
34. The method of claim 32, wherein said comparing further comprises generating one or more counts of said set of DMRs specific to said subject.
35. The method of claim 33, wherein said comparing further comprises generating one or more counts of said set of anti-DMRs specific to said subject.
36. The method of claim 34 or 35, further comprising normalizing said one or more counts of said set of DMRs specific to said subject to said one or more counts of said set of anti-DMRs specific to said subject.
37. The method of claim 36, wherein said normalizing generates another methylation score.
38. The method of claim 37, further comprising comparing said another methylation score to a threshold score, thereby generating said output indicative of said cancer in said subject.
39. The method of any one of claims 32-38, further comprising obtaining from said reference subject, a reference sample comprising one or more another reference nucleic acid molecules.
40. The method of claim 39, wherein said one or more another reference nucleic acid molecules are derived from one or more non-tissue samples and / or one or more tissue samples.
41. The method of claim 39 or 40, wherein said one or more another reference nucleic acid molecules comprise cell-free nucleic acid molecules.
42. The method of any one of claims 32-41, wherein said reference subject is same subject as said subject.
43. The method of any one of claims 32-42, wherein said reference sample is obtained prior to obtaining said sample.
44. The method of any one of claims 32-43, wherein said reference sample is obtained or derived from said subject subsequent to diagnosis with said cancer.
45. The method of any one of claims 32-44, wherein said reference subject is obtained or derived from said subject prior to treatment with a therapy.
46. The method of any one of claims 39-45, further comprising assaying said one or more another reference nucleic acid molecules to generate a reference data set comprising methylation states of one or more reference genomic regions.
47. The method of claim 46, wherein said assaying comprises enriching for said one or more reference genomic regions using said universal panel of genomic regions.
48. The method of claim 46, wherein said assaying comprises enriching for said one or more reference genomic regions without using said universal panel of genomic regions.
49. The method of claim 46 or 47, further comprising processing said reference data set to identify said set of DMRs specific to said reference subject.
50. The method of any one of claims 37-49, further comprising integrating said methylation score and said additional methylation score to generate a single score.
51. The method of claim 50, wherein said single score is indicative of said cancer.
52. The method of claim 50 or 51, wherein said integrating comprises Support Vector Machine, logistic regression, Bayesian Interference Model, weighted average, decision trees, and / or random forests.
53. The method of any one of claims 1-52, wherein said nucleic acid molecules are derived from one or more non-tissue samples.
54. The method of any one of claims 1-53, wherein said nucleic acid molecules are derived from one or more tissue samples.
55. The method of any one of claims 1-54, wherein said nucleic acid molecules comprise cell-free deoxyribonucleic acid (DNA) molecules.
56. The method of any one of claims 1-55, wherein said sample comprises a tissue sample, a blood sample, and / or a plasma sample.
57. The method of any one of claims 1-56, wherein said cancer is a late-stage cancer.
58. The method of any one of claims 1-57, wherein said cancer is an early-stage cancer.
59. The method of any one of claims 1-58, further comprising generating an output indicative of presence or absence of minimal residual disease in said subject.
60. The method of any one of claims 1-59, further comprising, based at least on said processing, treating said subject with a therapy capable of treating said cancer.
61. The method of claim 60, wherein said therapy comprises a chemotherapy, a radiation therapy, an immunotherapy, a targeted therapy, a surgical resection, or a combination thereof.
62. The method of any one of claims 1-60, further comprising, based at least on said processing, recommending a therapy regimen for said subject or changing a therapy regimen for said subject.
63. The method of any one of claims 1-62, wherein said assaying further comprises mixing said nucleic acid molecules with filler nucleic acid molecules.
64. The method of any one of claims 1-62, wherein said assaying does not comprise mixing said nucleic acid molecules with filler nucleic acid molecules.
65. The method of any one of claims 1-64, wherein said assaying further comprises enriching methylated nucleic acids.
66. The method of claim 65, wherein said enriching further comprises using a binder that binds to one or more methylated nucleotides.
67. The method of claim 66, wherein said binder comprises a protein comprising a methyl-CpG-binding domain.
68. The method of claim 67, wherein said protein is a MBD2 protein.
69. The method of any one of claims 65-68, wherein said binder comprises an antibody.
70. The method of claim 69, wherein said antibody is an anti 5-mC antibody.
71. The method of said 69, wherein said antibody is an anti 5 -hydroxymethyl cytosine antibody.
72. The method of any one of claims 65-71, wherein said binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of said cell free nucleic acid molecule or a sheared genomic nucleic acid molecule.
73. The method of any one of claims 65-72, wherein said assaying further comprises sequencing said nucleic acids, or derivatives thereof.
74. The method of claim 73, wherein said sequencing does not comprise bisulfite sequencing.
75. The method of claim 73, wherein said sequencing further comprises bisulfite sequencing with methylation specific PCR.
76. The method of any one of claims 73-75, wherein said sequencing further comprises targeted sequencing.
77. The method of any one of claims 73-76, wherein said sequencing further comprises using a plurality of capture probes.
78. The method of claim 77, wherein said plurality of capture probes comprises probes that are homologous or complementary to said one or more genomic regions.
79. The method of claim 77 or 78, wherein said plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation.
80. The method of any one of claims 77-79, wherein said plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation.
81. The method of any one of claims 77-80, wherein said sequencing generates sequencing reads corresponding to said one or more genomic regions.
82. The method of any one of claims 1-81, wherein said processing of (c) comprises counting a number of sequencing reads corresponding to a region of said one or more genomic regions.
83. A method for monitoring a subject for regression or progression of a disease or condition comprising: assaying a biological sample of said subject for one or more markers specific to said subject, using a universal panel of genomic regions.
84. The method of claim 84, wherein said assaying is without use of a primer or bait set specific to said one or more markers.
85. The method of claim 83 or 84, wherein said assaying further comprises sequencing nucleic acid molecules from said biological sample at a depth of at most 50 Million (M) single reads.
86. The method of any one of claims 83-85, wherein said assaying further comprises sequencing said nucleic acid molecules at a depth of at most 10 M single reads.
87. The method of any one of claims 83-86, wherein said biological sample is a blood sample or a plasma sample.
88. The method of any one of claims 83-87, wherein said biological sample is a tissue sample and / or a non-tissue sample.
89. The method of any one of claims 83-88, wherein said universal panel of genomic regions are derived from one or more control samples are derived from one or more control subjects without cancer.
90. The method of any one of claims 83-89, wherein said universal panel of genomic regions comprises non-m ethylated regions and regions that are amenable to methylation enrichment.
91. The method of any one of claims 83-90, wherein said one or more markers comprise differentially methylated regions (DMRs) specific to said subject.
92. The method of any one of claims 83-91, further comprising comparing one or more genomic regions of said biological sample to a set of anti-DMRs specific to said one or more reference samples to generate anti-DMRs specific to said subject.
93. The method of any one of claims 83-92, further comprising comparing said universal panel of genomic regions to one or more reference genomic regions to one or more reference samples to generate a set of DMRs specific to said one or more reference samples and / or a set of anti-DMRs specific to said one or more reference samples.
94. The method of claim 93, wherein said one or more reference samples are derived from one or more reference subj ects with cancer.
95. The method of claim 93 or 94, wherein said one or more reference samples comprise non-tissue samples or tissue samples.
96. The method of any one of claims 93-95, further comprising comparing said DMRs specific to said subject to said set of DMRs specific to said one or more reference samples to generate one or more counts of said DMRs specific to said subject.
97. The method of any one of claims 93-96, further comprising comparing said anti- DMRs specific to said subject to said set of anti-DMRs specific to said one or more reference samples to generate one or more counts of said anti-DMRs specific to said subject.
98. The method of claim 96 or 97, further comprising normalizing said one or more counts of said DMRs specific to said one or more counts of anti-DMRs specific to said subject, thereby generating a first methylation score.
99. The method of any one of claims 83-98, further comprising obtaining a reference sample.
100. The method of claim 99, wherein said reference sample is derived from said subject prior to obtaining said biological sample.
101. The method of claim 99 or 100, wherein said reference sample is a non-tissue sample and / or a tissue sample.
102. The method of any one of claims 99-101, further comprising comparing said universal panel of genomic regions to one or more genomic regions of said reference sample to generate a set of DMRs specific to said reference sample and / or a set of anti- DMRs specific to said reference sample.
103. The method of claim 102, further comprising comparing said DMRs specific to said subject to said set of DMRs specific to said reference sample to generate one or more counts of said DMRs specific to said subject.
104. The method of any one of claims 93-103, further comprising comparing said anti-DMRs specific to said subject to said set of anti-DMRs specific to said one or more reference samples to generate one or more counts of said anti-DMRs specific to said subject.
105. The method of claim 104, further comprising normalizing said one or more counts of said DMRs specific to said subject to said one or more counts of said anti- DMRs specific to said subject, thereby generating a second methylation score.
106. The method of any one of claims 83-105, wherein another biological sample is obtained at a time before or after obtaining said biological sample.
107. The method of claim 106, wherein said another biological sample comprises another one or more markers specific to said subject.
108. The method of claim 107, wherein said another one or more markers comprise another DMRs specific to said subject.
109. The method of any one of claims 106-108, further comprising comparing said another biological sample to said set of anti-DMRs specific to said one or more reference samples to generate another anti-DMRs specific to said subject.
110. The method of claim 109, further comprising comparing said another DMRs specific to said subject to said set of DMRs specific to said reference sample to generate one or more counts of said another DMRs specific to said subject.
111. The method of claim 109, further comprising comparing said another anti- DMRs specific to said subject to said set of anti -DMRs specific to said one or more reference samples to generate one or more counts of said another anti-DMRs specific to said subject.
112. The method of claim 110 or 111, further comprising normalizing said one or more DMR counts of said another DMRs specific to said subject to said one or more counts of said another anti-DMRs specific to said subject, thereby generating a third methylation score.
113. The method of any one of claims 105-112, further comprising comparing said second methylation score and said third methylation score, thereby generating an output indicative of said regression or said progression of said disease or said condition.
114. The method of any one of claims 105-113, further comprising integrating said first methylation score with said second methylation score to generate a single methylation score.
115. The method of claim 114, wherein said single methylation score is an output indicative of said regression or said progression of said disease or said condition.
116. The method of any one of claims 83-115, wherein said disease or condition comprises a cancer.
117. The method of claim 116, wherein said cancer is a late-stage cancer.
118. The method of claim 116, wherein said cancer is an early-stage cancer.
119. The method of claim 116, wherein said disease or condition is a pre-cancer.
120. The method of any one of claims 83-119, wherein said biological sample comprises nucleic acid molecules.
121. The method of claim 120, wherein said nucleic acid molecules are derived from one or more non-tissue samples.
122. The method of claim 120 or 121, wherein said nucleic acid molecules are derived from one or more tissue samples.
123. The method of any one of claims 120-122, wherein said nucleic acid molecules comprise cell-free nucleic acid molecules.
124. The method of any one of claims 83-123, wherein said biological sample comprises a blood sample and / or a plasma sample.
125. The method of any one of claims 120-124, wherein said assaying further comprises sequencing said nucleic acid molecules.
126. The method of claim 125, wherein said sequencing does not comprise bisulfite sequencing.
127. The method of claim 125, wherein said sequencing further comprises bisulfite sequencing with methylation specific PCR.
128. The method of any one of claims 125-127, wherein said sequencing further comprises generating sequencing reads corresponding to said one or more markers.
129. The method of any one of claims 125-128, wherein said sequencing further comprises targeted sequencing.
130. The method of any one of claims 125-129, wherein said sequencing further comprises using a plurality of capture probes.
131. The method of claim 130, said plurality of capture probes comprises probes that are homologous or complementary to said plurality of regions.
132. The method of claim 130 or 131, wherein said plurality of capture probes comprises probes that are homologous or complementary to regions with a known amount of methylation.
133. The method of any one of claims 130-132, wherein said plurality of capture probes comprises probes that are homologous or complementary to regions with no CpG methylation.
134. The method of any one of claims 83-133, wherein said assaying further comprises counting a number of sequencing reads corresponding to a marker of the one or more markers.
135. The method of any one of claims 83-134, wherein said assaying further comprises enriching methylated nucleic acids.
136. The method of claim 135, wherein the enriching further comprises using a binder that binds to one or more methylated nucleotides.
137. The method of claim 136, wherein said binder comprises a protein comprising a methyl -CpG-binding domain.
138. The method of claim 137, wherein said protein is a MBD2 protein.
139. The method of any one of claims 136-138, wherein said binder comprises an antibody.
140. The method of claim 139, wherein said antibody is an anti 5-mC antibody.
141. The method of claim 139 or 140, wherein said antibody is an anti 5- hydroxymethyl cytosine antibody.
142. The method of any one of claims 136-141, wherein said binder exhibits a reduced level of a non-specific binding to non-methylated nucleotides of said cell free nucleic acid molecule or a sheared genomic nucleic acid molecule.
143. The method of any one of claims 83-142, further comprising, based at least on said assaying, treating said subject with a therapy capable of treating said cancer.
144. The method of claim 143, wherein said therapy comprises a chemotherapy, a radiation therapy, an immunotherapy, a targeted therapy, a surgical resection, or a combination thereof.
145. The method of any one of claims 83-144, further comprising, based at least on said assaying, recommending a therapy regimen for said subject or changing a therapy regimen for said subject.
146. A method for analyzing a sample derived from a subject, comprising: assaying said sample for at least a portion of a set of differentially methylation regions (DMRs) specific to said subject to generate an output indicative of presence or absence of cancer, wherein said assaying comprises sequencing, wherein the sequencing has a depth of at most 50 Million (M) single reads.
147. The method of claim 146, wherein said sequencing has a depth of at most 10 M single reads.
148. The method of claim 146 or 147, further comprising assaying said sample to generate a data set comprising methylation states of one or more genomic regions.
149. The method of any one of claims 146-148, further comprising providing a universal panel of genomic regions.
150. The method of claim 149, further comprising comparing said universal panel of genomic regions to one or more reference genomic regions to generate a set of DMRs specific to a reference sample and / or a set of anti-DMRs specific to said reference sample.
151. The method of claim 149 or 150, further comprising comparing said universal panel of genomic regions to another one or more reference genomic regions to generate a set of DMRs specific to a one or more reference samples and / or a set of anti-DMRs specific to said one or more reference samples.
152. The method of any one of claims 146-151, wherein said sample comprises nucleic acid molecules.
153. The method of claim 152, wherein said nucleic acid molecules are derived from one or more tissue samples and / or one or more non-tissue samples.
154. The method of any one of claim 152 or 153, wherein said nucleic acid molecules comprise cell-free nucleic acid molecules.
155. The method of any one of claims 146-154, wherein said sample comprises a tissue sample.
156. The method of any one of claims 146-154, wherein said sample does not comprise a tissue sample.
157. The method of any one of claims 146-156, wherein said sample comprises a blood sample or a plasma sample.
158. The method of any one of claims 151-157, further comprising comparing said one or more genomic regions to said set of DMRs specific to said one or more reference samples to generate said set of DMRs specific to said subject.
159. The method of any one of claims 151-158, further comprising comparing said one or more genomic regions to said set of anti-DMRs specific to said one or more reference samples to generate said set of anti-DMRs specific to said subject.
160. The method of claim 158, wherein said comparing further comprises generating one or more counts of said set of DMRs specific to said subject.
161. The method of claim 59, wherein said comparing further comprises generating one or more counts of said set of anti-DMRs specific to said subject.
162. The method of claim 160 or 161, further comprising normalizing said one or more counts of said set of DMRs specific to said subject to said one or more counts of said set of anti-DMRs specific to said subject.
163. The method of claim 162, wherein said normalizing further comprises generating a methylation score.
164. The method of any one of claims 151-163, further comprising comparing said one or more genomic regions to said set of DMRs specific to said reference sample to generate said set of DMRs specific to said subject.
165. The method of claim 164, further comprising comparing said one or more genomic regions to said set of anti-DMRs specific to said one or more reference samples to generate said set of anti-DMRs specific to said subject.
166. The method of claim 165, wherein said comparing further comprises generating one or more counts of said set of DMRs specific to said subject.
167. The method of claim 165, wherein said comparing further comprises generating one or more counts of said set of anti-DMRs specific to said subject.
168. The method of claim 166 or 167, further comprising normalizing said one or more counts of said set of DMRs specific to said subject to said one or more counts of said set of anti-DMRs specific to said subject.
169. The method of claim 168, wherein said normalizing further comprises generating another methylation score.
170. The method of any one of claims 146-169, further using said output indicative of said presence or absence of cancer to determine progression of said cancer.
171. The method of any one of claims 146-171, further using said output indicative of said presence or absence of cancer to determine regression of said cancer.
172. The method of any one of claims 146-172, further using said output indicative of said presence or absence of cancer to determine therapy of said cancer.
173. A method for detecting Minimal Residual Disease (MRD), comprising: assaying a biological sample from a subject, wherein said assaying does not comprise analyzing a solid tumor sample of the subject, wherein said assaying comprises sequencing one or more genomic regions in said biological sample, wherein said sequencing has a depth of at most 50 million single reads.
174. A method comprising: assaying a sample for at least a portion of a set of differentially methylation regions (DMRs) specific to said subject to generate an output indicative of presence or absence of cancer at a specificity of at least 90%.
Citation Information
Patent Citations
Methods and compositions for noninvasive prenatal diagnosis of fetal aneuploidies
US20120282613A1
Cancer detection, classification, prognostication, therapy prediction and therapy monitoring using methylome analysis
US20210156863A1
Methods and systems for detecting methylation changes in DNA samples
US20220177956A1
Fragmentation for measuring methylation and disease
US20230374601A1
Methods and systems for generating sequencing libraries
WO2023107709A1
Cited By
Methods of capturing cell-free methylated DNA and uses of same
US12649915B2
Methods of capturing cell-free methylated DNA and uses of same
US12655417B2