Screening approaches for high cancer risk individuals
Non-invasive testing of cell-free nucleic acids with customized assays addresses the challenge of missed cancer detection in high-risk individuals, enabling early and personalized cancer detection in individuals with pathogenic germline variants or hereditary cancer syndrome.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
Individuals with pathogenic germline variants or hereditary cancer syndrome face an elevated risk of developing cancer, which often progresses quickly and may be missed by standard screening assays, leading to a need for improved monitoring and detection protocols.
A method involving non-invasive testing of cell-free nucleic acids for somatic and epigenomic variants, with customized assays based on the type of pathogenic germline variant, includes frequent monitoring and analysis of biomarkers to detect primary and secondary cancers using personalized frequency intervals and thresholds.
Enhances cancer detection by providing personalized and frequent monitoring, allowing for early identification of cancer development in high-risk individuals, thereby improving patient outcomes.
Smart Images

Figure IMGF000023_0001 
Figure IMGF000024_0001 
Figure IMGF000025_0001
Abstract
Description
SCREENING APPROACHES FOR HIGH CANCER RISK INDIVIDUALSCROSS-REFERENCE
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 712,939, filed October 28, 2024, which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Individuals with pathogenic germline variants or hereditary cancer syndrome (HCS) are at an elevated risk of developing cancer. For these individuals, there is typically a primary cancer associated with highest risk, and often there is also elevated risk of developing other cancers at a lower, secondary risk level. The variants or syndrome give the individual some of the molecular attributes of cancer, and they may likely show up as a “false positive” if ran through a non-invasive molecular screening assay (such as a liquid biopsy) that is intended for use on average risk population. If cancer does develop in these individuals, it tends to progress quickly, so there is a need for better screening protocols for such subjects.SUMMARY
[0003] The present disclosure generally provides improved screening approaches for individuals with an elevated risk of developing cancer or other disease, such as those subjects having pathogenic germline variants or hereditary cancer syndrome (HCS). Advantageously, by utilizing the screening methods described herein, greater monitoring of the subject can be carried out to determine if a primary or secondary cancer or other disease has developed. For example, the methods can comprise monitioring a level and / or rate of cancer signal increase in a sample from the subject using a non-invasive test (e.g., an assay analyzing cell-free nucleic acid for somatic and / or epigenomic variants) from an initial time point or baseline, performing testing at greater frequency intervals, and / or following standard screening guidelines for detecting the cancer of high risk (primary cancer) and utilizing a non-invasive test for detection of one or more secondary cancers with elevated risk. The methods can be further enhanced in some embedments by customizing subsequent assays based on the on the type of pathogenic germline variant(s) / disease suspected or determined to be present in the subject.
[0004] In certain aspects, the disclosure relates to a method for screening for cancer in a subject previously determined to be, or suspected of being, at an increased risk for developing cancer, comprising: (a) providing a first bodily fluid sample from the subject at a first time point and analyzing one or more analytes for one or more biomarkers from the first bodily fluid sample to determine a baseline cancer signal; (b) at one or more subsequent time points, providing an additional bodily fluid sample from the subject and analyzing one or more analytes for one or more biomarkers from the additional bodily fluid sample to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and (c) determining that the subject has cancer if the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold.
[0005] In some embodiments, the subject was previously determined to have one or more pathogenetic germline variants associated with increased cancer risk or was previously diagnosed with hereditary cancer syndrome (HCS). In some embodiments, the subject has a family history with a known germline pathogenic or likely pathogenic variant associated with increased cancer risk.
[0006] In some embodiments, wherein prior to providing the first bodily sample, the subject was previously determined to not have cancer.
[0007] In some embodiments, the analyzing in (a) and / or (b), above, comprises detecting for primary risk cancers. In some embodiments, the analyzing in (a) and / or (b), above, comprises detecting for secondary risk cancers. In some embodiments, the analyzing in (a) and / or (b), above, comprises detecting a non-cancer disease or to determine a risk of developing the non-cancer disease.
[0008] In some embodiment, the first and / or additional bodily fluid samples comprise blood, plasma, serum, urine, saliva, tears, or mucus. In some embodiments, the one or more analytes comprise cell-free deoxyribonucleic acid (cfDNA), cell-free ribonucleic acid (cfRNA), proteins, and / or exosomes. In some embodiments, the one or more analytes in the first bodily fluid sample and the additional bodily fluid sample are the same. In some embodiments, the one or more biomarkers in (a) and / or (b) comprise somatic variants, methylation status, fragmentomic patterns, transcription factor binding sites, and / or chromatin interaction sites.
[0009] In some embodiments, the subject was previously determined, or suspected of having, hereditary breast cancer, hereditary ovarian cancer, or hereditary breast and ovarian cancer(HBOC) syndrome. In some embodiments, the subject has germline mutations in BRCA1 and / or BRCA2.
[0010] In some embodiments, the subject was previously determined, or suspected of having Cowden syndrome. In some embodiments, the subject has a germline mutation in PTEN.
[0011] In some embodiments, the subject was previously determined, or suspected of having, Lynch syndrome (or hereditary non-polyposis colorectal cancer (HNPCC) syndrome) or Muir Torre syndrome. In some embodiments, the subject has a mutation in a mismatch repair gene, such as MLH1, MSH2, MSH6 or PMS2.
[0012] In some embodiments, the subject was previously determined, or suspected of having, hereditary leukemia or a hematologic malignancy syndrome. In some embodiments, the subject was previously determined, or suspected of having, Li-Fraumeni syndrome (LFS). In some embodiments, the subject has a germline mutation in TP53.
[0013] In some embodiments, the subject was previously determined, or suspected of having, Von Hippel-Lundau (VHL) disease. In some embodiments, the subject has a mutation in the VHL gene.
[0014] In some embodiments, the subject was previously determined, or suspected of having, multiple endocrine neoplasia (MEN) syndrome. In some embodiments, the subject has a mutation in the MEN1 and / or RET gene.
[0015] In some embodiments, the subject was previously determined to have familial adenomatous polyposis (FAP). In some embodiments, the subject has a mutation in the APC gene.
[0016] In some embodiments, the non-cancer disease is a neurological disease, autoimmune disease, cardiovascular disease, or cardiometabolic disorders. In some embodiments, the neurological disease is selected from the group consisting of Alzheimer’s disease, Parkinson’s disease, multiple sclerosis (MS), Amyotrophic Lateral Sclerosis (ALS), Huntington’s disease, stroke, or migraine. In some embodiments, the autoimmune disease is selected from the group consisting of rheumatoid arthritis, multiple sclerosis, lupus, Type 1 diabetes, psoriasis, celiac disease, Hashimoto’s thyroiditis, inflammatory bowel disease, Crohn’s disease or ulcerative colitis. In some embodiments, the cardiovascular disease is selected from the group consisting of Coronary Artery Disease (CAD), hypertension or high blood pressure, heart failure, arrhythmia, atherosclerosis, myocardial infarction, or stroke. In some embodiments, the cardiometabolic disorder is selected from the group consisting of: Type 2 Diabetes, hyperlipidemia, metabolic syndrome, obesity, hypercholesterolemia, or insulin resistance.
[0017] The time interval between the different analyses can be any days, weeks, months, years that a healthcare provider deems necessary to monitor the subject. In some embodiments, the one or more subsequent time points includes at least every week, 2 weeks, at least every 4 weeks, at least every 6 weeks, at least every 8 weeks, at least every 10 weeks, or at least every 12 weeks from the first time point. In some embodiments, the one or more subsequent time points includes at least every 6 months, at least every 12 months, at least every 18 months, at least every 24 months, at least every 30 months, or at least every 36 months from the first time point. In some embodiments, the one or more subsequent time points occurs every 3 months. In some embodiments, the one or more subsequent time points includes at least every 5 years or at least every 10 years from the previous time point.
[0018] In some embodiments, the subject was previously determined to be at an increased risk for developing cancer by undergoing germline genetic testing. In some embodiments, the subject was previously determined to not have cancer by undergoing a medical procedure. In some embodiments, the medical procedure comprises a colonoscopy, a stool-based test, a mammogram, a Pap test (or Pap smear), an HPV test, endoscopy, ultrasound, computed tomography (CT), magnetic resonance imaging (MRI), and / or other imaging method. In some embodiments, the subject was previously determined to not have cancer by analyzing a tissue sample from the subject. In some embodiments, the subject was previously determined to not have cancer by analyzing a bodily fluid sample from the subject. In some embodiments, the bodily fluid sample was analyzed by a multi-cancer detection (MCD) test. In some embodiments, the MCD test is a liquid biopsy assay. In some embodiments, wherein the analyzing in (a) and / or (b), above, comprises the use of a liquid biopsy assay. In some embodiments, a subject who was previously determined to be, or suspected of being, at an increased risk for developing cancer (e.g. who has HCS), will undergo testing for the primary cancer using the gold-standard tests, but at increased frequency relating to the standard screening guidelines.
[0019] In some embodiments, the analyzing in (a) and / or (b), above, comprises: (a) obtaining cfDNA molecules from the first and / or additional samples of the subject; (b) tagging, amplifying, and / or enriching the cfDNA molecules; and (c) sequencing a plurality of the cfDNA molecules or amplicons thereof.
[0020] In some embodiments, the enriching comprises contacting the cfDNA molecules of amplicons thereof with a probe set comprising oligonucleotide sequences configured to hybridize a plurality of target genomic regions within the cfDNA molecules or amplicons thereof to obtainenriched target genomic regions. In some embodiments, the enriching comprises contacting the cfDNA molecules or amplicons thereof with a set of oligonucleotide primers configured to amplify a plurality of target genomic regions within the cfDNA molecules or amplicons thereof to obtain enriched target genomic regions.
[0021] In some embodiments, each of the plurality of target genomic regions comprises members of the Homologous Recombination (HR), Non -Homologous End Joining (NHEJ), and / or PARP Pathways. In some embodiments, the target genomic regions comprise one or more sequences of BRCA1, BRCA2, TP53, MLH1, MSH2, MSH6, PMS2, PTEN, MUTYH, STK11, BMPR1A, SMAD4, PALB2, TSC1, TSC2, FLCN, RET, RBI, CDH1, APC, MUTYH, MEN1, CDKN1B, SDHA, SDHB, SDHC, SDHD, CTNNA1, VHL, FH, MET, FLCN, HOXB13, ATM, CHEK2, PTCHI, SUFU, PTCH2, CDKN2A, MITF, BAP1, and / or EPC AM. In some embodiments, the target genomic regions comprise a sequence in a gene body, a promoter region, a 3 or 5’ UTR or a regulatory element. In some embodiments, the regulatory element comprises a genetic or epigenetic element that modulates gene expression.
[0022] In some embodiments, the subject is determined to have an epigenetic alteration leading to gene inactivation. In some embodiments, the epigenetic alteration comprises promoter methylation, at least one regulatory variant that reduces gene expression, or a chromatin conformation change that reduces gene expression. In some embodiments, the subject is determined to have a somatic genetic variant comprising a single nucleotide variant (SNV), and insertion or deletion (indel), a gene fusion, and / or or copy number variation (CNV).
[0023] In some embodiments, the cancer is selected from the group consisting of breast cancer, ovarian cancer, prostate cancer, pancreatic cancer, bowel cancer, womb cancer, stomach cancer, gallbladder cancer, bladder cancer, bone cancer, acute myeloid leukemia (AML), soft tissue sarcoma, brain tumors, cancer of the adrenal gland, thyroid cancer, kidney cancer, skin cancer (melanoma), liver cancer, pancreatic neuroendocrine tumors (pNETs), renal cell carcinoma, retinoblastoma, medullary thyroid cancer, hereditary papillary renal cell carcinoma (HPRCC), and uterine leiomyomata.
[0024] In some embodiments, the cancer is determined to be a primary cancer. In some embodiments, the cancer is determined to be a secondary cancer.
[0025] In some embodiments, the tagging comprises attaching adapters to a plurality of the cfDNA molecules. In some embodiments, the adapters comprise molecular barcodes. In someembodiments, the tagging comprises incorporating molecular barcodes via ligation or PCR amplification.
[0026] In some embodiments, the baseline cancer signal and / or the level of cancer and / or rate of change in cancer level from the additional bodily fluid sample is determined by a method comprising quantifying somatic mutations, tumor fraction, mutant allele fraction, and / or methylation signals.
[0027] In some embodiments, the baseline cancer signal and / or the level of cancer and / or rate of change in cancer level from the additional bodily fluid sample is determined by a method comprising the use of any of the biomarkers disclosed above.
[0028] In some embodiments, the analysis in (a) and / or (b), above, is configured to be personalized for the subject based on the cancer or disease type the subject was previously determined to have or is suspected of having. In some embodiments, the method comprises targeting biomarkers associated with a primary and / or secondary cancer. In some embodiments, the targeting comprises the use of oligonucleotide capture probes or target amplification primers that hybridize or bind to nucleic acid sequences of cell-free nucleic acid molecules or amplicons thereof, wherein the oligonucleotide capture probes or target amplification primers are selected based the cancer or disease type the subject was previously determined to have or is suspected of having.
[0029] In some embodiments, the method further comprises administering a therapy to the subject to treat the cancer(s). In some embodiments, the therapy comprises PARP inhibitors, immunotherapies, a target-based therapy (such as kinase inhibitors), chemotherapy, or any combination thereof. In some embodiments, the therapy comprises one or more treatments selected from Table 1.
[0030] In certain aspects, the disclosure provides a method for screening for cancer in a subject previously determined to be, or suspected of being, at an increased risk for developing cancer or a non-cancerous disease, comprising: (a) providing a first sample from the subject at a first time point and performing a first test using one or more analytes from the sample to determine a baseline cancer or disease signal; (b) at one or more subsequent time points, providing an additional sample from the subject and performing an additional test using one or more analytes from the additional sample to determine the level of cancer or non-cancerous disease and / or rate of change in cancer or non-cancerous disease level from the baseline cancer or disease signal; and (c) determining that the subject has cancer or non-cancerous disease if the level of cancer or disease and / or rate of change in cancer or non-cancerous disease level from the additional test exceeds a predeterminedthreshold above the baseline cancer or disease signal or does not have cancer or disease if the level of cancer or disease and / or rate of change in cancer or disease level from the additional test is at or below the predetermined threshold.
[0031] In certain aspects, the disclosure provides a method for screening a subject with hereditary cancer syndrome (HCS), comprising: (a) obtaining an initial sample from the subject and performing an assay on the sample, wherein the assay comprises extracting a plurality of molecules from the sample, and sequencing the plurality of molecules to obtain a baseline molecular profile for the subject with HCS; (b) repeating the assay on a subsequent sample obtained from the subject at a later time point to obtain a progressive molecular profile; and (c) comparing the progressive molecular profile to the baseline molecular profile using multidimensional analysis, where each dimension corresponds to a specific feature within the molecular profiles to identify changes across the features, wherein changes in the distribution, levels and / or grouping in the feature space are analyzed to determine a change in cancer signal.
[0032] In certain aspects, the disclosure provides a method for screening an individual with hereditary cancer syndrome (HCS), comprising: (a) providing a sample comprising cell-free DNA molecules from the individual; (b) enriching the cfDNA molecules or amplicons thereof for a plurality of target genomic regions to obtain enriched target genomic regions; (c) sequencing the enriched target genomic regions to generate sequencing reads; and (d) analyzing a plurality of sequencing reads to detect at least one or more somatic alterations in the target genomic molecules, and determining that the subject has cancer based on the detection of the one or more somatic alterations.
[0033] In certain aspects, the disclosure provides a method for determining a presence or absence of cancer in a subject previously determined to be, or suspected of being, at an increased risk for developing cancer, comprising: (a) providing a first sample from the subject at a first time point and analyzing one or more biomarkers from one or more analytes in the sample to determine a baseline cancer signal; (b) at one or more subsequent time points, providing an additional sample from the subject and analyzing one or more biomarkers from one or more analytes from the additional sample to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and (c) determining that the subject has cancer if the level of cancer and / or rate of change in cancer level from the additional diagnostic test exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level from the additional diagnostic test is at or below the predetermined threshold,wherein the analysis in (a) and / or (b) are personalized to target biomarkers associated with one or more cancers. In certain embodiments, the analysis in (a) and / or (b) comprises enriching cfDNA molecules or amplicons thereof for genomic or epigenomic regions targeting the biomarkers associated with the one or more cancers. In certain embodiments, the one or more cancers is a primary cancer. In certain embodiments, the one or more cancers is, or also includes, a secondary cancer.
[0034] In certain aspects, the disclosure provides a method for determining that a subject is at an elevated risk for developing cancer, comprising: (a) providing a first bodily fluid sample from the subject and analyzing cell-free nucleic acid molecules or amplicons thereof from the sample to determine whether a genetic variant is somatic or germline in origin; and (b) determining that the genetic variant is a germline genetic variant, wherein the germline genetic variant is associated with hereditary cancer risk, thereby determining that the subject is at an elevated risk for developing cancer. In some embodiments, the subject does not exhibit any symptoms of cancer. In some embodiments, the method further comprises, based on determining a presence of the germline genetic variant in the first bodily fluid sample, performing a subsequent medical procedure to determine that the subject does not have a primary cancer. In some embodiments, the medical procedure comprises a colonoscopy, a stool-based test, a mammogram, a Pap test (or Pap smear), an HPV test, endoscopy, ultrasound, computed tomography (CT), magnetic resonance imaging (MRI), and / or other imaging method. In some embodiments, the method further comprises (c) providing a second bodily fluid sample from the subject at a time point after performing the medical procedure, and analyzing one or more biomarkers from the second bodily fluid sample to determine a baseline cancer signal; (d) providing one or more additional bodily fluid samples from the subject at one or more additional time points after the analyzing in (c), and analyzing one or more biomarkers from the one or more additional bodily fluid samples to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and (e) determining that the subject has cancer if the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold. In some embodiments, the analyzing of (c) and / or the analyzing in (d) comprises customizing the analysis to target one or more biomarkers associated with the hereditary cancer risk. In some embodiments, (c)-(e) are performed at intervals of about 1 to 6 month to monitor for one or more secondary cancers associated with the HCS. In some embodiments, if the subject isdetermined to have cancer, performing an additional medical procedure to confirm the presence of the cancer. In some embodiments, the method further comprises administering a therapy to the subject to treat the cancer. In some embodiments, the subject is determined to be at an increased risk of developing cancer by the methods disclosed herein.
[0035] In certain aspects, the disclosure provides a method of screening an individual with hereditary cancer syndrome (HCS) comprising: (a) determining a baseline cancer signal using a non-invasive molecular test from an initial time point; (b) repeating the non-invasive molecular test at frequency intervals to determine the levels and / or rate of change in the baseline cancer signal; (c) following standard screening guidelines or gold-standard test for detecting a cancer of highest risk or primary cancer when the level or rate of change of cancer signal in (b) exceeds a pre-set threshold; and (d) utilizing a follow-up non-invasive molecular screening test for detection of secondary cancers with elevated risk.
[0036] In certain aspects, the disclosure provides a method of screening an individual with hereditary cancer syndrome (HCS) or at high-risk for developing at least one cancer type due to germline variants, comprising: (a) performing at least one gold standard test to confirm the individual is cancer-free; (b) performing a non-invasive molecular screening test to determine a cancer signal baseline score for the individual; (c) monitoring the individual by repeating the non- invasive molecular screening test in (b) for at least one time point; and (d) determining the level or rate of change of cancer signal at each time point, wherein the level or rate of change above preset threshold is reported as a positive result that the individual has cancer.
[0037] In certain aspects, the disclosure provides a method of screening an individual for cancer, comprising: (a) diagnosing the individual as having hereditary cancer syndrome (HCS) or is at an increased risk for certain cancer types due to germline variants; (b) performing at least one gold standard test to confirm the individual is cancer-free; (c) performing a non-invasive molecular screening test to determine a cancer signal baseline score for the individual; (d) monitoring the individual by repeating the non-invasive molecular screening test in (b) for at least one time point; and (e) determining the level or rate of change of cancer signal at each time point, wherein the level or rate of change above a pre-set threshold is reported as a positive result that the individual has cancer.
[0038] In certain aspects, the disclosure provides a method of screening an individual for cancer, comprising: (a) determining that the individual as having hereditary cancer syndrome (HCS) or is at an increased risk for certain cancer types due to germline variants; (b) performing at least onegold standard test to screen for a primary risk cancer at one or more time points; and (c) performing a non-invasive molecular screening test to screen for a secondary risk cancer at multiple time points, wherein a cancer signal baseline score for the individual is established at a first time point and a level or rate of change of cancer signal is determined at a second time point, wherein the level or rate of change above a pre-set threshold is reported as a positive result that the individual has a secondary cancer.
[0039] In certain aspects, the disclosure provides a method for screening cancer in a subject, comprising: (a) providing a first sample from the subject at a first time point and analzying one or more analytes for one or more biomarkers from the first sample for determining (i) one or more germline variants in genes associated with hereditary cancer syndrome (HCS) to diagnose HCS, and (ii) one or more molecular features indicative of cancer to establish a baseline cancer signal for subsequent monitoring; (b) at one or more subsequent time points, providing a second sample from the subject and analyzing one or more analytes for one or more biomarkers from the second sample for determining (i) somatic alterations arising in the HCS gene(s) harboring the germline variant(s), and / or (ii) the one or more molecular features to determine changes relative to the baseline cancer signal established in (a); and (c) determining that the subject has cancer when the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold.
[0040] In some embodiments, genes associated with HCS comprise BRCA1 and / or BRCA2.In some embodiments, biomarkers are selected from the group consisting of mRNA expression, protein levels, methylation status, mutational signatures, and somatic variants in one or more genes in the homologous recombination deficiency (HRD) pathway. In some embodiments, the mutational signatures comprise Single Base Substitution Signature 3 (SBS3), Rearrangement Signature 3 (RS3), and / or Rearrangement Signature 5 (RS5).
[0041] In some embodiments, the determining in (c) further comprises: (i) preprocessing the first sample to generate a first set of standardized molecular feature values, (ii) inputting the first set of feature values into a trained model, (iii) receiving, from the model, an output comprising a baseline cancer signal score and / or a baseline feature vector, (iv) preprocessing the second sample to generate a second set of standardized molecular feature values; (v) inputting the second set of feature values into the same trained model to obtain a current cancer level score and / or a current feature vector, (vi) computing a change metric and / or rate of change between the current andbaseline outputs, (vii) determining cancer status by comparing the change metric and / or rate of change to a predetermined threshold, and (viii) outputting a result indicating a positive cancer status when the threshold is exceeded and a negative cancer status when it is not.
[0042] In some embodiments, determining that the subject has cancer further comprises: (i) preprocessing the first sample to generate a first set of standardized molecular feature values, (ii) inputting the first set of feature values into a trained model, (iii) receiving, from the model, an output comprising a baseline cancer signal score and / or a baseline feature vector, (iv) preprocessing the second sample to generate a second set of standardized molecular feature values; (v) inputting the second set of feature values into the same trained model to obtain a current cancer level score and / or a current feature vector, (vi) computing a change metric and / or rate of change between the current and baseline outputs, (vii) determining cancer status by comparing the change metric and / or rate of change to a predetermined threshold, and (viii) outputting a result indicating a positive cancer status when the threshold is exceeded and a negative cancer status when it is not.
[0043] In some embodiments, biomarkers comprise one or more of methylation status, methylation levels and / or patterns, fragmentomic features including fragment length distributions, end motifs and / or breakpoint density profiles, mRNA expression values, protein expression levels, or combinations thereof.
[0044] In some embodiments, wherein the determining in (c) further comprises: preprocessing the first and second samples to generate respective first and second sets of standardized molecular feature values, inputting the first and second sets of standardized molecular feature values into the trained model, and receiving, from the model, a cancer status for the subject comprising a binary call.
[0045] In some embodiments, determining that the subject has cancer further comprises: preprocessing the first and second samples to generate respective first and second sets of standardized molecular feature values, inputting the first and second sets of standardized molecular feature values into the trained model, and receiving, from the model, a cancer status for the subject comprising a binary call.
[0046] In any of the proceeding embodiments, the model comprises a neural network trained on datasets including subjects with confirmed cancer and cancer-free controls, optionally including HCS carriers.
[0047] In some embodiments, steps (b) and (c) are performed at intervals shorter than standard clinical guidelines, including intervals of about 1 to 3 months.
[0048] In some embodiments, the intervals between providing samples are shorter than standard clinical guidelines, including intervals of about 1 to 3 months.
[0049] In certain aspects, the disclosure provides a method for screening cancer in a subject, comprising: (a) establishing a baseline cancer signal for the subject by analyzing one or more biomarkers from a first sample obtained from the subject at a first time point; and (b) monitoring the subject at one or more subsequent time points by obtaining one or more additional samples and analyzing the same biomarkers to determine changes relative to the baseline cancer signal. In some embodiments, the method further comprises determining that the subject has cancer when the change relative to the baseline cancer signal exceeds a predetermined threshold. In some embodiments, the subject is determined to be, or suspected of being, at an increased risk for developing cancer.
[0050] In any of the proceeding embodiments, the subject is determined to be, or suspected of being, at an increase risk for developing cancer.
[0051] In any of the proceeding embodiments, monitoring is performed at intervals of about 1 to 6 months. In any of the proceeding embodiments, monitoring is performed at invtervals of about 1, 2, 3, 4, 5, or 6 months, In any of the proceeding embdoiments, monitoring is performed at intervals of 1 to 6 months, In any of the proceeding embodiments, monitoring is performed at invtervals of 1, 2, 3, 4, 5, or 6 months,
[0052] In any of the proceeding embodiments, the subject is screened for a primary cancer of highest risk using a clinical test and for one or more secondary cancers using steps (a)-(b). In any of the proceeding embodiments, the subject is screened for a primary cancer of highest risk using a clinical test and for one or more secondary cancers using the disclosed methods. In some embodiments, the clinical test is a gold standard clinical test. In some embodiments, the gold- standard clinical test comprises one or more of colonoscopy, stool-based test, mammogram, breast MRI, Pap smear, endoscopy, ultrasound, computed tomography (CT), or magnetic resonance imaging (MRI). In some embodiments, the subject is screened for a primary cancer of highest risk using a clinical test at interval shorter than standard clinical guidelines, including intervals of about 1 to 3 months, In some embodiments, the intervals are of 1 to 3 months. In some embodiments, the intervals are any one of 1, 2, 3 months.
[0053] In certain aspects, the disclosure provides a method for screening cancer in a subject who is identified or has been identified as having hereditary cancer syndrome, wherein the method comprises performing or having performed a clinical test on the subject for a primary cancerassociated with the hereditary cancer syndrome at intervals which are at an increased frequency relative to the standard screening guidelines.
[0054] In some embodiments, (i) the primary cancer is colorectal cancer and the clinical test is a colonoscopy; (ii) the primary cancer is breast cancer and the clinical test is a mammogram and / or breast MRI; (iii) the primary cancer is ovarian cancer and the clinical test is a transvaginal ultrasound and / or a serum CA-125 test; (iv) the primary cancer is endometrial cancer and the clinical test is an endometrial biopsy; or (v) the primary cancer is renal cell carcinoma and the clinical test is an abdominal MRI.
[0055] In some embodiments, the intervals are less than twelve months, less than six months, less than four months, less than three months, less than two months or less than one month.
[0056] In some embodiments, the method further comprises: providing a first bodily fluid sample from the subject at a first time point and analyzing one or more analytes for one or more biomarkers from the first bodily fluid sample to determine a baseline cancer signal; at one or more subsequent time points, providing an additional bodily fluid sample from the subject and analyzing one or more analytes for one or more biomarkers from the additional bodily fluid sample to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and determining that the subject has cancer if the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold. In some embodiments, the cancer determined to be present or absent via the analysis of the bodily fluid samples is a secondary cancer.
[0057] In some embodiments, the methods disclosed herein utilize a machine learning model that is trained on data from a population having similar health, genetic, and / or epigenomic profiles of the the subject (e.g., age, sex, weight, cancer type, somatic variant, methylation patterns). Such a model can be employed in the present methods to improve determining whether the subject has a primary or secondary cancer. In some embodiments, the machine learning model can be used to enhance the analysis of data from a non-invasive molecular assay (such as a liquid biopsy) for determining the presence of a secondary cancer, something that a gold-standard test may not catch if only screening for a primary cancer in the subject.
[0058] The various steps of the methods disclosed herein may be carried out at the same or different times, in the same or different geographical locations, e.g. countries, and / or by the same or different people.
[0059] In some embodiments, the results of the methods disclosed herein are used as an input to generate a report. The report may be in a paper or electronic format. For example, the quantitative measure indicative of a number of nucleic acids in a sample that map to a genomic region as obtained by the methods disclosed herein, or information derived therefrom, can be displayed directly in such a report. Alternatively, or additionally, diagnostic information or therapeutic recommendations which are at least in part based on the methods disclosed herein can be included in the report.
[0060] The various steps of the methods disclosed herein may be carried out at the same or different times, in the same or different geographical locations, e.g. countries, and / or by the same or different people.
[0061] Additional advantages will be set forth in part in the description which follows or may be learned by practice. The advantages will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.
[0062] This summary is provided for purposes of illustrating some exemplary embodiments, so as to provide a basic understanding of some aspects of the subject matter described herein. Accordingly, it will be appreciated that the above-described features are examples and should not be construed to narrow the scope or spirit of the subject matter described herein in any way. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following Detailed Description, Figures, and Claims.DETAILED DESCRIPTION
[0063] Reference will now be made in detail to certain embodiments of the invention. While the invention will be described in conjunction with such embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the invention as defined by the appended claims.
[0064] The section headings used herein are for organizational purposes and are not to be construed as limiting the disclosed subject matter in any way. In the event that any document or other material incorporated by reference contradicts any explicit content of this specification, including definitions, this specification controls.A. Definitions
[0065] Before describing the present teachings in detail, it is to be understood that the disclosure is not limited to specific compositions or process steps, as such may vary. It should be noted that, as used in this specification and the appended claims, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, reference to “a nucleic acid” includes a plurality of nucleic acids, reference to “a cell” includes a plurality of cells, and the like.
[0066] Numeric ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant digits and the error associated with the measurement. Also, the use of “comprise”, “comprises”, “comprising”, “contain”, “contains”, “containing”, “include”, “includes”, and “including” are not intended to be limiting. It is to be understood that both the foregoing general description and detailed description are exemplary and explanatory only and are not restrictive of the teachings.
[0067] Unless specifically noted in the above specification, embodiments in the specification that recite “comprising” various components are also contemplated as “consisting of’ or “consisting essentially of’ the recited components; embodiments in the specification that recite “consisting of’ various components are also contemplated as “comprising” or “consisting essentially of’ the recited components; and embodiments in the specification that recite “consisting essentially of’ various components are also contemplated as “consisting of’ or “comprising” the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0068] “Cell-free DNA,” “cfDNA molecules,” or simply “cfDNA” include DNA molecules that naturally occur in a subject in extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). While the cfDNA originally existed in a cell or cells in a large complex biological organism, e.g., a mammal, it has undergone release from the cell(s) into a fluid found in the organism, and may be obtained by obtaining a sample of the fluid without the need to perform an in vitro cell lysis step.
[0069] As used herein, a modification or other feature is present in “a greater proportion” in a first subsample or population of nucleic acid than in a second subsample or population when the fraction of nucleotides with the modification or other feature is higher in the first subsample or population than in the second population. For example, if in a first subsample, one tenth of the nucleotides are mC, and in a second subsample, one twentieth of the nucleotides are mC, then thefirst subsample comprises the cytosine modification of 5-methylation in a greater proportion than the second subsample.
[0070] As used herein, “machine learning model” (or “model”) refers to a collection of parameters and functions, where the parameters are trained on a set of training samples or individual data points or instances used to train a machine learning model. These samples are part of the dataset that provides the model with examples of input data along with the corresponding output (for supervised learning), or just input data (for unsupervised learning). The parameters and functions may be a collection of linear algebra operations, non-linear algebra operations, and tensor algebra operations. The parameters and functions may include statistical functions, tests, and probability models. The training samples can correspond to samples having measured properties of the sample (e.g., genomic, epigenomic, transcriptomic, metabolites etc. data and other subject data, such as histology, imaging data and / or electronic medical health records, or insurance claim data), as well as known patient / sample metadata including classifications or labels for example molecular phenotypes or specific cancer or disease therapies. Other phenotypes can include patient biomedical information including “cardiovascular phenotypes” or “cardiovascular risk factors” such as weight, height, Body Mass Index (BMI) , and other physical characteristics. Yet other phenotypes can include cancer risk factors including smoking, excessive alcohol consumption, poor diet, physical inactivity, obesity, genetic predispositions, exposure to harmful chemicals and radiation, chronic inflammation, certain infections (such as human papillomavirus, hepatitis B and C), hormonal imbalances, and advanced age. The model can learn from the training samples in a training process that optimizes the parameters (and potentially the functions) to provide an optimal quality metric (e.g., accuracy) for classifying new samples. A variety of advanced statistical and computational methods that can be employed as training functions including Expectation Maximization(EM) to find maximum likelihood estimates of parameters in probabilistic models, especially for models with latent variables, Maximum Likelihood Estimation (MLE) to estimate the parameters of a statistical model. MLE methods select the set of parameters that maximize the likelihood function i.e., the parameters under which the observed data is most probable. Bayesian Parameter Estimation Methods which incorporate prior knowledge in addition to the data at hand through the use of probability distributions. These include Markov Chain Monte Carlo (MCMC), Gibbs Sampling, Hamiltonian Monte Carlo (HMC), and Variational Inference (VI), or Gradient- Based Methods including Stochastic Gradient Descent (SGD) and the Broyden-Fletcher-Goldfarb- Shanno (BFGS) algorithm.
[0071] As used herein, “Deep learning models” refers to a collection of “architectures” or “algorithms” useful in scenarios where machine learning approaches may fall short, for example due to the complexity of the data (e.g., high dimensional genomics, epigenomics datasets used alone or in one or more combinations). Additional high dimensional data can include medical images such as MRIs, or histology reports. Deep learning models, especially CNNs, can automatically extract relevant features without manual feature engineering.
[0072] Exemplary applications for deep learning models include complex pattern recognition tasks for example, recognizing specific functional elements in genomics data, functional elements can include transcription factor binding sites (TFBS) or chromatin interaction sites for example promoter-enhancer interactions. Additionally, recognition of disease specific features in histology images including patters or specific features of PDL-1 expression.
[0073] As used herein, “without substantially altering base-pairing specificity” of a given nucleobase means that a majority of molecules comprising that nucleobase that can be sequenced do not have alterations of the base pairing specificity of the second nucleobase relative to its base pairing specificity as it was in the originally isolated sample. In some embodiments, 75%, 90%, 95%, or 99% of molecules comprising that nucleobase that can be sequenced do not have alterations of the base pairing specificity of the second nucleobase relative to its base pairing specificity as it was in the originally isolated sample.
[0074] As used herein, “base pairing specificity” refers to the standard DNA base (A, C, G, or T) for which a given base most preferentially pairs. Thus, for example, unmodified cytosine and 5- methylcytosine have the same base pairing specificity (i.e., specificity for G) whereas uracil and cytosine have different base pairing specificity because uracil has base pairing specificity for A while cytosine has base pairing specificity for G. The ability of uracil to form a wobble pair with G is irrelevant because uracil nonetheless most preferentially pairs with A among the four standard DNA bases.
[0075] As used herein, a “combination” comprising a plurality of members refers to either of a single composition comprising the members or a set of compositions in proximity, e.g., in separate containers or compartments within a larger container, such as a multiwell plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other form of storage.
[0076] The “capture yield” of a collection of probes for a given target set refers to the amount (e.g., amount relative to another target set or an absolute amount) of nucleic acid corresponding to the target set that the collection of probes captures under typical conditions. Exemplary typicalcapture conditions are an incubation of the sample nucleic acid and probes at 65°C for 10-18 hours in a small reaction volume (about 20 pL) containing stringent hybridization buffer. The capture yield may be expressed in absolute terms or, for a plurality of collections of probes, relative terms. When capture yields for a plurality of sets of target regions are compared, they are normalized for the footprint size of the target region set (e.g., on a per-kilobase basis). Thus, for example, if the footprint sizes of first and second target regions are 50 kb and 500 kb, respectively (giving a normalization factor of 0.1), then the DNA corresponding to the first target region set is captured with a higher yield than DNA corresponding to the second target region set when the mass per volume concentration of the captured DNA corresponding to the first target region set is more than 0.1 times the mass per volume concentration of the captured DNA corresponding to the second target region set. As a further example, using the same footprint sizes, if the captured DNA corresponding to the first target region set has a mass per volume concentration of 0.2 times the mass per volume concentration of the captured DNA corresponding to the second target region set, then the DNA corresponding to the first target region set was captured with a two-fold greater capture yield than the DNA corresponding to the second target region set.
[0077] “Capturing” one or more target nucleic acids refers to preferentially isolating or separating the one or more target nucleic acids from non-target nucleic acids.
[0078] A “captured set” of nucleic acids refers to nucleic acids that have undergone capture.
[0079] A “target-region set” or “set of target regions” refers to a plurality of genomic loci targeted for capture and / or targeted by a set of probes (e.g., through sequence complementarity).
[0080] “Corresponding to a target region set” means that a nucleic acid, such as cfDNA, originated from a locus in the target region set or specifically binds one or more probes for the target-region set.
[0081] “Specifically binds” in the context of an probe or other oligonucleotide and a target sequence means that under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence, or replicates thereof, to form a stable probe:target hybrid, while at the same time formation of stable probemon-target hybrids is minimized. Thus, a probe hybridizes to a target sequence or replicate thereof to a sufficiently greater extent than to a nontarget sequence, to enable capture or detection of the target sequence. Appropriate hybridization conditions are well-known in the art, may be predicted based on sequence composition, or can be determined by using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY,1989) at §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, particularly §§ 9.50-9.51, 11.12- 11.13, 11.45-11.47 and 11.55-11.57, incorporated by reference herein).
[0082] “Sequence-variable target region set” refers to a set of target regions that may exhibit changes in sequence such as nucleotide substitutions (i.e., single nucleotide variations), insertions, deletions, or gene fusions or transpositions in neoplastic cells (e.g., tumor cells and cancer cells).
[0083] “Epigenetic target region set” refers to a set of target regions that may show sequenceindependent changes in neoplastic cells (e g., tumor cells and cancer cells) or that may show sequence-independent changes in cfDNA from subjects having cancer relative to cfDNA from healthy subjects. Examples of sequence-independent changes include, but not limited to, changes in methylation (increases or decreases), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. For present purposes, loci susceptible to neoplasia-, tumor-, or cancer-associated focal amplifications and / or gene fusions may also be included in an epigenetic target region set because detection of a change in copy number by sequencing or a fused sequence that maps to more than one locus in a reference genome tends to be more similar to detection of exemplary epigenetic changes discussed above than detection of nucleotide substitutions, insertions, or deletions, e.g., in that the focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing because their detection does not depend on the accuracy of base calls at one or a few individual positions. In some embodiments, the epigenetic target region set includes one or more genomic regions, where the epigenetic state (e.g., methylation state) of cfDNA molecules in these regions is unchanged in cancer, but their presence / quantity in blood indicates increased, aberrant presentation of cfDNA from certain tissue (e.g. cancer origin) into circulation.
[0084] A nucleic acid is “produced by a tumor” or ctDNA or circulating tumor DNA, if it originated from a tumor cell. Tumor cells are neoplastic cells that originated from a tumor, regardless of whether they remain in the tumor or become separated from the tumor (as in the cases, e.g., of metastatic cancer cells and circulating tumor cells).
[0085] The term “methylation” or “DNA methylation” refers to addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to addition of a methyl group to a cytosine at a CpG site (cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in a 5’ - 3’ direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to addition of a methyl group to adenine, such as in N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon ofcytosine). In some embodiments, 5-methylation refers to addition of a methyl group to the 5C position of the cytosine to create 5-methylcytosine (5mC). In some embodiments, methylation comprises a derivative of 5mC. Derivatives of 5mC include, but are not limited to, 5- hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of cytosine). In some embodiments, 3C methylation comprises addition of a methyl group to the 3C position of the cytosine to generate 3 -methylcytosine (3mC). Methylation can also occur at non CpG sites, for example, methylation can occur at a CpA, CpT, or CpC site. DNA methylation can change the activity of methylated DNA region. For example, when DNA in a promoter region is methylated, transcription of the gene may be repressed. DNA methylation is critical for normal development and abnormality in methylation may disrupt epigenetic regulation. The disruption, e.g., repression, in epigenetic regulation may cause diseases, such as cancer. Promoter methylation in DNA may be indicative of cancer.
[0086] The term “hypermethylation” refers to an increased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypermethylated DNA can include DNA molecules comprising at least 1 methylated residue, at least 2 methylated residues, at least 3 methylated residues, at least 5 methylated residues, or at least 10 methylated residues.
[0087] The term “hypomethylation” refers to a decreased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules comprising 0 methylated residues, at most 1 methylated residue, at most 2 methylated residues, at most 3 methylated residues, at most 4 methylated residues, or at most 5 methylated residues.
[0088] The terms “or a combination thereof’ and “or combinations thereof’ as used herein refers to any and all permutations and combinations of the listed terms preceding the term. For example, “A, B, C, or combinations thereof’ is intended to include at least one of: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CAB ABB, and so forth. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.
[0089] “Or” is used in the inclusive sense, i.e., equivalent to “and / or,” unless the context requires otherwise.
[0090] ‘ ‘Derived from Cancer-Free Samples” as used herein refers to a threshold or cutoff value that is established based on data obtained from samples known to be free of cancer. By analyzing a broad range of cancer-free samples, one can identify what constitutes a "normal" range for various biomarkers, genetic sequences, or other measurable factors. This normal range can then serve as a baseline against which test results from potentially cancerous samples are compared.B. Methods
[0091] Individuals with pathogenic germline variants are at an increased risk of developing cancers and / or other diseases. A germline variant occurring in a sperm or egg cell can be copied into other cells in the body during human development. Since the variant is pathogenic in reproductive cells, it can be hereditary, passing from one generation to another.
[0092] While germline genetic testing is available to determine whether an individual has a pathogenic germline variant, family and personal histories may also be indicative of elevated cancer or disease risks. Such histories may include sex, age, high incidence of cancer with family members, early onset of cancer (e.g., in individuals younger than 50), rare cancers, certain pathologies (e.g., triple-negative breast cancer or microsatellite unstable cancer), more than one primary cancer in one person, evidence of autosomal-dominant inheritance (e.g. two or more generations affected, with both sexes affected), pattern of cancer associated with known cancer syndromes, and presence of premalignant conditions.
[0093] Identifying a pathogenic germline variant can impact and inform monitoring and treatment options for the individual. The approaches disclosed herein provide methods for screening for cancers or disease of individuals who have been determined to have a pathogenic germline variant or are suspected of having such variants. The methods can be used to monitor biomarkers collected from samples of an at-risk subject over time to determine whether a cancer or disease is present or absent.
[0094] In some embodiments, one or more biomarkers are collected from a sample from the subject and are analyzed to establish a baseline cancer signal. At a subsequent time point, an additional sample is taken from the subject where one or more biomarkers are analyzed to determine if the level of cancer or the rate of change of cancer levels has either increased above a pre-set threshold, indicating a presence of cancer, or is at or below the pre-set threshold, indicating that an absence of cancer. Multiple samples at multiple time points can be collected from thesubject. In certain embodiments, the threshold can distinguish not only the presence or absence of cancer, but also whether the disease is progressing or stable over time. Assessment of the rate of change in cancer levels thereby provides a dynamic measure of disease trajectory rather than a single static measurement.
[0095] In some embodiments, analysis of the one or more biomarkers is carried out using a non- invasive assay, such as a liquid biopsy. In some embodiments, the analysis is personalized or tailored to target specific biomarkers associated with hereditary cancer or disease risk. For example, in some embodiments, a subject with Lynch syndrome, biomarkers associated with colorectal cancer may be targeted, such as MLH1, MSH2, MSH6 PMS2, and EPCAM genes. In some embodiments, the analysis can be configured target biomarkers associated with one or more primary cancers and / or one or more secondary cancers.
[0096] In some embodiments, the one or more biomarkers analyzed in the first and / or additional bodily fluid samples include one or more genes associated with the cancer type(s) for which the subject is determined or suspected to be at increased risk of developing.
[0097] In certain embodiments, these biomarkers comprise hereditary cancer susceptibility genes, in which germline variants confer risk, and which may also be used for subsequent monitoring of disease progression through detection of somatic “second hits” alterations. Biomarkers associated with certain cancer types may include the following from Table 1:
[0098] Table 1: Cancer types and associated biomarkers.
[0099] In some embodiments, the subject is known or suspected of having constitutional mismatch repair deficiency (CMMRD) syndrome. Cancers that are common in individuals with CMMRD are colon, colorectal, brain and blood (leukemia or lymphoma). Mutations involved in CMMRD are PMS2, MLH1, MSH2 and MSH6.1. Sample and Subjects
[0100] The subject may be a human, a mammal, an animal, a primate, rodent (including mice and rats), or other common laboratory, domestic, companion, service or agricultural animal, for example a rabbit, dog, cat, horse, cow, sheep, goat or pig. The subject may in some cases have or be suspected of having a pathogenic germline variant. In other cases, the subject may not have cancer or a detectable cancer symptom. The subject may have been treated with one or more cancer therapies, e.g., any one or more of chemotherapies, antibodies, vaccines or biologies. The subject may be in remission, e.g. from a tumor, cancer, or neoplasia (e.g., following treatment such as chemotherapy, surgical resection, radiation, or a combination thereof). The subject may or may not be diagnosed as being susceptible to cancer or any cancer-associated genetic mutations / di sorders .
[0101] The sample can be any biological sample isolated from a subject. The sample can be a bodily sample. Samples can include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, bone marrow, pleural effusions, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, urine. Samples are preferably body fluids, particularly blood and fractions thereof, and urine. A sample can be in the form originally isolated from a subject or can have been subjected to further processing to remove or add components, such as cells, or enrich for one component relative to another. A sample can be isolated or obtained from a subject and transported to a site of sample analysis. The sample may be preserved and shipped at a desirable temperature, e.g., room temperature, 4°C, -20°C, or -80°C. A sample can be isolated or obtained from a subject at the site of the sample analysis.
[0102] The sample may be plasma. The volume of plasma can depend on the desired read depth for sequenced regions. Exemplary volumes are 0.4-40 ml, 5-20 ml, 10-20 ml. For examples, the volume can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, or 40 ml. A volume of sampled plasma may be 5 to 20 ml.
[0103] A sample can comprise various amount of nucleic acid that contains genome equivalents. For example, a sample of about 30 ng DNA can contain about 10,000 ( 104) haploid human genome equivalents and, in the case of cell free DNA (cfDNA), about 200 billion (2xlOn) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents and, in the case of cfDNA, about 600 billion (6 x 1011) individual molecules.
[0104] A sample can comprise nucleic acids from different sources, e.g., cellular DNA and cell- free DNA of the same subject, or cellular DNA and cell-free DNA of different subjects. A sample can comprise nucleic acids carrying mutations. For example, a sample can comprise DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations existing in germline DNA of a subject. Somatic mutations refer to mutations originating in somatic cells of a subject, e.g., cancer cells. A sample can comprise DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations). A sample can comprise an epigenetic variant (i.e. a chemical or protein modification), wherein the epigenetic variant associated with the presence of a genetic variant such as a cancer-associated mutation. In some embodiments, the sample comprises an epigenetic variant associated with the presence of a genetic variant, wherein the sample does not comprise the genetic variant.
[0105] The sample may comprise cell-free nucleic acids, such as cfDNA. The cfDNA may be obtained from a test subject, for example as described above. For example, the sample for analysis may be plasma or serum containing cell-free nucleic acids. “Cell-free DNA” “cfDNA molecules,” or “cfDNA”, for example, include DNA molecules that naturally occur in a subject in extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). While the cfDNA originally existed in a cell or cells in a large complex biological organism, e.g., a mammal, it has undergone release from the cell(s) in vivo into a fluid found in the organism, and may be obtained by obtaining a sample of the fluid without the need to perform an in vitro cell lysis step. In other words, cell-free nucleic acids or cfDNA are nucleic acids or cfDNA not contained within or otherwise bound to a cell. Cell-free nucleic acids include DNA, RNA, and hybrids thereof, including cfDNA derived from genomic DNA, mitochondrial DNA,siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi- interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or a hybrid thereof. A cell-free nucleic acid can be released into bodily fluid through secretion or cell death processes, e.g., cellular necrosis and apoptosis. Some cell-free nucleic acids are released into bodily fluid from cancer cells e g., circulating tumor DNA, (ctDNA). Others are released from healthy cells. Tn some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell free nucleic acids are produced by tumor cells. In some embodiments, cell free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.
[0106] Exemplary amounts of cell-free nucleic acids in a sample before amplification range from about 1 fg to about 1 pg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acids. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell- free nucleic acids. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acids. The method can comprise obtaining 1 femtogram (fg) to 200 ng cell-free nucleic acids from samples.
[0107] Cell-free nucleic acids have an exemplary size distribution of about 100-500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of molecules, with a mode of about 168 nucleotides and a second minor peak in a range between 240 to 440 nucleotides.
[0108] Cell-free nucleic acids can be isolated from bodily fluids through a fractionation step in which cell-free nucleic acids, as found in solution, are separated from intact cells and other nonsoluble components of the bodily fluid. Fractionation may include techniques such as centrifugation or filtration. Alternatively, cells in bodily fluids can be lysed and cell-free and cellular nucleic acids processed together. Generally, after addition of buffers and wash steps, nucleic acids can be precipitated with an alcohol. Further clean up steps may be used such as silica- based columns to remove contaminants or salts. Non-specific bulk carrier nucleic acids, such as Cl DNA, DNA or protein for hybridization may be added throughout the reaction to optimize certain aspects of the procedure such as yield.
[0109] After such processing, samples can include various forms of nucleic acid including double stranded DNA, single stranded DNA and single stranded RNA. In some embodiments, singlestranded DNA and RNA can be converted to double stranded forms so they are included in subsequent processing and analysis steps.
[0110] In some embodiments of the disclosed methods, the population of target nucleic acids comprises RNA and the method further comprises a cDNA synthesis step. RNA for use in the methods disclosed herein may be isolated from a blood sample or a sample comprising cells (such as a sample that includes immune and / or cancer-derived cells (e g., a blood sample such as a whole blood sample, a buffy coat sample, a leukapheresis sample, or a peripheral blood PBMC sample)). General methods for RNA extraction and isolation (such as mRNA extraction and isolation) are known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al., Current Protocols of Molecular Biology, John Wiley and Sons (1997). Methods for RNA extraction from paraffin embedded tissues are disclosed, for example, in Rupp and Locker, Lab Invest. 56:A67 (1987), and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using a purification kit, buffer set, and protease(s) from commercial manufacturers, such as PreAnalytix GmbH or Qiagen, according to the manufacturer’s instructions. For example, RNA can be extracted from whole blood samples using the PAXgene® Blood RNA Kit (PreAnalytix GmbH). Other commercially available RNA isolation kits include MasterPure Complete DNA and RNA Purification Kit (EPICENTRE, Madison, WI), and Paraffin Block RNA Isolation Kit (Ambion, Inc.). Total RNA from tissue samples can be isolated using RNA Stat-60 (Tel -Test). RNA prepared from tumor tissue can be isolated, for example, by cesium chloride density gradient centrifugation.
[0111] Following RNA extraction from a sample (such as a blood sample), a cDNA library is typically prepared in preparation for sequencing, e.g., as in RNA-Seq. In some embodiments, the cDNAs in a library, such as an RNA-Seq library, can comprise a cDNA insert flanked by adapter sequences, such as adapter sequences used for amplification and sequencing on a particular platform. Exemplary cDNA library preparation methods are discussed below; however, cDNA library preparation methods can vary depending on the RNA species under investigation, which can differ in size, sequence, structural features and abundance. One of ordinary skill in the art will be able to select cDNA library preparation methods suitable for cDNA library preparation using an RNA species of interest.
[0112] Ribosomal RNAs (rRNAs) are the most abundant RNA species in most cells. Globin mRNA is also abundant in certain cell types found in the blood. Thus, some embodiments of the present disclosure comprise a step of ribosomal RNA (rRNA) depletion and / or a step of globinmRNA depletion. Such steps can be performed, e.g., following RNA extraction from a sample, and prior to a step of RNA fragmentation or cDNA fragmentation, prior to a step preparing cDNA from the RNA, prior to a step of ligating adapters to the cDNA, and prior to a sequencing step. In some embodiments, the methods include a step of rRNA depletion. In other embodiments, the methods include a step of globin mRNA depletion. In yet other embodiments, the methods disclosed herein include both a step of rRNA depletion and a step of globin mRNA depletion.
[0113] Any suitable rRNA depletion and / or globin mRNA depletion methods are of use in the present disclosure. One approach is to eliminate rRNAs uses sequence-specific probes that can hybridize to rRNAs (Hrdlickova et al., Wiley Interdiscip Rev RNA. 2017; 8(1): 10.1002 / wrna.1364). Unwanted rRNAs or their cDNAs are hybridized with biotinylated DNA or locked nucleic acid (LNA) probes, followed by depletion with streptavidin beads. Alternatively, rRNAs can be targeted by anti-sense DNA oligos and digested by RNase H, a method also known as probe-directed degradation (PDD). Another approach for rRNA reduction uses specific, not-so-random (NSR) primers that bind to the RNA molecules of interest during reverse transcription, thus avoiding reverse transcription of the rRNAs. For example, a method known as Ovation RNA-Seq (NuGen) uses hexamer or heptamer primers whose sequences are not present in rRNAs. In addition to sequence-based approaches, some methods take advantage of certain features of rRNAs for their elimination. The COT-hybridization method is based on heat denaturation, re-annealing, and selective degradation by a duplex-specific nuclease (DSN). Double-stranded cDNAs from abundant sequences are preferentially degraded because of their more rapid annealing kinetics compared to less abundant ones. Selective degradation has also been achieved using the enzyme terminator 5 ’-phosphate-dependent exonuclease (TEX), which recognizes RNA molecules with 5 ’-monophosphate, as with rRNAs and tRNAs. Further, commercial kits are available for rRNA and globin mRNA depletion, including, e.g., the Watchmaker Genomics RNA Library Prep Kit with Polaris Depletion.
[0114] Other embodiments of the present disclosure comprise a step of poly(A) selection. Such a step can be performed, e.g., following RNA extraction from a sample, and prior to a step of RNA fragmentation or cDNA fragmentation, prior to a step preparing cDNA from the RNA, prior to a step of ligating adapters to the cDNA, and prior to a sequencing step. In eukaryotic organisms, most protein coding RNAs (mRNAs) and many long noncoding RNAs (IncRNAs) (>200 nt) comprise a poly(A) tail (“polyadenylated RNAs”). The poly(A) tail may be used to enrich for polyadenylated RNAs from total cellular RNA, in which polyadenylated RNAs may account forapproximately 1-5% of total cellular RNA (Hrdlickova et al., Wiley Interdiscip Rev RNA. 2017;8(l): 10.1002 / wrna.1364). Exemplary poly(A) selection methods include, but are not limited to, use of magnetic or cellulose beads coated with oligo-dT molecules. Alternatively, polyadenylated RNAs can be selected using oligo-dT priming for reverse transcription (RT). Poly(A) selection may be combined with globin mRNA depletion.
[0115] In some embodiments, one or more analytes from the same sample are analyzed. In some embodiments, the one or more analytes may comprise cell-free deoxyribonucleic acid (cfDNA), cell-free ribonucleic acid (cfRNA), proteins, and / or exosomes. In some embodiments, the one or more analytes are analyzed for different biomarkers or signatures, including somatic variants, methylation status, fragmentomic patterns, transcription factor binding sites, and / or chromatin interaction sites.2. Methylation Workflows
[0116] In some embodiments of the present disclosure, nucleic acid molecules in the sample are partitioned into two or more partitions. In some embodiments, the sample of nucleic acids has been subjected to a methylation-based partitioning assay. In some embodiments, the nucleic acid sample is partitioned based on the modification status of nucleic acids within the nucleic acid sample.
[0117] In such methods, different forms of DNA (e.g., hypermethylated and hypomethylated DNA) can be physically partitioned based on one or more characteristics of the DNA. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated. In some embodiments, the sample of parent nucleic acids are subjected to a methylation-based partitioning assay, wherein the methylation-based partitioning assay partitions nucleic acids using methyl-binding domain (MBD). In such methods, the methylation-based partitioning assay can form a hypermethylated partition and / or a hypomethylated partition. After partitioning, one or more of the resulting partitions can be analyzed by the methods disclosed herein. In some embodiments, the resulting partitions analyzed can include a hypermethylated partition obtained from the methylation-based partitioning assay. In some embodiments, the resulting partitions analyzed can include a hypomethylated partition obtained from the methylation-based partitioning assay. Such methods can further comprise detecting: (i) a quantitative measure indicative of a number of nucleic acids in the hypermethylated partition; and / or (ii) a quantitative measure indicative of a number of nucleic acids in the hypomethylated partition derived from a genomic region in the sample by determining a normalized quantitative measure at one or more genomic regions and determining a methylation level at that genomicregion based on the normalized quantitative measure. The methylation level can be determined, for example, by comparing the normalized quantitative measure in the hypermethylated partition to the normalized quantitative measure from the hypomethylated partition. The methylation level can be determined, for example, by comparing the normalized quantitative measure in the hypermethylated partition and / or the normalized quantitative measure from the hypomethylated partition to a reference value. The reference value may be, for example, derived from a normalized quantitative measure of a control genomic region from the same partition.
[0118] Partitioning may include physically partitioning nucleic acids into partitions based on the presence or absence of one or more methylated nucleobases. A sample may be partitioned into partitions based on a characteristic that is indicative of differential gene expression or a disease state. A sample may be partitioned based on a characteristic that provides a difference in signal between a normal and diseased state during analysis of nucleic acids, e.g., cell free DNA (cfDNA), non-cfDNA, tumor DNA, circulating tumor DNA (ctDNA) and cell free nucleic acids (cfNA).
[0119] In some instances, a nucleic acid sample is partitioned into two or more partitions (e g., at least 3, 4, 5, 6 or 7 partitions). The agents used to partition populations of nucleic acids within a sample can be affinity agents, such as antibodies with the desired specificity, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected e.g., by phage display to have specificity to a given target. In some embodiments, the agent used in the partitioning is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine, such as a methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is a product of a procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA of the sample. In some embodiments, the modified nucleobase may be a “converted nucleobase,” meaning that its base pairing specificity was changed by a procedure. For example, certain procedures convert unmethylated or unmodified cytosine to dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, resulting in uracil (considered a modified nucleobase in the context of DNA) or a further modified form of uracil. Examples of partitioning agents include antibodies, such as antibodies that recognize a modified nucleobase, which may be a modified cytosine, such as a methyl cytosine (e.g., 5-methylcytosine). In some embodiments, the partitioning agent is an antibody that recognizes a modified cytosine other than 5-methylcytosine, such as 5-carboxylcytosine (5caC). Alternative partitioning agentsinclude methyl binding domain (MBDs) and methyl binding proteins (MBPs), including proteins such as MeCP2.
[0120] Additional, non-limiting examples of partitioning agents are histone binding proteins which can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48 and SANT domain peptides.
[0121] In some embodiments, partitioning can comprise both binary partitioning and partitioning based on degree / level of modifications. For example, methylated fragments can be partitioned by methylated DNA immunoprecipitation (MeDIP), or all methylated fragments can be partitioned from unmethylated fragments using methyl binding domain proteins (e.g., Methyl Minder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning may involve eluting fragments having different levels of methylation by adjusting the salt concentration in a solution with the methyl binding domain and bound fragments. As salt concentration increases, fragments having greater methylation levels are eluted.
[0122] Various levels of methylation can be partitioned using sequential elutions. For example, a hypomethylated partition (no methylation) can be separated from a methylated partition by contacting the nucleic acid population with MBD, such as MBD attached to magnetic beads. The beads can be used to separate out the methylated nucleic acids from the nonmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids having different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After such methylated nucleic acids are eluted, magnetic separation can once again be used to separate higher level of methylated nucleic acids from those with lower level of methylation. The elution and magnetic separation steps can be repeated to create various partitions such as a hypomethylated partition (enriched in nucleic acids comprising no methylation), a methylated partition (enriched in nucleic acids comprising low levels of methylation), and a hypermethylated partition (enriched in nucleic acids comprising high levels of methylation). Any one or more partitions can then be analyzed using the methods disclosed herein.
[0123] In some methods, nucleic acids bound to an agent used for affinity separation-based partitioning are subjected to a wash step. The wash step washes off nucleic acids weakly bound to the affinity agent. Such nucleic acids can be enriched in nucleic acids having the modification toan extent close to the mean or median (i.e. intermediate between nucleic acids remaining bound to the solid phase and nucleic acids not binding to the solid phase on initial contacting of the sample with the agent).
[0124] For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0125] In some embodiments, the nucleic acids can be partitioned into different partitions based on the nucleic acids that are bound to a specific protein or a fragment thereof and those that are not bound to that specific protein or fragment thereof.
[0126] Nucleic acids can be partitioned based on DNA-protein binding. Protein-DNA complexes can be partitioned based on a specific property of a protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation) or enzymatic activity. Examples of proteins which may bind to DNA and serve as a basis for fractionation may include, but are not limited to, protein A and protein G. Any suitable method can be used to partition the nucleic acids based on protein bound regions. Examples of methods used to partition nucleic acids based on protein bound regions include, but are not limited to, SDS-PAGE, chromatin-immunoprecipitation (ChIP), heparin chromatography, and asymmetrical field flow fractionation (AF4).
[0127] In some embodiments, the partitioning is performed by contacting the nucleic acids with a methyl binding domain (“MBD”) of a methyl binding protein (“MBP”). In some such embodiments, the nucleic acids are contacted with an entire MBP. In some embodiments, an MBD binds to 5-methylcytosine (5mC), and an MBP comprises an MBD and is referred to interchangeably herein as a methyl binding protein or a methyl binding domain protein. In some embodiments, MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different extents of methylation can be performed by eluting fractions by increasing the NaCl concentration.
[0128] In some embodiments, bound DNA is eluted by contacting the antibody or MBD with a protease, such as proteinase K. This may be performed instead of or in addition to elution steps using NaCl as discussed above.
[0129] Examples of agents that recognize a modified nucleobase contemplated herein include, but are not limited to: (a) MeCP2 is a protein that preferentially binds to 5-methyl-cytosine over unmodified cytosine, (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 preferentially bind to 5- hydroxymethyl-cytosine over unmodified cytosine, (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind to 5-formyl -cytosine over unmodified cytosine (lurlaro et al.,Genome Biol. 14: R119 (2013)), and (d) antibodies specific to one or more methylated or modified nucleobases or conversion products thereof, such as 5mC, 5caC, or DHU.
[0130] In general, elution is a function of the number of modifications, such as the number of methylated sites per molecule, with molecules having more methylation eluting under increased salt concentrations. To elute the DNA into distinct populations based on the extent of methylation, one can use a series of elution buffers of increasing NaCl concentration. Salt concentration can range from about 100 nm to about 2500 mMNaCl. In one embodiment, the process results in three (3) partitions. Molecules are contacted with a solution at a first salt concentration and comprising a molecule comprising an agent that recognizes a modified nucleobase, which molecule can be attached to a capture moiety, such as streptavidin. At the first salt concentration a population of molecules will bind to the agent and a population will remain unbound. The unbound population can be separated as a “hypomethylated” population. For example, a first partition enriched in hypomethylated form of DNA is that which remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. A second partition (a residual partition) enriched in intermediate methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM concentration. This is also separated from the sample. A third partition enriched in hypermethylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.
[0131] In some embodiments, the partitioned nucleic acids can be contacted with a methylation sensitive restriction enzyme (MSRE) and / or a methylation dependent restriction enzyme (MORE). In one embodiment, a partition which is enriched for methylated nucleic acids (e.g. a hypermethylated partition and / or a residual partition) is treated with an MSRE such that unmethylated nucleic acids within the partition are digested. This can reduce the number of incorrectly partitioned nucleic acids in the partition enriched for methylated nucleic acids. In one embodiment, a partition which is unenriched for methylated nucleic acids (e.g. the hypomethylated partition) can be treated with an MORE such that methylated nucleic acids within the partition are digested. This can reduce the number of incorrectly partitioned nucleic acids in the partition enriched for unmethylated nucleic acids.
[0132] In some embodiments, a monoclonal antibody raised against 5-methylcytidine (5mC) is used to purify methylated DNA. DNA is denatured, e.g., at 95°C in order to yield single-stranded DNA fragments. Protein G coupled to standard or magnetic beads as well as washes following incubation with the anti-5mC antibody are used to immunoprecipitate DNA bound to the antibody.Such DNA may then be eluted. Partitions may comprise unprecipitated DNA and one or more partitions eluted from the beads.
[0133] In some embodiments, the nucleic acids of the nucleic acid sample may be exposed to methylation-sensitive restriction enzymes (MRSEs). Such restriction enzymes do not cleave methylated residues, leaving only the methylated nucleic acids of the nucleic acid sample intact. Exposure to such restriction enzymes would result in a sample of only methylated nucleic acids, which could then be analyzed by the disclosed methods. Exposure to differential concentrations of MRSEs would result in subsamples that contain nucleic acids increasingly enriched for hypermethylated nucleic acids.
[0134] In some embodiments, the nucleic acids may be subjected to a conversion-based procedure to identify the modification status of the nucleobases. Such conversion procedures can comprise subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA. In some embodiments, methods disclosed herein comprise a step of subjecting DNA, or a subsample thereof, to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, the procedure chemically converts the first or second nucleobase such that the base pairing specificity of the converted nucleobase is altered. In some embodiments, DNA is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA before library preparation using the DNA, before a first amplification of the DNA and / or before the ligation of adapters. In certain embodiments, the DNA is subjected to the procedure before or after contacting the DNA with a methylationsensitive nuclease.
[0135] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formyl cytosine (fC) or 5-carboxylcytosine (caC)) to uracil whereas other modified cytosines (e.g., 5- methylcytosine, 5-hydroxylmethylcystosine) are not converted. Thus, where bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, 5-formyl cytosine, 5- carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleobase may comprise one or more of 5-methyl cytosine (mC) and 5-hydroxymethylcytosine (hmC), such asmC and optionally hmC. Sequencing of bisulfite-treated DNA identifies positions that are read as cytosine as being mC or hmC positions. Meanwhile, positions that are read as T are identified as being T or a bisulfite-susceptible form of C, such as unmodified cytosine, 5-formyl cytosine, or 5- carboxylcytosine. Performing bisulfite conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing mC or hmC using the sequence reads obtained from the exemplary sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.
[0136] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to fC, which is bisulfite susceptible, followed by bisulfite conversion. Thus, when oxidative bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises mC. Sequencing of Ox-BS converted DNA identifies positions that are read as cytosine as being mC positions. Meanwhile, positions that are read as T are identified as being T, hmC, or a bisulfite-susceptible form of C, such as unmodified cytosine, fC, or hmC. Performing Ox-BS conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing mC using the sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, e.g., Booth et al., Science 2012; 336: 934-937.
[0137] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion and mC is oxidized in advance of bisulfite treatment, so that positions originally occupied by mC are converted to U while positions originally occupied by hmC remain as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, P-glucosyl transferase can be used to protect hmC (forming 5- glucosylhydroxymethyl cytosine (ghmC)), then a TET protein such as mTetl can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while ghmC remains unaffected. Alternatively, a carbamoyltransferase enzyme, such as 5- hydroxymethylcytosine carbamoyltransferase as described in Yang et al., Bio-protocol, 2023; 12(17): e4496, can be used to protect hmC (by converting hmC to 5-carbamoyloxymethylcytosine (5cmC)), then a TET protein such as mTetl can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while 5cmC remains unaffected. Thus, when TAB conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC,caC, mC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises hmC. Sequencing of TAB -converted DNA identifies positions that are read as cytosine as being hmC positions. Meanwhile, positions that are read as T are identified as being T, mC, or a bisulfite- susceptible form of C, such as unmodified cytosine, fC, or caC. Performing TAB conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing hmC using the sequence reads obtained from the sample.
[0138] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In Tet-assisted pic-borane conversion with a substituted borane reducing agent conversion, a TET protein is used to convert mC and hmC to caC, without affecting unmodified C. caC, and fC if present, are then converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting unmodified C. See, e.g., Liu et al., Nature Biotechnology’ 2019; 37:424-429 (e.g., at Supplementary Fig. 1 and Supplementary Note 7). DHU is read as a T in sequencing. Thus, when this type of conversion is used, the first nucleobase comprises one or more of mC, fC, caC, or hmC, and the second nucleobase comprises unmodified cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T, mC, fC, caC, or hmC. Performing TAP conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing unmodified C using the sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), described in further detail in Liu et al. 2019, supra.
[0139] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises chemical-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In chemical-assisted conversion with a substituted borane reducing agent, an oxidizing agent such as potassium perruthenate (KRuCL) (also suitable for use in ox-BS conversion) is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane converts fC and caC to DHU but does notaffect mC or unmodified C. Thus, when this type of conversion is used, the first nucleobase comprises one or more of hmC, fC, and caC, and the second nucleobase comprises one or more of unmodified cytosine or mC, such as unmodified cytosine and optionally mC. Sequencing of the converted DNA identifies positions that are read as cytosine as being either mC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T, fC, caC, or hmC. Performing this type of conversion, such as on a DNA sample as described herein, thus facilitates distinguishing positions containing unmodified C or mC on the one hand from positions containing hmC using the sequence reads obtained from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429. 5- hydroxymethylcytosine carbamoyltransferase is described in Yang et al., Bio-protocol, 2023; 12(17): e4496.
[0140] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises APOBEC-coupled epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme such as APOBEC3A (A3A) is used to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when ACE conversion is used, the first nucleobase comprises unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase comprises hmC. Sequencing of ACE-converted DNA identifies positions that are read as cytosine as being hmC, fC, or caC positions. Meanwhile, positions that are read as T are identified as being T, unmodified C, or mC. Performing ACE conversion on a DNA sample as described herein thus facilitates distinguishing positions containing hmC from positions containing mC or unmodified C using the sequence reads obtained from the sample. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0141] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv, DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692vl . For example, TET2 and T4- GT or 5 -hydroxymethyl cytosine carbamoyltransferase (described in Yang et al., Bio-protocol, 2023; 12(17): e4496) can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), and then a deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosines converting them to uracils.
[0142] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using a non-specific, modification-sensitive double-stranded DNA deaminase, e.g., as in SEM- seq. See, e.g., Vaisvila et al. (2023) Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA. bioRxiv, DOI: 10.1101 / 2023.06.29.547047, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047vl. SEM-Seq employs a nonspecific, modification-sensitive double-stranded DNA deaminase (MsddA) in a nondestructive single-enzyme 5-methylctyosine sequencing (SEM-seq) method that deaminates unmodified cytosines. Accordingly, SEM-seq does not require the TET2 and T4-PGT or 5- hydroxymethylcytosine carbamoyltransferase protection and denaturing steps that are of use, e.g., in APOEC3A-based protocols. Additionally, MsddA does not deaminate 5-formylated cytosines (5fC) or 5-carboxylated cytosines (5caC). In SEM-seq, unmodified cytosines in the DNA are deaminated to uracil and is read as “T” during sequencing. Modified cytosines (e.g., 5mC) are not converted and are read as “C” during sequencing. Cytosines that are read as thymines are identified as unmodified (e.g., unmethylated) cytosines or as thymines in the DNA. Performing SEM-seq conversion thus facilitates identifying positions containing 5mC using the sequence reads obtained. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using MsddA.
[0143] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase. In some such embodiments, the first nucleobase is hmC. DNA originally comprising the first nucleobase may be separated from other DNA using a labeling procedure comprising biotinylating positions that originally comprised the first nucleobase. In some embodiments, the first nucleobase is first derivatized with an azide-containing moiety, such as a glucosyl -azide containing moiety. The azide-containing moiety then may serve as a reagent for attaching biotin, e.g., through Huisgen cycloaddition chemistry. Then, the DNA originally comprising the first nucleobase, now biotinylated, can be separated from DNA not originally comprising the first nucleobase using a biotin-binding agent, such as avidin, neutravidin (deglycosylated avidin with an isoelectric point of about 6.3), orstreptavidin. An example of a procedure for separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase is hmC-seal, which labels hmC to form P-6-azide-glucosyl-5-hydroxymethylcytosine and then attaches a biotin moiety through Huisgen cycloaddition, followed by separation of the biotinylated DNA from other DNA using a biotin-binding agent. For an exemplary description of hmC-seal, see, e.g., Han et al., Mol. Cell 2016; 63 : 711 -719. This approach is useful for identifying fragments that include one or more hmC nucleobases.3. Tagging
[0144] In some embodiments, nucleic acids of the sample may be tagged with sample indexes, partition tags and / or molecular barcodes (referred to generally as “tags”). Tags can form part of an adapter.
[0145] Tags can be molecules, such as nucleic acids, containing information that indicates a feature of the molecule with which the tag is associated. For example, molecules can bear a sample tag or sample index (which distinguishes molecules in one sample from those in a different sample), a partition tag (which distinguishes molecules in one partition from those in a different partition) and / or a molecular barcode (which distinguishes different molecules from one another (in both unique and non-unique tagging scenarios). In certain embodiments, a tag can comprise one or a combination of barcodes.
[0146] Optionally, adapters may contain a partition-specific barcode and / or a molecular barcode. As used herein, the term “barcode” refers to a nucleic acid molecule having a particular nucleotide sequence, or to the nucleotide sequence, itself, depending on context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or can have sequences having a certain Hamming distance, as desired for the specific purpose. So, for example, a molecular barcode can be comprised of one barcode or a combination of two barcodes, each attached to different ends of a molecule. Additionally, or alternatively, for different partitions, different sets of molecular barcodes can be used such that the barcodes serve as a molecular tag through their individual sequences and also serve to identify the partition to which they correspond based the set of which they are a member.
[0147] Tags may be incorporated into or otherwise joined to adapters by chemical synthesis, ligation (e.g., as described above, e.g. by blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR), among other methods. Such adapters are ultimately joined to the parent nucleic acids. In other embodiments, one or more rounds of amplificationcycles (e.g., PCR amplification) may be applied to introduce sample indexes to a nucleic acid using conventional nucleic acid amplification methods. The amplifications may be conducted in one or more reaction mixtures (e.g., a plurality of microwells in an array). Molecular barcodes, partition tags and / or sample indexes may be introduced simultaneously, or in any sequential order. In some embodiments, molecular barcodes and / or sample indexes are introduced prior to and / or after a partitioning procedure. Tn some embodiments, molecular barcodes and / or sample indexes are introduced prior to and / or after sequence capturing steps, if present, are performed. In some embodiments, only the molecular barcodes are introduced prior to probe capturing and the sample indexes are introduced after sequence capturing steps are performed. In some embodiments, both the molecular barcodes and the sample indexes are introduced prior to performing probe-based sequence capturing steps, if present. In some embodiments, the sample indexes are introduced after sequence capturing steps are performed, if present. In some embodiments, sample indexes are incorporated through overlap extension polymerase chain reaction (PCR).
[0148] In some embodiments, the tags may be located at one end or at both ends of the sample nucleic acids. In some embodiments, tags are predetermined or random or semi-random sequences. In some embodiments, the tag(s) may together be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, or 5 nucleotides in length. Typically, tags are about 5 to 20 or 6 to 15 nucleotides in length. The tags may be linked to sample nucleic acids randomly or non-randomly.
[0149] In some embodiments, each sample is distinctly tagged with a sample index or a combination of sample indexes. In some examples, when multiple partitions are subsequently processed after the partitioning step, each partition can be distinctly tagged with a partition tag or a combination of partition tags. In some embodiments, each nucleic acid of a sample or subsample is uniquely tagged with a molecular barcode or a combination of molecular barcodes. In other embodiments, a plurality of molecular barcodes may be used such that molecular barcodes are not necessarily unique to one another in the plurality (e.g., non-unique molecular barcodes). In these embodiments, molecular barcodes are generally attached (e.g., by ligation) to individual nucleic acids such that the combination of the molecular barcode and the sequence of the sample nucleic acid that it is attached to creates a unique sequence that may be used for grouping the sequence reads into families, wherein a family corresponds to sequence reads derived from the same parent nucleic acid. Detection of non-unique molecular barcodes in combination with endogenous sequence information typically allows for the assignment of a unique identity to a particular molecule. Endogenous sequence information which can be used for grouping the sequence readsinto families includes the beginning (start) and / or end (stop) genomic location / position corresponding to the sequence of the parent nucleic acid in the sample, start and stop genomic positions corresponding to the sequence of the parent nucleic acid in the sample, the beginning (start) and / or end (stop) genomic location / position of the sequence read that is mapped to the reference sequence, start and stop genomic positions of the sequence read that is mapped to the reference sequence, sub-sequences of sequence reads at one or both ends, length of sequence reads, and / or length of the parent nucleic acids in the sample. In some embodiments, the beginning region comprises the first 5, the first 10, the first 15, the first 20, the first 25, the first 30 or at least the first 30 base positions at the 5' end of the sequencing read that align to the reference sequence. In some embodiments, the end region comprises the last 5, the last 10, the last 15, the last 20, the last 25, the last 30 or at least the last 30 base positions at the 3' end of the sequencing read that align to the reference sequence. The length, or number of base pairs, of an individual sequence read are also optionally used for grouping the sequence reads into families, wherein a family corresponds to sequence reads derived from the same parent nucleic acid. Methylation information comprises within sequence reads, for example after bisulfite sequencing, are also optionally used for grouping the sequence reads into families, wherein a family corresponds to sequence reads derived from the same parent nucleic acid.
[0150] In certain embodiments, the number of different tags used to uniquely identify a number of molecules, z, in a class can be between any of 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, 9*z, 10*z, 11 *z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z or 100*z (e.g., lower limit) and any of 100,000*z, 10,000*z, 1000*z or 100*z (e.g., upper limit). In some embodiments, molecular barcodes are introduced at an expected ratio of a set of identifiers (e.g., a combination of unique or non-unique molecular barcodes) to molecules in a sample. One example format uses from about 2 to about 1,000,000 different molecular barcode sequences, or from about 5 to about 150 different molecular barcode sequences, or from about 20 to about 50 different molecular barcode sequences, ligated to both ends of a target molecule. Alternatively, from about 25 to about 1,000,000 different molecular barcode sequences may be used. For example, 20-50 x 20-50 molecular barcode sequences (i.e., one of the 20-50 different molecular barcode sequences can be attached to each end of the parent nucleic acids) can be used. Such numbers of identifiers are typically sufficient for different molecules having the same start and stop points to have a high probability of receiving different combinations of identifiers.
[0151] In some embodiments, the assignment of unique or non-unique molecular barcodes in reactions is performed using methods and systems described in, for example, U.S. Patent Application Nos. 2001 / 0053519, 2003 / 0152490, and 2011 / 0160078, and U.S. Patent Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992. Alternatively, in some embodiments, grouping of sequence reads into families can be performed using only endogenous sequence information (e.g., start and / or stop positions, sub-sequences of one or both ends of a sequence, and / or lengths). Alternatively or additionally, in some embodiments, grouping of sequence reads into families can be performed using methylation status information, optionally in combination with other features. For example, conversion-based methylation sequencing (e.g. bisulfite sequencing) can change the base pairing specificity of bases in the parent nucleic acids depending on their methylation status, ultimately resulting in the sequence reads comprising a different nucleotide at the position of the converted base. This difference would be expected to be present all sequence reads derived from the same parent nucleic acid that had been subjected to the conversion procedure and hence can be used to group sequence reads into families. The addition of tags (e.g. sample indexes, partition and / or sub-partition tags and / or molecular barcodes) to parent nucleic acids can be done through amplification, wherein the tags are comprised in primers used for amplification.
[0152] In some embodiments, the nucleic acids are ligated to adapters comprising molecular barcodes. These molecular barcodes (optionally in combination with endogenous sequence information) can then be used for grouping the sequence reads into families, wherein a family corresponds to sequence reads derived from the same parent nucleic acid. The grouped sequence reads can then be analyzed, for example, to identify mutations.4. Amplification
[0153] In methods of the present disclosure, nucleic acids within the nucleic acid sample are amplified to provide progeny nucleic acids. Amplification of the nucleic acids can be used to maximize the likelihood that the assay will detect the target sequences present within the nucleic acid sample. In some embodiments, the method includes partitioning steps wherein amplification of the nucleic acid sample can be performed before or after the partitioning steps.
[0154] Before amplification, adapters can be ligated to the sample nucleic acids, wherein the adapters comprise primer binding sites. The sample nucleic acids flanked by adapters can then be amplified by PCR and / or other amplification methods primed by primers binding to the primer binding sites in the adapters. Amplification methods can involve cycles of denaturation, annealing and extension, resulting from thermocycling or can be isothermal as in transcription-mediatedamplification. Other amplification methods include the ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication.
[0155] DNA ligase can be used to ligate DNA molecules (e.g. cfDNA) in the sample with an adapter on one or both ends, i.e. to form adapted DNA. As used herein, “adapter” refers to short nucleic acids (e g., less than about 500, less than about 100 or less than about 50 nucleotides in length, or be 20-30, 20-40, 30-50, 30-60, 40-60, 40-70, 50-60, 50-70, 20-500, or 30-100 bases from end to end) that are typically at least partially double-stranded and can be ligated to the end of a given sample nucleic acid. In some instances, two adapters can be ligated to a single sample nucleic acid, with one adapter ligated to each end of the sample nucleic acid.
[0156] Ligation of adapters can comprise blunt end ligation or sticky-end ligation. In some embodiments, the present methods perform dsDNA ligations with T-tailed and C-tailed adapters when the sample nucleic acids have been subjected to A-tailing, e.g. using T4 polymerase or Klenow large fragment. This increases the efficiency of ligation and results in amplification of at least 50, 60, 70 or 80% of double stranded nucleic acids. Such methods can increase the amount or number of amplified molecules relative to control methods performed with T-tailed adapters alone by at least 10, 15 or 20%.
[0157] Adapters can include nucleic acid primer binding sites to permit amplification of a sample nucleic acid flanked by adapters at both ends, and / or a sequencing primer binding site, including primer binding sites for sequencing applications, such as various next generation sequencing (NGS) applications. Adapters can include a sequence for hybridizing to a solid support, e g., a flow cell sequence. Adapters can also include binding sites for capture probes, such as an oligonucleotide attached to a flow cell support or the like. Adapters can also include sample indexes and / or molecular barcodes. These are typically positioned relative to amplification primer and sequencing primer binding sites, such that the sample index and / or molecular barcode is included in amplicons and sequencing reads of a given nucleic acid. Adapters of the same or different sequence can be linked to the respective ends of a sample nucleic acid. In some cases, adapters of the same or different sequence are linked to the respective ends of the nucleic acid except that the sample index and / or molecular barcode differs in its sequence.
[0158] In some embodiments, primers relate to oligos which specifically target and enable amplification of amplicons within a set of amplicons. The primers may be of any suitable length depending on the particular needs and targeted sequences employed. In some embodiments, theprimers may at least 10 nucleotides in length. Longer primers are also within the scope of the present disclosure as well known in the art. In some embodiments, primers may be more than 30, more than 40, more than 50 nucleotides in length.
[0159] In some embodiments, the primers used for amplification can be designed by taking into consideration the melting point of hybridization thereof with its targeted sequence (Sambrook et al., 1989, Molecular Cloning — A Laboratory Manual, 2nd Edition, CSH Laboratories; Ausubel et al., 1994, in Current Protocols in Molecular Biology, John Wiley & Sons Inc., N.Y.). To enable hybridization to occur, primers may comprise an oligonucleotide sequence that has at least 70% (at least 71%, 72%, 73%, 74%), preferably at least 75% (75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%) and more preferably at least 90% (90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%) identity to a portion of their target sequence. In some embodiments, primers may have complete sequence identity to their target sequences.
[0160] In some embodiments, primers may contain high affinity RNA analogs such as locked nucleic acids (LNAs). LNA oligos exhibit much better thermal stability when hybridized to complementary nucleic acids compared to typical oligos. For each incorporated LNA within a primer, the melting point of the duplex increases by 2-8 °C. Incorporation of LNAs into primers can be used in the disclosed methods to improve the specificity and sensitivity of the amplification reaction. In some embodiments, primers may contain molecular barcodes.5. Enrichment
[0161] Nucleic acids may be subject to a sequence capture step, in which molecules having target sequences are captured for subsequent analysis. This allows nucleic acids derived from target regions of the genome to be isolated and analyzed, thus avoiding the need for whole genome analysis. Capture can be performed before or after the amplification step.
[0162] In some embodiments, target capture can involve use of a bait set comprising oligonucleotide baits labeled with a capture group, such as the examples noted below. The probes can have sequences selected to tile across a panel of regions, such as genes. Such bait sets are combined with a sample under conditions that allow hybridization of the target molecules with the baits. Then, captured molecules are isolated using the capture group. For example, a biotin capture group can be captured by bead-based streptavidin. Such methods are further described in, for example, U.S. 9,850,523.
[0163] Capture groups include, without limitation, biotin, avidin, streptavidin, a nucleic acid comprising a particular nucleotide sequence, a hapten recognized by an antibody, and magneticallyattractable particles. The capture group can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, a capture group that is attached to an analyte is captured by its binding pair which is attached to an isolatable moiety, such as a magnetically attractable particle or a large particle that can be sedimented through centrifugation. The capture group can be any type of molecule that allows affinity separation of nucleic acids bearing the capture group from nucleic acids lacking the capture group. An exemplary capture group are biotin which allows affinity separation by binding to streptavidin linked or linkable to a solid phase or an oligonucleotide, which allows affinity separation through binding to a complementary oligonucleotide linked or linkable to a solid phase.
[0164] In some embodiments, capturing comprises the use of target-specific amplification primers. In some embodiments, the target-specific primers are designed to hybridize one or more regions of nucleic acid molecules that are associated with the pathogenic germline variant and / or the hereditary cancer, including biomarkers associated with primary and secondary cancers. Thus, in some embodiments, the target-specific amplification primers are personalized for the subject based on the pathogenic germline variant and / or hereditary cancer type.
[0165] In some embodiments, the methods herein comprise capturing nucleic acids comprising epigenetic and / or sequence-variable target regions. In some embodiments, the methods herein comprise capturing nucleic acids comprising epigenetic target regions, such as differentially methylated regions. Such regions may be captured from a sample (e.g., a subsample) that has undergone attachment of adapters, derivatization, partitioning, and / or amplification. Enriching for or capturing DNA comprising epigenetic and / or sequence-variable target regions may comprise contacting the DNA with a set of target- specific probes. When the method comprises a partitioning step, capturing may be performed on one or more partitions. When capturing is performed on multiple partitions, the capture probes used for each partition may be different. In some embodiments, DNA is captured from the first partition and / or the second partition and / or the unbound partition. In some embodiments, the partitions are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.
[0166] The capturing step may be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on features of the probes such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given general knowledge in the art regarding nucleic acid hybridization. In some embodiments, complexes of target-specific probes and DNA are formed.
[0167] In some embodiments, methods described herein comprise capturing a plurality of sets of target regions. The target regions comprise intronic regions or VDJ regions that may comprise rearrangements. The target regions may comprise epigenetic target regions, which may show differences in methylation levels depending on whether they originated from a tumor or from healthy cells. The target regions may comprise sequence-variable regions, which may show differences in sequence, other than rearrangements, depending on whether they originated from a tumor or from healthy cells. The target regions may comprise both epigenetic target regions and sequence-variable regions. The capturing step produces a captured set of DNA molecules. In some embodiments, the DNA molecules corresponding to the sequence-variable target region set are captured at a greater capture yield in the captured set of DNA molecules than DNA molecules corresponding to the epigenetic target region set. In some embodiments, a method described herein comprises contacting DNA with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than DNA corresponding to the epigenetic target region set. For additional discussion of capturing steps, capture yields, and related aspects, see W02020 / 160414, incorporated herein by reference.
[0168] It can be beneficial to capture DNA corresponding to the sequence-variable target region set at a greater capture yield than DNA corresponding to the epigenetic target region set because a greater depth of sequencing may be necessary to analyze the sequence-variable target regions with sufficient confidence or accuracy than may be necessary to analyze the epigenetic target regions. The volume of data needed to determine fragmentation patterns (e.g., to test for perturbation of transcription start sites or CTCF binding sites) or methylation status is generally less than the volume of data needed to determine the presence or absence of genetic variants, such as cancer-related sequence mutations. Capturing the target region sets at different yields can facilitate sequencing the target regions to different depths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).
[0169] In some embodiments, amplification is performed before the capturing step. In some embodiments, amplification is performed after the capturing step. In some embodiments, an amplification step is performed before and after the capturing step. In some embodiments, the methods further comprise sequencing the captured DNA to different degrees of sequencing depth for the epigenetic and sequence-variable target region sets and for rearrangements, consistent with the discussion herein.
[0170] In some embodiments, a capturing step is performed with probes for a sequence-variable target region set and probes for an epigenetic target region set in the same vessel at the same time, e.g., the probes for the sequence-variable and epigenetic target region sets are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the sequence-variable target region set is greater that the concentration of the probes for the epigenetic target region set.
[0171] Alternatively, a capturing step is performed with a sequence-variable target region probe set in a first vessel and with an epigenetic target region probe set in a second vessel, or a contacting step is performed with a sequence-variable target region probe set at a first time and a first vessel and an epigenetic target region probe set at a second time before or after the first time. This approach allows for preparation of separate first and second compositions comprising captured DNA corresponding to a sequence-variable target region set and captured DNA corresponding to an epigenetic target region set. The compositions can be processed separately as desired. These can then be pooled in appropriate proportions to provide material for further processing and analysis such as sequencing.6. Sequencing
[0172] In general, sample nucleic acids can be subject to sequencing after amplification. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, Digital Gene Expression (Helicos), Next generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions can be performed in a variety of sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other mean of processing multiple sample sets substantially simultaneously. Sample processing unit can also include multiple sample chambers to enable processing of multiple runs simultaneously. In some embodiments, adapters are attached to one or both ends of the nucleic acid molecules to facilitate sequence. In some embodiments, the adapters may comprise a Y-shaped adapters or hairpin adapters.
[0173] Simultaneous sequencing reactions may be performed using multiplex sequencing. In some cases, cell-free nucleic acids may be sequenced with at least, for example, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other cases, cell-free nucleic acids may be sequenced with less than, for example, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. Sequencing reactions may be performed sequentially or simultaneously. Subsequent data analysis may be performed on all or part of the sequencing reactions. In some cases, data analysis may be performed on at least, for example, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other cases, data analysis may be performed on less than, for example, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is 1,000-50,000 or 1,000-10,000 or 1,000-20,000 reads per locus (base).
[0174] Sequence reads are aligned to a reference sequence, enabling the identification of reads that map to a genomic region. Prior to alignment, the sequence reads may undergo quality control analysis. Quality control analysis may perform quality control on the sequence read data from the sequencing pipeline. Only sequence reads that have passed a quality control analysis may be used by the sequence read mapper to align the sequence reads to the reference sequence. Quality control analysis of the sequence reads may include obtaining sequence reads that at least partially cover the locus of interest and analyzing coverage depth based on the sequence reads that align to a reference genome above a quality threshold. The quality threshold may vary depending on the particular locus of interest involved. Examples of quality thresholds include a minimum nucleotide overlap and / or minimum alignment identity or similarity. The minimum nucleotide overlap may include, without limitation, a minimum overlap of at least about 1 base, 2 bases, 4 bases, 4 bases, 5 bases, 10 bases, 15 bases, 40 bases, 25 bases, 40 bases, 45 bases, 40 bases, 45 bases, 50 bases, 55 bases, 60 bases, 65 bases, 70 bases, 75 bases, 80 bases, 85 bases, 90 bases, 95 bases, or 100 bases. The minimum alignment identity or similarity may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more.
[0175] In some embodiments, a sequence read mapper may be used to align the sequence reads to a reference sequence. The sequence read mapper may align sequence reads using various sequence alignment technique. The reference sequence is a known sequence used for purposes of comparison with experimentally determined sequences. For example, a known sequence can be an entire genome, a chromosome, or any segment thereof. A reference can typically include at leastIO1, IO3, 106, 109or more nucleotides. A reference sequence can align with a single contiguous sequence of a genome or chromosome or can include non-contiguous segments aligning with different regions of a genome or chromosome. The reference sequence may include a sequence, such as a whole genome sequence, of a species of the subject. For example, if the subject is human, the reference sequence may be the hG19 or hG38 whole genome sequence. In some examples, the reference sequence may be truncated to include only the sequences of interest. For example, the hG19 or hG38 whole genome sequence may be truncated to include only the sequence corresponding to chromosome 6.
[0176] As used herein, “genomic region” refers to any region (e.g., range of base pair locations) of a genome, e.g., a gene, or an exon. A genomic region may be a contiguous or a non-contiguous region. In some embodiments, a genomic region is less than 100 Mb, less than 50 Mb, less than 20 Mb, less than 10 Mb, less than 1 Mb, less than 500 kb, less than 250 kb, less than 100 kb, less than 50 kb, less than 25 kb, less than 10 kb, less than 5 kb, less than 1 kb, less than 500 bp or less than 200 bp. In some embodiments, a genomic region is less 500 bp. In some embodiments, a genomic region corresponds to the region of the genome to which a sequence read aligns, wherein the beginning and the end of the genomic region within 10 base pairs, within 5 base pairs, within 4 base pairs, within 3 base pairs, within 2 base pairs, or within 1 base pair of the terminal alignment positions of the sequence read. In some embodiments, a genomic region corresponds to the region of the genome to which a sequence read aligns, wherein the beginning and the end of the genomic region correspond to terminal alignment positions of the sequence read.7. Additional Workflows
[0177] In certain embodiments, the present methods can integrate genomic and / or epigenomic data with proteomic (proteins and their post-translational modifications), transcriptomic, fragmentomic, immunological, histological, and / or other analyte-specific data.
[0178] In some embodiments, analysis techniques may be utilized that include alternatives to nextgeneration sequencing. Such approaches may offer a cheaper alternative than next-generation sequencing, making it more cost-effective for the subject for monitoring the cancer syndrome or disease. For example, in some embodiments, the analysis comprises techniques that include polymerase chain reaction (PCR), quantitative PCR (qPCR), digital PCR (dPCR), droplet digital PCR (ddPCR), reverse-transcription PCR (RT-PCR) or other highly multiplex PCR methods (e g., UltraPCR). For example, in some embodiments, if the subject is known or suspected of havinghereditary cancer syndrome, relevant variants based on the type of hereditary cancer syndrome are targeted and qPCR is used to determine the level of cancer and / or rate of change in cancer level.
[0179] In some embodiments, single cell methylation profiling techniques can be applied to one or more of the methods in the disclosure, including for example, MLAD-seq, a technique for single-base resolution and quantitative detection of 5mC in DNA, EAC-seq that utilizes engineered proteins for bi sulfite-free, quantitative mapping of 5mC at single-base resolution, Di ital-scRRBS a microfluidics-based platform for single-cell methylation sequencing, msRRBS a scalable singlecell reduced representation bisulfite sequencing technology that allows pooling of cell-specific barcoded DNA fragments before bisulfite conversion which improves efficiency and reduces cost.
[0180] In some embodiments, the analysis comprises the detecting or quantifying macromolecules, such as proteins. In some embodiments, the analysis comprises immunoassays, immunohistocompatibility (IHC) detection, mass spectrometry, or proximity extension or proximity ligation assays.8. Computer Systems and Analysis
[0181] All methods of the present disclosure can be implemented using, or with the aid of, computer systems. In some embodiments, the present disclosure provides for a system comprising a sequencing system configured to receive and process a sample form a subject, wherein the sequencing system comprises a sequencing pipeline having one or more sequencing devices for generated sequence information from nucleic acid molecules from the sample, and a processor programmed to receive, via a sequence analysis pipeline, sequence reads from the sequencing system. The processor can be programmed to receive data from the sequence analysis pipeline. The processor can also be programed to store the data, train the data, and / or implement the data analysis for determine cancer signal levels as disclosed herein.
[0182] The system can include hardware, for example the hardware components of the system include a central processing unit (CPU), random access memory (RAM), storage devices (such as hard disk drives or solid-state drives), and input / output devices (such as keyboards, mice, and display monitors). The specific configuration of the hardware components may vary based on the system model and user requirements. The system can also include software including operating system software that manages the hardware resources and provides a platform for running application software. In addition to the operating system, the System may come with pre-installed application software designed to meet the needs of specific tasks or industries. Users may also install additional applications as required. Additional features in the computer system can includesecurity systems. For example the system can incorporate multiple layers of security measures, including firewalls, antivirus software, and encryption protocols, to protect against unauthorized access and ensure the confidentiality, integrity, and availability of data. The system can also include connectivity systems comprising various connectivity options, including wired and wireless network connections, to enable communication and data exchange with other systems and devices. Compatibility with standard networking protocols ensures the System can integrate seamlessly into existing network environments. The computer system can also include support and Maintenance systems. For example, a comprehensive support and maintenance services are provided to ensure the system operates efficiently and effectively. This includes technical support, software updates, and hardware repair or replacement services. Other features of the computer system can include compliance protocols. For example, the system can be designed to comply with relevant industry standards and regulatory requirements, ensuring reliability and safety in its operation.
[0183] Additional details relating to computer systems and networks, databases, and computer program products are also provided in, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7thEd. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11thEd. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety.
[0184] In some embodiments, the baseline cancer signal and / or the level of cancer and / or rate of change in cancer level from the additional bodily fluid sample is determined by a method comprising quantifying somatic mutations, tumor fraction, mutant allele fraction, and / or methylation signals. In some embodiments, the cancer signal determination comprises quantifying a number of molecules representing circulating tumor DNA in the sample (ctDNA). In some embodiments, the cancer level is determined by an aggregation of analyte signature data, including but not limited to somatic, methylation, fragmentomic, proteomic, histone modification, and / or extracellular vesicle (genomic content and / or cell surface protein) information.
[0185] In some embodiments, the “level of cancer” comprises one or more quantitative measures selected from: (i) the number of detected tumor-associated variants, (ii) the allele fraction of adetected tumor-associated variant, (iii) the frequency of a detected structural variant, (iv) the quantity of cell-free DNA fragments exhibiting aberrant methylation patterns, (v) the degree of hypermethylation or hypomethylation at specific genomic loci, (vi) the amount of circulating tumor DNA (ctDNA) in the sample, (vii) the tumor fraction or mutant allele fraction, or (viii) an aggregation of two or more analyte signature data types including somatic, methylation, fragmentomic, proteomic, histone modification, or extracellular vesicle-derived information. Tn some embodiments, the level of cancer comprises a quantitive measure based on a methylation signature. In some embodiments, the level of cancer comprises a quantitive measure based on a fragmentomic signature. In additional embodiments, the “change in cancer level” or “rate of change in cancer level” comprises the difference or temporal trend in one or more of the foregoing measures, relative to a baseline cancer signal established for the subject.
[0186] In some embodiments, the “change in cancer level” comprises differences in the level, presence / absence, and / or pattern of one or more biomarkers and / or molecular features measured relative to a baseline profile. The biomarkers and / or molecular features may include hereditary cancer syndrome (HCS) genes and / or molecular features associated with or modulated by cancer, such as methylation status, methylation levels, methylation patterns, degrees of hypermethylation or hypomethylation at specific genomic loci, transcription factor binding site occupancy, histone modification marks, fragmentomic features (including fragment length distributions, end motifs, or breakpoint density profiles), mRNA expression values, protein expression levels, or combinations thereof.
[0187] In some embodiments, the methods described herein involve integration of a plurality of sequencing datasets comprising a plurality of genetic and epigenetic states to monitor or screen individuals at high risk of developing one or more cancers due to one or more germline variants. Data integration can be achieved through a variety of artificial intelligence (Al), machine learning (ML), and deep learning methods. Combined, these methods can elucidate the genetic and / or epigenetic changes that occur as cancer initiates and evolves over time, thereby increasing the ability for the present methods to screen for cancer in the high-risk subject.
[0188] In some embodiments of the disclosure the molecular profiles comprise one or more of: genetic sequence data, DNA methylation data, histone modification data, chromatin conformation capture data, nucleosome positioning, histone variants (for example, replacement of the canonical histone H2A with the variant H2A.Z), RNA methylation, chromatin accessibility, DNA hydromethylation, DNA phosphorylation, acetylation, transcription factor binding sites, and / orchromatin looping or DNA-DNA interactions data. In some embodiments, of the disclosure the TFBS are directly assayed using methods such as ChlP-seq or CUT&Tag, in alternative embodiments the TFBS are inferred based on epigenetic data sets and / or machine learning algorithms.
[0189] In some embodiments, DNA methylation data is used to determine or monitor the cancer signal level. The following software tools can be used to perform DNA methylation analsyis: RnBeads is a software tool for large-scale analysis and interpretation of DNA methylation data. Msuite: is an analysis toolkit for DNA methylation profiling, specifically optimized for emerging bi sulfite-free methods. methylKit: An R package for the analysis of genome-wide DNA methylation profiles which also supports epigenome-wide association studies and biomarker discovery. B Smooth: Provides alignment, quality control, and analysis pipeline for whole-genome bisulfite sequencing. Similar open-source methods include MethLAB, MethCy and Methylation plotter. Additional methods include BEAT (BS-Seq Epimutation Analysis Toolkit), an RZBioconductor package for quantitative analysis of DNA methylation from bisulfite sequencing data, utilizing a binomial mixture model. SINBAD, designed for pre-processing, quality assessment, and analysis of single-cell methylation data, starting from multiplexed sequencing reads. CpGtools, a Python package for analyzing DNA methylation data, offering a comprehensive suite for analyzing, annotating, QC, and visualizing the data and MethTools a toolbox for visualizing and analyzing DNA methylation data generated by the Bisulfite sequencing.
[0190] In some embodiments, integration of various sources of cancer signal can be applied for cancer detection and monitoring. Multiple methods can be employed to integrate the various layers of epigenetic information such as methylation patterns, transcription factor binding sites (TFBS), and histone modifications described earlier in the disclosure to predict regional features like the presence of a TFB S or chromatin interaction sites (e.g., enhancer promoter interaction) at a specific genomic location. For example, a neural network approach can integrate multiple layers of epigenomic information to predict regional features, using deep learning to handle highdimensional data. Such model can undergo continuous iteration and validation against known biological insights can improve specificity and sensitivity of the methods by refining the model. In some embodiments, the integrated epigenetic features are employed to determine a baseline cancer signal or a level of cancer at a given time point. In further embodiments, repeated application of these methods across multiple time points may determine the rate of change in one or more cancer levels, patterns, or profiles (including but not limited to genetic, epigenetic,transcriptomic, proteomic, or metabolomic biomarkers), or in an overall cancer signal or composite measure derived therefrom.
[0191] In some embodiments, an initial sample from the subject may be analyzed for differential methylation at predefined cancer-associated loci using bisulfite sequencing and analyzed to determine the degree of hypermethylation and / or hypomethylation at these loci to establish a baseline methylation profile (e.g., a baseline cancer signal) for the subject. At subsequent time points, the same loci are re-analyzed, and differences in methylation levels from the baseline are calculated to determine a rate of change in the cancer level, wherein an increasing level of aberrant methylation beyond a threshold indicates cancer progression.
[0192] In additional embodiments, the baseline cancer signal may be determined by analyzing and integrating molecular features across multiple “omic” layers, including methylation levels and / or patterns, transcription factor binding sites, histone modification marks, and / or fragmentomic features including fragment length distributions, end motifs, and / or breakpoint density profiles. A machine learning model, such as a neural network trained on cancer reference datasets, may be employed to improve the present methods by integrating these multi-omic molecular features into a more refined baseline cancer signal as compared to analysis of a single molecular feature (e.g., such a genomic information). At subsequent time points, the same molecular features may be measured for one or more subsequent samples, and changes relative to the baseline profile may be assessed to determine a rate of change in the cancer level. A shift in the integrated multi-omic molecular features and / or a change in one or more individual molecular features beyond a predetermined threshold may be indicative of cancer initiation and / or progression.
[0193] The methods in the present disclosure provide datasets involving various types of epigenomic information, hence an input layer is designed to handle multiple data modalities. This can be achieved using separate input channels or sub-networks for each data type (methylation, TFBS, and histone modifications). Each channel can preprocess its respective data type, normalizing and encoding it in a form suitable for deep learning, that is including several steps to convert raw data into a format that neural networks can effectively process and learn from. These steps include data normalization / standardization, encoding, reshaping, handling missing values, feature engineering / selection, and data augmentation.
[0194] Feature extraction layers: For each data modality, a convolutional neural network (CNNs) or recurrent neural network (RNN) can be used to capture spatial dependencies and patterns within the genomic sequences. CNNs are particularly useful for identifying patterns in histonemodifications and TFBS, while RNNs or transformer-based models can effectively process sequential data, capturing long-range dependencies in methylation patterns for example.
[0195] Integration Layer: After feature extraction, the outputs of the separate channels are integrated. This can be achieved through concatenation, followed by dense layers, or by using more sophisticated integration techniques like attention mechanisms, which allow the model to weigh the importance of information from different epigenomic layers dynamically.
[0196] Prediction Layer: the integrated features can be fed into one or more dense layers with nonlinear activation functions to enable the prediction of regional features, for example the presence of a TFBS at a specific genomic location (this holds true for other feature like histone marks and interaction sites). The output layer is designed according to the specific prediction task, for example binary classification for predicting the presence / absence of a TFBS.
[0197] Training: The model should be trained on labeled datasets where the ground truth (e.g., presence or absence of TFBS in specific regions or other functional element or genomic feature) is known, a cross-entropy loss function can be used for classification tasks, and optimization of the model can be accomplished using gradient descent algorithms like Adam or SGD.
[0198] Regularization and Dropout: To prevent overfitting, given the complexity of epigenomic data and the deep architecture, we incorporate regularization techniques (L1 / L2 regularization) and dropout layers, particularly after dense layers in the network.
[0199] Evaluation and Fine-tuning: the model’s performance can be evaluated using standard metrics like accuracy, precision, recall, and Fl score. Depending on the results, fine-tune the model by adjusting the architecture, hyperparameters, or training procedure. In some embodiments we use techniques like cross-validation for a more robust evaluation. Similarly, for enhanced per Performance methods such as data augmentation techniques specific to genomic data to increase the diversity of the training set. Other techniques include multi-task learning when for example predicting multiple regional features simultaneously, under this framework the network is designed to make several predictions at once, sharing representations between tasks to improve learning efficiency and prediction accuracy.
[0200] Lastly, techniques such as transfer learning can be employed to leverage pre-trained models on related tasks to improve performance, this method is especially useful when labeled data are limited.
[0201] In some embodiments, the implementations described herein can integrate molecular data comprising transcription factor binding sites (TFBS), fragmentomic patterns, fragmentomic levels,fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, and / or H3S10ph with patient outcomes data, electronic health records and / or health insurance claims data to: (1) Identify common methylation changes associated with cancer types or subtypes and stages. The goal of this analysis is to reveal potential methylation biomarkers for cancer diagnosis (2) Correlate methylation changes with clinical outcomes. The goal of this analysis is to understand the impact of these methylation changes on cancer progression, resistance, and treatment response. (3) Integrate the identified methylation changes with existing biological pathways and networks. The goal of this analysis is to uncover how methylation changes disrupt cellular processes and contribute to cancer development, resistance, and response. (4) Compare methylation data across different cancer types or subtypes to identify unique and shared mechanisms. The goal of this analysis is to understand cancer heterogeneity and similarities across different cancers. (5) Develop a predictive model using genomic and / or epigenomic (e.g, methylation) data to predict treatment outcomes, recurrence, or drug resistance. The goal of this analysis is to improve and help guide personalized treatment strategies. Such analysis can be achieved using a variety of statistical and machine learning models including Linear Regression and Logistic Regression: For continuous and binary outcomes, respectively, to model the relationship between methylation levels at specific sites and the presence or severity of cancer. Cox Proportional Hazards Model, can be useful for survival analysis to correlate methylation levels with the time to event data, such as time to cancer recurrence or progression. Mixed Models are helpful when the data comprises multiple measurements or hierarchical structures, mixed models can account for the correlation within subjects or groups. Multivariate Additionally, analysis like principal component analysis (PCA) or partial least squares regression (PLSR) can reduce dimensionality and identify patterns in methylation data that correlate with cancer types.
[0202] Other methods employed to explore relationships or correlations between methylation data and cancer type / subtype include machine learning models. For example, to handle complex, nonlinear relationships we employ Decision Trees and Random Forests these are particularly helpful in situations where the association between the methylation status of certain genes (or CpG sites) and cancer characteristics does not follow a straight-line pattern. For example, to model interaction effects i.e., methylation at one site might affect the impact of methylation at another site. Alone orin one or more combinations these models can identify specific methylation sites that are important for classifying cancer types. In some embodiments, Support Vector Machines (SVM) can be employed for classification tasks, including distinguishing between different types of cancer based on methylation patterns. In yet other embodiments, for example when the data is sufficiently large, deep learning approaches (e.g., convolutional neural networks for structured data like methylation arrays) can capture complex patterns including interactions in the data. Methods such as Gradient Boosting Machines (GBM) including models like XGBoost, LightGBM, and CatBoost can provide robust predictive models for cancer classification based on methylation data. These models include sequential addition of weak learners (e.g., decision trees) in such a way that each new tree corrects the errors made by the previous ones. GBMs can handle various types of data, including categorical and continuous variables. In additional embodiments, Cluster Analysis are employed to identify subgroups within cancer types that share similar methylation patterns. These include unsupervised learning models, including K-means clustering or hierarchical clustering.
[0203] In some embodiments, the analysis of the methods provided herein comprises determining genetic and / or epigenetic alterations associated with a deficiency in a DNA repair pathway. In some embodiments, the DNA repair pathway is the DNA mismatch repair (MMR) system. A deficiency in the MMR pathway may lead to microsatellite instability (MSI). Thus, in some embodiments, the analysis comprises determining MSI by detecting genetic alterations and / or epigenetic signatures. In some embodiments, panels may be tailored to detect MSI, which may include MMR genes, such as MLH1 and MSH2.
[0204] Loss of heterozygosity (LOH) at the pathogenic germline variant locus / loci will initiate the cancer. In some embodiments, the methods disclosed herein are tailored to detect variants causing LOH at that loci only. For example, if the subject is a BRCA1 germline variant carrier, the method of the disclosure may comprise a multi-omic test that detects somatic variants (e.g., SNV, indel, large del, CNV), promoter DNA methylation, histone methylation (causing transcriptional silencing), RNA expression and / or protein expression all at or associated with BRCA1. This could also extend to loci involved with ‘BRCA1 inactivation’, and / or detecting phenotypes / methylation patterns corresponding to BRCA1 inactivation. For MLH1 -germline variant carrier, a similar type of test can be applied, albeit at / for MLH1 loci. The caller could be ‘static’ or involve longitudinal measurements, incorporating the rate of change of biomarker levels into caller.
[0205] As an example involving the effects of LOH, the methods comprise screening subjects carrying a germline TP53 mutation and predisposed to Li-Fraumeni syndrome (LFS), a hereditarycancer syndrome characterized by an elevated risk of developing multiple types of cancers at an early age. These individuals inherit one defective copy of the TP53 gene from birth, but this germline mutation alone may not immediately impair the full tumor-suppressive function of the p53 protein. The development of cancer typically requires a second hit, often through loss of heterozygosity (LOH) or additional mechanisms where the remaining functional copy of TP53 is lost or inactivated (for example, a somatic missense or nonsense mutations, chromosomal deletions and gene copy number variation, epigenetic silencing, loss of a cis or trans regulatory element such as an enhancer, histone marks, chromatin conformation, or transcription factor). In some embodiments of the disclosure, the screening methods are specifically tailored to detecting LOH or additional mechanisms that lead to a “second hit,” resulting in the somatic loss of the second allele.
[0206] Additionally, while the loss of TP53 tumor suppressor activity is commonly associated with cancers such as lung, breast, colorectal, and ovarian cancer, LFS patients often present with a diverse range of primary tumors, and the specific cancer type that will develop in these individuals is unpredictable. This variability in tumor types and origins can complicate diagnosis and management underscoring the need to incorporate cancer signal of origin (CSO) detection into clinical screening test. Thus, in certain embodiments, the methods described herein incorporate CSO detection or tissue of origin detection in subjects with hereditary cancer syndromes, such as LFS, where the cancer that will develop is unpredictable or that lead to a broad spectrum of possible cancers without a clear tissue of origin.
[0207] In some embodiments, the implementations described herein include methods and systems for determining a cell type, a tissue of origin, or both. In some embodiments, a cell type or the origin of a cell can be determined by genetic, transcriptomic, proteomic, and / or epigenetic states or methylation patterns of the DNA. In additional embodiments, fragments of DNA can be analyzed by methylation analysis to determine a methylation pattern or methylation levels, methylation signature, or one or more methylation sites or regions.
[0208] The methods described in this disclosure may further incorporate cancer of unknown primary (CUP) detection in individuals with LFS or other hereditary cancer syndromes where the cancer that will develop is unpredictable or that lead to a broad spectrum of possible cancers without a clear tissue of origin. CUP detection can involve obtaining a tumor sample from a patient, extracting a plurality of molecules including DNA, RNA, and proteins, DNA / protein interactions using chromatin immune purification followed by sequencing (ChlP-Seq), DNA-DNAinteractions using chromatin conformation capture approaches, enhancer RNA, non-coding RNA expression, miRNA levels, and transcript isoform diversity, and using at least one of these datatypes to a train machine learning algorithm by providing labeled datasets, allowing the model to adjust weights based on patterns identified in the data to predict the most likely tissue of origin. In cases where complex datasets are utilized and integrated ensemble methods can be employed to integrate a variety of biological data types, such as genomic, transcriptomic, and proteomic profiles, to enhance the robustness of the model’s predictions (i.e., tissue of origin, cell type, cancer of unknown priority, etc.). In additional embodiments, various characteristics of the plurality of molecules can further be derived from the data to improve the model’s performance, such as fragment length, fragment end motifs, fragment end density, fragment coverage distribution, sequence context at fragment boundaries, and other molecular features.
[0209] CUP, cancer signal of origin (CSO), tissue of origin, and / or cell type of origin detection in individuals with hereditary cancer syndromes as described in the methods of this disclosure can assist in identifying more precise, targeted treatments by revealing the underlying genetic or epigenetic mutations or molecular characteristics associated with hereditary cancer syndromes. Additionally, the methods described herein can help clinicians determine appropriate screening strategies, time points for other cancers, or secondary cancers associated with hereditary cancer syndromes.9. Therapies
[0210] In certain embodiments, the methods disclosed herein relate to identifying and / or administering therapies based on the determination of cancer signals in the subject. In some embodiments, the patient or subject has a given disease, disorder or condition, e.g., any of the cancers or other conditions described elsewhere herein. Essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, immunotherapy, and / or the like) may be included as part of these methods.
[0211] In certain embodiments, the therapy administered to a subject comprises at least one chemotherapy drug. In some embodiments, the chemotherapy drug may comprise alkylating agents (for example, but not limited to, Chlorambucil, Cyclophosphamide, Cisplatin and Carboplatin), nitrosoureas (for example, but not limited to, Carmustine and Lomustine), antimetabolites (for example, but not limited to, Fluorauracil, Methotrexate and Fludarabine), plant alkaloids and natural products (for example, but not limited to, Vincristine, Paclitaxel and Topotecan), anti- tumor antibiotics (for example, but not limited to, Bleomycin, Doxorubicin andMitoxantrone), hormonal agents (for example, but not limited to, Prednisone, Dexamethasone, Tamoxifen and Leuprolide) and biological response modifiers (for example, but not limited to, trastuzumab and bevacizumab, cetuximab and rituximab). In some embodiments, the chemotherapy administered to a subject may comprise FOLFOX or FOLFIRI. In certain embodiments, a therapy may be administered to a subject that comprises at least one PARP inhibitor. In certain embodiments, the PARP inhibitor may include OLAPARIB, TALAZOPARIB, RUCAPARIB, NIRAPARIB, among others. In some embodiments, the methods comprise administering a therapy comprising a PARP inhibitor, such as olaparib, to a subject determined to have homologous recombination repair (HRR) gene or deficiency (HRD), such as with BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and RAD54L alterations. In some embodiments, the subject has a metastatic castrate resistant prostate cancer (mCRPC). In some embodiments, the PARP inhibitor, such as olaprib is used to treat a subject having ovarian cancer, breast cancer, pancreatic cancer, or mCRPC, wherein the subject is determined to have alterations in BRCA1, BRCA2, and / or ATM.
[0212] In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, immunotherapy, and / or the like) may be included as part of these methods. Customized therapies can include at least one immunotherapy (or an immunotherapeutic agent). Immunotherapy refers generally to methods of enhancing an immune response against a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer. In some embodiments, the methods comprise administering a platinum compound to the subject to treat the cancer, such as cisplatin, carboplatin, or oxaliplatin.
[0213] In some embodiments, the immunotherapy or immunotherapeutic agent targets an immune checkpoint molecule. Certain tumors are able to evade the immune system by co-opting an immune checkpoint pathway. Thus, targeting immune checkpoints has emerged as an effective approach for countering a tumor’s ability to evade the immune system and activating anti-tumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012, 12:252-264.
[0214] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule that reduces a signal involved in the T cell response to antigen. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (aka B7.1) or CD86 (aka B7.2) on antigen presenting cells. PD-1 is another inhibitory checkpoint molecule that is expressed on T cells. PD-1 limits the activity of T cells in peripheral tissues during aninflammatory response. In addition, the ligand for PD-1 (PD-L1 or PD-L2) is commonly upregulated on the surface of many different tumors, resulting in the downregulation of anti-tumor immune responses in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for PD-1, such as PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for CTLA4, such as CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin like receptor (KIR), T cell membrane protein 3 (TIM3), gal ectin 9 (GAL9), or adenosine A2a receptor (A2aR).
[0215] Antagonists that target these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Accordingly, in certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the antagonist of the inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti-CTLA4, anti- PD-1, anti-PD-Ll, or anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal anti-PD-1 antibody. In some embodiments, the antibody is a monoclonal anti-PD-Ll antibody. In certain embodiments, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, an anti-CTLA4 antibody and an anti-PD-Ll antibody, or an anti-PD-Ll antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is one or more of pembrolizumab or nivolumab. In certain embodiments, the anti-CTLA4 antibody is ipilimumab. In certain embodiments, the anti-PD-Ll antibody is one or more of atezolizumab, avelumab, or durvalumab. In certain embodiments, immunotherapy, such as pembrolizumab, is used to treat a subject determined to have a high microsatellite instability status (MSLH). In certain embodiments, the immunotherapy, such as pembrolizumab, is used to treat a subject determined to have a high tumor mutational burden (TMB), for example, then the TMB status is greater than or equal to 10 mutations per megabase. In certain embodiments, the immunotherapy, such as pembrolizumab, is used to treat a subject determined to a have a mismatch repair deficiency (dMMR), such as in genes comprising MLH1, PMS2, MSH2 and MSH6.
[0216] In certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., antibody) against CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other embodiments,the antagonist is a soluble version of the inhibitory immune checkpoint molecule, such as a soluble fusion protein comprising the extracellular domain of the inhibitory immune checkpoint molecule and an Fc domain of an antibody. In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1, or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one embodiment, the soluble fusion protein comprises the extracellular domain of PD- L2 or LAG3.
[0217] In certain embodiments, the immune checkpoint molecule is a co-stimulatory molecule that amplifies a signal involved in a T cell response to an antigen. For example, CD28 is a costimulatory receptor expressed on T cells. When a T cell binds to antigen through its T cell receptor, CD28 binds to CD80 (aka B7.1) or CD86 (aka B7.2) on antigen-presenting cells to amplify T cell receptor signaling and promote T cell activation. Because CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 is able to counteract or regulate the co-stimulatory signaling mediated by CD28. In certain embodiments, the immune checkpoint molecule is a co- stimulatory molecule selected from CD28, inducible T cell co-stimulator (ICOS), CD 137, 0X40, or CD27. In other embodiments, the immune checkpoint molecule is a ligand of a co-stimulatory molecule, including, for example, CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.
[0218] Agonists that target these co-stimulatory checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Accordingly, in certain embodiments, the immunotherapy or immunotherapeutic agent is an agonist of a co-stimulatory checkpoint molecule. In certain embodiments, the agonist of the co-stimulatory checkpoint molecule is an agonist antibody and preferably is a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti -0X40, or anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80, anti-CD86, anti-B7RPl, anti-B7-H3, anti-B7-H4, anti-CD137L, anti-OX40L, or anti-CD70 antibody.
[0219] In certain embodiments, the status of a nucleic acid variant from a sample from a subject as being of somatic or germline origin may be compared with a database of comparator results from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the subject and / or patients who are receiving, or who have received, the same therapy as the subject. Acustomized or targeted therapy (or therapies) may be identified when the nucleic variant and the comparator results satisfy certain classification criteria (e.g., are a substantial or an approximate match).
[0220] In certain embodiments, the therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing an immunotherapeutic agent are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by any method known in the art, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraauricular, which administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like.
[0221] In certain embodiments, the present methods are also useful in determining the efficacy of particular treatment options. For example, the number of variations detected, irrespective of their precise identity, is a predictor of amenability to immunotherapy because the mutations create neoepitopes that can be subject of immune attack (see e.g., US20200370129).
[0222] In some embodiments, the therapy comprises one or more treatments from Table 2, below:Table 2; List of cancer types with associated biomarker target and drug
[0223] The present methods can be used to generate or profile, fingerprint or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation, nucleotide variation, epigenomic information, and / or tumor fraction. In some embodiments, the methods disclosed herein are used to monitor the efficacy or responsiveness of a treatment to the subject. In some embodiments, the methods disclosed herein can be used to determine whether the subject is a candidate for a therapy to treat the cancer or disease.
[0224] The present methods can be used to diagnose, prognose, monitor or observe cancers or other diseases of fetal origin. That is, these methodologies can be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other nucleic acids may co-circulate with maternal molecules.
[0225] In certain embodiments, the present methods can be used to determine minimal residual disease (MRD) of a subject, for example, based on a tumor fraction determination. In some embodiments, the methods may be directed to determining MRD by using a tissue-informed assay (i.e., using a tissue sample collected from a patient to determine a personalized panel to enrich for one or more genomic and / or epigenomic variants in a subsequent blood sample from the patient, optionally also using a buffy coat or normal sample from the subject) or a tissue-naive assay.
[0226] In certain embodiments, the present methods can integrate genomic and / or epigenomic data with proteomic (proteins and their post-translational modifications), transcriptomic,fragmentomic, immunological, histological, and / or other analyte-specific data to determine disease initiation, progression, malignant transformation, and therapeutic outcomes.
[0227] The intricate interplay of chemical modifications to DNA and histone proteins, known as the epigenome, plays a significant role in modulating gene expression in health and disease including cancer. Epigenetic changes in conjunction with genetic alterations contribute to the acquisition of cancer hallmarks such as sustaining proliferative signaling, evading growth suppressors, resisting cell death, enabling replicative immortality, inducing angiogenesis, and activating invasion and metastasis (Hanahan, D. 2022). Given the reversible nature of epigenetic modifications, understanding these mechanisms offers promising avenues for therapeutic intervention, with several epigenetic drugs already approved or in clinical trials for the treatment of cancer.
[0228] Methods of the disclosure can be applied to monitor patients treated with a PCV and / or an off-the-self vaccine, alone or in combination with other therapies including epigenetic drugs such as DNA methyltransferase (DNMT) inhibitors, Histone deacetylase (HDAC) inhibitors, Lysine methyltransferase inhibitors, Lysine demethylase inhibitors, and Bromodomain inhibitors.
[0229] DNA Methyltransferase Inhibitors (DNMTi)- inhibit DNA methyltransferases, enzymes that add methyl groups to DNA, typically silencing gene expression. DNMTi drugs include Azacitidine which was approved for the treatment of myelodysplastic syndromes (MDS) and Decitabine also approved for MDS. Histone Deacetylase Inhibitors (HDACi) inhibit histone deacetylases, enzymes that remove acetyl groups from histone proteins, typically leading to a closed chromatin structure and gene silencing. Inhibiting these enzymes can reactivate silenced genes beneficial in cancer treatment. HDACi drugs include Vorinostat approved for the treatment of cutaneous T cell lymphoma (CTCL), Romidepsin approved for CTCL and peripheral T-cell lymphoma (PTCL), Belinostat approved for PTCL, and Panobinostat approved for multiple myeloma in combination with bortezomib and dexamethasone. EZH2 Inhibitors, EZH2 is a component of the polycomb repressive complex 2 (PRC2) that methylates histone H3 on lysine 27 (H3K27me3), leading to gene silencing. EZH2 Inhibitors include Tazemetostat approved for the treatment of epithelioid sarcoma and follicular lymphoma. Additional epigenetic drugs include Bromodomain Inhibitors that target bromodomains, which recognize acetylated lysine residues on histone tails, influencing chromatin structure and gene expression; however, currently there are no bromodomain inhibitors approved by the FDA. For example, a subject with a germline BRCA1 or BRCA2 mutation may undergo an initial liquid biopsy to establish a baseline cancer signal. Theliquid biopsy test may comprise biomarkers associated with homologous recombination deficiency, including mutational signatures such as single-base substitution Signature 3 (SBS3) and / or rearrangement signatures such as RS3 and RS5. Additionally or alternatively, subsequent liquid biopsy assays may detect a “second hit,” such as loss of the wild-type allele, loss of heterozygosity, promoter methylation, or other somatic inactivation events in BRCA1 / 2.
[0230] In some embodiments primary cancer monitoring may be performed via breast MRI and / or mammography at intervals more frequent than standard screening guidelines. Secondary cancer monitoring may be performed via subsequent liquid biopsy assays, in which the emergence or increasing prevalence of the above-described mutational signatures and / or second-hit events is assessed. A determination that the level or rate of change in the cancer signal exceeds a predetermined threshold above the baseline indicates initiation or progression of a primary or secondary cancer, whereas stability of these measures indicates absence or non-progression of disease.
[0231] In some embodiments, the “standard screening guidelines” or “gold-standard tests” referenced herein may comprise established clinical practices for early detection of specific cancer types, for example, as set out in national and international medical guidelines (e.g., guidelines published by the U.S. Preventive Services Task Force (USPSTF), the American Cancer Society (ACS), the National Comprehensive Cancer Network (NCCN), among others. For example, clinical tests for breast cancer screening may include mammography, digital breast tomosynthesis (3D mammogram), contrast-enhanced mammography (CEM), breast ultrasound, clinical breast exam (CBE), and / or breast MRI; colorectal cancer screening may include colonoscopy, flexible sigmoidoscopy, CT coIonography (virtual colonoscopy), stool-based tests such as fecal immunochemical test (FIT), high-sensitivity guaiac-based fecal occult blood test (gFOBT), or multi -targeted stool DNA test, and, in some cases, blood-based colorectal cancer screening tests; cervical cancer screening may include Pap smear, human papillomavirus (HPV) testing, and / or co-testing (Pap plus HPV); endometrial cancer screening may include counseling at the time of menopause and, in selected cases, endometrial biopsy; lung cancer screening may include annual low-dose computed tomography (LDCT); prostate cancer screening may include prostate-specific antigen (PSA) testing with or without digital rectal examination; and skin cancer screening may include a full-body dermatological examination.
[0232] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way ofexample only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in practicing the invention. It is therefore contemplated that the disclosure shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0233] While the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be clear to one of ordinary skill in the art from a reading of this disclosure that various changes in form and detail can be made without departing from the true scope of the disclosure and may be practiced within the scope of the appended claims. For example, all the methods, systems, computer readable media, and / or component features, steps, elements, or other aspects thereof can be used in various combinations.
Claims
CLAIMSWhat is claimed is:
1. A method for screening for cancer in a subject previously determined to be, or suspected of being, at an increased risk for developing cancer, comprising:(a) providing a first bodily fluid sample from the subject at a first time point and analyzing one or more analytes for one or more biomarkers from the first bodily fluid sample to determine a baseline cancer signal;(b) at one or more subsequent time points, providing an additional bodily fluid sample from the subject and analyzing one or more analytes for one or more biomarkers from the additional bodily fluid sample to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and(c) determining that the subject has cancer if the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold.
2. The method of claim 1, wherein the subject was previously determined to have one or more pathogenetic germline variants associated with increased cancer risk or was previously diagnosed with hereditary cancer syndrome (HCS).
3. The method of claims 1 or 2, wherein the subject has a family history with a known germline pathogenic or likely pathogenic variant associated with increased cancer risk.
4. The method of any one of claims 1 to 3, wherein prior to (a), the subject was previously determined to not have cancer.
5. The method of any one of claims 1 to 4, wherein the analyzing in (a) and / or (b) comprises detecting for primary risk cancers.
6. The method of any one of claims 1 to 5, wherein the analyzing in (a) and / or (b) comprises detecting for secondary risk cancers.
7. The method of any one of claims 1 to 6, wherein the analyzing in (a) and / or (b) comprises detecting a non-cancer disease or to determine a risk of developing the non-cancer disease.
8. The method of any one of claims 1 to 7, wherein the first and / or additional bodily fluid samples comprise blood, plasma, serum, urine, saliva, tears, or mucus.
9. The method of any one of claims 1 to 8, wherein the one or more analytes in (a) comprise cell- free deoxyribonucleic acid (cfDNA), cell-free ribonucleic acid (cfRNA), proteins, and / or exosomes, and wherein the one or more analytes in (b) comprise the same analytes as in (a).
10. The method of any one of claims 1 to 9, wherein the one or more biomarkers in (a) and / or (b) comprise somatic variants, methylation status, fragmentomic patterns, transcription factor binding sites, and / or chromatin interaction sites.
11. The method of any one of claims 1 to 10, wherein the subject was previously determined, or suspected of having, hereditary breast cancer, hereditary ovarian cancer, or hereditary breast and ovarian cancer (HBOC) syndrome.
12. The method of claim 11, wherein the subject has germline mutations in BRCA1 and / or BRCA2.
13. The method of any one of claims 1 to 10, wherein the subject was previously determined, or suspected of having Cowden syndrome.
14. The method of claim 13, wherein the subject has a germline mutation in PTEN.
15. The method of any one of claims 1 to 10, wherein the subject was previously determined, or suspected of having, Lynch syndrome (or hereditary non-polyposis colorectal cancer (HNPCC) syndrome) or Muir Torre syndrome.
16. The method of claim 15, wherein the subject has a mutation in a mismatch repair gene, such as MLH1, MSH2, MSH6 or PMS2.
17. The method of any one of claims 1 to 10, wherein the subject was previously determined, or suspected of having, hereditary leukemia or a hematologic malignancy syndrome.
18. The method of any one of claims 1 to 10, wherein the subject was previously determined, or suspected of having, Li-Fraumeni syndrome (LFS).
19. The method of claim 18, wherein the subject has a germline mutation in TP53.
20. The method of any one of claims 1 to 10, wherein the subject was previously determined, or suspected of having, Von Hippel-Lundau (VHL) disease.
21. The method of claim 20, wherein the subject has a mutation in the VHL gene.
22. The method of any one of claims 1 to 10, wherein the subject was previously determined, or suspected of having, multiple endocrine neoplasia (MEN) syndrome.
23. The method of claim 22, wherein the subject has a mutation in the MEN1 and / or RET gene.
24. The method of any one of claims 1 to 10, wherein the subject was previously determined to have familial adenomatous polyposis (FAP).
25. The method of claim 24, wherein the subject has a mutation in the APC gene.
26. The method of claim 7, wherein the non-cancer disease is a neurological disease, autoimmune disease, cardiovascular disease, or cardiometabolic disorders.
27. The method of claim 26, wherein the neurological disease is selected from the group consisting of Alzheimer’s disease, Parkinson’s disease, multiple sclerosis (MS), Amyotrophic Lateral Sclerosis (ALS), Huntington’s disease, stroke, or migraine.
28. The method of claim 26, wherein the autoimmune disease is selected from the group consisting of rheumatoid arthritis, multiple sclerosis, lupus, Type 1 diabetes, psoriasis, celiac disease, Hashimoto’s thyroiditis, inflammatory bowel disease, Crohn’s disease or ulcerative colitis.
29. The method of claim 26, wherein the cardiovascular disease is selected from the group consisting of Coronary Artery Disease (CAD), hypertension or high blood pressure, heart failure, arrhythmia, atherosclerosis, myocardial infarction, or stroke.
30. The method of claim 26, wherein the cardiometabolic disorder is selected from the group consisting of Type 2 Diabetes, hyperlipidemia, metabolic syndrome, obesity, hypercholesterolemia, or insulin resistance.
31. The method of any one of claims 1 to 30, wherein the one or more subsequent time points includes at least every 2 weeks, at least every 4 weeks, at least every 6 weeks, at least every 8 weeks, at least every 10 weeks, or at least every 12 weeks from the first time point.
32. The method of any one of claims 1 to 30, wherein the one or more subsequent time points includes at least every 6 months, at least every 12 months, at least every 18 months, at least every 24 months, at least every 30 months, or at least every 36 months from the first time point.
33. The method of any one of claims 1 to 30, wherein the one or more subsequent time points occurs every 3 months.
34. The method of any one of claims 1 to 30, wherein the one or more subsequent time points includes at least every 5 years or at least every 10 years from the first time point.
35. The method of any one of claims 1 to 34, wherein the subject was previously determined to be at an increased risk for developing cancer by undergoing germline genetic testing.
36. The method of claim 4, wherein the subject was previously determined to not have cancer by undergoing a medical procedure.
37. The method of claim 36, wherein the medical procedure comprises a colonoscopy, a stoolbased test, a mammogram, a Pap test (or Pap smear), an HPV test, endoscopy, ultrasound,computed tomography (CT), magnetic resonance imaging (MRI), and / or other imaging method.
38. The method of claim 4, wherein the subject was previously determined to not have cancer by analyzing a tissue sample from the subject.
39. The method of claim 4, wherein the subject was previously determined to not have cancer by analyzing a bodily fluid sample from the subject.
40. The method of claim 39, wherein the bodily fluid sample was analyzed by a multi-cancer detection (MCD) test.
41. The method of claim 39, wherein the MCD test is a liquid biopsy assay.
42. The method of any one of claims 1 to 37, wherein the analyzing in (a) and / or (b) comprises the use of a liquid biopsy assay.
43. The method of any one of claims 1 to 42, wherein the analyzing in (a) and / or (b) comprises:(a) obtaining cfDNA molecules from the first and / or additional samples of the subject;(b) tagging, amplifying, and / or enriching the cfDNA molecules; and(c) sequencing a plurality of the cfDNA molecules or amplicons thereof.
44. The method of claim 43, wherein the enriching comprises contacting the cfDNA molecules of amplicons thereof with a probe set comprising oligonucleotide sequences configured to hybridize a plurality of target genomic regions within the cfDNA molecules or amplicons thereof to obtain enriched target genomic regions.
45. The method of claim 43, wherein the enriching comprises contacting the cfDNA molecules or amplicons thereof with a set of oligonucleotide primers configured to amplify a plurality of target genomic regions within the cfDNA molecules or amplicons thereof to obtain enriched target genomic regions.
46. The method of claims 44 or 45, wherein each of the plurality of target genomic regions comprises members of the Homologous Recombination (HR), Non -Homologous End Joining (NHEJ), and / or PARP Pathways.
47. The method of claims 44 or 45, wherein the target genomic regions comprise one or more sequences of BRCA1, BRCA2, TP53, MLH1, MSH2, MSH6, PMS2, PTEN, MUTYH, STK11, BMPR1A, SMAD4, PALB2, TSC1, TSC2, FLCN, RET, RBI, CDH1, APC, MUTYH, MEN1, CDKN1B, SDHA, SDHB, SDHC, SDHD, CTNNA1, VHL, FH, MET, FLCN, HOXB13, ATM, CHEK2, PTCHI, SUFU, PTCH2, CDKN2A, MITF, BAP1, and / or EPCAM.
48. The method of claims 44 or 45, wherein the target genomic regions comprise a sequence in a gene body, a promoter region, a 3 or 5’ UTR or a regulatory element.
49. The method of claim 48, wherein the regulatory element comprises a genetic or epigenetic element that modulates gene expression.
50. The method of any one of claims 1 to 46, wherein the subject is determined to have an epigenetic alteration leading to gene inactivation.
51. The method of claim 50, wherein the epigenetic alteration comprises promoter methylation, at least one regulatory variant that reduces gene expression, or a chromatin conformation change that reduces gene expression.
52. The method of any one of claims 1 to 51, wherein the subject is determined to have a somatic genetic variant comprising a single nucleotide variant (SNV), and insertion or deletion (indel), a gene fusion, and / or or copy number variation (CNV).
53. The method of any one of claims 1 to 52, wherein the cancer is selected from the group consisting of breast cancer, ovarian cancer, prostate cancer, pancreatic cancer, bowel cancer, womb cancer, stomach cancer, gallbladder cancer, bladder cancer, bone cancer, acute myeloid leukemia (AML), soft tissue sarcoma, brain tumors, cancer of the adrenal gland, thyroid cancer, kidney cancer, skin cancer (melanoma), liver cancer, pancreatic neuroendocrine tumors (pNETs), renal cell carcinoma, retinoblastoma, medullary thyroid cancer, hereditary papillary renal cell carcinoma (HPRCC), and uterine leiomyomata.
54. The method of claim 53, wherein the cancer is determined to be a primary cancer.
55. The method of claim 53, wherein the cancer is determined to be a secondary cancer.
56. The method of 43, wherein the tagging comprises attaching adapters to a plurality of the cfDNA molecules.
57. The method of claim 56, wherein the adapters comprise molecular barcodes.
58. The method of claim 43, wherein the tagging comprises incorporating molecular barcodes via ligation or PCR amplification.
59. The method of any one of claims 1 to 58, wherein the baseline cancer signal and / or the level of cancer and / or rate of change in cancer level from the additional bodily fluid sample is determined by a method comprising quantifying somatic mutations, tumor fraction, mutant allele fraction, and / or methylation signals.
60. The method of claim 59, wherein the baseline cancer signal and / or the level of cancer and / or rate of change in cancer level from the additional bodily fluid sample is determined by a method comprising the use of any of the biomarkers from claims 9 and / or 10.
61. The method of any one of claims 1 to 60, wherein the analysis in (a) and / or (b) is configured to be personalized for the subject based on the cancer or disease type the subject was previously determined to be, or suspected of being, at an increased risk of developing.
62. The method of claim 61, wherein the method comprises targeting biomarkers associated with a primary and / or secondary cancer.
63. The method of claim 62, wherein the targeting comprises the use of oligonucleotide capture probes or target amplification primers that hybridize or bind to nucleic acid sequences of cell- free nucleic acid molecules or amplicons thereof, wherein the oligonucleotide capture probes or target amplification primers are selected based the cancer or disease type the subject was previously determined to have or is suspected of having.
64. The method of any one of claims 1 to 63, wherein the method further comprises administering a therapy to the subject to treat the cancer(s).
65. The method of claim 64, wherein the therapy comprises P RP inhibitors, immunotherapies, a target -based therapy (such as kinase inhibitors), chemotherapy, or any combination thereof.
66. The method of claim 64, wherein the therapy comprises one or more treatments selected from Table 1.
67. A method for screening for cancer in a subject previously determined to be, or suspected of being, at an increased risk for developing cancer or a non-cancerous disease, comprising:(a) providing a first sample from the subject at a first time point and performing a first test using one or more analytes from the sample to determine a baseline cancer or disease signal;(b) at one or more subsequent time points, providing an additional sample from the subject and performing an additional test using one or more analytes from the additional sample to determine the level of cancer or non-cancerous disease and / or rate of change in cancer or non-cancerous disease level from the baseline cancer or disease signal; and(c) determining that the subject has cancer or non-cancerous disease if the level of cancer or disease and / or rate of change in cancer or non-cancerous disease level from the additional test exceeds a predetermined threshold above the baseline cancer or disease signal or does not have cancer or disease if the level of cancer or disease and / or rate ofchange in cancer or non-cancerous disease level from the additional test is at or below the predetermined threshold.
68. A method for screening a subject with hereditary cancer syndrome (HCS), comprising:(a) obtaining an initial sample from the subject and performing an assay on the sample, wherein the assay comprises extracting a plurality of molecules from the sample, and sequencing the plurality of molecules to obtain a baseline molecular profile for the subject with HCS;(b) repeating the assay on a subsequent sample obtained from the subject at a later time point to obtain a progressive molecular profile; and(c) comparing the progressive molecular profile to the baseline molecular profile using multidimensional analysis, where each dimension corresponds to a specific feature within the molecular profiles to identify changes across the features, wherein changes in the distribution, levels and / or grouping in the feature space are analyzed to determine a change in cancer signal.
69. A method for screening an individual with hereditary cancer syndrome (HCS), comprising:(a) providing a sample comprising cell-free DNA molecules from the individual;(b) enriching the cfDNA molecules or amplicons thereof for a plurality of target genomic regions to obtain enriched target genomic regions;(c) sequencing the enriched target genomic regions to generate sequencing reads; and(d) analyzing a plurality of sequencing reads to detect at least one or more somatic alterations in the target genomic molecules, and determining that the subject has cancer based on the detection of the one or more somatic alterations.
70. A method for determining a presence or absence of cancer in a subject previously determined to be, or suspected of being, at an increased risk for developing cancer, comprising:(a) providing a first sample from the subject at a first time point and analyzing one or more biomarkers from one or more analytes in the sample to determine a baseline cancer signal;(b) at one or more subsequent time points, providing an additional sample from the subject and analyzing one or more biomarkers from one or more analytes from the additional sample to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and(c) determining that the subject has cancer if the level of cancer and / or rate of change in cancer level from the additional diagnostic test exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level from the additional diagnostic test is at or below the predetermined threshold, wherein the analysis in (a) and / or (b) is personalized to target biomarkers associated with one or more cancers.
71. The method of claim 70, wherein the analysis in (a) and / or (b) comprises enriching cfDNA molecules or amplicons thereof for genomic or epigenomic regions targeting the biomarkers associated with the one or more cancers.
72. The method of claim 71, wherein the one or more cancers is a primary cancer.
73. The method of any one of claims 71 or 72, wherein the one or more cancers is, or also includes, a secondary cancer.
74. A method for determining that a subject is at an elevated risk for developing cancer, comprising:(a) providing a first bodily fluid sample from the subject and analyzing cell-free nucleic acid molecules or amplicons thereof from the sample to determine whether a genetic variant is somatic or germline in origin; and(b) determining that the genetic variant is a germline genetic variant, wherein the germline genetic variant is associated with hereditary cancer risk, thereby determining that the subject is at an elevated risk for developing cancer.
75. The method of claim 74, wherein the subject does not exhibit any symptoms of cancer.
76. The method of claims 74 or 75, further comprising, based on determining a presence of the germline genetic variant in the first bodily fluid sample, performing a subsequent medical procedure to determine that the subject does not have a primary cancer.
77. The method of claim 76, wherein the medical procedure comprises a colonoscopy, a stoolbased test, a mammogram, a Pap test (or Pap smear), an HPV test, endoscopy, ultrasound, computed tomography (CT), magnetic resonance imaging (MRI), and / or other imaging method.
78. The method of claims 76 or 77, further comprising:(c) providing a second bodily fluid sample from the subj ect at a time point after performing the medical procedure, and analyzing one or more biomarkers from the second bodily fluid sample to determine a baseline cancer signal;(d) providing one or more additional bodily fluid samples from the subject at one or more additional time points after the analyzing in (c), and analyzing one or more biomarkers from the one or more additional bodily fluid samples to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and(e) determining that the subject has cancer if the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold.
79. The method of claim 78, wherein the analyzing of (c) and / or the analyzing in (d) comprises customizing the analysis to target one or more biomarkers associated with the hereditary cancer risk.
80. The method of claim 78 or 79, wherein (c)-(e) are performed at intervals of about 1 to 6 month to monitor for one or more secondary cancers associated with the HCS.
81. The method of claims 78 or 79, wherein if the subject is determined to have cancer, performing an additional medical procedure to confirm the presence of the cancer.
82. The method of any one of claims 78 to 81, further comprising administering a therapy to the subject to treat the cancer.
83. The method of any one of claims 1-67 or 70-73, wherein the subject is determined to be at an increased risk of developing cancer by the method of any one of claims 74-77.
84. A method of screening an individual with hereditary cancer syndrome (HCS) comprising:(a) determining a baseline cancer signal using a non-invasive molecular test from an initial time point;(b) repeating the non-invasive molecular test at frequency intervals to determine the levels and / or rate of change from the baseline cancer signal;(c) following standard screening guidelines or gold-standard test for detecting a cancer of highest risk or primary cancer when the level or rate of change of cancer signal in (b) exceeds a pre-set threshold; and(d) utilizing a follow-up non-invasive molecular screening test for detection of secondary cancers with elevated risk.
85. A method of screening an individual with hereditary cancer syndrome (HCS) or at high-risk for developing at least one cancer type due to germline variants, comprising:(a) performing at least one gold standard test to confirm the individual is cancer-free;(b) performing a non-invasive molecular screening test to determine a cancer signal baseline score for the individual;(c) monitoring the individual by repeating the non-invasive molecular screening test in (b) for at least one time point; and(d) determining the level or rate of change of cancer signal at each time point, wherein the level or rate of change above pre-set threshold is reported as a positive result that the individual has cancer.
86. A method of screening an individual for cancer, comprising:(a) diagnosing the individual as having hereditary cancer syndrome (HCS) or is at an increased risk for certain cancer types due to germline variants;(b) performing at least one gold standard test to confirm the individual is cancer-free;(c) performing a non-invasive molecular screening test to determine a cancer signal baseline score for the individual;(d) monitoring the individual by repeating the non-invasive molecular screening test in (b) for at least one time point; and(e) determining the level or rate of change of cancer signal at each time point, wherein the level or rate of change above a pre-set threshold is reported as a positive result that the individual has cancer.
87. A method of screening an individual for cancer, comprising:(a) determining that the individual as having hereditary cancer syndrome (HCS) or is at an increased risk for certain cancer types due to germline variants;(b) performing at least one gold standard test to screen for a primary risk cancer at one or more time points; and(c) performing a non-invasive molecular screening test to screen for a secondary risk cancer at multiple time points, wherein a cancer signal baseline score for the individual is established at a first time point and a level or rate of change of cancer signal is determined at a second time point, wherein the level or rate of change above a pre-set threshold is reported as a positive result that the individual has a secondary cancer.
88. A method for screening cancer in a subject, comprising:(a) providing a first sample from the subject at a first time point and analzying one or more analytes for one or more biomarkers from the first sample for determining (i) one or more germline variants in genes associated with hereditary cancer syndrome (HCS) to diagnose HCS, and (ii) one or more molecular features indicative of cancer to establish a baseline cancer signal for subsequent monitoring;(b) at one or more subsequent time points, providing a second sample from the subject and analyzing one or more analytes for one or more biomarkers from the second sample for determining (i) somatic alterations arising in the HCS gene(s) harboring the germline variant(s), and / or (ii) the one or more molecular features to determine changes relative to the baseline cancer signal established in (a); and(c) determining that the subject has cancer when the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold.
89. The method of claim 88, wherein the genes associated with HCS comprise BRCA1 and / or BRCA2.
90. The method of claims 88 and 89, wherein the biomarkers are selected from the group consisting of mRNA expression, protein levels, methylation status, mutational signatures, and somatic variants in one or more genes in the homologous recombination deficiency (HRD) pathway.
91. The method of claim 90, wherein the mutational signatures comprise Single Base Substitution Signature 3 (SBS3), Rearrangement Signature 3 (RS3), and / or Rearrangement Signature 5 (RS5).
92. The method of any one of claims 88-91, wherein the determining in (c) further comprises: (i) preprocessing the first sample to generate a first set of standardized molecular feature values, (ii) inputting the first set of feature values into a trained model, (iii) receiving, from the model, an output comprising a baseline cancer signal score and / or a baseline feature vector, (iv) preprocessing the second sample to generate a second set of standardized molecular feature values; (v) inputting the second set of feature values into the same trained model to obtain a current cancer level score and / or a current feature vector, (vi) computing a change metric and / or rate of change between the current and baseline outputs, (vii) determining cancer statusby comparing the change metric and / or rate of change to a predetermined threshold, and (viii) outputting a result indicating a positive cancer status when the threshold is exceeded and a negative cancer status when it is not.
93. The method of claim 92, wherein the biomarkers comprise one or more of: methylation status, methylation levels and / or patterns, fragmentomic features including fragment length distributions, end motifs and / or breakpoint density profdes, mRNA expression values, protein expression levels, or combinations thereof.
94. The method of any one of claims 88-91, wherein the determining in (c) further comprises: preprocessing the first and second samples to generate respective first and second sets of standardized molecular feature values, inputting the first and second sets of standardized molecular feature values into the trained model, and receiving, from the model, a cancer status for the subject comprising a binary call.
95. The method of claim 92 or 94, wherein the model comprises a neural network trained on datasets including subjects with confirmed cancer and cancer-free controls, optionally including HCS carriers.
96. The method of any one of claims 88-95, wherein steps (b) and (c) are performed at intervals shorter than standard clinical guidelines, including intervals of about 1 to 3 months.
97. A method for screening cancer in a subject, comprising:(a) establishing a baseline cancer signal for the subject by analyzing one or more biomarkers from a first sample obtained from the subject at a first time point; and(b) monitoring the subject at one or more subsequent time points by obtaining one or more additional samples and analyzing the same biomarkers to determine changes relative to the baseline cancer signal.
98. The method of claim 97, further comprising determining that the subject has cancer when the change relative to the baseline cancer signal exceeds a predetermined threshold.
99. The method of any one of claims 1-98, wherein the subject is determined to be, or suspected of being, at an increased risk for developing cancer.
100. The method of any one of claims 88-99, wherein the monitoring is performed at intervals of about 1 to 6 months.
101. The method of any one of claims 1-100, wherein the subject is screened for a primary cancer of highest risk using a clinical test and for one or more secondary cancers using steps (a)-(b).
102. The method of claim 100, wherein the clinical test comprises one or more of colonoscopy, stool-based test, mammogram, breast MRI, Pap smear, endoscopy, ultrasound, computed tomography (CT), or magnetic resonance imaging (MRI).
103. A method for screening cancer in a subject who is identified or has been identified as having hereditary cancer syndrome, wherein the method comprises performing or having performed a clinical test on the subject for a primary cancer associated with the hereditary cancer syndrome at intervals which are at an increased frequency relative to the standard screening guidelines.
104. The method of claim 103 wherein:(i) the primary cancer is colorectal cancer and the clinical test is a colonoscopy;(ii) the primary cancer is breast cancer and the clinical test is a mammogram and / or breast MRI;(iii) the primary cancer is ovarian cancer and the clinical test is a transvaginal ultrasound and / or a serum CA-125 test;(iv) the primary cancer is endometrial cancer and the clinical test is an endometrial biopsy; or(v) the primary cancer is renal cell carcinoma and the clinical test is an abdominal MRI.
105. The method of claim 103 or claim 104, wherein the intervals are less than twelve months, less than six months, less than four months, less than three months, less than two months or less than one month.
106. The method of any one of claims 103-105, wherein the method further comprises:(a) providing a first bodily fluid sample from the subject at a first time point and analyzing one or more analytes for one or more biomarkers from the first bodily fluid sample to determine a baseline cancer signal;(b) at one or more subsequent time points, providing an additional bodily fluid sample from the subject and analyzing one or more analytes for one or more biomarkers from the additional bodily fluid sample to determine the level of cancer and / or rate of change in cancer level from the baseline cancer signal; and(c) determining that the subject has cancer if the level of cancer and / or rate of change in cancer level exceeds a predetermined threshold above the baseline cancer signal or does not have cancer if the level of cancer and / or rate of change in cancer level is at or below the predetermined threshold.
107. The method of claim 106, wherein the cancer determined to be present or absent via the analysis of the bodily fluid samples is a secondary cancer.
Citation Information
Patent Citations
Oligonucleotides
US20010053519A1
Method and apparatus for imaging a sample on a device
US20030152490A1
Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels
US20110160078A1
Methods and systems for adjusting tumor mutational burden by tumor fraction and coverage
US20200370129A1
Oligonucleotides
US6582908B2