Cell-free DNA blood-based test for cancer screening
A cell-free DNA-based CRC screening test using predictive models for tumor-derived nucleic acid analysis addresses invasiveness and adherence issues, offering a simpler and more effective screening option.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2024-05-06
- Publication Date
- 2026-05-19
AI Technical Summary
Current colorectal cancer (CRC) screening methods face barriers such as invasiveness, discomfort, and low adherence, necessitating a simpler and more accessible screening option like blood-based tests.
A CRC screening test using cell-free DNA (cfDNA) analysis, determining tumor-derived or non-tumor-derived nucleic acid samples through a predictive model based on tumor proportion regression (TFR) scores and epigenetic factors, including methylation patterns and genomic alterations.
Provides a less invasive and easier-to-administer screening method for CRC, potentially increasing adherence and early detection rates.
Smart Images

Figure 2026516012000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims the benefits based on U.S. Provisional Patent Application No. 63 / 500,480, filed on 5 May 2023, and U.S. Provisional Patent Application No. 63 / 614,350, filed on 22 December 2023, each of which is incorporated herein by reference in whole for all purposes. [Background technology]
[0002] background Colorectal cancer (CRC) is the third most common cancer diagnosed in men and women in the United States (US), and the second leading cause of cancer-related death. The lifetime risk of CRC in the US is approximately 4%, and 52,500 people are expected to die from the disease in 2023. Earlier detection of CRC impacts overall survival; the 5-year relative survival rate is 91% for patients with localized disease compared to 14% for patients with metastatic disease. Asymptomatic CRC screening reduces CRC incidence and mortality and is uniformly recommended in clinical guidelines published by major professional societies, including the US Preventive Services Taskforce (USPSTF), the US Multi-Society Taskforce, and the American Cancer Society (ACS). Numerous screening options are available, including direct visualization and stool-based tests. Despite the widely recognized benefits of CRC screening, the currently available options face significant barriers, resulting in approximately 59% of eligible individuals aged 45 and older being adherent. This falls far short of the 80% target set by the National Colorectal Cancer Roundtable (NCCRT), established by the Centers for Disease Control and Prevention (CDC) and the ACS. In addition, 76% of CRC-related deaths occur in individuals who were not re-screened. Therefore, there is an urgent need for CRC screening trials that are easier to administer and increase screening adherence.
[0003] Factors contributing to low support for CRC screening include the time required to perform the screening, challenges in scheduling colonoscopy, concerns about the invasiveness and pain of the test, fear of the test, discomfort or embarrassment associated with endoscopy, lack of insurance coverage, distance from the test provider, and lack of physician recommendations for screening. Integrating blood-based tests, which are drawn and completed as part of a routine medical encounter, into existing screening models would provide an additional screening option that is relatively simple to complete. [Overview of the Initiative] [Means for solving the problem]
[0004] Abstract A CRC screening test based on cell-free DNA (cfDNA) blood is disclosed herein. The CRC screening test includes determining whether a cell-free nucleic acid sample is tumor-derived or non-tumor-derived based on at least one of a cell-free nucleic acid score or a tumor proportion regression (TFR) score that meets respective thresholds, using a predictive model. The TFR score may be determined using a tumor proportion regression (TFR) model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples. The TFR score may include the proportion of molecules in several cell-free nucleic acid samples that indicate tumors. The cell-free nucleic acid score may indicate the presence of a tumor and may be based on at least one of epigenetic factors or genomic alterations in the cell-free nucleic acid sample. In some cases, epigenetic factors may include fragment mix data and logistic regression methylation data. In some cases, the cell-free nucleic acid score may also be based on the TFR score.
[0005] In addition, or alternatively, the method may include a step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0006] In addition, or or otherwise, the step of determining the quantification of observed tumor-associated abnormal methylation in each of multiple cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of multiple genomic regions.
[0007] In addition, or alternatively, the method may include a step of determining the epigenetic factors of each of several cell-free nucleic acid samples.
[0008] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
[0009] In addition, or otherwise, the method may include a step of using an LR model to determine whether a methylated LR model is cancerous or non-cancerous.
[0010] In addition, or otherwise, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0011] In addition, or alternatively, the method may include the step of determining cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples using a fragmentation model.
[0012] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0013] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0014] Additionally, or alternatively, the method may include determining genomic alterations in each of a plurality of cell-free nucleic acid samples.
[0015] Additionally, or alternatively, determining genomic alterations in each of a plurality of cell-free nucleic acid samples includes determining somatic variants observed in molecules each derived from a plurality of sequence fragments from the plurality of cell-free nucleic acid samples.
[0016] Additionally, or alternatively, the plurality of cell-free nucleic acid samples may be derived from a plurality of genomic regions. The plurality of genomic regions may include at least one of a genomic region known to be associated with a cancer type, a genomic region associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a treatment response.
[0017] Additionally, or alternatively, the plurality of genomic regions includes at least one genomic region known to be associated with colorectal cancer.
[0018] Additionally, or alternatively, determining the cell-free nucleic acid score is further based on the TFR score.
[0019] Additionally, or alternatively, the plurality of cell-free nucleic acid samples includes cell-free deoxyribonucleic acid (cfDNA) samples.
[0020] Additionally, or alternatively, the plurality of cell-free nucleic acid samples includes ribonucleic acid (RNA) samples.
[0021] Additionally, or alternatively, the plurality of cell-free nucleic acid samples includes cell-free ribonucleic acid (cfRNA) samples.
[0022] Additionally, or alternatively, the plurality of cell-free nucleic acid samples includes mitochondrial deoxyribonucleic acid (mtDNA) samples.
[0023] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
[0024] In addition, or separately, multiple cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
[0025] Another example of the method may include a step of determining a TFR score using a tumor proportion regression (TFR) model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples. The TFR score may represent the proportion of molecules in the several cell-free nucleic acid samples that indicate tumor. The method may further include a step of determining whether the several cell-free nucleic acid samples are tumor-derived or non-tumor-derived using a predictive model based on TFR scores that meet their respective thresholds.
[0026] In addition, or alternatively, the method may include a step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0027] In addition, or or otherwise, the step of determining the quantification of observed tumor-associated abnormal methylation in each of multiple cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of multiple genomic regions.
[0028] In addition, or alternatively, the step of determining whether multiple cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on a cell-free nucleic acid score indicating the presence of a tumor.
[0029] In addition, or alternatively, the method may include a step of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of several cell-free nucleic acid samples.
[0030] In addition, or alternatively, the method may include a step of determining the epigenetic factors of each of several cell-free nucleic acid samples.
[0031] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
[0032] In addition, or otherwise, the method may include a step of using an LR model to determine whether a methylated LR model is cancerous or non-cancerous.
[0033] In addition, or otherwise, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0034] In addition, or alternatively, the method may include the step of determining cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples using a fragmentation model.
[0035] In addition, or alternatively, the method may include a step of determining the genomic alteration of each of several cell-free nucleic acid samples.
[0036] In addition, or or the step of determining the genomic alterations of each of the multiple cell-free nucleic acid samples includes determining the somatic variants observed in the molecules derived from each of the multiple sequence fragments derived from the multiple cell-free nucleic acid samples.
[0037] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0038] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0039] In addition, or alternatively, multiple genomic regions may include at least one of the following: genomic regions known to be associated with cancer type, genomic regions known to be associated with a known methylation status, genomic regions known to be associated with hypomethylation, or genomic regions known to be associated with a therapeutic response.
[0040] In addition, or separately, multiple genomic regions include at least one genomic region known to be associated with colorectal cancer.
[0041] In addition, or separately, multiple cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
[0042] In addition, or separately, multiple cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
[0043] In addition, or separately, multiple cell-free nucleic acid samples include cell-free ribonucleic acid (cfRNA) samples.
[0044] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
[0045] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
[0046] In addition, or separately, multiple cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
[0047] Another example may include the steps of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of several cell-free nucleic acid samples, and determining whether the several cell-free nucleic acid samples are tumor-derived or non-tumor-derived using a predictive model based on the cell-free nucleic acid scores that meet their respective thresholds.
[0048] In addition, or otherwise, the step of determining whether multiple cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on a tumor proportion regression (TFR) score that meets a threshold. The TFR score may indicate the proportion of molecules in multiple cell-free nucleic acid samples that indicate tumors.
[0049] In addition, or alternatively, the method may include a step of determining a TFR score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0050] In addition, or alternatively, the method may include a step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0051] In addition, or or otherwise, the step of determining the quantification of observed tumor-associated abnormal methylation in each of multiple cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of multiple genomic regions.
[0052] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0053] In addition, or alternatively, the method may include a step of determining the epigenetic factors of each of several cell-free nucleic acid samples.
[0054] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
[0055] In addition, or otherwise, the method may include a step of using an LR model to determine whether a methylated LR model is cancerous or non-cancerous.
[0056] In addition, or otherwise, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0057] In addition, or alternatively, the method may include the step of determining cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples using a fragmentation model.
[0058] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0059] In addition, or alternatively, the method may include a step of determining that the cell-free nucleic acid score is further based on a tumor proportion regression (TFR) model. The TFR score may represent the proportion of molecules in multiple cell-free nucleic acid samples that indicate a tumor.
[0060] In addition, or alternatively, the method may include a step of determining the genomic alteration of each of several cell-free nucleic acid samples.
[0061] In addition, or or the step of determining the genomic alterations of each of the multiple cell-free nucleic acid samples includes determining the somatic variants observed in the molecules derived from each of the multiple sequence fragments derived from the multiple cell-free nucleic acid samples.
[0062] In addition, or otherwise, multiple cell-free nucleic acid samples may originate from multiple genomic regions. These multiple genomic regions may include at least one of the following: genomic regions known to be associated with oncology, genomic regions known to be associated with a known methylation status, genomic regions known to be associated with hypomethylation, or genomic regions known to be associated with a therapeutic response.
[0063] In addition, or alternatively, multiple genomic regions may include at least one genomic region known to be associated with colorectal cancer.
[0064] In addition, or alternatively, the step of determining whether multiple cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on methylation values that meet a threshold. The methylation score may indicate the quantity of molecules in multiple cell-free nucleic acid samples that indicate tumors.
[0065] In addition, or separately, multiple cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
[0066] In addition, or separately, multiple cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
[0067] In addition, or separately, multiple cell-free nucleic acid samples include cell-free ribonucleic acid (cfRNA) samples.
[0068] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
[0069] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
[0070] In addition, or separately, multiple cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
[0071] Predictive models can be trained using labeled cell-free nucleic acid sample data. Tumor prediction can be determined based on at least one of the cell-free nucleic acid score or TFR scores that meet their respective thresholds. Tumor prediction can be used to train predictive models along with tumor-derived or non-tumor-derived labels on the cell-free nucleic acid data. Using predictive models trained to detect CRC using cfDNA samples may be less invasive than traditional tests or screenings used to detect CRC.
[0072] In addition, or alternatively, the method may include a step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0073] In addition, or or otherwise, the step of determining the quantification of observed tumor-associated abnormal methylation in each of multiple cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of multiple genomic regions.
[0074] In addition, or alternatively, the method may include a step of determining the epigenetic factors of each of several cell-free nucleic acid samples.
[0075] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
[0076] In addition, or otherwise, the method may include a step of using an LR model to determine whether a methylated LR model is cancerous or non-cancerous.
[0077] In addition, or alternatively, the step of determining the epigenetic factors of each of the multiple cell-free nucleic acid samples is based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from the multiple cell-free nucleic acid samples.
[0078] In addition, or alternatively, the method may include the step of determining cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples using a fragmentation model.
[0079] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0080] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0081] In addition, or alternatively, the method may include a step of determining the genomic alteration of each of several cell-free nucleic acid samples.
[0082] In addition, or the step of determining the genomic alterations of each of the multiple cell-free nucleic acid samples includes determining the somatic variants observed in molecules derived from the multiple cell-free nucleic acid samples.
[0083] In addition, or otherwise, multiple cell-free nucleic acid samples are derived from multiple genomic regions. These multiple genomic regions may include at least one of the following: genomic regions known to be associated with oncology, genomic regions known to be associated with a known methylation status, genomic regions known to be associated with hypomethylation, or genomic regions known to be associated with a therapeutic response.
[0084] In addition, or alternatively, multiple genomic regions may include at least one genomic region known to be associated with colorectal cancer.
[0085] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0086] In addition, or separately, multiple cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
[0087] In addition, or separately, multiple cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
[0088] In addition, or separately, multiple cell-free nucleic acid samples include cell-free ribonucleic acid (cfRNA) samples.
[0089] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
[0090] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
[0091] In addition, or separately, multiple cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
[0092] Another example method may include the step of determining a TFR score using a tumor proportion regression (TFR) model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples. The TFR score may represent the proportion of molecules in the several cell-free nucleic acid samples that indicate tumor. Each of the several cell-free nucleic acid samples may be labeled with a tumor-derived label or a non-tumor-derived label. The example method may further include the step of determining a tumor prediction for each of the several cell-free nucleic acid samples based on the TFR score that satisfies each threshold. The example method may further include the step of generating a predictive model for predicting tumors in several cell-free nucleic acid samples based on the tumor-derived or non-tumor-derived labels and the tumor predictions for the several cell-free nucleic acid samples, and outputting the predictive model.
[0093] In addition, or alternatively, the method may include a step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0094] In addition, or or otherwise, the step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of several genomic regions.
[0095] In addition, or alternatively, the step of determining tumor prediction for each of multiple cell-free nucleic acid samples is further based on a cell-free nucleic acid score indicating the presence of a tumor.
[0096] In addition, or alternatively, the method may include a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of several cell-free nucleic acid samples.
[0097] In addition, or alternatively, the method may include a step of determining the epigenetic factors of each of several cell-free nucleic acid samples.
[0098] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
[0099] In addition, or otherwise, the method may include a step of using an LR model to determine whether a methylated LR model is cancerous or non-cancerous.
[0100] In addition, or otherwise, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0101] In addition, or alternatively, the method may include the step of determining cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples using a fragmentation model.
[0102] In addition, or alternatively, the method may include a step of determining the genomic alteration of each of several cell-free nucleic acid samples.
[0103] In addition, or or the step of determining the genomic alterations of each of the multiple cell-free nucleic acid samples includes determining the somatic variants observed in the molecules derived from each of the multiple sequence fragments derived from the multiple cell-free nucleic acid samples.
[0104] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0105] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0106] In addition, or otherwise, multiple cell-free nucleic acid samples may originate from multiple genomic regions. These multiple genomic regions may include at least one of the following: genomic regions known to be associated with oncology, genomic regions known to be associated with a known methylation status, genomic regions known to be associated with hypomethylation, or genomic regions known to be associated with a therapeutic response.
[0107] In addition, or separately, multiple genomic regions include at least one genomic region known to be associated with colorectal cancer.
[0108] In addition, or separately, multiple cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
[0109] In addition, or separately, multiple cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
[0110] In addition, or separately, multiple cell-free nucleic acid samples include cell-free ribonucleic acid (cfRNA) samples.
[0111] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
[0112] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
[0113] In addition, or separately, multiple cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
[0114] The example method may include the step of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of a group of cell-free nucleic acid samples. Each of the group of cell-free nucleic acid samples may be labeled with a tumor-derived label or a non-tumor-derived label. The example method may further include the step of determining a tumor prediction for each of the group of cell-free nucleic acid samples based on the cell-free nucleic acid scores that satisfy each threshold. The example method may further include the step of generating a predictive model for predicting tumors in the group of cell-free nucleic acid samples based on the tumor-derived label or the non-tumor-derived label and the tumor prediction for each of the group of cell-free nucleic acid samples, and outputting the predictive model.
[0115] In addition, or otherwise, the step of determining the tumor prediction for each of the multiple cell-free nucleic acids is further based on a tumor proportion regression (TFR) score that satisfies a threshold. The TFR score may represent the proportion of molecules in multiple cell-free nucleic acid samples that indicate a tumor.
[0116] In addition, or otherwise, the method includes the step of determining a TFR score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0117] In addition, or alternatively, the method may include a step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0118] In addition, or or otherwise, the step of determining the quantification of observed tumor-associated abnormal methylation in each of multiple cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of multiple genomic regions.
[0119] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0120] In addition, or alternatively, the method may include a step of determining the epigenetic factors of each of several cell-free nucleic acid samples.
[0121] In addition, or otherwise, the method may include the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
[0122] In addition, or otherwise, the method may include a step of using an LR model to determine whether a methylated LR model is cancerous or non-cancerous.
[0123] In addition, or otherwise, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0124] In addition, or alternatively, the method may include the step of determining cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples using a fragmentation model.
[0125] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0126] In addition, or alternatively, the step of determining the cell-free nucleic acid score may be further based on a tumor proportion regression (TFR) score. The TFR score may indicate the proportion of molecules in the multiple cell-free nucleic acid samples that represent a tumor.
[0127] In addition, or alternatively, the method may include a step of determining the genomic alteration of each of several cell-free nucleic acid samples.
[0128] In addition, or or the step of determining the genomic alterations of each of the multiple cell-free nucleic acid samples includes determining the somatic variants observed in the molecules derived from each of the multiple sequence fragments derived from the multiple cell-free nucleic acid samples.
[0129] In addition, or alternatively, multiple cell-free nucleic acid samples may originate from multiple genomic regions. These multiple genomic regions may include at least one of the following: genomic regions known to be associated with oncology, genomic regions known to be associated with a known methylation status, genomic regions known to be associated with hypomethylation, or genomic regions known to be associated with a therapeutic response.
[0130] In addition, or alternatively, multiple genomic regions include at least one genomic region known to be associated with colorectal cancer.
[0131] In addition, or alternatively, the step of determining whether multiple cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on methylation values that meet a threshold. The methylation score may indicate the quantity of molecules in multiple cell-free nucleic acid samples that indicate tumors.
[0132] In addition, or separately, multiple cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
[0133] In addition, or separately, multiple cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
[0134] In addition, or separately, multiple cell-free nucleic acid samples include cell-free ribonucleic acid (cfRNA) samples.
[0135] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
[0136] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
[0137] In addition, or separately, multiple cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
[0138] The example method may include the steps of detecting one or more biomarkers in a biological sample and determining a TFR score using a tumor proportion regression (TFR) model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples. The TFR score may represent the proportion of molecules in the several cell-free nucleic acid samples indicating a tumor. The example method may further include the steps of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations in each of the several cell-free nucleic acid samples and determining whether a biological sample is tumor-derived or non-tumor-derived based on at least one of the detected biomarkers, cell-free nucleic acid score, or TFR score that meet the respective thresholds.
[0139] In addition, or alternatively, the method may include a step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples.
[0140] In addition, or or otherwise, the step of determining the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of several genomic regions.
[0141] In addition, or alternatively, the method may include a step of determining the epigenetic factors of each of several cell-free nucleic acid samples.
[0142] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
[0143] In addition, or otherwise, the method may include the step of using the LR model to determine whether the methylated LR model is cancerous or non-cancerous.
[0144] In addition, or otherwise, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0145] In addition, or alternatively, the method may include the step of determining cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples using a fragmentation model.
[0146] In addition, or alternatively, the method may include the step of determining the epigenetic factors of each of several cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with multiple sequence fragments from multiple cell-free nucleic acid samples.
[0147] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0148] In addition, or alternatively, the method may include a step of determining the genomic alteration of each of several cell-free nucleic acid samples.
[0149] In addition, or or the step of determining the genomic alterations of each of the multiple cell-free nucleic acid samples includes determining the somatic variants observed in the molecules derived from each of the multiple sequence fragments derived from the multiple cell-free nucleic acid samples.
[0150] In addition, or otherwise, the method may include the possibility that multiple cell-free nucleic acid samples originate from multiple genomic regions. These multiple genomic regions may include at least one of the following: genomic regions known to be associated with oncology, genomic regions known to be associated with a known methylation status, genomic regions known to be associated with hypomethylation, or genomic regions known to be associated with a therapeutic response.
[0151] In addition, or separately, multiple genomic regions include at least one genomic region known to be associated with colorectal cancer.
[0152] In addition, or alternatively, the step of determining the cell-free nucleic acid score is further based on the TFR score.
[0153] In addition, or separately, multiple cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
[0154] In addition, or separately, multiple cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
[0155] In addition, or separately, multiple cell-free nucleic acid samples include cell-free ribonucleic acid (cfRNA) samples.
[0156] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
[0157] In addition, or separately, multiple cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
[0158] In addition, or separately, multiple cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
[0159] In addition, or alternatively, the biomarker may be one or more selected from proteins, exosomes, exomeres, microvesicles, apoptotic bodies, neutrophil extracellular traps (NETs), immune cells, tumor-educating platelets (TEPs), microbiome, pilomes, Toll-like receptors (TLRs), and mitochondrial DNA (mtDNA).
[0160] In addition, or or otherwise, the step of detecting one or more biomarkers includes detecting the presence or level of one or more biomarkers.
[0161] In addition, or otherwise, the step of determining whether a biological sample is tumor-derived or non-tumor-derived includes comparing the level of one or more biomarkers in the biological sample with a control.
[0162] In addition, or alternatively, the control is at a level present in the reference level or in healthy, non-cancer subjects.
[0163] Additional advantages of the disclosed methods and compositions are partially indicated in the description below, partially understood from the description, or can be learned through the practice of the disclosed methods and compositions. The advantages of the disclosed methods and compositions will be realized and achieved using the elements and combinations specifically indicated in the appended claims. It should be understood that both the above summary and the following detailed description are merely illustrative and descriptive of the claimed invention and not limiting.
[0164] The accompanying drawings incorporated herein and constituting part thereof illustrate several embodiments of the disclosed methods and compositions and, together with the description, serve to illustrate the principles of the disclosed methods and compositions. [Brief explanation of the drawing]
[0165] [Figure 1] Figure 1 is a flowchart illustrating an example of an artificial intelligence (e.g., machine learning) technique for generating a classifier configured to differentiate or classify tumor and non-tumor-derived nucleic acid variants in cell-free nucleic acid (cfDNA) samples obtained from a test subject.
[0166] [Figure 2] Figure 2 illustrates an example of a system for determining whether a test specimen is tumor-derived, according to one embodiment of the present disclosure.
[0167] [Figure 3A] Figure 3A illustrates a method for sequencing cfDNA molecules to obtain a methylation state vector.
[0168] [Figure 3B] Figure 3B is a diagrammatic representation of an example environment 307 for identifying nucleic acids corresponding to the classification region of a reference sequence, where the classification region has at least a threshold number of CpGs according to one or more implementations.
[0169] [Figure 4] Figure 4 shows an example of a terminal motif according to an embodiment of the present disclosure.
[0170] [Figure 5] Figure 5 illustrates an example of how the degree of overhang of cell-free DNA molecules (i.e., the overhang index) can be determined.
[0171] [Figure 6]Figure 6 shows an example of calculating methylation levels along DNA molecules after mapping to the human reference genome.
[0172] [Figure 7] Figure 7 shows how to determine the overhang index.
[0173] [Figure 8] Figure 8 is a flowchart illustrating an example method for generating a predictive model.
[0174] [Figure 9] Figure 9 is a flowchart illustrating an example training method for generating the ML module shown in Figure 8 using the training module shown in Figure 8.
[0175] [Figure 10] Figure 10 illustrates an exemplary process flow for classifying sequence fragments / reads and / or variants as tumor-derived or non-tumor-derived using a machine learning-based classifier.
[0176] [Figure 11] Figure 11 shows an illustrative process flow for a method of classifying nucleic acid samples as either tumor-derived or non-tumor-derived.
[0177] [Figure 12] Figure 12 shows an illustrative process flow for a method of classifying nucleic acid samples as either tumor-derived or non-tumor-derived.
[0178] [Figure 13] Figure 13 shows an illustrative process flow for a method of classifying nucleic acid samples as either tumor-derived or non-tumor-derived.
[0179] [Figure 14]Figure 14 shows an illustrative process flow for training a predictive model to classify nucleic acid samples as either tumor-derived or non-tumor-derived.
[0180] [Figure 15] Figure 15 illustrates an exemplary process flow for training a predictive model to classify nucleic acid samples as either tumor-derived or non-tumor-derived.
[0181] [Figure 16] Figure 16 shows an illustrative process flow for training a predictive model to classify nucleic acid samples as either tumor-derived or non-tumor-derived.
[0182] [Figure 17] Figure 17 is a chart showing the sensitivity of colorectal cancer according to the diagnostic stage. [Modes for carrying out the invention]
[0183] Detailed explanation The methods and compositions disclosed may be more readily understood by referring to the following detailed description of specific embodiments and the examples contained herein, as well as the drawings and the above and below descriptions relating thereto.
[0184] It should be understood that the methods and compositions disclosed are not limited to specific synthesis methods, specific analytical techniques, or specific reagents unless otherwise specified, and such may vary. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting.
[0185] Materials, compositions, and components that can be used in, in conjunction with, or in the preparation thereof, or are products thereof, are disclosed. These and other materials are disclosed herein, and where combinations, subsets, interactions, groups, etc., of these materials are disclosed, it is understood that each is specifically intended and described herein, even though specific references to various individual and collective combinations and rearrangements of each of these compounds may not be explicitly disclosed. For example, where peptides are disclosed and discussed, and several modifications that can be made to several molecules containing amino acids are discussed, all possible combinations and rearrangements of peptides and modifications are specifically intended unless the opposite is explicitly indicated. Thus, where classes A, B, and C of molecules are disclosed, along with classes D, E, and F and examples A-D of combined molecules, each is individually and collectively intended, even if each is not individually enumerated. Therefore, in this example, each of the combinations A-E, A-F, B-D, B-E, B-F, C-D, C-E, and C-F should be considered to be specifically contemplated and disclosed from the disclosures of A, B and C; D, E and F; and the exemplary combination A-D. Similarly, any subset or combination of these should also be considered specifically contemplated and disclosed. Thus, for example, the subgroups A-E, B-F, and C-E should be considered to be specifically contemplated and disclosed from the disclosures of A, B and C; D, E and F; and the exemplary combination A-D. This concept applies to all embodiments of the present application, including but not limited to steps in a method for preparing and using the disclosed compositions. Thus, where there are various additional steps that can be taken, each of these additional steps may be taken in conjunction with any specific embodiment or combination of embodiments of the disclosed method, and each such combination should be considered to be specifically contemplated and disclosed.
[0186] definition To facilitate understanding of this disclosure, certain terms are first defined below. Further definitions of the following terms and other terms may be provided throughout this specification. If any definition of a term provided below conflicts with a definition in a patent application or granted patent incorporated by reference, the definition provided here should be used to understand the meaning of the term.
[0187] Where used herein and in the appended claims, the singular forms “a,” “an,” and “the” include multiple subjects unless explicitly indicated otherwise by the context. Thus, for example, a reference to “a method” includes one or more methods and / or steps of the type described herein and / or of the type that would be apparent to those skilled in the art by reading this disclosure, etc. It will also be understood that there is an implicit “about” before temperatures, concentrations, times, number of bases or base pairs, coverage, etc., discussed herein, and therefore only a few, non-substantial equivalents are also within the scope of this disclosure. In this application, the use of the singular form includes the plural unless specifically stated otherwise. Also, “comprise,” “comprises,” “comprising,” “contain,” “contains,” “containing,” “include,” “includes,” and “including” are not intended to be limiting.
[0188] It should also be understood that the terminology used herein is intended solely to describe specific embodiments and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure pertains. The following terms, and their grammatical variants, will be used in descriptions of methods, computer-readable media, and systems, and in the claims, according to the definitions set forth below.
[0189] About: As used herein, “about” or “approximately” when applied to one or more values or elements of interest means a value or element that is similar to the reference value or element being referred to. In certain embodiments, unless otherwise stated or the context makes it clear that the term “about” or “approximately” means a range of values or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less than 1% of the reference value or element being referred to (except where such a number exceeds 100% of the possible values or elements).
[0190] Adapter: As used herein, “adapter” refers to a short nucleic acid (e.g., less than approximately 500 nucleotides, less than approximately 100 nucleotides, or less than approximately 50 nucleotides in length) that is typically at least partially double-stranded and used to ligate to either or both ends of a given sample nucleic acid molecule. An adapter may include a nucleic acid primer binding site for enabling amplification of the nucleic acid molecule whose ends are adjacent to the adapter, and / or a sequencing primer binding site, including a primer binding site for sequencing applications such as various next-generation sequencing (NGS) applications. An adapter may also include a binding site for a capture probe, such as an oligonucleotide bound to a flow cell support or similar. An adapter may also include nucleic acid tags as described herein. Nucleic acid tags are typically positioned relative to amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequencing read of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of a nucleic acid molecule. In certain embodiments, the same adapter is attached to each end of a nucleic acid molecule, except that the nucleic acid tag differs in its sequence. In some embodiments, the adapter is a Y-shaped adapter having one end blunt-ended or having a tail as described herein, for binding to a nucleic acid molecule that is also blunt-ended, or to a nucleic acid molecule having a tail with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter including an end with a blunt end or tail for binding to the nucleic acid to be analyzed. Other exemplary adapters include an adapter having a T tail and an adapter having a C tail.
[0191] To administer: As used herein, “to administer” or “to administer” a therapeutic agent (e.g., an immunological therapeutic agent, a DNA damage response (DDR) inhibitor (e.g., a poly(ADP-ribose) polymerase (PARP) inhibitor (PARPi))) to a subject means to give, apply, or bring into contact with the subject the composition. Administration may be carried out by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal routes.
[0192] Align: As used herein, “align,” “alignment,” and “to align” in the context of nucleic acids refer to aligning DNA or RNA sequences to identify regions of similarity. Similarity may relate to functional, structural, and / or evolutionary relationships between sequences. DNA sequence alignment includes alignment of the genomic DNA of one sequence with the genomic DNA of at least one other sequence. Such alignment may exclude non-genomic DNA, such as molecular barcodes, padding bases, and similar elements. For example, the genomic DNA of a sequence read may be aligned with the genomic DNA of a reference DNA sequence, excluding any molecular tags that may be bound to the sequence read.
[0193] Allele: As used herein, “allele” or “allele variant” refers to a specific genetic variant at a defined gene location or genomic locus. Allele variants are typically expressed with a frequency of 50% (0.5) or 100%, depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and typically have a frequency of 0.5 or 1. Somatic variants, however, are acquired variants and typically have a frequency of <0.5. Major and minor alleles of a gene locus refer to nucleic acids that have loci occupied by a reference sequence nucleotide and a variant nucleotide different from the reference sequence, respectively. Measurements at a gene locus may take the form of allele proportions (AF), which are measures of the frequency of the allele found in a sample.
[0194] Amplification: As used herein, "amplification" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or a portion of a polynucleotide, usually starting with a small amount of polynucleotide (e.g., a single polynucleotide molecule), and the amplified product or amplicon is generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes.
[0195] Barcode: As used herein, “barcode” in the context of nucleic acids refers to a nucleic acid molecule having a sequence that can function as a molecular identifier (molecular barcode), a partition identifier (partition barcode), or a sample identifier (sample barcode or sample index). For example, individual “barcode” sequences are typically added to DNA fragments during next-generation sequencing (NGS) library preparation so that each read can be identified and sorted before final data analysis.
[0196] Breakpoint: As used herein, “breakpoint” in the context of a nucleic acid fusion molecule or a corresponding sequencing read refers to a terminal nucleotide position at the junction between fused subsequences of a nucleic acid fusion, or a terminal nucleotide position represented by the corresponding sequencing read. For example, a given split sequence read contains a first subsequence that is contiguous with a second subsequence in the split sequence read and is located 5' to the second subsequence, and this first subsequence is located at a first gene locus in a reference sequence, which is discontinuous with a second gene locus in the reference sequence where the second subsequence is located. In this example, the first subsequence of the split sequence read contains a breakpoint at its 3' terminal nucleotide, while the second subsequence of the split sequence read contains a breakpoint at its 5' terminal nucleotide. In certain applications, breakpoints, for example, these breakpoints, are referred to as a “breakpoint pair.”
[0197] Oncology type: As used herein, “cancer,” “oncology type,” or “tumor type” refers to a type or subtype of cancer as defined, for example, by histopathological diagnosis. Oncology type is defined by any conventional criteria, for example, by the presence in a given tissue (e.g., hematological cancer, central nervous system (CNS) cancer, brain cancer, lung cancer (small cell and non-small cell), skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer). Cancer can be defined based on cancers that present cancer markers, such as solid cancers, heterologous cancers, allologous cancers, cancers of unknown cause, and similar cancers, as well as cancers of the same cell lineage (e.g., carcinomas, sarcomas, lymphomas, cholangiocarcinomas, leukemias, mesotheliomas, melanomas, or glioblastomas), and cancers that present cancer markers, such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, KRAS, BRAF, NRAS, hormone receptors, and NMP-22. Cancer can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether it is primary or secondary.
[0198] Cell-free nucleic acids: As used herein, “cell-free nucleic acids” refers to nucleic acids that are not contained within cells and are not otherwise bound to cells. In some embodiments, “cell-free nucleic acids” refers to nucleic acids that, at the time of isolation from the subject, are not contained within cells and are not otherwise bound to cells. Cell-free nucleic acids may include, for example, all unencapsulated nucleic acids supplied from bodily fluids from the subject (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, nucleolar small RNA (snoRNA), Piwi-binding RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids may be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids by secretion or cell death processes, such as cell necrosis, apoptosis, or similar processes. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. ctDNA can be unencapsulated tumor-derived fragmented DNA. Another example of cell-free nucleic acids is fetal DNA that circulates freely in the maternal bloodstream, also known as cell-free fetal DNA (cffDNA). Cell-free nucleic acids may have one or more epigenetic modifications; for example, cell-free nucleic acids may be acetylated, 5-methylated, ubiquitinated, phosphorylated, SUMOylated, ribosylated, and / or citrullinated.
[0199] Cellular origin: As used herein, “cellular origin” in the context of cell-free nucleic acids means the cell type from which a given cell-free nucleic acid molecule originates or otherwise arises (e.g., by an apoptotic process, a necrotic process, or similar). In certain embodiments, for example, a given cell-free nucleic acid molecule may originate from tumor cells (e.g., cancerous lung cells) or from non-tumor cells or normal cells (e.g., non-cancerous lung cells).
[0200] Classification Region: As used herein, “classification region” refers to a genomic region that can exhibit sequence-independent changes in neoplastic cells (e.g., tumor cells and cancer cells) or sequence-independent changes in cfDNA from a cancerous subject compared to cfDNA from a subject without cancer. Examples of sequence-independent changes include, but are not limited to, changes in methylation rates (increase or decrease), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. In one or more examples, sequence-independent changes in a classification region can indicate the presence of a single form of cancer in a subject. In one or more additional examples, sequence-independent changes in a classification region can correspond to the presence of multiple forms in a subject. A classification region can be enriched by one or more probes. In addition, a classification region can be defined by a pair of primer-binding sites. Furthermore, a classification region can be defined by a predetermined beginning genomic locus and a predetermined ending genomic locus. A classification region may consist of approximately 25 to 250 nucleotides, approximately 50 to 200 nucleotides, or approximately 75 to 150 nucleotides. For example, a classification region may be a variable methylation region. A "variable methylation region" or "DMR" refers to a region of DNA that has a detectably different degree of methylation in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type; or that has a detectably different degree of methylation in at least one cell or tissue type obtained from a diseased or impaired subject compared to the degree of methylation in the same region of DNA from the same cell or tissue type obtained from a healthy subject. In some embodiments, a variable methylation region has a detectably higher degree of methylation (e.g., a hypermethylated region / hypermethylated target region) in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type contributing to cfDNA in a healthy individual or from the same cell or tissue type obtained from a healthy subject.In some embodiments, the variable methylation region has a detectably lower degree of methylation (e.g., a hypomethylated region / hypomethylated target region) in at least one cell or tissue type that contributes to cfDNA in a healthy individual, for example, the degree of methylation in the same region of DNA from other immune cell types and / or cell types or from the same cell or tissue type derived from a healthy subject. In some embodiments, the classification region includes a hypermethylated target region and / or a hypomethylated target region.
[0201] Classifier: As used herein, “classifier” generally refers to an algorithmic computer code that receives test data as input data and produces output data of the classification of the input data as belonging to one or another class (e.g., tumor DNA or non-tumor DNA, having or not having DNA damage repair deficiency (DDRD)).
[0202] Contiguous Sequence: As used herein, “contiguous sequence” or “contig” refers to a set of overlapping nucleic acid segments that together represent a consensus region of nucleic acids.
[0203] Copy number variant: As used herein, “copy number variant,” “CNV,” or “copy number diversity” refers to the phenomenon in which genomic compartments are repeated and the number of repeats within the genome differs among individuals in the population under consideration.
[0204] Coverage: As used herein, the terms “coverage,” “total molecular count,” and “total allele count” are used synonymously. These terms refer to the total number of DNA molecules at a specific genomic location in a given sample.
[0205] Deoxyribonucleic acid or ribonucleic acid: As used herein, “deoxyribonucleic acid” or “DNA” refers to a natural or modified nucleotide having a hydrogen group at the 2' position of the sugar moiety. Typically, DNA refers to a nucleotide chain containing a deoxyribonucleoside with one of four types of nucleic acid bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, “ribonucleic acid” or “RNA” refers to a natural or modified nucleotide having a hydroxyl group at the 2' position of the sugar moiety. Typically, RNA refers to a nucleotide chain containing a ribonucleoside with one of four types of nucleic acid bases: A, uracil (U), G, and C. As used herein, the term “nucleotide” refers to a natural or modified nucleotide. Certain nucleotide pairs bind specifically to each other in a complementary manner (this is called complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand is joined to a second nucleic acid strand composed of nucleotides complementary to those in the first strand, these two strands join to form a double helix. As used herein, “nucleic acid sequencing data,” “nucleic acid sequencing information,” “sequence information,” “nucleic acid sequence,” “nucleotide sequence,” “genome sequence,” “gene sequence,” or “fragment sequence,” or “nucleic acid sequencing read” means any information or data indicating the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a nucleic acid molecule such as DNA or RNA (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment).It should be understood that this instruction intends to describe sequence information obtained using any available type of technique, platform, or technology, including but not limited to capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion or pH-based detection stems, and digital signature-based systems.
[0206] To detect: As used herein, “to detect,” “to detect,” or “detect” means the act of determining the existence or presence of one or more target nucleic acids (e.g., nucleic acids having a target mutation or other marker) in a sample.
[0207] Concentrated Sample: As used herein, “concentrated sample” refers to a sample in which a specific region of interest has been concentrated. A sample can be concentrated by amplifying the region of interest or by using a single-stranded DNA / RNA probe or a double-stranded DNA probe (e.g., SureSelect® probe, Agilent Technologies) that can hybridize to the nucleic acid molecule of interest. In some embodiments, a concentrated sample refers to a subset or portion of the processed sample to be concentrated, and this subset or portion of the processed sample to be concentrated contains nucleic acid molecules from a cell-free polynucleotide sample or a polynucleotide sample.
[0208] Epigenetic Information: As used herein, “epigenetic information” in the context of a DNA polymer means one or more epigenetic patterns or signatures exhibited in that polymer.
[0209] Epigenetic loci: As used herein, “epigenetic loci” or “epigenetic site” means a fixed location on a chromosome that exhibits a different state or status without alteration or modification of the nucleotide sequence. To avoid misunderstanding, a given epigenetic loci may coincide with a given nucleotide location or genomic region that also exhibits genetic or sequence diversity (e.g., mutation). For example, a given epigenetic loci may or may not be acetylated, methylated (e.g., modified with 5-methylcytosine (5mC), modified with 5-hydroxymethylcytosine (5hmC), and / or similarly modified), ubiquitinated, phosphorylated, SUMOylated, ribosylated, citrullinated, have post-translational histone modifications or other histone diversity, and / or similar.
[0210] Epigenetic signature: As used herein, “epigenetic signature” means an epigenetic state or status indicated by one or more epigenetic loci in a given DNA molecule. For example, a DNA molecule or cfDNA fragment containing a given genomic region or locus (e.g., a CTCF-binding region) may exhibit an epigenetic pattern in which a certain number of epigenetic loci are methylated in some of these DNA molecules, while in other cases the corresponding epigenetic loci in other DNA molecules or cfDNA fragments containing the same genomic region are not methylated. “Methylation signature” means an epigenetic signature associated with a methylation state or status indicated by one or more epigenetic loci in a given DNA molecule.
[0211] Fusion event: As used herein, “fusion event” refers to the fusion of at least two separate genes at a specific location. Examples of causes of fusion events include translocations, intermediate deletions, or chromosomal inversions.
[0212] Gene: As used herein, “gene” refers to any segment of DNA associated with a biological function. Thus, a gene includes coding sequences and, if necessary, regulatory sequences required for their expression. A gene also includes, if necessary, non-expressed DNA segments that form recognition sequences for, for example, other proteins.
[0213] Genomic region: As used herein, “genomic region” means a fixed location or compartment on a chromosome, such as the location of a gene or genomic marker. Exemplary genomic markers include transcription factor binding regions (e.g., CTCF binding regions), distal regulatory elements (DREs), repeat elements (e.g., microsatellites), intron-exon or exon-intron junctions, transcription start sites (TSSs), and similar structures.
[0214] Germline mutation: As used herein, “germline mutation” means a mutation in germ cells, and therefore a mutation that can be passed on to offspring.
[0215] Indel: As used herein, “indel” refers to a mutation that involves the insertion or deletion of a nucleotide in the genome in question.
[0216] Machine Learning Algorithms: As used herein, “machine learning algorithms” generally refers to algorithms performed by a computer that automate the construction of analytical models, such as clustering, classification, or pattern recognition. Machine learning algorithms may be supervised or unsupervised. Examples of learning algorithms include artificial neural networks (e.g., backpropagation networks), discriminant analysis (e.g., Bayesian classifiers or Fisher analysis), support vector machines, decision trees (e.g., recursive partitioning processes, e.g., CART classification and regression trees, or random forests), linear classifiers (e.g., linear multiple regression (MLR), partial least squares (PLS) regression, and principal component regression), hierarchical clustering, and cluster analysis. The datasets on which a machine learning algorithm learns may be called “training data.”
[0217] Match: As used herein, “match” means that at least one value or element is at least substantially equal to at least one second value or element. In certain embodiments, for example, the cellular origin of at least a subset of DNA molecules from a cfDNA sample is determined when there is at least a substantial or approximate match between the test sample distribution of cfDNA fragment properties and the reference sample distribution of cfDNA fragment properties.
[0218] Minor allele frequency: As used herein, “minor allele frequency” refers to the frequency at which a minor allele (e.g., not the most frequent allele) is present in a given nucleic acid population, such as a sample obtained from a subject. Genetic variants with low minor allele frequencies are typically relatively infrequent in the sample.
[0219] Mutant Allele Proportion: As used herein, “mutant allele proportion” or “MAF” refers to the proportion of nucleic acid molecules in a given sample that contain an allele modification or mutation relative to a given genomic location. MAF is generally expressed as a proportion or percentage. For example, MAF is typically about 0.5, 0.1, 0.05, or less than 0.01 (i.e., about 50%, 10%, 5%, or less than 1%) of all somatic variants or alleles present at a given locus.
[0220] Maximum Mutagenic Allele Proportion: As used herein, “Maximum Mutagenic Allele Proportion,” “Maximum MAF,” or “MAX MAF” refers to the maximum or largest MAF of all somatic variants present or observed in a given sample.
[0221] Mutation: As used herein, “mutation,” “nucleic acid variant,” “variant,” or “genetic abnormality” refers to a variation from a known reference sequence, including, for example, single nucleotide variants (SNVs), copy number variants or variations (CNVs) / abnormalities, insertions or deletions (indels), shortenings, gene fusions, transversions, translocations, frameshifts, duplications, repeat extensions, and epigenetic variants. Mutations can be germline or somatic mutations. In some embodiments, the reference sequence for comparison is the wild-type genome sequence of the species under consideration, typically the human genome, which provides the test sample. In certain cases, the mutation or variant is a “tumor-associated genetic variant” that causes tumorigenesis or at least contributes to tumorigenesis.
[0222] Negative control region: As used herein, “negative control region” refers to a genomic region that is expected to be unmethylated or hypomethylated in all samples, regardless of whether the DNA originates from cancer cells or normal cells.
[0223] Next-generation sequencing: As used herein, “next-generation sequencing” or “NGS” refers to sequencing techniques that offer increased throughput compared to conventional Sanger and capillary electrophoresis-based methods, such as the ability to simultaneously generate hundreds of thousands of relatively short sequence reads. Some examples of next-generation sequencing techniques include, but are not limited to, single-nucleotide synthesis, ligation sequencing, and hybridization sequencing.
[0224] Nucleic acid tags: As used herein, “nucleic acid tags” means short nucleic acids (e.g., less than approximately 500, 100, 50, or 10 nucleotides in length) used to label nucleic acid molecules to distinguish nucleic acids from different samples of different types or those that have undergone different treatments (e.g., representing a sample index), or nucleic acids used to label nucleic acid molecules to distinguish different nucleic acid molecules in the same sample of different types or those that have undergone different treatments (e.g., representing a molecular tag). Nucleic acid tags may be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags may have the same length or a variety of lengths, as necessary. Nucleic acid tags may contain a double-stranded molecule with one or more blunt ends, a 5' or 3' single-stranded region (e.g., an overhang), and / or one or more other single-stranded regions at other positions within a given molecule. Nucleic acid tags may be conjugated to one or both ends of another nucleic acid (e.g., a sample nucleic acid to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information about a given nucleic acid, such as its origin, morphology, or processing. Nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different nucleic acid tags and / or sample indices, which are then deconvoluted by reading the nucleic acid tags. Nucleic acid tags are sometimes called molecular identifiers or tags, sample identifiers, index tags, and / or barcodes. In addition or alternatively, nucleic acid tags can be used to distinguish different molecules within the same sample. This includes, for example, uniquely tagging different nucleic acid molecules within a given sample, or non-uniquely tagging such molecules. In the case of non-unique tagging applications, nucleic acid molecules can be tagged using tags with a limited number of different sequences, and thus different molecules can be distinguished, for example, based on their start and / or stop locations in a selected reference genome in combination with at least one nucleic acid tag.Typically, a sufficient number of different nucleic acid tags are used, and therefore, the probability that any two molecules will have the same start / stop position and also have the same nucleic acid tag is low (e.g., less than approximately 10%, less than approximately 5%, less than approximately 1%, or less than approximately 0.1%). Some nucleic acid tags include multiple molecular identifiers to label a sample, the morphology of the nucleic acid molecule in the sample, and the nucleic acid molecule in a morphology that has the same start and stop position. Such nucleic acid tags can be referred to using the exemplary morphology "A1i", where uppercase letters indicate the sample type, Arabic numerals indicate the morphology of the molecule in the sample, and lowercase Roman numerals indicate the molecule in a particular morphology.
[0225] Polynucleotides: As used herein, “polynucleotides,” “nucleic acids,” “nucleic acid molecules,” or “oligonucleotides” refer to linear polymers of nucleosides linked by nucleoside linkages (deoxyribonucleosides, ribonucleosides, or analogs thereof). Typically, polynucleotides contain at least three nucleosides. Oligonucleotides often range in size from a small number of monomer units, e.g., 3-4, to several hundred. Whenever polynucleotides are represented by a sequence of letters such as “ATGCCTG,” it will be understood that, unless otherwise noted, the nucleotides are in 5'→3' order from left to right, and in the case of DNA, “A” represents deoxyadenosine, “C” represents deoxycytidine, “G” represents deoxyguanosine, and “T” represents deoxythymidine. The letters A, C, G, and T may be used to refer to the base itself, to refer to a nucleoside, or to refer to a nucleotide containing a base, as is common in the art.
[0226] Positive control region: As used herein, “positive control region” refers to a genomic region that is expected to be methylated or hypermethylated in all samples, regardless of whether the DNA originates from cancer cells or normal cells.
[0227] Prevalence: As used herein, “prevalence” in the context of nucleic acid variants means the extent, prevalence, or frequency to which a given nucleic acid variant is found or observed in a given sample (e.g., a given body fluid sample, a given non-body fluid sample, etc.) or other population (e.g., a given population of body fluid samples, a given population of non-body fluid samples, etc.).
[0228] Reference Sample: As used herein, “reference sample” or “reference cfDNA sample” refers to a sample of known composition and / or known to have, or to have or lack, certain characteristics (e.g., known nucleic acid variants, known cellular origin, known tumor percentage, known coverage, and / or similar) that is analyzed together with or compared to a test sample to evaluate the accuracy of an analytical procedure. A reference sample dataset typically contains at least about 25 to at least about 30,000 or more reference samples. In some embodiments, the reference sample dataset includes approximately 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000, or more reference samples.
[0229] Reference Sequence: As used herein, “reference sequence” or “reference genome” refers to a known sequence used for comparison with an experimentally determined sequence. For example, a known sequence may be an entire genome, a chromosome, or any segment thereof. A reference sequence typically contains at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1,000, at least about 10,000, at least about 100,000, at least about 1,000,000, at least about 10,000,000, at least about 1,000,000, or more nucleotides. A reference sequence may be aligned with a single contiguous sequence of genome or chromosome, or it may contain discontinuous segments that align with different regions of genome or chromosome. Examples of reference sequences include, for instance, the human genome, specifically hG19 and hG38.
[0230] Sample: As used herein, “Sample” means any biological sample that can be analyzed by the methods and / or systems disclosed herein. In certain aspects of this disclosure, Sample is a fluid sample, in particular whole blood or a fraction thereof, lymph, urine, and / or cerebrospinal fluid, among the types of fluids supplied with cell-free (circulating, not contained within cells or otherwise bound to cells) nucleic acids. In certain implementations, a fluid sample is a plasma sample, which is the fluid portion of whole blood excluding cells such as red blood cells and white blood cells. In some implementations, a fluid sample is a serum sample, i.e., fibrinogen-free plasma. In certain aspects of this disclosure, Sample is a “non-fluid sample” or “non-plasma sample,” i.e., a biological sample other than a “fluid sample,” such as a cell and / or tissue sample supplied with nucleic acids other than cell-free nucleic acids.
[0231] Sensitivity: As used herein, “sensitivity” in the context of a given assay or method refers to the ability of the assay or method to detect and distinguish between target (e.g., nucleic acid variant) analytes and non-target analytes.
[0232] Sequence fragment: As used herein, “sequence fragment” refers to a portion of a nucleic acid molecule that may vary in length and may contain sequence information (or sequence data) of the nucleic acid molecule. Sequence information may be derived from sequencing reads obtained from sequencing of the sequence fragment.
[0233] Sequence Read: As used herein, “sequence read” refers to a sequence of base pairs corresponding to all or part of a sequence fragment.
[0234] Sequencing: As used herein, “sequencing” refers to any of several techniques used to determine the sequence (e.g., identity and order of monomeric units) of a biomolecule, such as nucleic acids, such as DNA or RNA. Exemplary sequencing methods include targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, sungardideoxytermination sequencing, whole-genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, and single-nucleotide extension sequencing. Examples of sequencing methods include, but are not limited to, solid-phase sequencing, high-throughput sequencing, large-scale parallel signature sequencing, emulsion PCR, co-amplification-PCR at lower denaturation temperatures (COLD-PCR), multiplex PCR, sequencing by reversible dye terminators, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, single-nucleotide synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD® sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed using a gene analyzer, for example, a gene analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among others.
[0235] Sequence information: As used herein, “sequence information” in the context of nucleic acid polymers means the order and / or identity of monomer units (e.g., nucleotides, etc.) within the polymer.
[0236] Sequence motif: As used herein, “sequence motif” may refer to a short, repeating pattern of bases in a DNA fragment (e.g., a cell-free DNA fragment). Sequence motifs may be located at the ends of a fragment and may therefore be part of or contain a termination sequence. “Terminal motif” may refer to a sequence motif of a termination sequence that is preferentially located at the ends of a DNA fragment, potentially in a particular type of tissue. Terminal motifs may also correspond to a termination sequence by being located immediately before or immediately after the end of the fragment. Nucleases may have a specific cleavage preference for a particular terminal motif, as well as a second-highest cleavage preference for a second terminal motif.
[0237] Single nucleotide variant: As used herein, “single nucleotide variant” or “SNV” means a variation or diversity of a single nucleotide at a specific location within the genome.
[0238] Somatic mutation: As used herein, “somatic mutation” means a given mutation in the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0239] Specificity: As used herein, “specificity” in the context of a diagnostic analysis or assay refers to the degree to which the analysis or assay detects the intended target analyte while excluding other components of a given sample.
[0240] Status: As used herein, “status” in the context of the subject means one or more conditions of a given subject, for example, whether or not the subject has cancer.
[0241] Subject: As used herein, “subject” or “test subject” means an animal, e.g., a mammalian species (e.g., human) or a bird species (e.g., bird), or another organism, e.g., a plant. More specifically, a subject may be a vertebrate, e.g., a mammal, e.g., a mouse, a primate, a monkey, or a human. Animals include livestock (e.g., production cattle, dairy cows, poultry, horses, pigs, and similar animals), sports animals, and companion animals (e.g., pets or support animals). A subject may be a healthy individual, an individual with or suspected to have a disease or disease predisposition, or an individual in need of therapy or suspected to need therapy. The terms “individual” or “patient” are intended to be synonymous with “subject.” In some embodiments, a subject is a human being who has or is suspected to have cancer. For example, a subject may be an individual diagnosed with cancer, an individual who will undergo cancer therapy, and / or an individual who has received at least one cancer therapy. A subject may also be in remission from cancer. In another example, a subject may be an individual diagnosed with an autoimmune disease. In another example, the subject may be a pregnant woman who has been diagnosed with, or is suspected of having, a disease such as cancer or an autoimmune disease, or a woman who is planning to become pregnant. A “reference subject” refers to a subject that is known to have or lack certain characteristics (e.g., known cancer or disease status, known nucleic acid variant, known cellular origin, known tumor rate, known coverage, and / or similar).
[0242] Threshold: As used herein, “threshold” refers to a separately determined value used to characterize or classify experimentally determined values. In certain embodiments, for example, “threshold value” refers to a selected value from which a quantitative value is compared to determine whether a given target nucleic acid variant is absent at a given gene locus.
[0243] Tumor Percentage: As used herein, “tumor percentage” refers to an estimate of the proportion of nucleic acid molecules in a given sample that originate from tumors. For example, the tumor percentage of a sample may be the maximum variant allele frequency (MAX MAF) of the sample, or the coverage of the sample, or a measure derived from the length, epigenetic state, or other properties of cfDNA fragments in the sample, or any other selected feature of the sample. The term “MAX MAF” refers to the maximum or greatest MAF of all somatic variants present in a given sample. In some embodiments, the tumor percentage of a sample is equal to the MAX MAF of the sample.
[0244] Value: As used herein, “Value” or “Score” generally refers to any entry in a dataset that may characterize the feature pointed to by the Value. This includes, but is not limited to, numbers, words or phrases, symbols (e.g., + or -) or degrees. B. Introduction
[0245] Methods and systems for differentiating or classifying tumor and non-tumor-derived nucleic acid variants in nucleic acid samples obtained from test subjects are provided herein. In some embodiments, the methods and systems link genomic modification data (e.g., somatic genomic data) with epigenetic data (e.g., methylation data, fragment mix data). In some embodiments, the nucleic acid sample may be cell-free nucleic acid (cfNA), genomic DNA, or RNA, but is not limited to these.
[0246] Figure 1 is a schematic flowchart illustrating an example of an artificial intelligence (e.g., machine learning) technique for generating a classifier configured for differentiating or classifying tumor-derived and non-tumor-derived nucleic acid variants in cell-free nucleic acid (cfDNA) samples obtained from test subjects. As shown, Method 100 may include, in step 102, obtaining data, for example, in the form of cancer (e.g., tumor)-derived and non-tumor-derived sequence data from cell-free nucleic acid (cfDNA) samples of multiple subjects. Method 100 may also include obtaining epigenetic data and / or genomic modification data related to or otherwise derived from the sequence data. Both epigenetic data and genomic modification data can be determined from genomic regions in the cfDNA samples. Epigenetic data may include, for example, DNA methylation, histone state or modification, inflammation-mediated cytosine damage products, protein binding, fragment mix data (information on fragment size, nucleotide motifs at fragment ends, single-strand jagged ends, and / or gene location at fragment endpoints), or the state of other molecules reflected in the analyzed nucleic acid fragment that cannot be determined solely from the nucleotide sequence, such as information on the methylation state of a given base or set base. In some embodiments, epigenetic data and genome modification data of nucleic acid sequences known to be tumor-derived may be labeled as tumor-derived, and epigenetic data and genome modification data of nucleic acid sequences known to be non-tumor-derived may be labeled as non-tumor-derived. Furthermore, additional labels, such as cancer type, histological type, and similar may be assigned.
[0247] In some embodiments, the methods disclosed herein, as well as associated systems and computer-readable media implementations, involve identifying a set of DNA molecules or cfDNA fragments from a cfDNA sample, wherein each member cfDNA fragment of a given set contains a genomic region common to one another. Essentially any genomic region can be used if the cfDNA fragment containing a given genomic region exhibits different characteristics between at least two cell or tissue types (e.g., cfDNA fragment length, offset of the midpoint of the cfDNA fragment relative to the midpoint of the genomic region contained in the cfDNA fragment, epigenetic state, and / or similar). In certain embodiments, for example, the genomic region contains a region of differential chromatin composition between at least two cell or tissue types. More specifically, the fragmentation pattern of DNA molecules in a cfDNA sample holds information about the chromatin composition of the cell or tissue from which the cfDNA fragment originates. In particular, DNA fragments released into the bloodstream are often fragmented or cleaved around nucleosomes and / or other DNA-binding proteins in the originating cell or primary tissue. Furthermore, the arrangement of nucleosomes and the position of DNA-binding proteins are highly tissue-specific and are therefore used herein to amplify signals originating from cells or tissues (e.g., tumor cells, as well as cells in the tumor microenvironment and cells involved in immune responses) from which cfDNA fragments originate. In certain embodiments, genomic regions include transcription factor-binding regions, distal regulatory elements (DREs), repeat elements, intron-exon or exon-intron junctions (splice junctions), transcription start sites (TSSs), and / or similar.
[0248] In some embodiments, the methods disclosed herein, as well as related systems and computer-readable media implementations, include determining the cellular origin of DNA molecules from a cfDNA sample using the characteristics of these DNA molecules, e.g., epigenetic patterns indicated by these molecules or fragments. As described herein, epigenetic changes in genomic regions often occur in conjunction with changes in chromatin composition and nucleosome positioning within these genomic regions. Therefore, the methods and related embodiments of this disclosure use these signal sources in combination to increase the ability to detect the presence of target cells in a cfDNA sample (e.g., diseased cells, e.g., tumor cells or similar, fetal cells, transplant donor cells, and similar).
[0249] The methods and related embodiments of this disclosure can be carried out using any epigenetic site or locus that exhibits differential modification (e.g., post-replication modification or similar) between at least two cell or tissue types. Examples of such sites include methylation sites, acetylation sites, ubiquitination sites, phosphorylation sites, SUMOylation sites, ribosylation sites, citrullination sites, histone post-translational modification sites, histone variant sites, and / or similar sites. Examples of post-replication modifications include, among many others, 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxylcytosine, and 5-formylcytosine. Further details regarding epigenetic sites or loci can be found, for example, in Jin et al., "DNA Methylation: Superior or Subordinate in the Epigenetic Hierarchy?," Genes Cancer, 2(6):607-617 (2011)、Javaid et al., "Acetylation- and Methylation-Related Epigenetic Proteins in the Context of Their Target," Genes (Basel), 8(8):196 (2017)、Cao et al., "Histone Ubiquitination and Deubiquitination in Transcription, DNA Damage Response, and Cancer," Front Oncol, 2:26 (2012)、Rossetto et al., "Histone phosphorylation: A chromatin modification involved in diverse nuclear event," Epigenetics, 7(10):1098-1108 (2012), Vranych et al., "SUMOylation and deimination of proteins: two epigenetic modifications involved in Giardia encystation," Biochim Biophys Acta, 1843(9):1805-17 (2014), Sadakierska-Chudy et al., "A Comprehensive View of the Epigenetic Landscape. Part II: Histone Post-translational Modification, Nucleosome Level, and Chromatin Regulation by ncRNAs," Neurotox Res, 27:172-197 (2015), Fuhrmann et al., "Protein Arginine Methylation and Citrullination in Epigenetic Regulation," ACS Chem Biol, 11(3):654-668 (2016), Fan et al., “Metabolic regulation of Histone post-translational modifications, ACS Chem Biol, 10(1):95-108 (2015), and Henikoff et al., "Histone Variants and Epigenetics," Cold Spring Harb Perspect Biol, 7(1) (2015), are described herein and are incorporated by reference.
[0250] Epigenetic information can be obtained from cfDNA fragments using any technique known to those skilled in the art. In some embodiments, for example, DNA molecules from a given cfDNA sample are physically fractionated to generate partitions (e.g., fractionation with methyl-binding domain protein ("MBD") beads to stratify cfDNA fragments to various degrees of methylation or similar). In these embodiments, differential molecular tags and NGS-compatible adapters are applied to each of two or more partitions to generate molecularly tagged partitions. In addition, these embodiments also include assaying the molecularly tagged partitions with an NGS instrument to generate sequence data for deconvoluting the sample into differentially partitioned molecules to generate epigenetic information. In some embodiments, bisulfite sequencing techniques are also used to generate epigenetic information from cfDNA samples. Further details relating to the analysis of epigenetic modifications, which may be adapted as necessary for use in carrying out the methods disclosed herein, are described, for example, in WO2018 / 119452, filed on 22 December 2017, which is incorporated herein by reference.
[0251] In some embodiments, the methods and related systems and computer-readable media implementations disclosed herein include determining the cellular origin of DNA molecules from nucleic acid samples, e.g., cfDNA samples, using another form of epigenetic data, by characteristics of sequences (e.g., sequence fragments / reads) identified by a sequencing process, e.g., the fragment mix pattern indicated by those molecules or fragments. Human plasma DNA contains a mixture of DNA fragments of different sizes, and therefore the size of sequence fragments can form part of the fragment mix signature. The mode of size is approximately 166 base pairs (bp), which may be related to nucleosome structure. Cellless tumor-derived DNA in the plasma of cancer patients has a shorter mode of size of approximately 143 bp. The size profile of ctDNA may have a shorter median length and be more variable in subjects with cancer than in subjects without cancer. In addition, tumor sequence fragments and non-tumor sequence fragments can be distinguished using the pattern of cellless DNA size peaks.
[0252] Cell-free tumor-derived DNA may exhibit different ends compared to cell-free non-tumor-derived DNA, and therefore, terminal motifs can form part of the fragment mix signature. Termination sequences may reveal overpresentation of certain motifs that can feature various nucleotides, such as 2-nucleotide oligomer (2-mer) or 4-mer motifs. Many human cancers exhibit downregulation of DNASE1L3 expression, and this downregulation results in a reduction of plasma DNA with DNASE1L3-related terminal motifs. Plasma DNA terminal motifs demonstrate their advantage in that their maximum diagnostic power can be obtained by analyzing a relatively small number of DNA molecules. For example, based on computer simulations, with a 10% tumor DNA percentage, it would require only 50,000 plasma DNA molecules (the DNA content of each cell is fractionated into approximately 20,000,000 cell-free DNA molecules) to differentiate patients with hepatocellular carcinoma from those without, while at least 7,500,000 DNA molecules would be required to detect a 1 megabase (Mb) copy number anomaly. The detection of tumor-derived single-nucleotide variants in plasma DNA has been shown to require a much greater sequencing depth (e.g., >200 times the coverage of the haploid human genome).
[0253] Double-stranded cell-free DNA may have blunt or jagged ends, and therefore the presence and / or degree of jagged ends may form part of the fragment mix signature. Nucleases differ in their preference for generating blunt-ended double-stranded DNA fragments over those with protruding or jagged ends. Jagged ends can be repaired with either methylated or unmethylated cytosine, and therefore the abundance of jagged ends can be measured by changes in methylation levels from those of the genome. The frequency of jagged ends has been found to be increased in ctDNA of cancer patients. The frequency of jagged ends can be related to the relative activity between DNASE1 and DNASE1L3, with the former increasing the frequency of jagged ends and the latter decreasing it.
[0254] Plasma DNA fragmentation is a non-random process in which certain genomic regions are more likely to be cleaved and found at the ends of plasma DNA fragments, called "preferred end sites," and such sites can therefore form part of the fragment mix signature. These sites can differ in DNA molecules from different tissue sources. When cell-free DNA is aligned to the human genome, their ends tend to cluster at gene locations (preferred end sites), and this tendency can differ among DNA molecules depending on their tissue of origin. A window protection score, which can be calculated as the number of complete fragments minus the number of fragment endpoints within a given window size, can convey information about DNA protection from digestion, and this can be used to infer nucleosome arrangement. Genomic coverage and directional information at cell-free DNA termination sites—i.e., upstream or downstream ends—reflect the chromatin structure of the primary tissue (e.g., TFs, transcription factors).
[0255] The major local locations of nucleosomes across the human genome within tissues that contribute to cfDNA can be inferred by comparing the distribution of aligned fragment endpoints with one or more reference maps or by mathematical transformation thereof. An example of a value that may be used in fragment mix analysis is the Windowed Protection Score ("WPS"), as described in PCT application WO2016 / 015058, which was developed to represent such arrangements and therefore can form part of the fragment mix signature. Specifically, cfDNA fragment endpoints are expected to cluster adjacent to nucleosome boundaries and simultaneously be depleted within the nucleosomes themselves. WPS values correlate with the location of nucleosomes in tightly arranged arrays, as mapped by other groups using in vitro methods or ancient DNA. In other areas, WPS correlates with genomic features, e.g., DNase I high-sensitivity (DHS) sites (e.g., corresponding to nucleosome rearrangements adjacent to distal regulatory elements). Fragment mix analysis typically involves determining a value (single or multiple) based on the number of fragment endpoints located at a specific gene location (one or more bases), normalizing it to the amount of sequence data at or near that gene location. Therefore, fragment mix values can be input into models comparing healthy and diseased individuals to determine the likelihood of disease presence or absence in a test subject. For example, if 10,000 paired end reads have ends located within a 500 bp genomic region, and 100 ends are located at single base locations within that 500 bp region, then a value of 100 / 1000 might be the fragment mix value for that single base location. While not theoretically constrained, fragment mix values appear to indicate the presence or absence of proteins bound to the matched genomic region, such as histones or transcription factors. The presence or absence of such bound proteins is thought to affect the accessibility of nucleases to DNA protected by the bound proteins.
[0256] In one embodiment, in the feature engineering step 104, input features for the machine learning step may be generated by analyzing, for example, sequence data, epigenetic data, genome modification data, combinations thereof, and similar data. Additional or other data types may be used in the feature engineering step as needed. Method 100 may also include one or more transformations and / or cleanups in the data normalization step 106, e.g., sample retention cleanup (e.g., adjusting for samples with fewer nucleic acid variants, samples with fewer samples, etc.), logarithmic transformations (e.g., Log(x+1) or Np.log1p), and normalizations (e.g., Yeo-Johnson normalization, minimax normalization, z-score normalization, and / or similar normalizations) (step 108).
[0257] Method 100 may include a machine learning step 108 that generates a machine learning model (e.g., a classifier) according to a training dataset generated from the data obtained in step 102 (e.g., by creating a training dataset) and input features from step 104. The machine learning model may be configured to provide, classify, predict, or otherwise determine one or more possibilities that the origin of a given nucleic acid variant present in a test sample is tumorous or non-tumorous. Machine learning step 108 may use any machine learning technique, e.g., logistic regression or deep learning technique. Exemplary models that may be used for training and classification may include, but are not limited to, one or more ensembles of logistic regression, probit regression, decision trees, random forests, gradient boosting, support vector machines, k nearest neighbors, neural networks, or one or more of these methods. An ensemble method is a meta-algorithm that combines several machine learning techniques into a single predictive model to reduce variance (bagging), reduce bias (boosting), or improve prediction (stacking). Most ensemble methods use a single base learning algorithm to produce a uniform base learner, i.e., learners of the same type, resulting in a uniform ensemble. Some methods, however, use heterogeneous learners, i.e., learners of different types, to produce a heterogeneous ensemble. For an ensemble to be more accurate than any of its individual members, the base learners must be as accurate and as diverse as possible.
[0258] Method 100 may output a machine learning model / classifier configured in step 110 to classify or otherwise predict the origin of a sample when epigenetic data and / or genomic modification data related to the sample are provided.
[0259] Machine learning models / classifiers can be used to determine the origin of newly presented sequence fragments in a test sample. The origin may be tumor-derived or non-tumor-derived. Sequence fragments classified as tumor-derived by the machine learning model / classifier can be used to direct the treatment of the subject. It may be unknown beforehand whether the subject has a disease, or it may be known that the subject has a disease. The disease may be cancer. The method may include a step of administering one or more therapies to the subject to treat the disease. Therapies may include administering chemotherapy, administering radiotherapy, or performing surgery to remove all or part of the tumor. The method may include a step of assisting in communicating the determination of origin as tumor-derived to the subject associated with the test sample.
[0260] C. Example Systems and Methods The systems and methods described herein concern blood-based assays for cfDNA detection of CRCs. These methods match epigenetic factors (abnormal methylation status and fragment mix patterns) with cfDNA genomic alterations. The results are integrated into a binary "detected abnormal signal" ("positive") or "detected normal signal" ("negative"). The following is a description of each of the cancer screening assays and their components.
[0261] Figure 2 shows an example of a system 200 for determining whether a sample of a test subject 211 is tumor-derived, according to embodiments of the present disclosure. The system 200 can process one or more samples 201 from the test subject 211 to generate sequence reads. The system 200 may include a laboratory system 202, a computer system 210, and / or other components. Note that the laboratory system 202 and the computer system 210 may be geographically separated from each other and may be connected to each other by a computer network (not shown). The laboratory system 202 may include a sample collection and preparation pipeline 203, a sequencing pipeline 205, a sequence read data store 209, and / or other components. The sequencing pipeline 205 may include one or more sequencing devices 207 (illustrated in Figure 2 as sequencing devices 207a...n).
[0262] The methods of this disclosure can be used in a wide variety of applications in the manipulation, preparation, identification, quantification, and / or analysis of cell-free nucleic acids. As shown in Figure 2, a sample collection and preparation pipeline 203 may include obtaining a cfDNA reference sample 201 from one or more reference subjects and a cfDNA test sample 211 from a test subject. As described herein, polynucleotides may include any type of nucleic acid, such as DNA and / or RNA. For example, if the polynucleotide is DNA, it may be genomic DNA, complementary DNA (cDNA), or any other deoxyribonucleic acid. Polynucleotides may also be cell-free nucleic acids, such as cell-free DNA (cfDNA). For example, a polynucleotide may be circulating cfDNA. Circulating cfDNA may include DNA that is expelled from cells of the body by apoptosis or necrosis. cfDNA expelled by apoptosis or necrosis may originate from normal (e.g., healthy) cells of the body. Tumor DNA may be expelled if there is abnormal tissue growth, such as abnormal tissue growth for cancer. Circulating cfDNA may include circulating tumor DNA (ctDNA). 1. Sample
[0263] Cell-free polynucleotides can be isolated and extracted by collecting samples using various techniques. The sample may be any biological sample isolated from the subject. Samples may include body tissues, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsy material (e.g., biopsy material from known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial or extracellular fluid (e.g., fluid from the interstitial space), gingival exudate, gingival crevicular exudate, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Preferably, the sample is a body fluid, particularly blood and its fractions, as well as urine. Such samples may contain nucleic acids excreted from tumors. Nucleic acids may include DNA and RNA and may be in double-stranded or single-stranded form. A sample may be the form in which it was initially isolated from the subject, or it may have undergone further processing to remove or add components such as cells, to concentrate one component compared to another, or to convert one form of nucleic acid to another, for example, RNA to DNA, or single-stranded nucleic acid to double-stranded nucleic acid. Therefore, for example, a bodily fluid sample for analysis may be plasma or serum containing cell-free nucleic acid, such as cell-free DNA (cfDNA).
[0264] In some embodiments, the sample volume of bodily fluids taken from the subject depends on the desired reading depth for the region being sequenced. Exemplary volumes are approximately 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, volumes are approximately 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, 40 ml, or more milliliters. The volume of plasma sampled is typically between approximately 5 ml and 20 ml.
[0265] A sample may contain varying amounts of nucleic acids. Typically, the amount of nucleic acid in a given sample is considered equivalent to several genome equivalents. For example, a sample with approximately 30 ng of DNA is equivalent to approximately 10,000 (10) 4 ) In the case of haploid human genome equivalents and cfDNA, approximately 200,000,000,000 (2 × 10⁻¹⁶).11 ) may contain individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, about 600,000,000,000 individual molecules.
[0266] In some embodiments, the sample includes nucleic acids from different sources, e.g., from cells and from cell-free sources (e.g., blood samples). Typically, the sample includes nucleic acids that carry mutations. For example, the sample may include DNA that carries germline mutations and / or somatic mutations, as is appropriate. Typically, the sample includes DNA that carries cancer-related mutations (e.g., cancer-related somatic mutations). In some embodiments of this disclosure, cell-free nucleic acids in the subject may be derived from a tumor. For example, cell-free DNA isolated from the sample may include ctDNA.
[0267] Exemplary amounts of cell-free nucleic acids in a sample before amplification are typically in the range of about 1 femtogram (fg) to about 1 microgram (μg), for example, about 1 picogram (pg) to about 200 nanograms (ng), about 1 ng to about 100 ng, or about 10 ng to about 1000 ng. In some embodiments, the sample contains cell-free nucleic acid molecules in amounts of about 600 ng or less, about 500 ng or less, about 400 ng or less, about 300 ng or less, about 200 ng or less, about 100 ng or less, about 50 ng or less, or about 20 ng or less. If necessary, the amounts are at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the amounts are approximately 1 fg, 10 fg, 100 fg, 1 pg, 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or less than 200 ng of cell-free nucleic acid molecules. In some embodiments, the method includes the step of obtaining approximately 1 fg to approximately 200 ng of cell-free nucleic acid molecules from a sample.
[0268] Cell-free nucleic acids typically have a size distribution between approximately 100 and 500 nucleotides in length, with molecules between approximately 110 and 230 nucleotides accounting for about 90% of the molecules in the sample. The mode is approximately 168 nucleotides, and a second small peak is in the range of approximately 240 to 440 nucleotides. In certain embodiments, cell-free nucleic acids are approximately 160 to 180 nucleotides, or approximately 320 to 360 nucleotides, or approximately 440 to 480 nucleotides.
[0269] In some embodiments, cell-free nucleic acids are isolated from body fluids by a fractionation step, in which cell-free nucleic acids, if present in solution, are separated from intact cells and other insoluble components in the body fluids. In some of these embodiments, fractionation includes techniques such as centrifugation or filtration. Alternatively, cells in the body fluids are lysed, and cell-free nucleic acids and cellular nucleic acids are processed together. Generally, after the addition of buffers and washing steps, cell-free nucleic acids are precipitated, for example, with alcohol. In certain embodiments, additional purification steps are used, such as silica-based columns to remove impurities or salts. For example, non-specific bulk carrier nucleic acids are added as needed through the reaction to optimize certain aspects of the exemplary procedure, such as yield. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. If necessary, single-stranded DNA and / or single-stranded RNA are converted to double-stranded form for inclusion in subsequent processing and analysis steps. Further details regarding cfDNA fractionation and the relevant analysis of epigenetic modifications as necessary for use in carrying out the methods disclosed herein are described, for example, in WO2018 / 119452, filed on 22 December 2017, which is incorporated herein by reference.
[0270] 2. Array data Sequence information can be obtained from cfDNA. This sequence information can be used for further analysis of epigenetic factors and genomic modifications. Several components may be involved in obtaining the sequence data described herein.
[0271] An example overview of the disclosed workflow is as follows. In some embodiments, some of the steps, particularly some of the tagging of the nucleic acid sample, can be performed in a different order. After obtaining the cfDNA sample, the sample can be partitioned based on methylation status. An adapter containing a molecular barcode can be ligated to the sample. Methylation-dependent restriction enzyme (MSRE) treatment can be performed on the hypermethylated partition to remove inaccurately partitioned molecules. The hypomethylated partition can optionally be treated with MDRE to perform the step of removing methylated molecules from the hypo partition. After MSRE digestion, the partitions can be pooled and PCR amplification can be performed. The target region can be enriched using a probe (e.g., an RNA probe or a DNA probe). After enrichment, further PCR amplification can be performed. The nucleic acid can be tagged with a sample index via a primer during PCR (which can be either the first PCR before enrichment or the PCR after enrichment). Next, the samples can be pooled and sequenced using an NGS instrument. Next, the generated sequencing reads can be aligned to the human genome. The molecular barcode (and optionally, along with the alignment position) can be used to group the sequencing reads into families of individual cfDNA molecules, and then this family can be used to estimate the count of molecules at one or more loci (and in genomic regions). Next, the raw molecule count can be normalized using a positive control region, and then one of a model, LR or TFR model can be applied. A final score for the presence or absence of cancer in the subject from whom the cfDNA was obtained can be generated (e.g., based on ctDNA) using an LR and / or TFR model with or without biomarker analysis. Each of these steps is described in further detail throughout. i. Partitioning; Analysis of Epigenetic Features
[0272] In certain embodiments described herein, a population of different forms of nucleic acids (e.g., highly methylated and hypomethylated DNA in a sample from a subject, e.g., tagged DNA or aliquots thereof) can be physically partitioned based on one or more characteristics of the nucleic acids before analysis, e.g., sequencing, or tagging and sequencing. This technique can be used to determine, for example, whether highly methylated variable epigenetic target regions indicate the highly methylated characteristics of tumor cells, or whether hypomethylated variable epigenetic target regions indicate the hypomethylated characteristics of tumor cells, or whether they indicate the presence of disease in a different way. In addition, partitioning heterogeneous nucleic acid populations can increase the presence of rare signals, for example, by concentrating rare nucleic acid molecules that are more abundant in one fraction (or partition) of the population. For example, genetic diversity present in highly methylated DNA but less (or none) in hypomethylated DNA can be more easily detected by partitioning the sample into highly methylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of single gene loci or nucleic acid species in the genome can be performed, thus achieving higher sensitivity.
[0273] In some embodiments, the partitions are differentially tagged and then rearranged, after which the sample is divided into a first aliquot and a second aliquot, followed by subsequent steps of the method described herein. In some embodiments, the sample divided into a first aliquot and a second aliquot is a partition such as a low-methylated partition, and the second aliquot is combined with at least one other partition, such as a high-methylated partition, before being subjected to the concentration and / or other steps of the method.
[0274] In some cases, a heterogeneous nucleic acid sample is partitioned into two or more partitions (e.g., at least 3, 4, 5, 6, or 7 partitions). In some embodiments, each partition is differentially tagged. The tagged partitions can then be pooled together for population sample preparation and / or sequencing. The partitioning-tagging-pool steps can be performed more than once, and each partitioning round is tagged and performed based on different features (examples provided herein) and using differential tags that distinguish it from other partitions and partitioning means.
[0275] Examples of features that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitions can contain one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids having one or more epigenetic modifications and nucleic acids having no epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; methylation level; type of methylation (e.g., 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and association and level of association with one or more proteins such as histones. Alternatively, or in addition, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules with nucleosomes and nucleic acid molecules lacking nucleosomes. Alternatively, or in addition, a heterogeneous population of nucleic acids can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or in addition, a heterogeneous population of nucleic acids can be partitioned based on nucleic acid length (e.g., molecules of 160 bp or less and molecules having a length greater than 160 bp).
[0276] In some cases, each partition (representing a different nucleic acid morphology) is differentially labeled, and these partitions are pooled together before sequencing. In other cases, different morphologies are sequenced separately.
[0277] The samples may contain nucleic acids that differ in terms of modification, including post-replication modifications to nucleotides and binding to one or more proteins, usually non-covalently.
[0278] In one embodiment, the nucleic acid population is obtained from serum, plasma, or blood samples from subjects suspected of having a neoplasm, tumor, or cancer, or from subjects previously diagnosed with a neoplasm, tumor, or cancer. The nucleic acid population includes nucleic acids having varying levels of methylation. Methylation may result from any one or more post-replication or post-transcriptional modifications. Post-replication modifications include modifications of nucleotide cytosines, particularly modifications at the 5-position of the nucleic acid base, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0279] In some embodiments, nucleic acids in the original population may be single-stranded and / or double-stranded. Partitioning based on single-strandedness versus double-strandedness of nucleic acids can be achieved, for example, by using labeled capture probes to partition ssDNA and double-stranded adapters to partition dsDNA.
[0280] Affinity agents may be antibodies with desired specificity, their natural binding partners or variants (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or they may be artificial peptides selected, for example, by phage display to have specificity for a given target.
[0281] Examples of the capture portions intended herein include the methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein.
[0282] Similarly, partitioning of different nucleic acid forms can be performed using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that may be used in the methods disclosed herein include RBBP4 (RbAp48) and SANT domain peptides.
[0283] For some affinity agents and modifiers, binding to the drug may occur in an all-encompassing manner, depending on whether the nucleic acid has the modification, but separation may be a matter of degree. In such cases, nucleic acids that are over-presented in the modifier bind to the drug to a greater extent than nucleic acids that are under-presented in the modifier. Alternatively, nucleic acids with modifications may bind all or nothing. Nevertheless, various levels of modifications can be successively eluted from the binder.
[0284] For example, in some embodiments, partitioning may be binary or based on degree / level of modification. For instance, a methyl-binding domain protein (e.g., MethylMiner Methylated DNA Enrichment Kit (Thermo Fisher Scientific)) can be used to partition all methylated fragments from unmethylated fragments. Further partitioning may then involve eluting fragments with different methylation levels by adjusting the salt concentration in the solution containing the fragments bound to the methyl-binding domain. As the salt concentration increases, fragments with higher methylation levels are eluted.
[0285] In some cases, the final partition represents nucleic acids with varying degrees of modification (over-representation or under-representation of modification). Over-presentation and under-presentation can be defined by the number of modifications a nucleic acid has relative to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is 2, nucleic acids containing more than 2 5-methylcytosine residues are over-presented in this modification, and nucleic acids with 1 or zero 5-methylcytosine residues are under-presented. The effect of affinity separation is to enrich nucleic acids that are over-presented in the modification in the bound phase and nucleic acids that are under-presented in the modification in the unbound phase (i.e., in soluble state). Nucleic acids in the bound phase can be eluted before subsequent processing.
[0286] When using the MethylMiner Methylated DNA Enrichment Kit (Thermo Fisher Scientific), sequential elution can be used to partition nucleic acids at various methylation levels. For example, by contacting the MBD from the kit, which is bound to magnetic beads, with the nucleic acid population, a low-methylated partition (e.g., unmethylated) can be separated from the methylated partition. The beads are used to separate and remove methylated nucleic acids from unmethylated nucleic acids. Then, one or more elution steps are performed sequentially to elute nucleic acids with different methylation levels. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After such methylated nucleic acids have been eluted, magnetic separation is used again to separate nucleic acids with higher methylation levels from those with lower methylation levels. The elution and magnetic separation steps themselves can be repeated to produce various partitions, such as low-methylated partitions (e.g., representing unmethylated areas), methylated partitions (representing low methylation levels), and high-methylated partitions (representing high methylation levels).
[0287] In some methods, nucleic acids bound to the affinity agent used for affinity separation are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity agent. In such nucleic acids, the modified nucleic acids may be concentrated to a degree close to the mean or median (i.e., an intermediate value between nucleic acids that remained bound to the solid phase and nucleic acids that were not bound to the solid phase at the time of the initial contact between the sample and the agent).
[0288] Affinity separation yields at least two partitions of nucleic acids with different degrees of modification, and sometimes three or more. The partitions are further separated, but at least one, usually two or three (or more) partitions of nucleic acids are ligated to nucleic acid tags, typically provided as components of the adapter, so that nucleic acids in different partitions are given different tags, thereby distinguishing members of one partition from members of another. Tags ligated to nucleic acid molecules in the same partition may be the same or different from each other. However, even if they are different, those tags have a common part in their coding, and therefore can identify the molecule to which they are ligated when they belong to a particular partition.
[0289] For further details regarding the portioning of nucleic acid samples based on characteristics such as methylation, please refer to WO2018 / 119452, which is incorporated herein by reference.
[0290] In some embodiments, nucleic acid molecules can be fractionated into different partitions based on whether they are bound to a specific protein or fragment thereof, or whether they are not bound to that specific protein or fragment thereof.
[0291] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that can bind to DNA and serve as a basis for fractionation include, but are not limited to, protein A and protein G. Nucleic acid molecules can be fractionated based on protein-binding regions using any preferred method. Examples of methods used to fractionate nucleic acid molecules based on protein-binding regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric field flow fractionation (AF4).
[0292] In some embodiments, nucleic acid partitioning is performed by contacting the nucleic acid with the methylation-binding domain ("MBD") of a methylation-binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is linked to paramagnetic beads such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different degrees of methylation can be performed by eluting the fractions by increasing the NaCl concentration.
[0293] Examples of MBPs intended herein, but not limited to these, include: (a) MeCP2 is a protein that preferentially binds to 5-methylcytosine rather than unmodified cytosine. (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethylcytosine rather than unmodified cytosine. (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferentially bind to 5-formyl-cytosine rather than unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)). (d) Antibodies specific for one or more methylated nucleotide bases.
[0294] Generally, elution is correlated with the number of methylation sites per molecule, and molecules with a higher degree of methylation elute at higher salt concentrations. To elute DNA into separate populations based on the degree of methylation, a series of elution buffers with increasing NaCl concentrations can be used. The salt concentration can range from about 100 mM to about 2500 mM NaCl. In one embodiment, three partitions are obtained as a result of this process. The molecules are contacted with a solution having a first salt concentration, which contains molecules comprising a methyl binding domain, and which molecules can be bound to a capture moiety such as streptavidin. At the first salt concentration, a population of molecules will bind to the MBD and a population will remain unbound. The unbound population can be separated as the "low methylation" population. For example, the first partition representative of the low methylation form of DNA is the partition that remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. The second partition representative of moderately methylated DNA is eluted using a moderate salt concentration, e.g., a concentration between 100 mM and 2000 mM. This is also separated from the sample. The third partition representative of the highly methylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.
[0295] In some embodiments, for example, when an epigenetic target region set is to be captured, sample DNA (e.g., about 1–300 ng) is mixed with an appropriate amount of methyl-binding domain (MBD) buffer (the amount of MBD buffer depends on the amount of DNA used) and magnetic beads conjugated with MBD protein, and incubated overnight. Methylated DNA (highly methylated DNA) binds to the MBD protein on the magnetic beads during this incubation. Unmethylated (lowly methylated DNA) or not-so-methylated DNA (medium-methylated DNA) is washed away from the beads with a buffer containing gradually increasing concentrations of salt. For example, one, two or more fractions containing unmethylated, low-methylated and / or medium-methylated DNA can be obtained from such washing. Finally, a high-salt buffer is used to elute heavily methylated DNA (highly methylated DNA) from the MBD protein. In some embodiments, these washes result in three partitions of DNA with progressively increasing methylation levels: a low-methylation partition, a medium-methylation fraction, and a high-methylation partition.
[0296] In some embodiments, these three partitions of DNA are desalted and concentrated during preparation for the enzymatic steps of library preparation.
[0297] In some embodiments, the methylation signature of a molecule can be determined by methods such as MeDIP-seq, MBD-seq, BS-seq, Ox-BS-seq, TAP-seq, ACE-seq, hmC-seal, and TAB-seq. For example, Schutsky, EK et al. Nondestructive, base-resolution sequencing of 5-hydroxymethylcytosine using a DNA deaminase. Nature Biotech, 2018; doi.10.1038 / nbt.4204 (ACE-Seq);Yu, Miao et al. Base-resolution analysis of 5-hydroxymethylcytosine in the Mammalian Genome. Cell, 2012; 149(6):1368-80 (TAB-Seq);Han, D. A highly sensitive and robust method for genome-wide 5hmC profiling of rare cell populations. Mol Cell. 2016; 63(4):711-719 (5hmC-Seal);Shen, SY et al. Sensitive tumor detection and classification using plasma cell-free DNA methylomes. Nature. 2018; See 563(7732):579-583 (cfMeDIP); Nair, SS et al. Comparison of methyl-DNA immunoprecipitation (MeDIP) and methyl-CpG binding domain (MBD) protein capture for genome-wide DNA. Epigenetics. 2011; 6(1):34-44. In some embodiments, the methylation signature of a molecule can be determined by treating the sample with one or more methylation-sensitive restriction enzymes (MSREs) and / or methylation-dependent restriction enzymes (MDREs).In some embodiments, any of the above methods can be used alone or in combination to determine the methylation signature of a molecule. ii. Nucleic acid tags
[0298] In some embodiments, nucleic acid molecules (from polynucleotides obtained from a sample) can be tagged with a sample index and / or molecular barcode (commonly referred to as a “tag”). Among several methods, the tag can be incorporated into or otherwise conjugated to an adapter by chemical synthesis, ligation (e.g., blunt-end ligation or adherent-end ligation), or overlap-extension polymerase chain reaction (PCR). Such an adapter can then be ultimately conjugated to a target nucleic acid molecule. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are commonly applied to introduce the sample index to the nucleic acid molecule using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., multiple microwells in an array). The molecular barcode and / or sample index can be introduced simultaneously or in any order. In some embodiments, the molecular barcode and / or sample index are introduced before and / or after the sequence capture step. In some embodiments, only the molecular barcode is introduced before probe capture, and the sample index is introduced after the sequence capture step. In some embodiments, both the molecular barcode and sample index are introduced before the probe-based capture step. In some embodiments, the sample index is introduced after the sequence capture step has been performed. In some embodiments, the molecular barcode is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample via an adapter by ligation (e.g., blunt-end ligation or adherent-end ligation). In some embodiments, the sample index is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample by overlap extension polymerase chain reaction (PCR). Typically, the sequence capture protocol involves introducing a single-stranded nucleic acid molecule complementary to a target nucleic acid sequence, such as the coding sequence of a genomic region, where mutations in such a region are associated with cancer type.
[0299] In some embodiments, the tag may be located at one end of the sample nucleic acid molecule, or at both ends. In some embodiments, the tag is an oligonucleotide of a predetermined sequence, a random sequence, or a semi-random sequence. In some embodiments, the tag may be approximately 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or less than 1 nucleotide in length. The tag may be randomly or intentionally ligated to the sample nucleic acid.
[0300] In some embodiments, each sample is uniquely tagged with a sample index or a combination of sample indices. In some embodiments, each nucleic acid molecule in a sample or secondary sample is uniquely tagged with a molecular barcode or a combination of molecular barcodes. In other embodiments, multiple barcodes (e.g., non-unique molecular barcodes) may be used, such that the molecular barcodes among them are not necessarily unique to one another. In these embodiments, molecular barcodes are generally attached to individual molecules (e.g., by ligation), resulting in a unique sequence that can be individually tracked by the combination of molecular barcodes and sequences that can be attached to them. By detecting a non-unique molecular barcode in combination with endogenous sequence information (e.g., the first (start) and / or last (stop) gene / genomic location corresponding to the sequence of the original nucleic acid molecule in the sample, the start and stop genomic locations corresponding to the sequence of the original nucleic acid molecule in the sample, the first (start) and / or last (stop) gene / genomic locations of the sequence read mapped to the reference sequence, the start and stop gene locations of the sequence read mapped to the reference sequence, a subsequence at one or both ends of the sequence read, the length of the sequence read, and / or the length of the original nucleic acid molecule in the sample), it is typically possible to assign a unique identity to a particular molecule. In some embodiments, the first region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions of the 5' end of the sequencing read aligned to the reference sequence. In some embodiments, the last region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least last 30 base positions of the 3' end of the sequencing read aligning to the reference sequence. The length or number of base pairs of individual sequence reads are also used, if necessary, to assign a unique identity to a given molecule. As described herein, a fragment from a single strand of nucleic acid to which a unique identity has been assigned may then enable the subsequent identification of fragments from the parent and / or complementary strands.
[0301] In certain embodiments, the number of different tags used to uniquely identify a number, z, of molecules in a certain class may be between 2×z, 3×z, 4×z, 5×z, 6×z, 7×z, 8×z, 9×z, 10×z, 11×z, 12×z, 13×z, 14×z, 15×z, 16×z, 17×z, 18×z, 19×z, 20×z, or 100×z (e.g., lower limit) and 100,000×z, 10,000×z, 1000×z, or 100×z (e.g., upper limit). In some embodiments, molecular barcodes are introduced in an expected ratio of a set of identifiers (e.g., combinations of unique or non-unique molecular barcodes) to the molecules in the sample. One exemplary form uses approximately 2 to approximately 1,000,000 different molecular barcode sequences, or approximately 5 to approximately 150 different molecular barcode sequences, or approximately 20 to approximately 50 different molecular barcode sequences, ligated to both ends of the target molecule. Alternatively, approximately 25 to approximately 1,000,000 different molecular barcode sequences may be used. For example, 20 to 50 × 20 to 50 molecular barcode sequences (i.e., one of the 20 to 50 different molecular barcode sequences may be ligated to each end of the target molecule) may be used. Typically, such a number of identifiers is sufficient to ensure that different molecules having the same start and end points have a high probability of receiving different combinations of identifiers (e.g., at least 94%, 99.5%, 99.99%, or 99.999%). In some embodiments, approximately 80%, approximately 90%, approximately 95%, or approximately 99% of molecules have the same combination of molecular barcodes.
[0302] In some embodiments, the assignment of unique or non-unique molecular barcodes in a reaction is carried out using methods and systems described, for example, in U.S. Patent Applications Publications 20010053519, 20030152490, and 20110160078, and in U.S. Patents 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is thus incorporated herein by reference in its entirety. Alternatively, in some embodiments, different nucleic acid molecules in a sample can be identified using only endogenous sequence information (e.g., start and / or stop positions, partial sequences of one or both ends of the sequence, and / or length).
[0303] In certain embodiments described herein, populations of different forms of nucleic acids (e.g., highly methylated and hypomethylated DNA in a sample) can be physically partitioned before analysis, e.g., sequencing, or tagging and sequencing. This technique can be used to determine, for example, whether highly methylated variable epigenetic target regions indicate the highly methylated characteristics of tumor cells, or whether hypomethylated variable epigenetic target regions indicate the hypomethylated characteristics of tumor cells. In addition, partitioning heterogeneous nucleic acid populations can increase the presence of rare signals, for example, by enriching rare nucleic acid molecules that are more abundant in one fraction (or partition) of the population. For example, genetic diversity present in highly methylated DNA but less (or none at all) in hypomethylated DNA can be more easily detected by partitioning the sample into highly methylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of a single locus of genome or a species of nucleic acid can be performed, and therefore, higher sensitivity can be achieved.
[0304] In some cases, heterogeneous nucleic acid samples are partitioned into two or more partitions (e.g., at least three, four, five, six, or seven partitions). In some embodiments, each partition is differentially tagged—that is, each partition may have a different set of molecular barcodes. The tagged partitions can then be pooled together for population sample preparation and / or sequencing. The partitioning-tagging-pooling step can be performed more than once, with each partitioning round being tagged based on different characteristics (examples provided herein) and using differential tags that distinguish them from other partitions and partitioning means.
[0305] In some cases, each partition (representing a different nucleic acid form) is differentially tagged with a molecular barcode, and these partitions are pooled together before sequencing. In other cases, the different forms are sequenced separately. In some embodiments, a single tag may be used to label a particular partition. In some embodiments, multiple different tags may be used to label a particular partition. In embodiments utilizing multiple different tags to label a particular partition, the set of tags used to label one partition can be easily differentiated from the set of tags used to label other partitions. In some embodiments, the tag can be multifunctional—that is, it can simultaneously function as a molecular identifier (i.e., molecular barcode), a partition identifier (i.e., partition tag), and a sample identifier (i.e., sample index). For example, if there are four DNA samples and each DNA sample is partitioned into three partitions, then the DNA molecules in each of the 12 partitions (i.e., 12 partitions in total for the four DNA samples) can be tagged with separate sets of tags such that the identity of the DNA molecule, the partition to which it belongs, and the sample from which it originates are revealed by the tag sequence attached to the DNA molecule. In some embodiments, the tags can be used as both molecular barcodes and partition tags. For example, if a DNA sample is partitioned into three partitions, the DNA molecules in each partition are tagged with a separate set of tags such that the identity of the DNA molecule and the partition to which it belongs are revealed by the tag sequence attached to the DNA molecule. In some embodiments, the tags can be used as both molecular barcodes and sample indexes. For example, if there are four DNA samples, the DNA molecules in each sample are tagged with a separate set of tags that may be distinguishable from each sample such that the tag sequence attached to the DNA molecule functions as both a molecular identifier and a sample identifier.
[0306] In one embodiment, partition tagging involves tagging molecules within each partition with partition tags. After recombining the partitions and sequencing molecules, the source partition is identified by the partition tags. In another embodiment, different partitions are tagged with different sets of molecular tags, for example, consisting of a pair of barcodes. Thus, each molecular barcode is useful not only for distinguishing molecules within a partition but also for indicating the source partition. For example, molecules in a first partition can be tagged using a first set of 35 barcodes, while molecules in a second partition can be tagged using a second set of 35 barcodes.
[0307] In some embodiments, molecules can be pooled for sequencing in a single run after partitioning and tagging with partition tags. In some embodiments, sample tags are attached to molecules, for example, in a step after the step of attaching partition tags and the step of pooling. Sample tags can facilitate pooling materials generated from multiple samples for sequencing in a single sequencing run.
[0308] Alternatively, in some embodiments, partition tags can be correlated not only with partitions but also with samples. As a simple example, a first tag may indicate a first partition of a first sample, a second tag may indicate a second partition of a first sample, a third tag may indicate a first partition of a second sample, and a fourth tag may indicate a second partition of a second sample.
[0309] While tags can be attached to molecules that have already been partitioned based on one or more epigenetic features, the final tagged molecule in the library may no longer possess those epigenetic features. For example, single-stranded DNA molecules can be partitioned and tagged, but the final tagged molecule in the library is likely to be double-stranded. Similarly, DNA can be partitioned based on different methylation levels, but the tagged molecules derived from these molecules in the final library are likely to be unmethylated. Therefore, tags attached to molecules in a library typically exhibit the features of the "parent molecule," and the final tagged molecule originates from this parent molecule, but the parent molecule itself does not necessarily exhibit the features of the tagged molecule.
[0310] For example, barcodes 1, 2, 3, 4, etc., are used to tag and label molecules in the first partition; barcodes A, B, C, D, etc., are used to tag and label molecules in the second partition; and barcodes a, b, c, d, etc., are used to tag and label molecules in the third partition. Differentially tagged partitions can be pooled before sequencing. Differentially tagged partitions can be sequenced separately or together in parallel, for example, in the same flow cell of an Illumina sequencer.
[0311] In some embodiments, tags are introduced into the microwells in a ratio of expected identifiers (e.g., combinations of unique and / or non-unique barcodes). For example, identifiers can be loaded so that approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or more identifiers are loaded per genome sample. In some embodiments, identifiers are loaded such that each genome sample is loaded with approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or fewer than 1,000,000,000 identifiers. In certain embodiments, the average number of identifiers loaded per sample genome is less than or greater than approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers per genome sample. Identifiers are generally unique and / or non-unique.
[0312] One exemplary form uses approximately 2 to 1,000,000 different tags, or approximately 5 to 150 different tags, or approximately 20 to 50 different tags, ligated to both ends of a target nucleic acid molecule. In the case of 20 to 50 × 20 to 50 tags, a total of 400 to 2,500 tags are produced. Typically, such a number of tags is sufficient to ensure that different molecules with the same start and stop points have a high probability of receiving different combinations of tags (e.g., at least 94%, 99.5%, 99.99%, 99.999%).
[0313] After sequencing, read analysis for detecting genetic variants can be performed at the partition level and at the whole nucleic acid population level. Tags are used to sort reads from different partitions. Analysis may include in silico analysis to determine genetic and epigenetic diversity (one or more of the following, such as methylation and chromatin structure) using sequence information, genomic coordinate length, coverage, and / or copy number. In some embodiments, higher coverage may correlate with higher nucleosome occupancy in genomic regions, while lower coverage may correlate with lower nucleosome occupancy or nucleosome-depleted regions (NDRs). iii. Conversion procedure
[0314] As described in the methods disclosed herein, the use of quality control nucleosides in adapters can be advantageously used in conjunction with enzymatic conversion procedures that alter the base pairing specificity of modified nucleosides (e.g., DM-seq conversion including the addition of a protecting group (such as a carboxymethyl group) to unmodified cytosine and 5mC deamination, such as the use of APOBEC enzymes), or enzymatic conversion procedures that alter the base pairing specificity of unmodified nucleosides. For example, in some embodiments, if a molecule containing an adapter with two or more quality control nucleosides is exposed to a conversion procedure selected to alter the base pairing specificity of the quality control nucleosides, the base pairing specificity of a first portion (e.g., at least one) of the quality control nucleosides is altered, but the base pairing specificity of a second portion (e.g., at least one) of the quality control nucleosides in the adapter remains unaffected, which can represent a suboptimal conversion. As described in the methods disclosed herein, the use of quality control nucleosides in adapters can be advantageously used to predict / infer / indicate false negative detection and / or identification of modified nucleosides in DNA samples (i.e., inaccurate identification of a base as unmodified), and / or false positive detection and / or identification of modified nucleosides in DNA samples (i.e., inaccurate identification of a base as modified). A quality control nucleoside described herein for use in detecting the occurrence of false positive detection of modified nucleosides may be referred to as a “false positive quality control nucleoside.” A quality control nucleoside described herein for use in detecting the occurrence of false negative detection of modified nucleosides may be referred to as a “false negative quality control nucleoside.” A nucleoside having a modified status, meaning that its base pairing specificity is not altered when exposed to a particular conversion procedure, may in some cases be referred to as a “protected” nucleoside, or having a “protected modified status,” or similarly.
[0315] When detecting false negatives using a conversion procedure that transforms the base pairing specificity of modified nucleosides, the quality control nucleosides in the adapter may include modified nucleosides so that the conversion efficiency of the conversion procedure / suboptimal conversion can be measured and thus the frequency of false negatives can be predicted. A suboptimal conversion refers to the conversion of fewer than all types of nucleosides that would normally be converted by the reagents used in the conversion procedure; for example, a suboptimal conversion by deaminase, as in DM-seq, results in the conversion of some, but not all, 5mC thymine. The terms suboptimal and suboptimal have equivalent meanings. A suboptimal conversion may also be called an incomplete conversion, meaning that some nucleosides (modified or unmodified) that would have been converted by the conversion procedure in a complete reaction were not actually converted.
[0316] When detecting false positives using a conversion procedure that alters the base pairing specificity of unmodified nucleosides, the quality control nucleosides in the adapter may contain modified nucleosides so that the frequency of incorrect conversions of modified nucleosides can be measured, and thus the frequency of false positives can be predicted. Incorrect conversion refers to the conversion of nucleosides other than those typically converted by the conversion procedure. The conversion of methylated cytosine by a conversion method that typically converts only unmodified cytosine is an example of incorrect conversion.
[0317] When detecting false positives using a conversion procedure that alters the base pairing specificity of modified nucleosides, the quality control nucleosides in the adapter may include unmodified nucleosides so that the frequency of incorrect conversion of unmodified nucleosides can be measured, and thus the frequency of false positives can be predicted.
[0318] When detecting false positives using a conversion procedure that alters the base pairing specificity of unmodified nucleosides, the quality control nucleosides in the adapter may include unmodified nucleosides so that the conversion efficiency of the conversion procedure / suboptimal conversion can be measured and thus the frequency of false positives can be predicted.
[0319] Various methods exist for detecting and / or identifying modified nucleosides, relying on transformation procedures that alter the base pairing specificity of nucleosides based on their modification status. Such changes in base pairing specificity can then be detected by sequencing, and thus the modification status of the nucleoside can be inferred.
[0320] In some cases, the conversion procedure used in the methods of this disclosure alters the base pairing specificity of a modified nucleoside (e.g., methylated cytosine) but does not alter the base pairing specificity of the corresponding unmodified nucleoside (e.g., cytosine), or does not alter the base pairing specificity of any of the unmodified nucleosides (e.g., cytosine, adenosine, guanosine, and thymidine (or uracil)). Advantages of methods that do not alter the base pairing specificity of unmodified nucleosides include reduced loss of sequence complexity, higher sequencing efficiency, and reduced alignment loss. In addition, methods such as DM-seq may be preferred in some cases over methods such as bisulfite sequencing and EM-seq, as they are less destructive (particularly important for low-yield samples such as cfDNA) and do not require denaturation, which theoretically means that non-conversion errors are more likely to be random. In methods that require denaturation for conversion, failure to denaturate the DNA molecule would result in the non-conversion of all bases in the DNA molecule. Because biological changes in methylation are predominantly coordinated towards the desired localization region, such non-random (localized) conversions can appear as false negatives (unmethylated regions). Random non-conversion methods can maximize the effect on a low percentage of bases within a region, and thus the specificity of methylation change detection can be maximized (reducing false positives) by setting a threshold on the percentage of bases within a region that are methylated / unmethylated. Therefore, in some cases, conversion procedures that do not involve denaturation are preferred.
[0321] In some embodiments, the adapter includes a first quality control nucleoside having a first modification status (e.g., modified, e.g., methylated) and a second quality control nucleoside having a second modification status (e.g., unmodified). Using such an adapter, both suboptimal and erroneous conversions can be detected.
[0322] Figure 1 illustrates an embodiment of a quality control method for monitoring false-negative and / or false-positive detection of DNA subjected to a DM-seq base-transformation procedure, using optional protection of 5hmC (e.g., by glucosylation). An adapter containing unmethylated C (e.g., in a molecular barcode) is ligated to the DNA and then subjected to a DM-seq transformation procedure that alters the base pairing of methylated cytosine (the sequence read as "T") but does not alter the base pairing of unmethylated cytosine (still read as "C"). Each strand is sequenced. Molecules that have undergone suboptimal transformation are identified and can be filtered out, at least for the purpose of determining methylation. In such examples, quality control bases in the adapter at the 5' end of the dsDNA molecule strand, quality control bases in the adapter at the 3' end of the dsDNA molecule strand, or both, can be evaluated to determine whether the methylated cytosine in the molecule has been successfully deaminated. In some embodiments, suboptimal conversions include ssDNA molecules in which methylated C is not deaminated but converted to T, or ssDNA molecules in which 0 / 2 barcode 5mCs are converted to T. In such examples, the quality control bases at the 5' end adapter of the ssDNA molecule, the quality control bases at the 3' end adapter of the ssDNA molecule, or both can be evaluated to determine whether the methylated cytosine in the molecule was successfully deaminated. The sample conversion rate can be calculated by dividing all converted barcode 5mCs by the total number of barcode 5mCs.
[0323] In other cases, the transformation procedure used in the method of the present disclosure is a procedure that alters the base pairing specificity of the unmodified nucleoside (e.g., cytosine) but does not alter the base pairing specificity of the corresponding modified nucleoside (e.g., methylated cytosine).
[0324] Those skilled in the art can select appropriate methods as needed, including which nucleoside modifications to detect and / or identify.
[0325] In some embodiments, the conversion procedure converts the modified nucleoside. In some embodiments, the conversion procedure for converting the modified nucleoside includes enzymatic conversion, such as DM-seq, as described in WO2023 / 288222A1. In DM-seq, unmodified cytosine in DNA is enzymatically protected from the subsequent deamination step, and 5mC in 5mCpG is converted to T. Enzymatically protected unmodified (e.g., unmethylated) cytosine is not converted and is read as "C" during sequencing. Cytosine read as thymine (in the CpG context) is identified as methylated cytosine in DNA.
[0326] Therefore, when this type of conversion is used, the first nucleic acid base contains unmodified (non-methylated, etc.) cytosine, and the second nucleic acid base contains modified (methylated, etc.) cytosine. Sequencing of the converted DNA identifies the positions read as cytosine as unmodified C positions. On the other hand, positions read as T are identified as T or 5mC. Therefore, performing DM-seq conversion facilitates the identification of 5mC-containing positions using the resulting sequence reads. Accordingly, in these embodiments, the quality control nucleoside in the adapter used in the method contains unmodified (non-methylated) cytosine.
[0327] The exemplary cytosine deaminases for use herein include APOBEC enzymes, e.g., APOBEC3A. Generally, AID / APOBEC family DNA deaminase enzymes such as APOBEC3A (A3A) are used to deaminate (unprotected) unmodified cytosine and 5mC. For an exemplary description of APOBEC conversion, see, for example, Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0328] Enzymatic protection of unmodified cytosine in DNA involves the addition of a protecting group to unmodified cytosine. Such protecting groups may include alkyl groups, alkyne groups, carboxyl groups, carboxyalkyl groups, amino groups, hydroxymethyl groups, glucosyl groups, glucosylhydroxymethyl groups, isopropyl groups, or dyes. For example, DNA can be treated with a methyltransferase, such as a CpG-specific methyltransferase, which adds a protecting group to unmodified cytosine. The term methyltransferase is used herein to refer to an enzyme capable of transferring methyl or substituted methyl (e.g., carboxymethyl) to a substrate (e.g., cytosine in nucleic acids). In some embodiments, DNA is contacted with a CpG-specific DNA methyltransferase (MTase), such as a CpG-specific carboxymethyltransferase (CxMTase), and a substituted methyl donor, such as a carboxymethyl donor (e.g., carboxymethyl-S-adenosyl-L-methionine). See, for example, WO2021 / 236778A2. In certain embodiments, CxMTase can facilitate the addition of a protective carboxymethyl group to unmethylated cytosine. In some embodiments, unmethylated cytosine is unmodified cytosine. The carboxymethyl group can prevent deamination of cytosine in a deamination step (such as a deamination step using an APOBEC enzyme such as A3A). Useful substituted methyl or carboxymethyl donors in the disclosed methods include, but are not limited to, S-adenosyl-L-methionine (SAM) analogs, and if necessary, the SAM analog is carboxy-S-adenosyl-L-methionine (CxSAM). SAM analogs are described, for example, in WO2022 / 197593A1. MTase may be, for example, CpG methyltransferase (M.SssI), DNA-methyltransferase 1 (DNMT1), DNA-methyltransferase 3-alpha (DNMT3A), DNA-methyltransferase 3-beta (DNMT3B), or DNA adenine methyltransferase (Dam) derived from the Spiroplasma strain MQ1.CxMTase may be CpG methyltransferase (M.MpeI) derived from Mycoplasma penetrans. In certain embodiments, the methyltransferase enzyme is a variant of M.MpeI, where the amino acid corresponding to position 374 is R or K, or a sequence that is at least 90%, at least 92%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% identical to them, and where the amino acid corresponding to position 374 is R or K, if necessary.
[0329] In one embodiment, the methyltransferase enzyme is a variant of M.MpeI having an N374R substitution or an N374K substitution. The methyltransferase having an N374R substitution or an N374K substitution may further include one or both of the residues T300 and E305 by S, A, G, Q, D, or N; b) one or more of the residues A323, N306, and Y299 by a positively charged amino acid selected from K, R, or H; and / or c) one or more amino acid substitutions selected from the substitution of S323 by A, G, K, R, or H, which can enhance the activity of the enzyme.
[0330] If necessary, the conversion procedure further includes enzymatic protection of 5hmC in the DNA prior to deamination of the unprotected modified cytosine, such as by glucosylation of 5hmC (e.g., using βGT) or carbamoylation of 5hmC (e.g., using 5-hydroxymethylcytosine carbamoyltransferase). In this method, 5hmC can be protected from conversion, for example, by glucosylation using β-glucosyltransferase (βGT) to form (5-glucosylhydroxymethylcytosine)5ghmC, or by carbamoylation using 5-hydroxymethylcytosine carbamoyltransferase to form 5cmC. Examples of these are described, for example, in Yu et al., Cell 2012; 149: 1368-80 and Yang et al., Bio-protocol, 2023; 12(17): e4496. Glucosylation or carbamoylation of 5hmC can reduce or eliminate deamination of 5hmC by deaminases such as APOBEC3A. Next, treatment with MTase or CxMTase adds a protecting group to the unmodified (unmethylated) cytosine in the DNA. Then, 5mC (but not the protected, unmodified cytosine, nor 5hmC or 5cmC) is deaminated by treatment with a deaminase, e.g., APOBEC enzyme (such as APOBEC3A) (converted to T in the case of 5mC). Sequencing of the converted DNA identifies the positions read as cytosine as either 5hmC or unmodified C positions. On the other hand, positions read as T are identified as either T or 5mC. Therefore, performing DM-seq conversion using the glucosylation of 5hmC in the samples described herein facilitates the distinction, using the obtained sequence reads, between positions containing 5mC and positions containing either unmodified cytosine or 5hmC. Accordingly, in these embodiments, the quality control nucleoside in the adapter used in the method can contain both unmodified cytosine and 5hmC. This allows the efficiency of each of the two steps to be determined separately.For example, if the sequencing of the adapter shows that both 5mC and 5hmC nucleosides have the converted base pairing specificity, this indicates that the 5hmC protection step was ineffective. If the 5mC nucleosides do not have the converted base pairing specificity, this indicates that the DM-seq process was ineffective (at least). If the base pairing of the 5mC nucleosides in the adapter, but not the 5hmC nucleosides, has the converted base pairing specificity, then both steps were effective.
[0331] In addition to controlling for near-optimal conversion of modified nucleosides, a quality control nucleoside in the adapter can also be used to predict false positives (i.e., nucleosides that are incorrectly classified as modified). In this case, the quality control nucleoside in the adapter contains unmodified C for proper conversion procedures. If sequencing of the adapter indicates that the quality control nucleoside has altered the base pairing specificity, this indicates that the unmodified base (e.g., unmodified C) has been incorrectly converted. This information can then be used to predict false positive detection of modified nucleosides (e.g., modified C) in the DNA sample.
[0332] In certain embodiments, the methods of the present disclosure are useful in providing a quality control method for identifying methylated cytosines (i.e., CpG and CpH cytosines) that do not exist in any sequence context. Methylated CpH or non-CpG cytosines are rare and therefore require a high level of sensitivity for reliable detection. In addition, methylated CpG co-locates with methylated non-CpG and cannot be detected by methods that use the methylation status of non-CpG cytosines as an indicator of suboptimal molecular transformation. The methods of the present disclosure achieve this by providing quality control nucleosides known to have specific modification states and thus provide a reliable measurement of the frequency of erroneous and / or suboptimal transformations.
[0333] In some embodiments, the methods of the present disclosure include analysis of sequence diversity and / or fragmentation patterns, and do not exclude adapted DNA having suboptimal or incorrect conversions of quality control nucleosides from the analysis of sequence diversity and / or fragmentation patterns. For example, the method may include a step of detecting the presence or absence of sequence diversity and / or a step of determining a fragmentation pattern, and adapted DNA containing a quality control nucleoside exhibiting suboptimal or incorrect conversions of the quality control nucleoside is included in the step of detecting the presence or absence of sequence diversity and / or determining a fragmentation pattern. In this way, the method can reduce the likelihood of false negatives and / or false positives in the detection of modified nucleosides (e.g., 5mC) by excluding adapted DNA that is unsuitable for this purpose due to suboptimal or incorrect conversions, while retaining such adapted DNA for the analysis of sequence diversity and / or fragmentation patterns (which are not affected by suboptimal or incorrect conversions) and thus avoiding an impact on sensitivity.
[0334] Some embodiments of the disclosed quality control methods are (a) A step of ligating DNA to an oligonucleotide adapter, wherein the adapter contains a quality control nucleoside, the quality control nucleoside has the same nucleoside identity and the same or different modification status as the modified nucleoside to be detected in the DNA, and the modification status of the quality control nucleoside is known; (b) A step of subjecting adapted DNA or a subsample to a conversion procedure that alters the base pairing specificity of a quality control nucleoside, depending on the modification status of the nucleoside, wherein the conversion procedure includes deamination of unmodified cytosine, and the conversion procedure is selected to (i) alter the base pairing specificity of adapted DNA nucleosides having the same nucleoside identity and modification status as the quality control nucleoside in the adapter, and not alter the base pairing specificity of adapted DNA nucleosides having the same nucleoside identity but a different modification status as the quality control nucleoside in the adapter; and / or (ii) not alter the base pairing specificity of adapted DNA nucleosides having the same nucleoside identity and modification status as the quality control nucleoside in the adapter, and alter the base pairing specificity of adapted DNA nucleosides having the same pairing identity but a different modification status as the quality control nucleoside in the adapter; (c) A step of sequencing the DNA adapted after the conversion step (b); (d) Using the sequence data obtained in step (c), determine the base pairing specificity conversion of the quality control nucleoside in the adapter; and (e) A step in which a base pairing specificity conversion of a quality control nucleoside in an adapter is used as a quality control measure for conversion step (b), wherein a suboptimal conversion of the adapter quality control nucleoside according to the conversion procedure of step (b)(i) and / or an incorrect conversion of the adapter quality control nucleoside according to the conversion procedure of step (b)(ii) predicts false negative and / or false positive detection of the modified nucleoside in the DNA sample. Includes.
[0335] In some embodiments of the disclosed method, the quality control conversion procedure is selected to alter the base pairing specificity of the unmodified quality control nucleoside in the adapter, but not alter the base pairing specificity of DNA sample nucleosides having the same nucleoside identity but a different modification status. In some such embodiments, the suboptimal conversion of the unmodified quality control nucleoside predicts false negative detection of DNA sample nucleosides having the same nucleoside identity and modification status as the quality control nucleoside, or having a different modification status and the same change in base pairing specificity upon exposure to the conversion procedure. In some such embodiments, the suboptimal conversion of the unmodified quality control nucleoside predicts false positive detection of DNA sample nucleosides having the same nucleoside identity and a different modification status as the quality control nucleoside, or having a different modification status and the same change in base pairing specificity upon exposure to the conversion procedure.
[0336] In other embodiments of the disclosed method, the quality control conversion procedure is selected to alter the base pairing specificity of a DNA sample nucleoside that has the same nucleoside identity but no modification, without altering the base pairing specificity of the modified quality control nucleoside in the adapter. In some such embodiments, an incorrect conversion of the modified quality control nucleoside predicts a false negative detection of a DNA sample nucleoside that has the same nucleoside identity and modification status as the quality control nucleoside, or the same change in base pairing specificity upon exposure to a different modification status and conversion procedure. In some such embodiments, an incorrect conversion of the modified quality control nucleoside predicts a false positive detection of a DNA sample nucleoside that has the same nucleoside identity but no modification as the quality control nucleoside, or the same change in base pairing specificity upon exposure to a different modification status and conversion procedure.
[0337] In some embodiments, the quality control nucleoside in the adapter contains unmodified cytosine. In some embodiments, the quality control nucleoside in the adapter contains modified cytosine. In some such embodiments, the quality control nucleoside in the adapter contains 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC). In some embodiments, the quality control nucleoside in the adapter contains 5-methylcytosine (5mC). In some embodiments, the quality control nucleoside in the adapter contains 5-hydroxymethylcytosine (5hmC).
[0338] Therefore, methods in which the conversion procedure includes the deamination of unmodified nucleosides such as unmodified cytosine are also provided herein. In some embodiments, the conversion procedure includes enzymatic conversion of unmodified nucleosides such as unmodified cytosine using a nonspecific, modification-sensitive double-stranded DNA deaminase, for example, as in SEM-seq. See, for example, Vaisvila et al. (2023) Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA. bioRxiv; DOI: 10.1101 / 2023.06.29.547047, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047v1. SEM-Seq is a non-destructive single-enzyme 5-methylcytosine sequencing (SEM-seq) method that uses a non-specific, modification-sensitive double-stranded DNA deaminase (MsddA) to deaminate unmodified cytosine. Therefore, SEM-seq does not require the TET2 and T4-βGT or 5-hydroxymethylcytosine carbamoyltransferase protection and denaturation steps used in protocols based on APOEC3A, for example. In addition, MsddA does not deaminate 5-formylated cytosine (5fC) or 5-carboxylated cytosine (5caC). In SEM-seq, unmodified cytosine in DNA is deaminated to uracil and read as "T" during sequencing. Modified cytosine (e.g., 5mC) is not converted and read as "C" during sequencing. Cytosine read as thymine is identified in DNA as either unmodified (e.g., unmethylated) cytosine or thymine.Therefore, by performing SEM-seq conversion, it becomes easier to identify the positions containing 5mC using the obtained sequence reads. In some embodiments, the procedure that affects the first nucleic acid base in DNA, unlike the second nucleic acid base in DNA, includes enzymatic conversion of the first nucleic acid base using MsddA. However, in some embodiments of the disclosed method in which the conversion procedure deaminates an unmodified nucleoside (such as unmodified cytosine), the method further includes enzymatic protection of at least one type of modified nucleoside (such as modified cytosine, 5mC and / or 5hmC) in DNA prior to deamination of the unprotected unmodified nucleoside (such as unprotected unmodified cytosine). In some embodiments, the at least one type of modified nucleoside is 5mC. In some embodiments, the enzymatic protection of 5mC includes converting 5mC to carboxylcytosine. For example, the conversion of 5mC to carboxylcytosine may involve contacting 5mC with a TET enzyme, such as TET1, TET2, or TET3, or any other suitable TET enzyme disclosed herein. In some embodiments, at least one type of modified nucleoside is 5hmC. In some embodiments, the enzymatic protection of 5hmC in DNA prior to the deamination of unmodified cytosine is glucosylation of 5hmC, such as as described herein.
[0339] Methods employing alternative base conversion schemes are also provided herein. For example, unmethylated cytosine can be left unchanged (for example, by being protected using the methods disclosed herein), while methylated cytosine and hydroxymethylcytosine are converted to bases read as thymine (e.g., uracil, thymine, or dihydrouracil).
[0340] In some embodiments, converting a modified cytosine (such as methylation or hydroxymethylation) in at least one first or second chain to thymine or a base read as thymine involves oxidizing hydroxymethylcytosine, for example, hydroxymethylcytosine being oxidized to formylcytosine. In some embodiments, oxidizing hydroxymethylcytosine to formylcytosine involves contacting hydroxymethylcytosine with a ruthenate, such as potassium ruthenate (KRuO4).
[0341] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil. In any such embodiment, the amplification method may include a uracil and / or dihydrouracil-resistant amplification method, such as PCR using a uracil and / or dihydrouracil-resistant DNA polymerase.
[0342] In some embodiments, the method includes a step of converting formylcytosine and / or methylcytosine to carboxylcytosine as part of a step of converting modified cytosine in at least one first or second chain to thymine or a base read as thymine. For example, the step of converting formylcytosine and / or methylcytosine to carboxylcytosine may include contacting formylcytosine and / or methylcytosine with a TET enzyme, e.g., TET1, TET2, or TET3. In some embodiments, the method includes a step of reducing carboxylcytosine as part of a step of converting modified cytosine in at least one first or second chain to thymine or a base read as thymine, and / or the carboxylcytosine is reduced to dihydrouracil. In some embodiments, the step of reducing carboxylcytosine includes contacting carboxylcytosine with a borane or borohydride reducing agent.
[0343] In some embodiments, the borane or borohydride reducing agent includes pyridineborane, 2-picolineborane, borane, tert-butylamineborane, ammoniaborane, sodium borohydride, sodium cyanoborohydride (NaBH3CN), lithium borohydride (LiBH4), ethylenediamineborane, dimethylamineborane, sodium triacetoxyborohydride, morpholineborane, 4-methylmorpholineborane, trimethylamineborane, dicyclohexylamineborane, or salts thereof. In other embodiments, the reducing agent includes lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof.
[0344] Various TET enzymes can be used in the disclosed methods as appropriate. In some embodiments, one or more TET enzymes include TETv. TETv is described in U.S. Patent No. 10,260,088, and its sequence is Sequence ID No. 1 in that document. In some embodiments, one or more TET enzymes include TETcd. TETcd is described in U.S. Patent No. 10,260,088, and its sequence is Sequence ID No. 3 in that document. In some embodiments, one or more TET enzymes include TET1. In some embodiments, one or more TET enzymes include TET2. TET2 can be expressed and used as a fragment containing TET2 residues 1129-1480 linked by a linker to TET2 residues 1844-1936, for example, as described in U.S. Patent No. 10,961,525. In some embodiments, one or more TET enzymes include TET1 and TET2. In some embodiments, one or more TET enzymes include V1900 TET variants such as V1900A, V1900C, V1900G, V1900I, or V1900P TET variant. In some embodiments, one or more TET enzymes include V1900A, V1900C, V1900G, V1900I, or V1900P TET2 variant. Since 5-carboxylcytosine (5-caC) is not a substrate for enzymatic deamination by APOBEC enzymes such as APOBEC3A, it may be beneficial to use a TET enzyme that maximizes 5-caC formation compared to modified cytosines with lower oxidation levels, particularly 5-formylcytosine. Therefore, maximizing 5-caC formation reduces the risk of false calls where bases are identified as unmethylated due to deamination, even if they were methylated (or hydroxymethylated) in the original sample. Therefore, in some embodiments, the TET enzyme contains a mutation that increases 5-caC formation. An exemplary mutation is shown above. "Mutation that increases 5-caC formation" means that the TET enzyme having this mutation produces more 5-caC than the TET enzyme lacking the mutation but otherwise identical.5-caC production can be measured, for example, as described in Liu et al., Nat Chem Biol 13:181-187 (2017) (see online methods section, TET reaction in vitro subsection, "driving" conditions). Any variants and / or mutants described in Liu et al. (2017) may be used in the disclosed method as appropriate.
[0345] In some embodiments, one or more TET enzymes include TET2 enzymes containing T1372S mutations, such as TET2-CS-T1372S and TET2-CD-T1372S. TET2 containing the T1372S mutation is described in U.S. Patent No. 10,961,525 and can be expressed and used as a fragment containing TET2 residues 1129-1480 linked by a linker to TET2 residues 1844-1936. Position 1372 of TET2 corresponds to position 258 of Sequence ID No. 21 (wild-type TET2 catalytic domain) in U.S. Patent No. 10,961,525. Thus, the sequence of the T1372S TET2 catalytic domain can be obtained by changing the threonine to serine at position 258 of Sequence ID No. 21 in U.S. Patent No. 10,961,525. TET2 containing the T1372S mutation is also described in Liu et al., Nat Chem Biol. 2017 February; 13(2): 181-187. As demonstrated in Liu et al., TET2 containing the T1372S mutation can oxidize 5mC more efficiently to produce 5-carboxylcytosine (5caC) than other versions of TET2, such as TET2 lacking the T1372S mutation.
[0346] A method is provided herein that includes the steps of contacting DNA with a TET2 enzyme containing the T1372S mutation to oxidize 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC) present in the DNA to 5-carboxycytosine (5caC); subsequently contacting at least a portion of the DNA with a substituted borane reducing agent to convert 5-caC in the DNA to dihydrouracil (DHU) to produce treated DNA; and sequencing at least a portion of the treated DNA. iv. Nucleic acid amplification
[0347] The sample nucleic acid adjacent to the adapter is typically amplified by PCR and other amplification methods, using the binding of nucleic acid primers to primer binding sites in the adapter adjacent to the DNA molecule to be amplified, as part of the sample collection and preparation pipeline 203. In some embodiments, the amplification method may involve a cycle of extension, denaturation, and annealing resulting from thermal circulation, or it may be isothermal, for example, in the case of transcription-mediated amplification. Other exemplary amplification methods that may be used as needed include, among many other techniques, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and auto-persistent sequence-based replication.
[0348] Conventional nucleic acid amplification methods typically involve one or more rounds of amplification cycles to introduce molecular tags and / or sample index / tags into nucleic acid molecules. Amplification is usually performed using one or more reaction mixtures. Molecular tags and sample index / tags are introduced simultaneously or in any order, as needed. In some embodiments, molecular tags and sample index / tags are introduced before and / or after the sequence capture step. In some embodiments, only the molecular tag is introduced before probe capture, and the sample index / tag is introduced after the sequence capture step. In certain embodiments, both the molecular tag and sample index / tag are introduced before the probe-based capture step. In some embodiments, the sample index / tag is introduced after the sequence capture step. Typically, the sequence capture protocol involves introducing a single-stranded nucleic acid molecule complementary to a target nucleic acid sequence, e.g., the coding sequence of a genomic region, and the coding sequence of a mutation in such a region associated with an oncogenic type. Typically, the amplification reaction produces multiple nucleic acid amplicons non-uniquely or uniquely tagged with molecular tags and sample index / tags, ranging in size from approximately 200 nucleotides (nt) to approximately 700 nt, 250 nt to approximately 350 nt, or approximately 320 nt to approximately 550 nt. In some embodiments, the amplicons have a size of approximately 300 nt. In some embodiments, the amplicons have a size of approximately 500 nt. a. Nucleic acid enrichment
[0349] In some embodiments, sequences are enriched as part of a sample collection and preparation pipeline 203 before sequencing the nucleic acids. The enrichment may be performed on specific target regions or non-specifically (on “target sequences”), as required. In some embodiments, the target regions of interest can be enriched using nucleic acid capture probes (“baits”) selected for one or more bait set panels using a differential tiling and capture scheme. A differential tiling and capture scheme typically involves using bait sets of different relative concentrations to tile differentially (e.g., at different “resolutions”) across genomic regions associated with the baits, imposing a set of constraints (e.g., sequencer constraints, e.g., sequencing load, utilization of each bait, etc.) to capture the target nucleic acids at the desired level for downstream sequencing. These target genomic regions of interest may optionally include native or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, the target sequences may be captured using biotin-labeled beads having probes on one or more compartments of interest, and these compartments may subsequently be amplified to enrich the regions of interest, as required.
[0350] Sequence capture typically involves the use of oligonucleotide probes that hybridize to a target nucleic acid sequence. In certain embodiments, a probe set strategy involves tiling probes across a desired compartment. Such probes may be, for example, about 60 to about 120 nucleotides in length. This set may have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x, or more. Generally, the effectiveness of sequence capture depends, in part, on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the probe sequence. v. Nucleic acid sequencing
[0351] As shown in Figure 2, after extraction and isolation of cfDNA from the sample by the sample collection and preparation pipeline 203, the cfDNA can be sequenced by a sequencing pipeline 205 including one or more sequencing devices 207. Sample nucleic acids, pre-amplified or unamplified, and optionally adjacent to adapters, are generally subjected to sequencing. Sequencing methods or commercially available formats used as needed include, for example, Sanger sequencing, high-throughput sequencing, bisulfite sequencing, pyrosequencing, single-nucleotide synthesis, single-molecule sequencing, nanopore-based sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), large-scale parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking; and sequencing using PacBio, SOLiD, Ion Torrent, or nanopore platforms. Sequencing reactions can be carried out in various sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. The sample processing unit may also include multiple sample chambers to enable simultaneous processing of multiple runs.
[0352] In some embodiments, sequencing includes the detection and / or differentiation of unmodified and modified nucleic acid bases. For example, long-read sequencing (also referred to herein as third-generation sequencing) methods include methods that can produce longer sequencing reads, such as reads exceeding 10 kilobases, compared to short-read sequencing methods that typically produce reads up to about 600 bases in length. Compared to short reads, long reads can improve de novo assembly, transcript isoform identification, and detection and / or mapping of structural variants. Furthermore, long-read sequencing of native DNA or RNA molecules reduces amplification bias and preserves base modifications such as methylation status. Useful long-read sequencing techniques herein include any preferred long-read sequencing method, including but not limited to Pacific Biosciences (PacBio) single-molecule real-time (SMRT) sequencing, Oxford Nanopore Technologies (ONT) nanopore sequencing, and synthetic long-read sequencing approaches, such as ligated reads, proximity ligation strategies, and optical mapping. The synthetic long-read approach involves assembling short reads derived from the same DNA molecule to generate synthetic long reads and can be used in conjunction with "true" long-read sequencing techniques such as SMRT and nanopore sequencing methods.
[0353] Single-molecule real-time (SMRT) sequencing can facilitate the direct detection of unmodified cytosine, for example, along with 5-methylcytosine and 5-hydroxymethylcytosine (Weirather JL, et al., “Comprehensive comparison of Pacific Biosciences and Oxford Nanopore Technologies and their applications to transcriptome analysis,” F1000Research, 6:100, 2017). While next-generation sequencing methods detect amplified signals from a clonal population of amplified DNA fragments, SMRT sequencing captures single DNA molecules and preserves base modifications during sequencing. Because the signal-to-noise ratio from single DNA molecules is not high, the error rate of raw PacBio SMRT sequencing-generated data is approximately 13–15%. To increase accuracy, the platform uses a circular DNA template by ligating hairpin adapters to both ends of the target double-stranded DNA. Because polymerase replicates by repeatedly moving back and forth through circular molecules, the DNA template is sequenced multiple times to produce continuous long reads (CLRs). CLRs can be split into multiple reads ("subreads") by removing adapter sequences, and these subreads generate circular consensus sequence ("CCS") reads with higher accuracy. The average length of a CLR is >10kb and up to 60kb, and the length depends on the polymerase lifetime. Therefore, the length and accuracy of CCS reads depend on the fragment size. PacBio sequencing is used for genomic (e.g., de novo assembly, structural variant detection, and haplotyping) and transcriptome (e.g., gene isoform reconstruction and novel gene / isoform discovery) studies.
[0354] SMRT sequencing relies on single-nucleotide synthesis, where the sequence of a circular DNA template is determined from a sequence of fluorescence pulses resulting from the addition of a single labeled nucleotide by a polymerase fixed at the bottom of the well. Base modifications do not affect the base call sequence but do affect the polymerase dynamics. By considering the inter-pulse duration (IPD), base modifications can be inferred from an in silico model or from a comparison of the unmodified and modified templates. Thus, such methods can use the pulse width of the signal from base sequencing, the base inter-pulse duration (IPD), and base identity to detect modifications at a base or adjacent bases (see, e.g., Weirather et al., F1000Research, 6:100, 2017). Therefore, SMRT sequencing can be used to detect base modifications such as 5-caC, 4mC, 5mC, 5hmC, 6mA, and 8oxoG (Gouil & Keniry Essays in Biochemistry (2019) 63 639-648). Thus, in some embodiments, sequencing includes SMRT sequencing. In such embodiments, end repair can be performed using dNTPs including 5-caC, 4mC, 5mC, 5hmC, 6mA, and / or 8oxoG.
[0355] Some sequencing reactions involve the use of enzymes to control the passage of nucleic acids through nanopores, in which case the reaction data may include both enzyme dynamics and other behaviors, as well as fluctuations in the current passing through the nanopore. For example, ratchet proteins, helicases, or motor proteins can be used to push or pull nucleic acid molecules through pores in biological or synthetic membranes. The dynamics of these proteins may vary depending on the context of the nucleic acid sequence on which they act. For example, they may slow down or pause at modified bases, and this behavior, captured as part of the reaction data, indicates the presence of modified bases even if the modified bases are not within the sensing region of the nanopore.
[0356] One example of a nanopore-based single-molecule sequencing system is the system commercialized by Oxford Nanopore Technologies (ONT) (Weirather JL, et al., F1000Research, 6:100, 2017). ONT directly sequences native single-stranded DNA (ssDNA) molecules by measuring characteristic current changes as bases pass through nanopores by molecular motion proteins. ONT uses a hairpin library structure similar to the PacBio circular DNA template: the DNA template and its complement are bound by a hairpin adapter. Thus, the DNA template passes through the nanopore, followed by the hairpin, and finally the complement. The raw read can be split into two “1D” reads (“template” and “complement”) by removing the adapter. The consensus sequence of the two “1D” reads is a “2D” read with higher precision.
[0357] Nanopore sequencing can be used to detect base modifications including 5-caC, 5mC, 5hmC, 6mA, BrdU, FldU, IdU, and EdU (see, for example, Gouil & Keniry Essays in Biochemistry (2019) 63 639-648; Kutyavin, Biochemistry (2008), 47, 51, 13666-1367; Muller et al., Nature Methods (2019), volume 16, pages 429-436; Hennion et al., Genome Biology (2020), volume 21, Article number: 125). Therefore, in some embodiments, sequencing includes nanopore sequencing. In such embodiments, end repair can be performed using dNTPs including 5-caC, 4mC, 5mC, 5hmC, 6mA, BrdU, FldU, IdU, and / or EdU.
[0358] The 5-letter and 6-letter sequencing methods include whole-genome sequencing methods capable of sequencing A, C, T, and G in addition to 5mC and 5hmC, providing 5-letter (A, C, T, G and either 5mC or 5hmC) or 6-letter (A, C, T, G, 5mC, and 5hmC) digital readouts in a single workflow. DNA sample processing is entirely enzymatic, avoiding DNA degradation and genomic coverage bias associated with bisulfite treatment. In an exemplary 5-letter sequencing method developed by Cambridge Epigenetix, the sample DNA is first fragmented by sonication and then ligated at both ends to short synthetic DNA hairpin adapters (Fullgrabe, et al. 2022, bioRxiv doi: https: / / doi.org / 10.1101 / 2022.07.08.499285). The constructs are then split to separate the sense and antisense sample strands. For each original sample strand, a complementary copy strand is synthesized by DNA polymerase extension at the 3' end, generating a hairpin construct with the original sample DNA strand attached to its complementary strand via a synthesis loop, lacking epigenetic modifications. A sequencing adapter is then ligated to the end. Modified cytosines are protected by the enzyme. Unprotected C is then deaminated to uracil, which is subsequently read as thymine. In any such embodiment, the amplification method may include a uracil and / or dihydrouracil-resistant amplification method, such as PCR using a uracil and / or dihydrouracil-resistant DNA polymerase (i.e., a DNA polymerase capable of reading and amplifying templates containing uracil and / or dihydrouracil bases). The deaminated construct is no longer fully complementary and has substantially reduced double-strand stability, thus readily opening the hairpin and allowing amplification by PCR. The construct can be sequenced in paired-end form, where lead 1 (P1 prime) is the original strand and lead 2 (P2 prime) is the copy strand.Since read data are pairwise aligned, read 1 is aligned to its complementary read 2. Cognate residues from both reads are degraded by computer to produce a single genetic or epigenetic letter. Cognate base pairings other than acceptable ones are the result of imperfect fidelity at some stage, including incorrect base calling in sample preparation, amplification, or sequencing. Since these errors occur independently of cognate bases in each strand, substitutions result in unacceptable pairs. Unacceptable pairs are masked (marked as N) within the degraded read, preserving the read itself and resulting in minimal information loss and high read-level accuracy. The degraded reads are aligned to a reference genome. Genetic variants and methylation counts are produced by read counting at the base level.
[0359] 5hmC has been shown to have value as a marker of biological status and disease, including early cancer detection from cell-free DNA. In the adaptation of 5-letter to 6-letter sequencing, 5mC is deambiguated from 5hmC without compromising the genetic base calling within the same sample fragment. The first three steps of the workflow are identical to the 5-letter sequencing described above to generate an adapter-ligated sample fragment with a synthetic copy strand. Methylation at 5mC is enzymatically copied to C across CpG units in the copy strand, while 5hmC is enzymatically protected from such copying. Thus, the unmodified C, 5mC, and 5hmC in each of the original CpG units are distinguished by a unique 2-base combination. Next, the unmodified cytosine is deaminated to uracil, which is then read as thymine. The DNA is subjected to PCR amplification and sequencing as previously described. Reads are pairwise aligned and degraded using the 2-base code. Unmodified C, 5mC, and 5hmC can each be decomposed because these three types of CpG units are separate sequencing environments for two-base coding.
[0360]
[0361] Sequencing can be performed on one or more nucleic acid fragment types or compartments known to contain markers for cancer or other diseases. Sequencing can also be performed on any nucleic acid fragment present in the sample. Sequencing can provide genome sequence coverage of at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome. In other cases, genome sequence coverage may be less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome.
[0362] Simultaneous sequencing reactions can be performed using multiplex sequencing techniques. In some embodiments, cell-free polynucleotides are sequenced in at least about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. In other embodiments, cell-free polynucleotides are sequenced in less than about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. Sequencing reactions are typically performed sequentially or simultaneously. Subsequent data analysis is generally performed with respect to all or some of the sequencing reactions. In some embodiments, data analysis is performed for at least about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. In other embodiments, data analysis may be performed for fewer than about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. An example read depth is about 1,000 to about 50,000 reads per locus (base position).
[0363] In some embodiments, nucleic acid populations are prepared for sequencing by enzymatically forming blunt ends on double-stranded nucleic acids having single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Examples of enzymes or catalytic fragments of which may be used as needed include Krenow large fragments and T4 polymerase. At the 5' overhang, the enzyme typically extends the recessed 3' end on the opposite strand until its 3' end aligns perfectly with the 5' end to produce a blunt end. At the 3' overhang, the enzyme generally digests from the 3' end to the 5' end of the opposite strand, and may also digest beyond the 5' end. If this digestion proceeds beyond the 5' end of the opposite strand, an enzyme with the same polymerase activity used for the 5' overhang can fill the gap. The formation of blunt ends on double-stranded nucleic acids facilitates, for example, adapter binding and subsequent amplification.
[0364] In some embodiments, the nucleic acid population is subjected to additional processing, such as the conversion of single-stranded nucleic acids to double-stranded nucleic acids and / or RNA to DNA. These forms of nucleic acids are also ligated to adapters and amplified as needed.
[0365] Sequenced nucleic acids can be produced by sequencing nucleic acids subjected to the blunt-end formation process described above, whether previously amplified or not, and, if necessary, other nucleic acids in the sample. Sequentially, sequenced nucleic acids may refer to the sequence of a nucleic acid (i.e., sequence information), or to a nucleic acid whose sequence has been determined. Sequence data for individual nucleic acid molecules in a sample can be obtained, directly or indirectly, from the consensus sequences of the amplified products of individual nucleic acid molecules in the sample.
[0366] In some embodiments, double-stranded nucleic acids with single-stranded overhangs in a sample are blunt-ended, and both ends are ligated to an adapter containing a barcode. Sequencing then determines not only the nucleic acid sequence but also the inline barcode induced by the adapter. The blunt-ended DNA molecule is ligated, as necessary, to the blunt ends of an adapter that is at least partially double-stranded (e.g., Y-shaped or bell-shaped). Alternatively, complementary nucleotide tails can be attached to the blunt ends of the sample nucleic acid and the adapter to facilitate ligation (e.g., adherent-end ligation).
[0367] Typically, a nucleic acid sample is brought into contact with a sufficient number of adapters, and therefore, the probability that any two copies of the same nucleic acid will receive the same combination of adapter barcodes from adapters ligated to both ends is low (e.g., <1 or 0.1%). Using adapters in this way allows for the identification of families of nucleic acid sequences that have the same start and stop points on the reference nucleic acid and are ligated to the same combination of barcodes. Such families represent the sequences of the amplified products of the nucleic acids in the sample before amplification. When modified by blunt end formation and adapter ligation, the sequences of family members can be compiled to derive the consensus nucleotide or complete consensus sequence of the nucleic acid molecule in the original sample. In other words, a nucleotide occupying a specific position in the nucleic acid in the sample is determined to be the consensus of the nucleotide occupying the corresponding position in the family member sequence. A family may include sequences from one or both strands of a double-stranded nucleic acid. If a family member includes sequences from both strands of a double-stranded nucleic acid, the sequence from one strand is converted to its complementary strand for the purpose of compiling all sequences to derive a consensus nucleotide or sequence. Some families contain only a single member sequence. In this case, this sequence can be considered the nucleic acid sequence in the sample before amplification. Alternatively, families containing only a single member sequence can be excluded from subsequent analyses.
[0368] For further details regarding nucleic acid sequencing, including the forms and applications described herein, see, for example, Levy et al., Annual Review of Genomics and Human Genetics, 17: 95-115 (2016), Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012), Voelkerding et al., Clinical Chem., 55: 641-658 (2009), MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009), and Astier et al., J Am Chem Soc., 128(5):1705-10. (2006), U.S. Patent Nos. 6,210,891, 6,258,568, 6,833,246, 7,115,400, 6,969,488, 5,912,148, 6,130,073, 7,169,560, 7,282,337, 7,482,120, and 7,501,245 These references are also provided in U.S. Patent Nos. 6,818,395, 6,911,345, 7,501,245, 7,329,492, 7,170,050, 7,302,146, 7,313,308, and 7,476,503, and these references are incorporated herein by reference in their entirety. a. Sequencing panel
[0369] To increase the likelihood of detecting a target genomic region and, if necessary, the likelihood of detecting a tumor exhibiting the mutation, the DNA parcels to be sequenced may include a panel of genes or genomic parcels containing known genomic regions. By selecting a limited number of parcels (e.g., a limited panel) for sequencing, the total sequencing required (e.g., the total amount of nucleotides to be sequenced) can be reduced. The sequencing panel may target multiple different genes or regions, for example, to detect a single cancer, a set of cancers, or all cancers. Alternatively, DNA can be sequenced without using a sequencing panel by whole-genome sequencing (WGS) or other unbiased sequencing methods. Examples of suitable panels and targets for use in panels can be found in the epigenetic targets described in international application WO2020160414, filed on 31 January 2020, the aforementioned reference patent document, is incorporated herein by reference in its entirety.
[0370] In some embodiments, a panel is selected that targets multiple different genes or genomic regions (e.g., CHIP genes, transcription factor binding regions, distal regulatory elements (DREs), repeat elements, intron-exon junctions, transcription start sites (TSSs), and / or similar) such that a predetermined proportion of subjects with cancer exhibit genetic variants or tumor markers in one or more different genes within the panel. The panel may be selected to limit the region to be sequenced to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA. The panel may further be selected to achieve a desired sequence reading depth. The panel may be selected to achieve a desired sequence reading depth or sequence reading coverage with respect to the amount of base pairs being sequenced. The panel may be selected to achieve theoretical sensitivity, theoretical specificity, and / or theoretical precision for detecting one or more genetic variants in a sample.
[0371] The genes included in this panel are ATM, ATR, BAP1, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, HDAC2, MRE11, NBN, PALB2, RAD50, RAD51, RAD51B, RAD51C, RAD51D, RAD54L, XRCC2, XRCC3, DNMT3A, TP53, LRP1B, KRAS, MARCH11, TAC1, TCF21, SHOX2, p16, Casp8, CDH13, MGMT, MLH1, MSH2, TSLC1, APC, DKK1, DKK3, LKB1, WIF1, RUN This may include one or more of the following: X3, GATA4, GATA5, PAX5, E-cadherin, H-cadherin, VIM, SEPT9, CYCD2, TFPI2, GATA4, RARB2, p16INK4a, APC, NDRG4, HLTF, HPP1, hMLH1, RASSF1A, IGFBP3, ITGA4, PIK3CA, ERBB2 (HER2), BRCA1 / 2, NTRK1 / 2 / 3, MSI-High, ESR1, ATM, HRR, FGFR2 / 3, IDH1, KRAS, NRAS, BRAF, KIT, PDGFRA, EGFR, ALK, ROS1, MET, TMB, or RET.
[0372] Probes for detecting a panel of regions can include not only those for detecting target genomic regions (hotspot regions) but also nucleosome recognition probes (e.g., KRAS codons 12 and 13), and such probes can be designed to optimize capture based on an analysis of cfDNA coverage and fragment size variations influenced by nucleosome binding patterns and GC sequence composition. In this case, the regions used may also include non-hotspot regions optimized based on nucleosome location and GC model. The panel may include multiple subpanels, including a subpanel for identifying primary tissue (e.g., using published literature to define 50-100 baits representing genes (not necessarily promoters) with the most diverse transcriptional profiles across tissue), a subpanel for identifying whole-genome scaffolds (e.g., for identifying hyperconservative genomic content and sparsely tiling across chromosomes with a small number of probes for copy-number-based lining), and a subpanel for identifying transcription start sites (TSS) / CpG islands (e.g., for capturing differential methylation regions (e.g., variable methylation regions (DMRs)) in the promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer)). In some embodiments, the marker for primary tissue is a tissue-specific epigenetic marker.
[0373] Some examples of the list of target gene locations can be found in Tables 1 and 2. In some embodiments, the gene locations used in the methods of the disclosure include at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or 97 of the genes in Table 1. In some embodiments, the gene locations used in the methods of the disclosure include at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 1. In some embodiments, the gene locations used in the methods of the disclosure include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the CNVs in Table 1. In some embodiments, the gene locations used in the methods of the disclosure include at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 1. In some embodiments, the gene locations used in the methods of the disclosure include at least one, at least two, or at least a portion of three of the indels in Table 1. In some embodiments, the gene locations used in the methods of this disclosure include at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, or 115 of the genes in Table 2.In some embodiments, the gene locations used in the methods of the disclosure include at least five, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 2. In some embodiments, the gene locations used in the methods of the disclosure include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the CNVs in Table 2. In some embodiments, the gene locations used in the methods of the disclosure include at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 2. In some embodiments, the gene locations used in the methods of the present disclosure include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or at least 18 of the indels in Table 2. Each of these target gene locations can be identified as a skeletal region or hotspot region for a given baitset panel. An example of a list of target hotspot gene locations can be found in Table 3. In some embodiments, the gene locations used in the methods of the present disclosure include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 3.Each hotspot gene location is listed along with several characteristics, including the associated gene, the chromosome in which it resides, the genomic start and stop locations representing the gene's locus, the length of the gene's locus in base pairs, the exons covered by the gene, and several other crucial features that the given gene location of interest may attempt to acquire (e.g., the type of mutation). [Table 1] [Table 2-1] [Table 2-2] [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4] [Table 3-5]
[0374] In some embodiments, one or more regions within the panel include one or more loci from one or more genes to detect residual cancer after surgery. This detection may be earlier than that possible with existing cancer detection methods. In some embodiments, one or more gene locations within the panel include one or more loci from one or more genes to detect cancer in high-risk patient populations. For example, smokers have a much higher rate of lung cancer than the general population. Furthermore, smokers may develop other lung conditions that make cancer detection more difficult, such as the development of irregular nodules in the lungs. In some embodiments, the methods described herein can detect cancer in high-risk patients earlier than that possible with existing cancer detection methods.
[0375] Gene locations to be included in the sequencing panel can be selected based on the number of subjects with cancer in which the tumor marker is located in that gene or region. Gene locations to be included in the sequencing panel can also be selected based on the prevalence of cancer and subjects with cancer in which the tumor marker is located in that gene. The presence of a tumor marker within a region may indicate a subject with cancer.
[0376] In some cases, information from one or more databases can be used to select a panel. Information about cancer may be derived from cancer tumor biopsies or cfDNA assays. A database may contain information describing a population of sequenced tumor samples. A database may contain information about mRNA expression in tumor samples. A database may contain information about regulatory elements or genomic regions in tumor samples. Information about sequenced tumor samples may include the frequencies of various genetic variants and may describe the genes or regions in which these genetic variants exist. Genetic variants may be tumor markers. A non-limiting example of such a database is COSMIC, a catalog of somatic mutations found in various cancers. For a particular cancer, COSMIC ranks genes based on their mutation frequency. Genes with high mutation frequencies among given genes can be selected for inclusion in a panel. For example, COSMIC shows that 33% of a population of sequenced breast cancer samples have the TP53 mutation, and 22% of a population of sampled breast cancers have the KRAS mutation. Other ranked genes, including APC, have mutations found in only about 4% of the population of sequenced breast cancer samples. TP53 and KRAS can be included in the sequencing panel based on their relatively high frequency in the sampled breast cancers (compared to APC, for example, which is present at a frequency of about 4%). While COSMIC is provided as a non-limiting example, any database or set of information that associates cancer with tumor markers located in genes or genetic regions can be used. In another example, as provided by COSMIC, 380 out of 1156 biliary tract cancer samples (33%) carry the TP53 mutation. Several other genes, such as APC, have mutations in 4–8% of all samples. Therefore, TP53 can be selected for inclusion in the panel based on its relatively high frequency in the population of biliary tract cancer samples.
[0377] A gene or genomic compartment may be selected for a panel if the frequency of a tumor marker is significantly higher in the sampled tumor tissue or circulating tumor DNA than in a given background population. A combination of gene locations may be selected to include in the panel such that at least the majority of subjects with cancer have a tumor marker or genomic region located in at least one gene location or gene within the panel. For a particular cancer or set of cancers, a combination of gene locations may be selected based on data showing that the majority of subjects have one or more tumor markers in one or more selected regions. For example, to detect cancer 1, a panel including regions A, B, C, and / or D may be selected based on data showing that 90% of subjects with cancer 1 have tumor markers in regions A, B, C, and / or D of the panel. Alternatively, a tumor marker may be shown to be independently present in two or more regions of subjects with cancer, and therefore, combined, tumor markers in two or more regions are present in the majority of the population of subjects with cancer. For example, to detect cancer 2, a panel including regions X, Y, and Z can be selected based on data showing that 90% of subjects have tumor markers in one or more regions, and that in 30% of such subjects the tumor marker is detected only in region X, while in the remaining subjects where the tumor marker is detected, it is detected only in regions Y and / or Z. Tumor markers located at one or more gene locations previously proven to be associated with one or more cancers can indicate or predict that a subject has cancer if the tumor marker is detected in one or more regions in 50% or more of those regions at that time. Computational methods, such as models that utilize the conditional probability of detecting cancer based on the frequency of cancer for a set of tumor markers in one or more regions, can be used to predict which regions, individually or in combination, may predict cancer.Other methods for panel selection include the use of databases describing information from studies using comprehensive genomic profiling and / or whole-genome sequencing (WGS, RNA-seq, Chip-seq, bisulfite sequencing, ATAC-seq, etc.) of tumors in large panels. Information gathered from the literature may also describe pathways that are generally affected and mutate in certain cancers. The use of ontologs describing genetic information can provide further information for panel selection.
[0378] The genes included in the sequencing panel may include the complete transcription region, promoter region, enhancer region, regulatory elements, and / or downstream sequences. To further increase the likelihood of detecting tumors exhibiting mutations, only exons may be included in the panel. The panel may include all exons of a selected gene, or one or more exons of a selected gene. The panel may include exons from each of several different genes. The panel may also include at least one exon from each of several different genes.
[0379] In some embodiments, a panel of exons from each of several different genes is selected such that a predetermined proportion of subjects having cancer exhibits a genetic variant in at least one exon within the panel of exons.
[0380] At least one whole exon from each of the different genes within a gene panel can be sequenced. The panel being sequenced may contain exons from multiple genes. The panel may contain exons from 2 to 100 different genes, 2 to 70 genes, 2 to 50 genes, 2 to 30 genes, 2 to 15 genes, or 2 to 10 genes.
[0381] The selected panel may contain a variety of exons. The panel may contain 2 to 3000 exons. The panel may contain 2 to 1000 exons. The panel may contain 2 to 500 exons. The panel may contain 2 to 100 exons. The panel may contain 2 to 50 exons. The panel may contain 300 or fewer exons. The panel may contain 200 or fewer exons. The panel may contain 100 or fewer exons. The panel may contain 50 or fewer exons. The panel may contain 40 or fewer exons. The panel may contain 30 or fewer exons. The panel may contain 25 or fewer exons. The panel may contain 20 or fewer exons. The panel may contain 15 or fewer exons. The panel may contain 10 or fewer exons. The panel may contain 9 or fewer exons. The panel may contain 8 or fewer exons. The panel may contain 7 or fewer exons.
[0382] The panel may contain one or more exons from multiple different genes. The panel may contain one or more exons from each of a certain proportion of multiple different genes. The panel may contain at least two exons from each of at least 25%, 50%, 75%, or 90% of different genes. The panel may contain at least three exons from each of at least 25%, 50%, 75%, or 90% of different genes. The panel may contain at least four exons from each of at least 25%, 50%, 75%, or 90% of different genes.
[0383] The size of a sequencing panel can vary. For example, the size of a sequencing panel can be larger or smaller (in terms of nucleotide size) depending on several factors, including the total amount of nucleotides being sequenced, or the number of unique molecules being sequenced in a particular region of the panel. Sequencing panels can be 5kb to 50kb in size. Sequencing panels can be 10kb to 30kb in size. Sequencing panels can be 12kb to 20kb in size. Sequencing panels can be 12kb to 60kb in size. The sequencing panel may be at least 10kb, 12kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 110kb, 120kb, 130kb, 140kb, or 150kb in size. The sequencing panel may also be less than 100kb, 90kb, 80kb, 70kb, 60kb, or 50kb in size.
[0384] The panel selected for sequencing may contain at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 gene locations (e.g., each containing a genomic region of interest). In some cases, gene locations within the panel that are relatively small in size are selected. In some cases, the regions within the panel have a size of approximately 10kb or less, approximately 8kb or less, approximately 6kb or less, approximately 5kb or less, approximately 4kb or less, approximately 3kb or less, approximately 2.5kb or less, approximately 2kb or less, approximately 1.5kb or less, or approximately 1kb or less or less. In some cases, gene locations within a panel have sizes ranging from approximately 0.5kb to 10kb, 0.5kb to 6kb, 1kb to 11kb, 1kb to 15kb, 1kb to 20kb, 0.1kb to 10kb, or 0.2kb to 1kb. For example, regions within a panel may have sizes ranging from approximately 0.1kb to 5kb.
[0385] The panels selected herein may enable deep sequencing sufficient to detect low-frequency genetic variants (e.g., in cell-free nucleic acid molecules obtained from a sample). The amount of a genetic variant in a sample may be referred to in terms of the minor allele frequency for a given genetic variant. Minor allele frequency may refer to the frequency at which a minor allele (e.g., not the most frequent allele) is present in a given nucleic acid population, such as a sample. Genetic variants with low minor allele frequencies may be relatively infrequent in a sample. In some cases, the panel may enable the detection of genetic variants with minor allele frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, or 0.5%. The panel may enable the detection of genetic variants with minor allele frequencies of 0.001% or greater. The panel may enable the detection of genetic variants with minor allele frequencies of 0.01% or greater. The panel enables the detection of genetic variants present in the sample at frequencies as low as 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel enables the detection of tumor markers present in the sample at frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel may enable the detection of tumor markers present in the sample at frequencies as low as 1.0%. The panel may enable the detection of tumor markers present in the sample at frequencies as low as 0.75%. The panel can enable the detection of tumor markers at frequencies as low as 0.5% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.25% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.1% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.075% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.05% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.025% in the sample.The panel may enable the detection of tumor markers at frequencies as low as 0.01% in a sample. The panel may enable the detection of tumor markers at frequencies as low as 0.005% in a sample. The panel may enable the detection of tumor markers at frequencies as low as 0.001% in a sample. The panel may enable the detection of tumor markers at frequencies as low as 0.0001% in a sample. The panel may enable the detection of tumor markers in sequenced cfDNA at frequencies as low as 1.0% to 0.0001% in a sample. The panel may enable the detection of tumor markers in sequenced cfDNA at frequencies as low as 0.01% to 0.0001% in a sample.
[0386] Genetic variants can be expressed as a percentage of the target population having a disease (e.g., cancer). In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the population with cancer exhibit one or more genetic variants in at least one region within the panel. For example, at least 80% of the population with cancer may exhibit one or more genetic variants in at least one genomic location within the panel.
[0387] The panel may include one or more locations containing the target genomic region from each of one or more genes. In some cases, the panel may include one or more locations containing the target genomic region from each of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the panel may include one or more locations containing the target genomic region from each of at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the panel may include one or more locations containing the target genomic region from each of about 1 to about 80, 1 to about 50, about 3 to about 40, 5 to about 30, or 10 to about 20 different genes.
[0388] A location within the panel containing a genomic region can be selected so that one or more epigenetically modified regions are detected. These epigenetically modified regions may be acetylated, methylated, ubiquitinated, phosphorylated, SUMOylated, ribosylated, and / or citrullinated. For example, a region within the panel can be selected so that one or more methylated regions are detected. In some embodiments, the genomic region of the panel may contain one or more of the following genes: DNMT3A, TP53, LRP1B, KRAS, MARCH11, TAC1, TCF21, SHOX2, p16, Casp8, CDH13, MGMT, MLH1, MSH2, TSLC1, APC, DKK1, DKK3, LKB1, WIF1, RUNX3, GATA4, GATA5, PAX5, E-cadherin, H-cadherin, VIM, SEPT9, CYC D2, TFPI2, GATA4, RARB2, p16INK4a, APC, NDRG4, HLTF, HPP1, hMLH1, RASSF1A, IGFBP3, ITGA4, PIK3CA, ERBB2(HER2), BRCA1 / 2, NTRK1 / 2 / 3, MSI-High, ESR1, ATM, HRR, FGFR2 / 3, IDH1, KRAS, NRAS, BRAF, KIT, PDGFRA, EGFR, ALK, ROS1, MET, TMB, or RET.
[0389] Regions within a panel can be selected to contain sequences that are differentially transcribed across one or more tissues. In some cases, locations containing genomic regions may contain sequences that are transcribed at a higher level in certain tissues compared to others. For example, locations containing genomic regions may contain transcribed sequences in certain tissues but not in others.
[0390] Gene locations within a panel may include coding and / or non-coding sequences. For example, gene locations within a panel may include one or more sequences in exons, introns, 3' untranslated regions, 5' untranslated regions, regulatory elements, transcription start sites, and / or splice sites. In some cases, regions within a panel may include other non-coding sequences, including pseudogenes, repetitive sequences, transposons, viral elements, and telomeres. In some cases, gene locations within a panel may include sequences in non-coding RNA, such as ribosomal RNA, transfer RNA, Piwi-binding RNA, orphan non-coding RNA, and microRNAs.
[0391] Gene locations within the panel can be selected to detect (diagnose) cancer at a desired sensitivity level (e.g., by detecting one or more genetic variants). For example, regions within the panel can be selected to detect cancer with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% (e.g., by detecting one or more genetic variants). Gene locations within the panel can also be selected to detect cancer with 100% sensitivity.
[0392] Gene locations within the panel can be selected to detect (diagnose) cancer at a desired specificity level (e.g., by detecting one or more genetic variants). For example, gene locations within the panel can be selected to detect cancer with a specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% (e.g., by detecting one or more genetic variants). Gene locations within the panel can also be selected to detect one or more genetic variants with 100% specificity.
[0393] Genetic locations within a panel can be selected to detect (diagnose) cancer with a desired positive predictive value. Positive predictive values can be increased by improving sensitivity (the chance of detecting an actual positive) and / or specificity (the chance of not mistaking an actual negative for a positive). As a non-limiting example, genetic locations within a panel can be selected to detect one or more genetic variants with positive predictive values of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within a panel can be selected to detect one or more genetic variants with a 100% positive predictive value.
[0394] Gene locations within the panel can be selected to detect (diagnose) cancer with the desired accuracy. As used herein, the term “accuracy” refers to the ability of a test to discriminate between a diseased state (e.g., cancer) and a healthy state. Accuracy can be quantified using measures such as sensitivity and specificity, predictive value, likelihood ratio, area under the ROC curve, Joden index, and / or diagnostic odds ratio.
[0395] Accuracy can be presented as a percentage, which refers to the ratio of the number of trials that yielded correct results to the total number of trials performed. Regions within the panel can be selected to detect cancer with an accuracy of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Gene locations within the panel can be selected to detect cancer with 100% accuracy.
[0396] Panels can be selected to achieve high sensitivity and to detect low-frequency genetic variants. For example, a panel can be selected to detect genetic variants or tumor markers present at frequencies as low as 0.01%, 0.05%, or 0.001% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Gene locations within the panel can be selected to detect tumor markers present at frequencies of 1% or less in a sample with a sensitivity of 70% or higher. The panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.1%, with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.01%, with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.001% with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0397] Panels can be selected to achieve high specificity and to detect low-frequency genetic variants. For example, a panel can be selected to detect genetic variants or tumor markers present at frequencies as low as 0.01%, 0.05%, or 0.001% in a sample with a specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Gene locations within the panel can be selected to detect tumor markers present at frequencies of 1% or less in a sample with a specificity of 70% or higher. A panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.1%, with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.01%, with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.001%, with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0398] Panels can be selected to achieve high accuracy and to detect low-frequency genetic variants. Panels can be selected to detect genetic variants or tumor markers present in the sample at frequencies as low as 0.01%, 0.05%, or 0.001% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Gene locations within the panel can be selected to detect tumor markers present in the sample at frequencies of 1% or less with 70% or higher accuracy. Panels can be selected to detect tumor markers present in the sample at frequencies as low as 0.1% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The panel can be selected to detect tumor markers present in the sample at frequencies as low as 0.01% with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers present in the sample at frequencies as low as 0.001% with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0399] Panels can be selected to provide high predictive capabilities and to detect low-frequency genetic variants. Panels can be selected so that genetic variants or tumor markers present at frequencies as low as 0.01%, 0.05%, or 0.001% in the sample can have positive predictive values of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0400] The concentration of the probe or bait used in the panel can be increased (from 2 to 6 ng / μL) to capture more nucleic acid molecules in the sample. The concentration of the probe or bait used in the panel may be at least 2 ng / μL, 3 ng / μL, 4 ng / μL, 5 ng / μL, 6 ng / μL, or higher. The probe concentration may be approximately 2 ng / μL to 3 ng / μL, approximately 2 ng / μL to 4 ng / μL, approximately 2 ng / μL to 5 ng / μL, or approximately 2 ng / μL to 6 ng / μL. The concentration of the probe or bait used in the panel may be 2 ng / μL or higher to 6 ng / μL or lower. In some cases, this allows for the analysis of more molecules in the biologic, thereby enabling the detection of lower-frequency alleles.
[0401] In one embodiment, the sequencing pipeline 205 can be used to subject the panel to one or more of the following: whole-genome bisulfite sequencing (WGBS); whole-genome sequencing (WGS); and / or targeted sequencing methods that examine copy number variants (CNVs) and single nucleotide variants (SNVs).
[0402] By combining genetic and / or epigenetic information obtained from the subject's DNA, it is possible to determine whether the subject has cancer or the likelihood of the subject having cancer. A detailed description of a method for analyzing cell-free human DNA for both cancer-related genetic and epigenetic variants can be found in U.S. Provisional Patent Application No. 62 / 799637, which is incorporated herein by reference in its entirety. Further guidance for analyzing cell-free DNA to detect cancer can be found in several places, among others, U.S. Patent No. 9,834,822, PCT Application WO2018064629A1, and PCT Application WO2017106768A1.
[0403] Various embodiments include the step of sequencing DNA (e.g., cfDNA) for the purpose of detecting genetic variants of cancer-related genes. Various embodiments also include the step of sequencing DNA (e.g., cfDNA) for the purpose of detecting epigenetic variants of cancer-related genes, including, but not limited to, DNA sequences and nucleosome fragmentation patterns that are differentially methylated in cancerous and non-cancerous cells, e.g., those described in U.S. Patent Application Publication 2017 / 0211143.
[0404] In some embodiments, a captured set of nucleic acids, for example, nucleic acids including DNA (e.g., cfDNA), is provided. With respect to the disclosed method, the captured set of DNA may be provided, for example, after the capture and / or separation steps described herein. The captured set may include DNA corresponding to one or both of the sequence variable target region set and the epigenetic target region set. In some embodiments, the captured set includes DNA corresponding to the sequence variable target region set and the epigenetic target region set. In all embodiments described herein, including the sequence variable target region set and the epigenetic target region set, the sequence variable target region set includes regions not present in the epigenetic target region set, and conversely, the epigenetic target region set includes regions not present in the sequence variable target region set, although in some cases a fraction of these regions may overlap (e.g., a fraction of genomic locations may be shown in both target region sets). (A) Methylation target region set
[0405] In some embodiments, a set of epigenetic target regions is captured. This set of epigenetic target regions may include DNA and one or more types of target regions that are likely to differentiate neoplastic (e.g., tumor or cancer) cells from healthy cells, e.g., non-neoplastic circulating cells. The set of epigenetic target regions can be analyzed in a variety of ways, including methods that do not rely on high precision in sequencing specific nucleotides within the target. Exemplary types of such regions are discussed in detail herein. In some embodiments, the methods according to this disclosure include the step of determining whether the cfDNA molecule corresponding to the set of epigenetic target regions contains or exhibits cancer-related epigenetic modifications (e.g., hypermethylation in one or more hypermethylated variable target regions; one or more perturbations of CTCF binding; and / or one or more perturbations of transcription start sites) and / or copy number diversity (e.g., local amplification). Such analysis can be performed by sequencing and may require less data (e.g., number of sequence reads or depth of sequencing coverage) to determine the presence or absence of sequence mutations, such as base substitutions, insertions, or deletions. The epigenetic target region set may also include one or more control regions, such as those described herein.
[0406] In some embodiments, the epigenetic target region set has a footprint of at least 100kb, for example, at least 200kb, at least 300kb, or at least 400kb. In some embodiments, the epigenetic target region set has a footprint in the range of 100 to 1000kb, for example, 100 to 200kb, 200 to 300kb, 300 to 400kb, 400 to 500kb, 500 to 600kb, 600 to 700kb, 700 to 800kb, 800 to 900kb, and 900 to 1,000kb. (B) Highly methylated variable target region
[0407] In some embodiments, the epigenetic target region set includes one or more hypermethylated variable target regions. Generally, a hypermethylated variable target region refers to a region where an observed increase in methylation level indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by neoplastic cells such as tumor or cancer cells. For example, hypermethylation of tumor suppressor gene promoters has been repeatedly observed. See, for example, Kang et al., Genome Biol. 18:53 (2017) and the references cited therein.
[0408] A comprehensive discussion of methylation-variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866:106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylation-variable target regions containing genes or portions thereof, based on colorectal cancer (CRC) studies, is provided in Table 4. Many of these genes are likely to be associated with cancers other than colorectal cancer; for example, TP53 is widely recognized as a very important tumor suppressor, and inactivation based on hypermethylation of this gene may be a common carcinogenic mechanism.
[0409] [Table 4]
[0410] In some embodiments, the highly methylated variable target region includes multiple genes or portions thereof listed in Table 4, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the genes or portions thereof listed in Table 4. For example, each locus included in the target region may have one or more probes having a hybridization site that binds between the transcription start site of the gene and the stop codon (the last stop codon of the gene to be spliced alternatively). In some embodiments, one or more probes bind within 300 bp upstream and / or downstream of the genes or portions thereof listed in Table 4, for example, within 200 bp or 100 bp.
[0411] Methylation-variable target regions in various types of lung cancer are, for example, Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77:453-74 (2015); Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017); Shi et al., BMC Genomics 18:901 (2017); Schneider et al., BMC Cancer. 11:102 (2011); Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016); Skvortsova et al., Br. J. Cancer. 94(10):1492-1495 (2006); Kim et al., Cancer Res. 61:3419-3424 (2001);Furonaka et al., Pathology International 55:303-309 (2005);Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014);Kim et al., Oncogene. 20:1765-70 (2001);Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003);Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005);Heller et al., Oncogene 25:959-968 (2006);Licchesi et al., Carcinogenesis. 29:895-904 (2008);Guo et al., Clin. Cancer Res. 10:7917-24 (2004);Palmisano et al., This is discussed in detail in Cancer Res. 63:4620-4625 (2003) and Toyooka et al., Cancer Res. 61:4556-4560, (2001).
[0412] Table 5 provides an exemplary set of hypermethylated variable target regions, including genes or portions thereof, based on lung cancer research. Many of these genes are likely to be associated with cancers other than lung cancer; for example, Casp8 (caspase 8) is a key enzyme in programmed cell death, and inactivation based on hypermethylation of this gene may be a general carcinogenic mechanism not limited to lung cancer. In addition, several genes are listed in both Tables 4 and 5, which present a general concept.
[0413] [Table 5-1] [Table 5-2]
[0414] Any of the above embodiments relating to the target regions identified in Table 2 may be combined with any of the above embodiments relating to the target regions identified in Table 1. In some embodiments, the highly methylated variable target region comprises a plurality of genes or portions thereof listed in Table 1 or Table 2, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the genes or portions thereof listed in Table 1 or Table 2.
[0415] Further hypermethylation target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a probabilistic method called Cancer Locator using hypermethylation target regions from breast, colon, kidney, liver, and lung. In some embodiments, hypermethylation target regions may be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylation target regions include one, two, three, four, or five subsets of hypermethylation target regions that collectively exhibit hypermethylation in one, two, three, four, or five of the following cancers: breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0416] Low-methylation variable target region
[0417] Overall hypomethylation is a common phenomenon in various cancers. See, for example, Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (a review article describing the observation of hypomethylation in colon, ovarian, prostate, leukemia, hepatocyte, and cervical cancers). For example, regions such as repeating elements, e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA, as well as intergeneric regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Therefore, in some embodiments, a set of epigenetic target regions includes hypomethylation-variable target regions in which the observed decrease in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by neoplastic cells such as tumor or cancer cells.
[0418] In some embodiments, the low-methylation variable target region includes repeat elements and / or intergenetic regions. In some embodiments, the repeat elements include one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and / or satellite DNA.
[0419] Exemplary specific genomic regions exhibiting cancer-related hypomethylation include, for example, nucleotides 8403565–8953708 and 151104701–151106035 of human chromosome 1, as defined by the HG19 or HG38 human genome constructs. In some embodiments, the hypomethylation variable target region overlaps with or includes one or both of these regions. (C)CTCF binding region
[0420] CTCF is a DNA-binding protein that contributes to chromatin structure and often colocalizes with cohesin. Perturbations of CTCF binding sites have been reported in various different cancers. See, for example, Katainen et al., Nature Genetics, doi:10.1038 / ng.3335, published online on June 8, 2015; and Guo et al., Nat. Commun. 9:1520 (2018). CTCF binding results in recognizable patterns in cfDNA, which can be detected by sequencing, for example, by fragment length analysis. For example, details on sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and U.S. Patent Application Publication No. 20170211143A1, each of which is incorporated herein by reference.
[0421] As a result, perturbations to CTCF binding lead to variations in the fragmentation pattern of cfDNA. Therefore, CTCF binding sites represent a certain type of fragmentation-variable target region.
[0422] There are many known CTCF binding sites. For example, see CTCFBSDB (CTCF Binding Site Database), available on the internet at insulatordb.uthsc.edu / ; see Cuddapah et al., Genome Res. 19:24-32 (2009); Martin et al., Nat. Struct. Mol. Biol. 18:708-14 (2011); and see Rhee et al., Cell. 147:1408-19 (2011), each of which is incorporated herein by reference. Exemplary CTCF binding sites are located, for example, on nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13, according to the hg19 or hg38 human genome construct.
[0423] Therefore, in some embodiments, the epigenetic target region set includes CTCF-binding regions. In some embodiments, the CTCF-binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF-binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF-binding regions, such as the CTCF-binding regions mentioned above, or CTCF-binding regions described in one or more of the CTCFBSDB cited above or in the papers by Cuddapah et al., Martin et al., or Rhee et al.
[0424] In some embodiments, at least a portion of the CTCF site may or may not be methylated, and this methylation status correlates with whether or not the cell is a cancer cell. In some embodiments, the epigenetic target region set includes at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, and at least 1000 bp upstream and / or downstream regions of the CTCF binding site. (D) Transcription start site
[0425] Transcription start sites can also undergo perturbations in neoplastic cells. For example, the nucleosome composition at various transcription start sites in healthy hematopoietic cells—which significantly contributes to cfDNA in healthy individuals—may differ from that at the same sites in neoplastic cells. This can result in different cfDNA patterns that can be detected by sequencing, as commonly discussed, for example, in Snyder et al., Cell 164:57-68 (2016);WO2018 / 009723; and U.S. Patent Application Publication No. 20170211143A1.
[0426] As a result, perturbations at the transcription start site also lead to variations in the fragmentation pattern of cfDNA. Therefore, the transcription start site also represents a certain type of fragmentation-variable target region.
[0427] Human transcription start sites are available from the DBTSS (Database of Human Transciption Start Sites), which can be accessed via the internet at dbtss.hgc.jp, and are described in Yamashita et al., Nucleic Acids Res. 34 (Database issue): D86-D89 (2006), which is incorporated herein by reference.
[0428] Therefore, in some embodiments, the epigenetic target region set includes transcription start sites. In some embodiments, the transcription start sites include at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those listed in DBTSS. In some embodiments, at least a portion of the transcription start sites may or may not be methylated, and this methylation status correlates with whether or not the cell is a cancer cell. In some embodiments, the epigenetic target region set includes at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, and at least 1000 bp upstream and / or downstream regions of the transcription start site. (E) Methylation control region
[0429] Including control regions can be useful to facilitate data validation. In some embodiments, the epigenetic target region set includes a control region that is expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA originates from cancer cells or normal cells. In some embodiments, the epigenetic target region set includes a control hypomethylated region that is expected to be hypomethylated in essentially all samples. In some embodiments, the epigenetic target region set includes a control hypermethylated region that is expected to be hypermethylated in essentially all samples. (F) Copy number diversity; local amplification
[0430] Copy number diversity, such as local amplification, is a somatic mutation, but it can be detected based on read frequency by sequencing in a manner similar to that used to detect certain epigenetic changes, such as changes in methylation. For this reason, regions that may exhibit copy number diversity, such as local amplification, in cancer can be included in an epigenetic target region set, and these regions may include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, the epigenetic target region set includes at least two, three, four, five, six, seven, eight, nine, ten, 11, 12, 13, 14, 15, 16, 17, or 18 of the aforementioned targets. vi. Sequence Analysis Pipeline
[0431] In one embodiment, after sequencing, sequence reads and any associated data can be stored in a sequence data store 209. Sequence reads can be stored in any format. The sequence data store 209 can be local and / or remote from the location where sequencing is performed. As shown in Figure 2, the stored reads can be supplied to a sequence analysis pipeline 212. a. Sequence alignment
[0432] The sequence analysis pipeline 230 may include an alignment component 236 configured to align sequence fragments / reads from the laboratory system 102 to identify regions of similarity and to place the sequences in the sequence data store 209. Similarity may relate to functional, structural, and / or evolutionary relationships between sequences. For DNA sequences, the alignment by the alignment component 236 may include the alignment of the genomic DNA of one sequence to the genomic DNA of at least one other sequence. Such alignments may exclude non-genomic DNA such as molecular barcodes, padding bases, and others. For example, the genomic DNA of a sequence read may be aligned to the genomic DNA of a reference DNA sequence, excluding any molecular tags that may be bound to the sequence read. b. Sequence quality control
[0433] The sequence analysis pipeline 230 may include a sequence quality control (QC) component 231 that can filter sequence fragments / reads from the laboratory system 102. The sequence QC component 231 can assign a quality score to one or more sequence fragments / reads. The quality score may be an indication of the sequence fragments / reads, indicating whether those sequence fragments / reads may be useful for subsequent analysis based on a threshold. In some cases, some sequence fragments / reads are neither of sufficient quality nor length to perform subsequent mapping steps. Sequence fragments / reads with a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the sequence fragment / read dataset. In other cases, sequence fragments / reads that have been assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the dataset.
[0434] Sequence fragments / reads that meet a specific quality score threshold can be mapped to a reference genome by the sequence QC component 231. After mapping alignment, a mapping score can be assigned to the sequence fragments / reads. The mapping score may be an indication of the remapped sequence fragments / reads to the reference sequence, showing whether each position is uniquely mappable. Sequence fragments / reads with a mapping score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the dataset. In other cases, sequencing fragments / reads assigned a mapping score of less than 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the dataset. c. Epigenetic factors
[0435] Epigenetic factors that may be used in the systems and methods described herein are disclosed throughout.
[0436] In some embodiments, the epigenetic component 232 can perform sequence fragment / read analysis to determine epigenetic data. Epigenetic data may include, for example, information about DNA methylation, histone state or modification, inflammation-mediated cytosine damage products, protein binding, fragment mix (fragment size, nucleotide motifs at fragment ends, single-strand jagged ends, and / or genomic location of fragment endpoints), or the state of other molecules reflected in the analyzed nucleic acid fragment that cannot be determined solely from the nucleotide sequence. Epigenetic data can be used as an epigenetic signature. Epigenetic data can be determined by any means known in the art. Epigenetic data can be based on methylation data determined by the LR methylation component 233 and fragment mix data determined by the fragment mix component 234. In some examples, epigenetic data can also be based on fragment mix data determined via methylation data from the TFR methylation component 235. Epigenetic data can be stored in the analysis data store 240. (A) Methylation status
[0437] As described herein, methylation status can be determined in cfDNA and can be used to determine whether a sample is tumor-derived. Generally, cfDNA can be separated into methylated and unmethylated partitions based on the overall methylation status of each molecule. cfDNA can be partitioned based on the differential binding affinity of methylated nucleic acid molecules to a binder (i.e., a binder that binds to methylated nucleotides). In some embodiments, bisulfite conversion is not used. Next, the DNA in each partition can be tagged with a separate set of dual barcodes that uniquely identify the partition associated with the entire molecule and help in the identification of unique cfDNA molecules after sequencing. Next, the DNA molecules in the methylated partition can be treated with restriction enzymes to deplete samples of partially methylated molecules. Next, all partitions can be PCR amplified and enriched by hybridization to oligonucleotides representing the target genomic region targeting approximately 1 Mb of the human genome. The enriched partitions can be pooled and tagged with an index that uniquely identifies each sample before pooling multiple enriched samples into a sequencing pool. The sequencing pool was sequenced using a NovaSeq 6000 instrument. In addition, cfDNA fragments from sample 201 and / or subject 211 can be treated in the sample collection and preparation pipeline 203, for example, by converting unmethylated cytosine to uracil, and then sequenced according to the sequencing pipeline 205.
[0438] In accordance with this specification, sequence fragments / reads can be compared to a reference genome using LR methylation component 233 and / or tumor proportion regression (TFR) methylation component 235 to identify the methylation status at specific CpG sites within the sequence fragments / reads. Each CpG site may or may not be methylated. Identifying fragments that are abnormally methylated compared to healthy individuals provides clues to understanding the cancer status of a subject. Abnormal DNA methylation (compared to healthy controls) can cause a variety of effects that may contribute to cancer. Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5-methylcytosine. In particular, methylation tends to occur at cytosine and guanine dinucleotides, which are referred to herein as “CpG sites.” Abnormal DNA methylation can be identified as hypermethylation or hypomethylation, both of which may indicate cancer status. Throughout this disclosure, hypermethylation and hypomethylation can be characterized for a sequence fragment / read if the sequence fragment / read contains more CpG sites methylated than a threshold percentage or more unmethylated CpG sites than a threshold percentage. Examples of thresholds for the number of CpG sites include values greater than 3, 4, 5, 6, 7, 8, 9, 10, etc. Examples of percentage thresholds for methylation or unmethylation include percentages greater than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%. Those skilled in the art will understand that the principles described herein are similarly applicable to the detection of methylation in non-CpG contexts, including non-cytosine methylation.
[0439] In one embodiment, the LR methylation component 233 and / or the TFR methylation component 235 may be configured to determine the location and methylation state for each CpG site based on alignment to a reference genome. The LR methylation component 233 and / or the TFR methylation component 235 can generate a methylation state vector for each fragment, specifying the location of the fragment in the reference genome (e.g., the location of the first CpG site in each fragment, or as identified by another similar metric), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, whether it is methylated (e.g., indicated as M), unmethylated (e.g., indicated as U), or undetermined (e.g., indicated as I). The observed state is either methylated or unmethylated, while the unobserved state is undetermined. Undetermined methylation states may arise from sequencing errors and / or mismatches between the methylation states of the complementary strands of the DNA fragment. The methylation state vectors can be stored in the analytical data store 240 for subsequent use and processing. Furthermore, the LR methylation component 233 and / or the TFR methylation component 235 can remove duplicate reads or duplicate methylation state vectors from a single sample. The LR methylation component 233 and / or the TFR methylation component 235 can determine that certain fragments having one or more CpG sites have an undetermined methylation status exceeding a threshold number or percentage, and such fragments can be excluded.
[0440] Figure 3A is an explanatory diagram of method 300 for sequencing a cfDNA molecule to obtain a methylation state vector. Method 300 may involve single-site methylation. As an example, laboratory system 202 receives cfDNA molecule 301, which in this example contains three CpG sites. As shown, the first and third CpG sites of cfDNA molecule 301 are methylated 302. As part of a sample collection and preparation pipeline 203, cfDNA molecule 301 is converted to produce a converted cfDNA molecule 303. The second CpG site, which was not methylated, has its cytosine converted to uracil, while the first and third CpG sites were not converted.
[0441] In one or more examples, methylated cytosine can be determined using at least one of the following: sodium bisulfite conversion and sequencing, Tet-assisted bisulfite sequencing (TAB-Seq), differential enzymatic cleavage, treatment with MSRE and / or MDRE, MBD partitioning, ACE-Seq, Ox-BS, Tet-assisted pyridineborane sequencing (TAPS); EM-Seq; SEM-Seq, DM-Seq, and TrueMethyl oxidative bisulfite sequencing.
[0442] After conversion, sequencing pipeline 205 is used to generate sequence fragments / reads 304. LR methylation components 233 and / or TFR methylation components 235 may be configured to align sequence fragments / reads 304 to a reference genome 305. The reference genome 305 provides context about the location within the human genome where the fragment cfDNA originates. In this simplified example, LR methylation components 233 and / or TFR methylation components 235 align sequence reads 304 so that three CpG sites correlate with CpG sites 1, 2, and 3. Thus, LR methylation components 233 and / or TFR methylation components 235 can generate information about both the methylation status of all CpG sites on the cfDNA molecule 301 and the location within the human genome where the CpG sites are located. As shown, the CpG sites on the methylated sequence reads 304 are read as cytosine. In this example, cytosine appears only at the first and third CpG sites of sequence read 304, from which it can be inferred that the first and third CpG sites in the original cfDNA molecule were methylated. On the other hand, the second CpG site is read as thymine (U is converted to T during the sequencing process), and therefore it can be inferred that the second CpG site in the original cfDNA molecule was not methylated. Using these two pieces of information, methylation status and location, the LR methylation component 233 and / or TFR methylation component 235 can generate a methylation state vector 306 for fragment cfDNA 301. In this example, the resulting methylation state vector 306 is:<M1、U2、M3> Here, M corresponds to a methylated CpG site, U corresponds to a non-methylated CpG site, and the subscript corresponds to the location of each CpG site in the reference genome.
[0443] In another embodiment, after sequencing and alignment, the methylation status of individual CpG sites can be inferred from the counts of methylated sequence reads "M" (methylated) and unmethylated sequence reads "U" (unmethylated) at cytosine residues in the CpG context. The mean methylated CpG density (also called methylation density m) at a particular locus in plasma can be calculated using the equation: m = M / (M + U), where M is the count of methylated reads at the CpG site within the gene locus, and U is the count of unmethylated reads at the CpG site within the gene locus. If there is more than one CpG site within a locus, M and U correspond to the counts across that site.
[0444] In addition to sequencing, other techniques can be used to determine information regarding DNA methylation. In one embodiment, methylation profiling can be performed by methylation-specific PCR, PCR following methylation-sensitive restriction enzyme digestion, or PCR following ligase chain reaction. In yet another embodiment, PCR may be in the form of single-molecule or digital PCR (B. Vogelstein et al. 1999 Proc Natl Acad Sci USA; 96: 9236-9241). In yet another embodiment, PCR may be real-time PCR. In yet another embodiment, PCR may be multiplex PCR.
[0445] Using methylation status and location, the TFR methylation component 235 can be used with a TFR model to quantify the proportion of tumor-derived cfDNA in a sample (e.g., tumor proportion) based on the quantification of observed tumor-associated abnormal methylation of cfDNA molecules. This quantification can be based on the observed number of unique methylated molecules that map to each of the targeted classification regions. These molecular counts are normalized to the total number of unique methylated molecules observed in the normalized region of the panel. After normalization, the dependence of the classification region feature values (normalized molecular counts) on the total number of measured molecules and the input cfDNA amount for the sample is minimized. Region-level normalized molecular counts can be used as input features to the TFR model. The predicted tumor proportion can be used as a TFR model score for evaluating the cancer status of individual samples.
[0446] Figure 3B is a diagrammatic representation of an example environment 307, identifying nucleic acids corresponding to the classification region of a reference sequence, where the classification region has at least a threshold number of CpGs according to one or more implementations. In one or more examples, the disease under consideration is a certain type of cancer.
[0447] The environment 307 may include a sample 308. The sample 308 may be derived from a biological fluid obtained from a subject. For example, the sample 308 may be derived from blood obtained from a subject. In one or more further examples, the sample 308 may be derived from the tissue of the subject. In various examples, the sample 308 may be derived from multiple sources. For illustrative purposes, the sample 308 may be derived from one or more bodily fluids and / or tissues of the subject. In one or more exemplary examples, the subject may be a mammal. In one or more further exemplary examples, the subject may be a human. In one or more exemplary examples, the subject may be a non-human mammal.
[0448] Sample 308 may contain several nucleic acids 309. Each nucleic acid 309 may contain several regions having at least a threshold number of cytosine and guanine molecules. In one or more examples, each nucleic acid 309 may contain a region having at least a threshold number of cytosine-guanine dinucleotides. In various examples, at least a portion of the cytosine-guanine pairs contained in a region may be sequentially positioned in the sequence of nucleic acid 309. In one or more exemplary examples, a region of nucleic acid having at least a threshold number of cytosine-guanine pairs may be referred to herein as a "CG region" or "CpG region". In one or more examples, a CG region may contain at least 200 CpG dinucleotides. In one or more exemplary examples, a CG region may contain 200 to 5000 CpG dinucleotides, 300 to 3000 CpG dinucleotides, 200 to 2500 CpG dinucleotides, or 500 to 1500 CpG dinucleotides. In addition, a CG region may have at least a 50% GC percentage and at least a 60% observed-to-expected CpG ratio. The observed-to-expected CpG ratio can be calculated, where observed CpG is the number of CpGs identified in a given genomic region, and expected CpG is the number of cytosines multiplied by the number of guanines divided by the number of bases in the genomic region. Expected CpG can also be calculated by the following formula: ((number of cytosine + number of guanine) / 2)² / length of genome region
[0449] For example, CG regions can be determined using the techniques described in Gardiner-Garden M, Frommer M (1987). "CpG islands in vertebrate genomes". Journal of Molecular Biology. 196 (2): 261-282., and / or Saxonov S, Berg P, Brutlag DL (2006). "A genome-wide analysis of CpG dinucleotides in the human genome distinguishes two distinct classes of promoters". Proc Natl Acad Sci USA. 103 (5): 1412-1417.
[0450] In the exemplary example shown in Figure 3B, a portion of the sequence of the example nucleic acid 309 may include a first CG region 310, a second CG region 311, and a third CG region 312. While the exemplary example in Figure 3B illustrates a portion of the sequence of nucleic acid 309 having three CG regions, the nucleic acid 309 contained in sample 308 may have a different number of CG regions. For example, each individual nucleic acid 309 contained in sample 308 may contain at least one CG region, at least five CG regions, at least ten CG regions, at least 25 CG regions, at least 50 CG regions, at least 100 CG regions, at least 250 CG regions, at least 500 CG regions, or at least 1000 CG regions.
[0451] Individual CG regions can correspond to several molecules having one or more methylated cytosines. In the exemplary example in Figure 3B, CG region 310 may include a molecule having methylated cytosine 313. In the exemplary example in Figure 3B, the molecule having methylated cytosine 313 is 5-methylcytosine. Individual CG regions can also correspond to several molecules having unmethylated cytosines. For example, CG region 310 may include a molecule having unmethylated cytosine 316. In various examples, at least a portion of the CG region of nucleic acid 309 can correspond to a taxonomic region of a reference genome. The taxonomic region can correspond to a genomic region of the reference genome that corresponds to a non-sequence difference consistent with one or more biological conditions, such as one or more types of cancer. In at least some examples, the non-sequence difference may include one or more mutations consistent with one or more biological conditions. In one or more examples, the taxonomic region can correspond to a genomic region of a reference sequence from which the molecule originates from a subject having at least one form of cancer. In at least some examples, nucleic acid molecules (e.g., highly methylated molecules) having at least a threshold amount of methylated cytosine in at least one CG region may originate from subjects where cancer is present and correspond to a classification region.
[0452] In addition to the classification region, the CG region may include one or more positive control regions, for example, positive control region 318. The positive control region 311 has at least a threshold number of methylated cytosine molecules in at least one CG region and can be mapped to nucleic acid molecules from both cancer-free and cancer-present subjects. In various examples, the positive control region 310 may be hypermethylated in cells from cancer-free subjects and in cells from cancer-present subjects. The CG region may also include one or more negative control regions, for example, negative control region 320. The negative control region 320 has fewer than a threshold number of methylated cytosine molecules in at least one CG region and can be mapped to nucleic acid molecules from both cancer-free and cancer-present subjects. In one or more exemplary examples, the negative control region 320 may be hypomethylated in both cancer-free and cancer-present subjects. In various examples, the positive and negative control regions can be used to perform normalization calculations. Normalization calculations can be performed to generate input data for one or more models to be executed to determine tumor metrics for a given sample 308.
[0453] A first molecular separation process 322 can be performed. The first molecular separation process 322 can separate nucleic acids 309 contained in sample 308 based on the amount of methylated cytosine in each nucleic acid 309. In one or more examples, the first molecular separation process can separate nucleic acids 309 contained in sample 308 based on the amount of methylated cytosine contained in the CG region of each nucleic acid 309. In various examples, the first molecular separation process 322 can separate nucleic acids 309 into multiple groups, each group corresponding to the respective amount of methylated cytosine in nucleic acids 309.
[0454] In the exemplary example shown in Figure 3B, the first molecular separation process 322 can be performed in relation to a first methylation threshold 324. Performing the first molecular separation process 322 with respect to the first methylation threshold 324 can produce a first partition 326 of nucleic acid. In one or more examples, the first methylation threshold 324 may represent a first threshold number of molecules having methylated cytosine located in the CG region of nucleic acid 309. The first molecular separation process 322 can identify several nucleic acids 309 having molecules with methylated cytosine in the CG region that are less than the first methylation threshold 324. In various examples, the first methylation threshold 324 may correspond to a first methylation rate.
[0455] The first molecular separation process 322 can also be performed with respect to a second methylation threshold 328. The second methylation threshold 328 may represent the amount of methylated cytosine in one or more genomic regions of nucleic acid 309 that is greater than the amount of methylated cytosine in one or more regions corresponding to the first methylation threshold 324. The second methylation threshold 324 may represent the number of molecules having methylated cytosine per number of nucleic acids. In one or more additional examples, the second methylation threshold 324 may correspond to a methylation rate of nucleic acid that is greater than the methylation rate corresponding to the first methylation threshold 324. Performing the first molecular separation process 322 with respect to the second methylation threshold 328 can produce a second partition 330 of the nucleic acid. In one or more examples, the first molecular separation process 322 can identify nucleic acids 309 having a greater amount of methylated cytosine than a first methylation threshold 324 and a less amount of methylated cytosine than a second methylation threshold 328, thereby producing a second partition 330 of the nucleic acid.
[0456] In addition, the first molecular separation process 322 can also be performed with respect to a third methylation threshold 332. The third methylation threshold 332 can represent the amount of methylated cytosine in one or more genomic regions of nucleic acid 309 that is greater than the amount of methylated cytosine in one or more regions corresponding to the first methylation threshold 324 and greater than the amount of methylated cytosine in one or more regions corresponding to the second methylation threshold 328. The third methylation threshold 332 can represent the number of molecules having methylated cytosine per number of nucleic acids. In one or more additional examples, the third methylation threshold 332 can correspond to a methylated cytosine rate that is greater than the methylation rate corresponding to the first methylation threshold 324 and greater than the methylation rate corresponding to the second methylation threshold 328. Performing the first molecular separation process 322 with respect to the third methylation threshold 332 can produce a third partition 334 of the nucleic acid. In one or more examples, the first molecular separation process 322 can identify nucleic acid 309 having a greater amount of methylated cytosine than nucleic acid 309 contained in the second partition 328 of the nucleic acid. In this way, the amount of methylated cytosine in the nucleic acids contained in the first partition 322, the second partition 326, and the third partition 330 increases from the first partition 322 to the second partition 326, and from the second partition 326 to the third partition 330. In one or more exemplary examples, the first partition 326 of the nucleic acid may be referred to as the low-methylated partition, the second partition 330 of the nucleic acid may be referred to as the intermediate partition, and the third partition 334 of the nucleic acid may be referred to as the high-methylated partition.
[0457] In one or more examples, the amount of methylated cytosine in the nucleic acid can correspond to the strength of binding to the methyl-binding domain (MBD). In these scenarios, the first partition 326, the second partition 330, and the third partition 334 can be produced based on different strengths of binding to the MBD for nucleotides having different amounts of methylated cytosine. In one or more examples, the first molecular separation process 322 may include a series of washes in which the nucleic acid 309 is contacted with solutions having different concentrations of sodium chloride (NaCl).
[0458] Nucleic acid partitioning can be performed by contacting the nucleic acid with a modified nucleotide-specific binding agent, such as MBD from MBP. The modified nucleotide-specific binding agent can bind to 5-methylcytosine (5mC). Modified nucleotide-specific binding agents, such as MBD, may be linked to paramagnetic beads, such as Dynabeads® M-280 streptavidin, via a biotin linker. Partitioning into fractions with different degrees of methylation can be performed by increasing the NaCl concentration in a series of washes. Sequences eluted from the modified nucleotide-specific binding agent are partitioned into two or more fractions (e.g., low, high), depending on which wash (e.g., NaCl concentration) eluted the sequence. The resulting partitions may contain one or more of the following nucleic acid forms: double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments.
[0459] The binding of modified nucleotide-specific binding reagents to nucleic acids can be a function of the number of methylation (or modification) sites per molecule, with molecules having more methylation that elute under increasing salt concentrations. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into separate populations based on the degree of methylation. Salt concentrations can range from approximately 100 nM to approximately 2500 mM NaCl in one or more implementations. In various implementations, the process results in three partitions. Molecules are contacted with the solution at the first salt concentration and include molecules containing methyl-binding domains, which can be bound to a capture site such as streptavidin. At the first salt concentration, some populations of molecules will bind to the MBD, while others will remain unbound. The unbound populations can be separated as a “low-methylated” population (low partition). For example, the first partition 326 can represent the low-methylated form of DNA in that it remains unbound at low salt concentrations. In one or more exemplary examples, the concentration of NaCl in the solution used to produce the first partition 326 may be about 100 nM, about 120 nM, about 140 nM, about 160 nM, about 180 nM, about 200 nM, or about 250 nM. The second partition 330 may be referred to as the “residue partition” or “intermediate partition” and may represent intermediate methylated DNA eluted using an intermediate salt concentration between 100 mM and 2000 mM, for example. In one or more additional exemplary examples, the concentration of NaCl in the solution used to produce the second partition 330 may be about 100 mM to about 500 mM, about 100 mM to about 1000 mM, about 100 mM to about 1500 mM, about 250 mM to about 1000 mM, about 250 mM to about 1500 mM, about 500 mM to about 1500 mM, about 250 mM to about 2000 mM, about 500 mM to about 2000 mM, or about 1000 mM to about 2000 mM. This is also separated from the sample. The third partition 334 may represent a hypermethylated form of DNA (hyperpartition) and is eluted using a high salt concentration, e.g., at least about 2000 mM.In one or more further exemplary examples, the concentration of NaCl in the solution used to produce the third partition 334 may be about 2000 mM to about 5000 mM, about 2000 mM to about 4000 mM, about 2000 mM to about 3500 mM, about 2000 mM to about 3000 mM, or about 2500 mM to about 4000 mM.
[0460] In various examples, the first partition 326 may correspond to a first range of binding strength of nucleic acids to MBD and a first range of methylated CG regions, and the second partition 330 may correspond to a second range of binding strength of nucleic acids to MBD and a second range of methylated CG regions. The first range of binding strength may be less than the second range of binding strength. In one or more scenarios, a first solution having a first NaCl concentration can separate a first group of nucleic acids having a first range of binding strength from MBD, and a second solution having a second NaCl concentration can separate a second group of nucleic acids having a second range of binding strength from MBD, where the second NaCl concentration is greater than the first NaCl concentration. In addition, the third partition 334 may correspond to a third range of binding strength and a third range of methylated CG regions. The third range of binding strength may be greater than the first range of binding strength and the second range of binding strength. In one or more examples, a third solution having a third NaCl concentration can separate a third group of nucleic acids having a third range of binding strength from NaCl. The third NaCl concentration may be greater than the first and second NaCl concentrations.
[0461] In one or more exemplary examples, a nucleic acid-MBD solution can be produced by combining a plurality of nucleic acids derived from at least one of the blood or tissues of interest with a solution containing a certain amount of MBD. A first wash of the nucleic acid-MBD solution can be performed using a first solution containing a first NaCl concentration to produce a first nucleic acid fraction and a first residue solution. The first nucleic acid fraction may contain a first portion of the plurality of nucleic acids, and the first residue solution may contain a second portion of the plurality of nucleic acids. In one or more examples, the first portion of the plurality of nucleic acids may have a first range of binding strength to MBD that is less than a second range of binding strength to MBD of the second portion of the plurality of nucleic acids.
[0462] In addition, a second wash of the first residue solution can be performed using a second solution containing a second NaCl concentration greater than the first NaCl concentration required to produce a second nucleic acid fraction and a second residue solution. The second nucleic acid fraction may contain a first subset of the second parts of multiple nucleic acids, and the second residue solution may contain a second subset of the second parts of multiple nucleic acids. The first subset of the second parts of multiple nucleic acids may have a third range of binding strength to MBD, which is less than a fourth range of binding strength to MBD for the second subset of the second parts of multiple nucleic acids. Furthermore, a third wash of the second residue solution can be performed using a third solution containing a third NaCl concentration greater than the second NaCl concentration required to produce a third nucleic acid fraction containing a second subset of the second parts of multiple nucleic acids.
[0463] After the first, second, and third washes, it can be determined that a first portion of a plurality of nucleic acids is associated with a first partition 326. The first portion of the plurality of nucleic acids can be coupled to molecular barcodes derived from a first set of molecular barcodes indicating the first partition 326. In this way, a sequencing read corresponding to the first partition 326 can be identified based on the determination that the sequencing read contains the first molecular barcode. In addition, it can be determined that a first subset of a second portion of a plurality of nucleic acids is associated with an additional partition of the plurality of partitions. In these situations, a second set of molecular barcodes, different from the first set of molecular barcodes, can be coupled to a second portion of the plurality of nucleic acids having a second molecular barcode indicating the additional partition. As a result, a sequencing read corresponding to the additional partition can be identified based on the determination that the sequencing read contains one or more molecular barcodes from the second set of molecular barcodes. Furthermore, it can be determined that a second subset of a second portion of a plurality of nucleic acids is associated with a second partition 330. Next, a third set of molecular barcodes, distinct from the first and second sets of molecular barcodes, can be combined with a second subset of the second portion of multiple nucleic acids, where the third set of molecular barcodes represents the second partition 330. In these cases, the sequencing read corresponding to the second partition 330 can be identified based on the determination that the sequencing read contains a third molecular barcode from the third set of molecular barcodes.
[0464] In at least some examples, the first molecular separation process 322 may result in nucleic acids present in at least one of the first partition 326, the second partition 330, or the third partition 334 having a different amount of methylation than the amount of methylation of other nucleic acids in each partition. For example, the first partition 326 may contain a number of nucleic acids having an amount of methylation corresponding to the amount of methylation of nucleic acids contained in at least one of the second partition 330 or the third partition 334. In addition, at least one of the second partition 330 or the third partition 334 may contain nucleic acids having an amount of methylation corresponding to the amount of methylation of nucleic acids contained in the first partition 326. The presence of nucleic acids in at least one of the first partition 326, the second partition 330, or the third partition 334 that does not correspond to at least the majority of the amount of methylation of other nucleic acids contained in each partition may cause data noise when performing computer operations on sequence reads produced from the nucleic acids contained in the first partition 326, the second partition 330, and the third partition 334. Data noise can lead to inaccuracies in calculations based on sequence reads derived from nucleic acids contained in the first partition 326, the second partition 330, and the third partition 334.
[0465] A second molecular separation process 336 can be performed after the first molecular separation process 322 to reduce or eliminate data noise associated with nucleic acids present in at least one of the first partition 326, the second partition 330, or the third partition 334, which have an amount of methylation inconsistent with the amount of methylation of at least the majority of other molecules present in each partition. The second molecular separation process 336 can be performed with respect to nucleic acids present in the first partition 326, the second partition 330, and the third partition 334. In one or more examples, the second molecular separation process 336 may include digesting the nucleic acids present in the first partition 326 using methylation-dependent restriction enzymes (MDREs), while the nucleic acids present in the second partition 330 and the third partition 334 may be digested using methylation-sensitive restriction enzymes (MSREs). Digestion of nucleic acids contained in the first partition 326 by MDRE can result in the separation of nucleic acids contained in the first partition having the amounts of methylation corresponding to the second partition 330 and the third partition 334 from nucleic acids having the amounts of methylation corresponding to the first partition. In addition, digestion of nucleic acids contained in the second partition 330 and the third partition 334 by MDRE can result in the separation of nucleic acids having the amount of methylation corresponding to the first partition 326 from nucleic acids having the amounts of methylation corresponding to the nucleic acids in the second partition 330 and the third partition 334. An additional group of nucleic acids 338 can be produced by removing nucleic acids from the first partition 326 having the amounts of methylation corresponding to the second partition 330 and the third partition 334, and by removing nucleic acids from the second partition 330 and the third partition 334 having the amounts of methylation corresponding to the first partition 326.The additional group 338 of nucleic acids may include nucleic acids corresponding to the amounts of methylation in the second partition 330 and the third partition 334, with a minimum amount or none at all of nucleic acids having the amount of methylation corresponding to the first partition 326. For example, less than 50% of the nucleic acids in the additional group 338 may have the amounts of methylation corresponding to the second partition 330 and the third partition 334, or at least 50% of the nucleic acids in the additional group 338 may have the amounts of methylation corresponding to the second partition 330 and the third partition 334, or at least 60% of the nucleic acids in the additional group 338 may have the amounts of methylation corresponding to the second partition 330 and the third partition 334, or at least 70% of the nucleic acids in the additional group 338 may have the amounts of methylation corresponding to the second partition 330 and the third partition 334, or at least 90% of the nucleic acids in the additional group 338 may have the amounts of methylation corresponding to the second partition 330 and the third partition 334, or additional At least 95% of the nucleic acids in group 338 may have the amount of methylation corresponding to the second partition 330 and the third partition 334, or at least 97% of the nucleic acids in the additional group 338 may have the amount of methylation corresponding to the second partition 330 and the third partition 334, or at least 99% of the nucleic acids in the additional group 338 may have the amount of methylation corresponding to the second partition 330 and the third partition 334, or at least 99.5% of the nucleic acids in the additional group 338 may have the amount of methylation corresponding to the second partition 330 and the third partition 334, or at least 99.9% of the nucleic acids in the additional group 338 may have the amount of methylation corresponding to the second partition 330 and the third partition 334.
[0466] Architecture 307 may include a sequencing machine 340. In one or more examples, the sequencing machine 340 may be one of several sequencing machines capable of performing one or more sequencing operations to amplify nucleic acids present in the sample 309. In various examples, the sequencing machine 340 may perform next-generation sequencing operations. In one or more examples, the sample 309 may include a volume of at least one body fluid extracted from the subject. In one or more additional examples, the sample 309 may include a tissue sample obtained from the subject.
[0467] In one or more examples, prior to sequencing, the extracted polynucleotides may be partitioned into two or more partitions based on the binding strength of the polynucleotides to the MBD. Blunt-end ligation can be performed on the partitioned polynucleotides and adapters, and tags (e.g., molecular barcodes) can be attached to the partitioned polynucleotides. Tagged polynucleotides in one or more partitions (e.g., high and / or intermediate partitions) can be treated with one or more methylation-sensitive restriction enzymes (MSREs). In some examples, the low partition can be treated with one or more methylation-dependent restriction enzymes (MDREs). After MSRE and / or MORE treatment, the molecules can also be enriched by inducing hybridization between the extracted polynucleotides and probes corresponding to target regions of the reference sequence. The enrichment process can identify thousands, hundreds of thousands, or up to millions of polynucleotides corresponding to on-target regions associated with the probes.
[0468] After and / or before the enrichment process, the molecules may be amplified according to one or more amplification processes. One or more amplification processes can produce thousands to millions of copies of individual nucleic acid molecules. In one or more cases, a portion of the unenriched polynucleotide may be amplified, but in some cases, it may not be amplified to the extent that the enriched polynucleotide is amplified. One or more amplification processes can produce amplified products that undergo one or more sequencing operations. After one or more sequencing operations are performed on sample 309, the sequencing machine can generate sequencing data 342.
[0469] The sequencing data 342 may include alphanumeric representations of the nucleic acids contained in the amplification product. For example, for each nucleic acid in the amplification product, the sequencing data 342 may include data corresponding to a string of characters representing each nucleotide chain corresponding to the individual nucleic acid.
[0470] The sequencing data 342 may be stored in one or more data files. For example, the sequencing data 342 may be stored in a FASTQ file which includes a text-based sequencing data file format that stores raw sequence data and quality scores. In one or more additional examples, the sequencing data 342 may be stored in a data file conforming to the Binary-Based Call (BCL) sequence file format. In one or more further examples, the sequencing data 342 may be stored in a BAM file. In one or more examples, the sequencing data 342 may include at least about 1 gigabyte (GB), at least about 2 GB, at least about 3 GB, at least about 4 GB, at least about 5 GB, at least about 8 GB, or at least about 10 GB. Individual sequence representations contained in the sequencing data 310 may be referred to herein as “reads” or “sequencing reads.” In various examples, individual first nucleic acids contained in pool 338 may correspond to multiple sequence representations contained in the sequencing data 342 as a result of amplification of the individual first nucleic acids. In one or more additional examples, each second nucleic acid contained in pool 338 can correspond to a single sequence representation contained in sequencing data 342 as a result of no amplification of the individual second nucleic acids.
[0471] Using methylation status and location as well, the LR methylation component 233 can include an LR model to differentiate tumor-associated methylation signatures of cfDNA molecules from those observed in non-tumor subjects. The methylation LR model can use the same input feature space as the TFR model of the TFR methylation component 235 (e.g., reg...
Claims
1. A step of determining a tumor proportion regression (TFR) score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples, wherein the TFR score includes the proportion of molecules in the several cell-free nucleic acid samples that indicate a tumor; A step of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of the plurality of cell-free nucleic acid samples; and A step in which a predictive model is used to determine whether the plurality of cell-free nucleic acid samples are tumor-derived or non-tumor-derived, based on at least one of the cell-free nucleic acid scores or TFR scores that satisfy the respective thresholds. A method that includes this.
2. The method according to claim 1, further comprising the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
3. The method according to claim 2, wherein the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of the plurality of genomic regions.
4. The method according to claim 1, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples.
5. The method according to claim 4, further comprising the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
6. The method according to claim 5, further comprising the step of using an LR model to determine whether the methylated LR model is cancerous or non-cancer.
7. The method according to claim 4, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
8. The method according to claim 7, further comprising the step of determining the cancer signal derived from the cell-free nucleic acid fragmentation pattern associated with the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples using a fragment mix model.
9. The method according to claim 4, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
10. The method according to claim 9, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
11. The method according to claim 1, further comprising the step of determining the genome modification of each of the plurality of cell-free nucleic acid samples.
12. The method according to claim 11, wherein the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples includes determining the somatic variant observed in the molecule derived from each of the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
13. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one of the following: a genomic region known to be associated with cancer type, a genomic region known to be associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapeutic response.
14. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one genomic region known to be associated with colorectal cancer.
15. The method according to claim 1, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
16. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
17. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
18. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples include a cell-free ribonucleic acid (cfRNA) sample.
19. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples include a mitochondrial deoxyribonucleic acid (mtDNA) sample.
20. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
21. The method according to claim 1, wherein the plurality of cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
22. A step of determining a tumor proportion regression (TFR) score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples, wherein the TFR score includes the proportion of molecules in the several cell-free nucleic acid samples that indicate a tumor; and Based on the TFR scores that satisfy each threshold, a predictive model is used to determine whether the plurality of cell-free nucleic acid samples are tumor-derived or non-tumor-derived. A method that includes this.
23. The method according to claim 22, further comprising the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
24. The method according to claim 23, wherein the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of the plurality of genomic regions.
25. The method according to claim 22, wherein the step of determining whether the plurality of cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on a cell-free nucleic acid score indicating the presence of a tumor.
26. The method according to claim 25, further comprising the step of determining the cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of the plurality of cell-free nucleic acid samples.
27. The method according to claim 26, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples.
28. The method according to claim 27, further comprising the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
29. The method according to claim 28, further comprising the step of using an LR model to determine whether the methylated LR model is cancerous or non-cancer.
30. The method according to claim 27, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
31. The method according to claim 30, further comprising the step of determining the cancer signal derived from the cell-free nucleic acid fragmentation pattern associated with the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples using a fragment mix model.
32. The method according to claim 26, further comprising the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples.
33. The method according to claim 32, wherein the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples includes determining the somatic variant observed in the molecule derived from each of the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
34. The method according to claim 26, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
35. The method according to claim 34, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
36. The method according to claim 22, wherein the plurality of genomic regions include at least one of genomic regions known to be associated with cancer type, genomic regions known to be associated with known methylation status, genomic regions known to be associated with hypomethylation, or genomic regions known to be associated with therapeutic response.
37. The method according to claim 22, wherein the plurality of genomic regions include at least one genomic region known to be associated with colorectal cancer.
38. The method according to claim 22, wherein the plurality of cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
39. The method according to claim 22, wherein the plurality of cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
40. The method according to claim 22, wherein the plurality of cell-free nucleic acid samples include a cell-free ribonucleic acid (cfRNA) sample.
41. The method according to claim 22, wherein the plurality of cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
42. The method according to claim 22, wherein the plurality of cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
43. The method according to claim 22, wherein the plurality of cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
44. A step of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of several cell-free nucleic acid samples; and A step in which a predictive model is used to determine whether the multiple cell-free nucleic acid samples are tumor-derived or non-tumor-derived, based on the cell-free nucleic acid scores that satisfy each threshold. A method that includes this.
45. The method according to claim 44, wherein the step of determining whether the plurality of cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on tumor proportion regression (TFR) scores that satisfy a threshold, wherein the TFR score indicates the proportion of molecules in the plurality of cell-free nucleic acid samples that indicate tumors.
46. The method according to claim 45, further comprising the step of determining the TFR score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
47. The method according to claim 46, further comprising the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
48. The method according to claim 47, wherein the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of the plurality of genomic regions.
49. The method according to claim 45, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
50. The method according to claim 44, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples.
51. The method according to claim 50, further comprising the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
52. The method according to claim 51, further comprising the step of using an LR model to determine whether the methylated LR model is cancerous or non-cancer.
53. The method according to claim 50, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
54. The method according to claim 53, further comprising the step of determining the cancer signal derived from the cell-free nucleic acid fragmentation pattern associated with the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples using a fragment mix model.
55. The method according to claim 50, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
56. The method according to claim 55, wherein the step of determining the cell-free nucleic acid score is further based on a tumor proportion regression (TFR) score, the TFR score indicating the proportion of molecules of the plurality of cell-free nucleic acid samples that represent tumors.
57. The method according to claim 44, further comprising the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples.
58. The method according to claim 57, wherein the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples includes determining the somatic variant observed in the molecule derived from each of the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
59. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one of the following: a genomic region known to be associated with cancer type, a genomic region known to be associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapeutic response.
60. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one genomic region known to be associated with colorectal cancer.
61. The method according to claim 44, wherein the step of determining whether the plurality of cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on methylation values that satisfy a threshold, and the methylation score indicates the number of molecules of the plurality of cell-free nucleic acid samples that indicate tumor.
62. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
63. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
64. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples include a cell-free ribonucleic acid (cfRNA) sample.
65. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples include a mitochondrial deoxyribonucleic acid (mtDNA) sample.
66. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
67. The method according to claim 44, wherein the plurality of cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
68. A step of determining a tumor proportion regression (TFR) score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples, wherein the TFR score comprises the proportion of molecules in the several cell-free nucleic acid samples that indicate tumor, and each of the several cell-free nucleic acid samples is labeled with a tumor-derived label or a non-tumor-derived label; A step of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of the plurality of cell-free nucleic acid samples; A step of determining a tumor prediction for each of the plurality of cell-free nucleic acid samples based on at least one of the cell-free nucleic acid scores or the TFR score that satisfies the respective thresholds; A step of generating a predictive model for predicting tumors in the plurality of cell-free nucleic acid samples based on the tumor-derived label or the non-tumor-derived label and the tumor prediction for each of the plurality of cell-free nucleic acid samples; and step of outputting the prediction model A method that includes this.
69. The method according to claim 68, further comprising the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
70. The method according to claim 69, wherein the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of the plurality of genomic regions.
71. The method according to claim 68, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples.
72. The method according to claim 71, further comprising the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
73. The method according to claim 72, further comprising the step of using an LR model to determine whether the methylated LR model is cancerous or non-cancer.
74. The method according to claim 71, wherein the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples is based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
75. The method according to claim 74, further comprising the step of determining the cancer signal derived from the cell-free nucleic acid fragmentation pattern associated with the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples using a fragment mix model.
76. The method according to claim 71, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
77. The method according to claim 76, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
78. The method according to claim 68, further comprising the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples.
79. The method according to claim 78, wherein the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples includes determining the somatic variant observed in the molecules derived from the plurality of cell-free nucleic acid samples.
80. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one of the following: a genomic region known to be associated with cancer type, a genomic region known to be associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapeutic response.
81. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one genomic region known to be associated with colorectal cancer.
82. The method according to claim 68, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
83. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
84. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
85. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples include a cell-free ribonucleic acid (cfRNA) sample.
86. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
87. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
88. The method according to claim 68, wherein the plurality of cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
89. A step of determining a tumor proportion regression (TFR) score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples, wherein the TFR score comprises the proportion of molecules in the several cell-free nucleic acid samples that indicate tumor, and each of the several cell-free nucleic acid samples is labeled with a tumor-derived label or a non-tumor-derived label; A step of determining a tumor prediction for each of the plurality of cell-free nucleic acid samples based on the TFR scores that satisfy each threshold; A step of generating a predictive model for predicting tumors in the plurality of cell-free nucleic acid samples based on the tumor-derived label or the non-tumor-derived label and the tumor prediction of the plurality of cell-free nucleic acid samples; and step of outputting the prediction model A method that includes this.
90. The method according to claim 89, further comprising the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
91. The method according to claim 90, wherein the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of the plurality of genomic regions.
92. The method according to claim 89, wherein the step of determining the tumor prediction for each of the plurality of cell-free nucleic acid samples is further based on a cell-free nucleic acid score indicating the presence of a tumor.
93. The method according to claim 92, further comprising the step of determining the cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of the plurality of cell-free nucleic acid samples.
94. The method according to claim 93, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples.
95. The method according to claim 94, further comprising the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
96. The method according to claim 95, further comprising the step of using an LR model to determine whether the methylated LR model is cancerous or non-cancer.
97. The method according to claim 94, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
98. The method according to claim 97, further comprising the step of determining the cancer signal derived from the cell-free nucleic acid fragmentation pattern associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples using a fragment mix model.
99. The method according to claim 93, further comprising the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples.
100. The method according to claim 99, wherein the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples includes determining the somatic variant observed in the molecule derived from each of the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
101. The method according to claim 93, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
102. The method according to claim 101, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
103. The method according to claim 89, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one of the following: a genomic region known to be associated with cancer type, a genomic region known to be associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapeutic response.
104. The method according to claim 89, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one genomic region known to be associated with colorectal cancer.
105. The method according to claim 89, wherein the plurality of cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
106. The method according to claim 88, wherein the plurality of cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
107. The method according to claim 89, wherein the plurality of cell-free nucleic acid samples include a cell-free ribonucleic acid (cfRNA) sample.
108. The method according to claim 89, wherein the plurality of cell-free nucleic acid samples include mitochondrial deoxyribonucleic acid (mtDNA) samples.
109. The method according to claim 89, wherein the plurality of cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
110. The method according to claim 89, wherein the plurality of cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
111. A step of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of a plurality of cell-free nucleic acid samples, wherein each of the plurality of cell-free nucleic acid samples is labeled with a tumor-derived label or a non-tumor-derived label; A step of determining a tumor prediction for each of the plurality of cell-free nucleic acid samples based on the cell-free nucleic acid scores that satisfy each threshold; A step of generating a predictive model for predicting tumors in the plurality of cell-free nucleic acid samples based on the tumor-derived label or the non-tumor-derived label and the tumor prediction for each of the plurality of cell-free nucleic acid samples; and step of outputting the prediction model A method that includes this.
112. The method according to claim 111, wherein the step of determining the tumor prediction for each of the plurality of cell-free nucleic acids is further based on a tumor proportion regression (TFR) score that satisfies a threshold, the TFR score indicating the proportion of molecules of the plurality of cell-free nucleic acid samples that indicate a tumor.
113. The method according to claim 112, further comprising the step of determining the TFR score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
114. The method according to claim 113, further comprising the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
115. The method according to claim 114, wherein the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of the plurality of genomic regions.
116. The method according to claim 112, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
117. The method according to claim 111, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples.
118. The method according to claim 117, further comprising the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
119. The method according to claim 118, further comprising the step of using an LR model to determine whether the methylated LR model is cancerous or non-cancer.
120. The method according to claim 117, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
121. The method according to claim 120, further comprising the step of determining the cancer signal derived from the cell-free nucleic acid fragmentation pattern associated with the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples using a fragment mix model.
122. The method according to claim 117, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
123. The method according to claim 122, wherein the step of determining the cell-free nucleic acid score is further based on a tumor proportion regression (TFR) score, the TFR score indicating the proportion of molecules of the plurality of cell-free nucleic acid samples that represent tumors.
124. The method according to claim 111, further comprising the step of determining the genome modification of each of the plurality of cell-free nucleic acid samples.
125. The method according to claim 124, wherein the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples includes determining the somatic variant observed in the molecule derived from each of the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
126. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one of the following: a genomic region known to be associated with cancer type, a genomic region known to be associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapeutic response.
127. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one genomic region known to be associated with colorectal cancer.
128. The method according to claim 111, wherein the step of determining whether the plurality of cell-free nucleic acid samples are tumor-derived or non-tumor-derived is further based on methylation values that satisfy a threshold, and the methylation score indicates the number of molecules of the plurality of cell-free nucleic acid samples that indicate tumor.
129. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
130. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
131. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples include a cell-free ribonucleic acid (cfRNA) sample.
132. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples include a mitochondrial deoxyribonucleic acid (mtDNA) sample.
133. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
134. The method according to claim 111, wherein the plurality of cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
135. A step of detecting one or more biomarkers in a biological sample; A step of determining a tumor proportion regression (TFR) score using a TFR model based on the quantification of observed tumor-associated abnormal methylation in each of several cell-free nucleic acid samples, wherein the TFR score includes the proportion of molecules in the several cell-free nucleic acid samples that indicate a tumor; A step of determining a cell-free nucleic acid score indicating the presence of a tumor based on at least one of the epigenetic factors or genomic alterations of each of the plurality of cell-free nucleic acid samples; and A step of determining whether the biological sample is tumor-derived or non-tumor-derived based on at least one of the detected biomarkers, the cell-free nucleic acid score, or the TFR score that meets the respective thresholds. A method that includes this.
136. The method according to claim 135, further comprising the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples.
137. The method according to claim 136, wherein the step of determining the quantification of the observed tumor-associated abnormal methylation in each of the plurality of cell-free nucleic acid samples includes quantifying the number of unique methylation molecules that map to each of the plurality of genomic regions.
138. The method according to claim 135, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples.
139. The method according to claim 138, further comprising the step of determining the epigenetic factors of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification.
140. The method according to claim 139, further comprising the step of using an LR model to determine whether the methylated LR model is cancerous or non-cancer.
141. The method according to claim 138, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
142. The method according to claim 141, further comprising the step of determining the cancer signal derived from the cell-free nucleic acid fragmentation pattern associated with the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples using a fragment mix model.
143. The method according to claim 138, further comprising the step of determining the epigenetic factor of each of the plurality of cell-free nucleic acid samples based on a methylated logistic regression (LR) model for cancer or non-cancer classification, and based on cancer signals derived from cell-free nucleic acid fragmentation patterns associated with a plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
144. The method according to claim 143, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
145. The method according to claim 135, further comprising the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples.
146. The method according to claim 145, wherein the step of determining the genomic modification of each of the plurality of cell-free nucleic acid samples includes determining the somatic variant observed in the molecule derived from each of the plurality of sequence fragments derived from the plurality of cell-free nucleic acid samples.
147. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one of the following: a genomic region known to be associated with cancer type, a genomic region known to be associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapeutic response.
148. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples are derived from a plurality of genomic regions, and the plurality of genomic regions include at least one genomic region known to be associated with colorectal cancer.
149. The method according to claim 135, wherein the step of determining the cell-free nucleic acid score is further based on the TFR score.
150. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples include cell-free deoxyribonucleic acid (cfDNA) samples.
151. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples include ribonucleic acid (RNA) samples.
152. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples include cell-free ribonucleic acid (cfRNA) samples.
153. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples include a mitochondrial deoxyribonucleic acid (mtDNA) sample.
154. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples include mitochondrial ribonucleic acid (mtRNA) samples.
155. The method according to claim 135, wherein the plurality of cell-free nucleic acid samples include extracellular vesicle-bound deoxyribonucleic acid (evDNA) samples.
156. The method according to any one of claims 135 to 155, wherein the biomarker is one or more selected from proteins, exosomes, exomers, microvesicles, apoptotic bodies, neutrophil extracellular traps (NETs), immune cells, tumor-educating platelets (TEPs), microbiomes, pilomes, Toll-like receptors (TLRs), and mitochondrial DNA (mtDNA).
157. The method according to any one of claims 135 to 156, wherein the step of detecting one or more biomarkers includes detecting the presence or level of the one or more biomarkers.
158. The method according to any one of claims 135 to 156, wherein the step of determining whether the biological sample is tumor-derived or non-tumor-derived includes comparing the level of one or more biomarkers in the biological sample with a control.
159. The method according to claim 158, wherein the control is a reference level or a level present in a healthy non-cancer subject.