Method for monitoring esophageal squamous cell carcinoma treatment effect based on multi-omics dynamic ctDNA
By employing multi-omics analysis and machine learning models, this study addresses the shortcomings in assessing the response and prognosis of neoadjuvant immunotherapy for esophageal squamous cell carcinoma in existing technologies. It enables precise monitoring and prediction of the efficacy of treatment for esophageal squamous cell carcinoma, thereby improving the accuracy of treatment monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2024-04-26
- Publication Date
- 2026-07-24
AI Technical Summary
Current technologies are not yet able to effectively utilize ctDNA to assess the response and prognosis of patients with locally advanced esophageal squamous cell carcinoma to neoadjuvant immunotherapy, and there is a lack of accurate monitoring methods.
Through multi-omics analysis, including ctDNA targeted sequencing, cfMeDIP-seq, WGBS differential methylation region identification, TCGA 450K methylation chip differential methylation probe identification, and cfMeDIP-seq differential methylation region identification, a methylation risk score and immune index were constructed. Combined with a logistic regression model, the efficacy of treatment for esophageal squamous cell carcinoma was predicted.
It enabled accurate prediction and prognostic assessment of the response to neoadjuvant immunotherapy for esophageal squamous cell carcinoma, improved the monitoring accuracy of treatment effects, and achieved a classification model accuracy of 88%.
Smart Images

Figure CN118448038B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cancer cell therapy monitoring technology, specifically, it relates to a method for monitoring the efficacy of esophageal squamous cell carcinoma based on multi-omics dynamic ctDNA. Background Technology
[0002] The main histological subtype of esophageal cancer is esophageal squamous cell carcinoma (ESCC). The standard treatment for locally advanced esophageal cancer recommended by the National Comprehensive Cancer Network guidelines is neoadjuvant chemoradiotherapy combined with esophagectomy. According to the CROSS and NEOCRTEC5010 studies, the recurrence rate after neoadjuvant chemoradiotherapy for esophageal cancer is between 30% and 50%, especially with distant metastasis. Therefore, the clinical need to improve treatment outcomes remains unmet.
[0003] Neoadjuvant immunotherapy has shown effective pathological responses in early-stage tumors such as melanoma, lung cancer, bladder cancer, and colorectal cancer, improving patient prognosis to some extent. Neoadjuvant immunotherapy provides an excellent opportunity for real-time monitoring of tumor response and evaluation of drug efficacy. Currently, an increasing number of clinical trials are evaluating the role of neoadjuvant immunotherapy in ESCC. Therefore, accurately predicting patient response to neoadjuvant immunotherapy is a key issue in this regard.
[0004] Circulating tumor DNA (ctDNA) testing in the blood is a non-invasive, repeatable, and minimally invasive method that allows for real-time monitoring, making it an effective means of disease monitoring and assessing treatment response. Clinical studies have shown that ctDNA may play a role in detecting molecular residual disease (MRD) and emerging treatment resistance (i.e., molecular recurrence (MR)) in early-stage breast cancer, as well as in monitoring disease progression in patients with advanced cancer. In high-risk early-stage breast cancer, personalized ctDNA monitoring during neoadjuvant chemotherapy can be used to assess treatment response, help fine-tune pathological complete remission, and serve as a surrogate endpoint for improved prognosis.
[0005] Circulating cell-free DNA 5' motifs combined with MRI tumor regression grade (mrTRG) may predict the response of locally advanced rectal cancer to neoadjuvant chemoradiotherapy. Longitudinal tracking of resectable esophageal adenocarcinoma using serial plasma sampling has revealed that postoperative ctDNA can identify patients at risk of recurrence who could benefit from intensive adjuvant chemotherapy. Several studies have also assessed aberrant DNA methylation in tumors to obtain diagnostic or prognostic biomarkers. However, the role of ctDNA in assessing tumor response and prognosis in patients with locally advanced ESCC to neoadjuvant immunotherapy remains unclear. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides a method for monitoring the efficacy of esophageal squamous cell carcinoma treatment based on multi-omics dynamic ctDNA.
[0007] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows: A method for monitoring the efficacy of esophageal squamous cell carcinoma treatment based on multi-omics dynamic ctDNA includes the following steps: S1. Analyze ctDNA targeted sequencing data and cfMeDIP-seq data; Identification of S2 and WGBS differentially methylated regions; S3 and TCGA 450K methylation chip differential methylation probe identification; S4, cfMeDIP-seq differential methylation region identification; S5, methylation risk score; S6. Define the methylation immune index as an evaluation indicator to assess the prognosis of neoadjuvant immunotherapy; S7. Using the Logistic Regression method, construct a multi-omics hybrid model to obtain the model risk score.
[0008] Further, analysis of ctDNA targeted sequencing data in step S1: The raw sequencing FASTQ data was filtered using Fastp software to remove adapter sequences, poly-N sequences, and low-quality reads, resulting in clean reads. The Burrows-Wheeler Aligner was used to align clean reads to the Ensembl GRCh37 / hg19 reference genome; PCR replicates were processed using Picard to obtain uniquely aligned reads; Local alignment and base recalibration were performed using the Genome Analysis Toolkit (GATK). GATK MuTect2, VarDict, and VarScan were used to identify single nucleotide variants (SNVs), small insertions, and deletions (Indels); ANNOVAR was used for variant annotation based on multiple databases, including population frequency databases (1000G and ExAC), disease or phenotype databases (COSMIC, InterVar and ClinVar), and variant prediction tools (PolyPhen2 and SIFT).
[0009] Further analysis of cfMeDIP-seq data: FastQC was used to evaluate the quality of the raw cfMeDIP-seq FastQ data; Fastp software was used to remove connector sequences, poly-N sequences, and low-quality reads to obtain clean reads; The Burrows-Wheeler Aligner was used to align clean reads to the Ensembl GRCh37 / hg19 reference genome; Use Picard MarkDuplicates to remove PCR duplicate reads; The peak region for each sample is obtained using the callpeak function in the MACS2 software.
[0010] Further, identification of differentially methylated regions in WGBS: Raw data of whole-genome bisulfite sequencing of ESCC tumors and matched healthy tissues were downloaded from the GEO database; Fastp software was used to remove connector sequences, poly-N sequences, and low-quality reads to obtain clean reads; The clean reads were aligned to the Ensembl GRCh37 / hg19 reference genome using Bismark software, and PCR duplicate reads were removed. The reference genome was divided into 300bp non-overlapping windows. The methylation level of the WGBS data in each window was defined as the total number of methylated reads in that window divided by the sum of methylated and non-methylated reads. After calculating the methylation level for each window, the windows are filtered, requiring that at least 70% of the samples have a methylation level higher than 0.7 or lower than 0.25. For the filtered window, the differentially methylated regions were identified using ROTS software as WGBS-DMRs; Finally, a total of 182,338 WGBS-DMRs were obtained.
[0011] Furthermore, differential methylation probes from the TCGA 450K methylation chip were used for identification: Download the infinite human methylation 450 Bead Chips data for esophageal squamous cell carcinoma from the TCGA data portal, and download the relevant clinical information; Microarray data of a 450K human whole blood cohort containing 656 individuals were downloaded from the Gene Expression Omnibus; ChAMP software was used to perform differential methylation analysis on TCGA tumors and adjacent normal samples or blood-derived normal samples, resulting in 135,653 differential methylation probes.
[0012] Further, cfMeDIP-seq differential methylation region identification: Before performing DMRs, the reference genome of the Ensembl GRCh37 / hg19 version was split into 300bp non-overlapping windows; For each sample, the read count for each bin is calculated using featureCounts; The bins are filtered as follows: only bins with more than 5 reads in at least 80% of the samples are retained; Use the total number of reads as the library size to calculate the RPKM for each bin; The ROTS method was used to identify DMRs with a threshold of pvalue < 0.05, and the DMR count data was obtained. All DMRs were further filtered. The following criteria were used for further filtering: (1) overlap with DMRs from WGBS or differentially methylated probes from the TCGA 450K chip; (2) complete overlap with the promoter regions of immune-related genes (https: / / www.innatedb.com / ), ESCC driver genes, or ESCC methylation-related genes. In the end, a total of 13 DMRs were obtained.
[0013] Furthermore, the detailed process for methylation risk scoring is as follows: Calculate the number of RPKMs in the DMRs region of the cfMeDIP-seq obtained in the above steps; Using a random forest model, the importance of these DMRs to the prognostic effect was evaluated in the training set queue at time T0, and the top 60% were selected as candidate DMR regions. LASSO regression models were constructed based on candidate DMR regions in the training set queue at time T0, and methylation risk scores were obtained.
[0014] Furthermore, the formula for calculating the methylation immune index is as follows: Meti = ∑PIIP / ∑NI , PIIP represents the number of peaks that overlap with the promoter of immune genes, and NI represents the number of immune genes.
[0015] Furthermore, for the training set samples, the difference in methylation risk scores between time points T1 and T0 (∆T1-T0Mets) and the ctDNA status at time point T1 are used as inputs. The model risk score is used to predict the efficacy of the patient, and the Youden index is used as the threshold for the risk score. The risk scoring formula is as follows:
[0016] In the formula, α and β are the coefficients of ctDNA state and ∆T1-T0 Mets in the mixture model, respectively, and b is the intercept of the mixture model. α, β, and b are equal to 47.8, 48.7, and -24.5, respectively.
[0017] Compared with the prior art, the present invention has the following advantages: By integrating dynamic genomic and epigenetic features from multiple time points, a machine learning model was constructed to predict the response to neoadjuvant immunotherapy for esophageal squamous cell carcinoma, and to evaluate primer and probe combinations and classification models for esophageal cancer prognosis. Attached Figure Description
[0018] Figure 1 This is an overall flowchart of a method for monitoring the efficacy of esophageal squamous cell carcinoma based on multi-omics dynamic ctDNA in an embodiment of the present invention; Figure 2 This is a schematic diagram of the performance verification process in the experimental procedure of this invention. Detailed Implementation
[0019] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and accompanying drawings. The content mentioned in the embodiments is not intended to limit the present invention.
[0020] like Figure 1 As shown, this embodiment provides a method for monitoring the efficacy of esophageal squamous cell carcinoma treatment based on multi-omics dynamic ctDNA, including the following steps: S1. Analyze ctDNA targeted sequencing data and cfMeDIP-seq data.
[0021] ctDNA targeted sequencing data analysis: For the raw sequencing FASTQ data, we used Fastp software to filter the data, removing adapter sequences, poly-N sequences, and low-quality reads (default parameters) to obtain clean reads. The clean reads were then aligned to the Ensembl GRCh37 / hg19 reference genome using Burrows-WheelerAligner (BWA). PCR repeats were processed using Picard to obtain uniquely aligned reads. Local alignment and base recalibration were performed using the Genome Analysis Toolkit (GATK). Single nucleotide variants (SNVs), small insertions, and deletions (Indels) were identified using GATK MuTect2, VarDict, and VarScan. Variant annotation was performed using ANNOVAR based on multiple databases, including population frequency databases (1000G and ExAC), disease or phenotype databases (COSMIC, InterVar, and ClinVar), and variant prediction tools (PolyPhen2 and SIFT). To remove clonal hematopoiesis, paired leukocytes from each sample were also sequenced. SNVs annotated as benign or possibly benign, or with PopFreqMax > 0.005, were excluded. Non-synonymous SNVs with VAF > 5‰ in non-hotspot regions and supported by at least four high-quality reads, and mutations in driver genes or cancer hotspot regions in the database with VAF > 3‰ and supported by at least two high-quality reads, were retained for subsequent analysis. To ensure the reliability of the mutations, we preprocessed a large number of internal control samples and established a blacklist database, which was used to further filter the aforementioned mutations. If more than one variant was detected in plasma, it was considered ctDNA positive; otherwise, it was ctDNA negative. The concentration of ctDNA in plasma was determined as haplogenous genomic equivalents per milliliter of plasma (hGE / mL). The CNV Kit was used to detect copy number variations.
[0022] Analysis of cfMeDIP-seq data: For the raw cfMeDIP-seq FastQ data, we used FastQC (http: / / www.bioinformatics.babraham.ac.uk / projects / fastqc / ) to assess data quality. Fastp software was used to remove adapter sequences, poly-N sequences, and low-quality reads (default parameters) to obtain clean reads. Burrows-WheelerAligner (BWA) was used to align the clean reads to the Ensembl GRCh37 / hg19 reference genome (default parameters). Picard MarkDuplicates (http: / / picard.sourceforge.net) was used to remove PCR duplicate reads. Then, the callpeak function in MACS2 software was used to obtain the peak region for each sample (qvalue < 0.05).
[0023] Identification of S2 and WGBS differentially methylated regions; Raw whole-genome bisulfite sequencing (WGBS) data from ESCC tumors and matched healthy tissues were downloaded from the GEO database, including 10 tumor samples and 9 healthy samples. Adaptor sequences, poly-N sequences, and low-quality reads were removed using Fastp software [PMID: 30423086] (default parameters) to obtain clean reads. Then, Bismark software was used to align the clean reads to the Ensembl GRCh37 / hg19 reference genome, and PCR duplicate reads were removed. The reference genome was split into 300 bp non-overlapping windows (bins). The methylation level of the WGBS data in each bin was defined as the total number of methylated reads in that bin divided by the sum of methylated and unmethylated reads. After calculating the methylation level for each bin, the bins were filtered, requiring that at least 70% of the samples have a methylation level higher than 0.7 or lower than 0.25. For the filtered bins, differentially methylated regions (p<0.05) were identified using ROTS software and termed WGBS-DMRs. Finally, a total of 182,338 WGBS-DMRs were obtained. These WGBS-DMRs were used for subsequent filtering of cfMeDIP-seqDMRs.
[0024] S3 and TCGA 450K methylation chip differential methylation probe identification; We downloaded infinite human methylation 450 BeadChips (450K) data (based on hg19) for esophageal squamous cell carcinoma (ESCC) from the TCGA data portal (https: / / tcga-data.nci.nih.gov / tcga / ), including 96 tumor samples and 16 adjacent normal samples, along with relevant clinical information. Additionally, we downloaded microarray data of a 450K human whole blood cohort containing 656 individuals from the GeneExpression Omnibus (GSE40279). Differential methylation analysis was performed on TCGA tumors and adjacent or blood-derived normal samples using ChAMP software (FDR < 0.05, |deltaBeta| > 0.1). Finally, a total of 135,653 differentially methylated probes (DMPs) were obtained. These DMPs were used for subsequent filtering of cfMeDIP-seq DMRs.
[0025] S4, cfMeDIP-seq differential methylation region identification; The training set at time point T0 contained 4 pCR and 9 non-pCR patients. We identified differentially methylated regions (DMGs) between the pCR and non-pCR groups. Before performing DMRs, the Ensembl GRCh37 / hg19 version of the reference genome was split into 300bp non-overlapping windows (bins). For each sample, read counts for each bin were calculated using featureCounts, and then the bins were filtered as follows: only bins with at least 80% of the samples having more than 5 reads were retained. After filtering, a total of 424,561 bins were retained, and the RPKM for each bin was calculated using the total read count as the library size. Next, the ROTS method was used to identify DMRs with a pvalue < 0.05, resulting in 23,297 DMRs. All of these DMRs were further filtered using the following criteria: (1) overlap with DMRs from WGBS or differentially methylated probes from the TCGA 450K chip; (2) complete overlap with the promoter regions of immune-related genes (https: / / www.innatedb.com / ), ESCC driver genes, or ESCC methylation-related genes. Ultimately, 13 DMRs were obtained.
[0026] S5, methylation risk score; First, the RPKM of the DMRs regions obtained from the above steps for the 13 cfMeDIP-seq datasets was calculated. Then, using a random forest model, the importance of these DMRs to the prognostic effect was evaluated in the training set at time T0, and the top 60% were selected as candidate DMRs regions. Similarly, a LASSO regression model was constructed based on the candidate DMRs regions in the training set at time T0, and a methylation risk score (Mets) was obtained.
[0027] S6. Define the methylation immune index as an evaluation indicator to assess the prognosis of neoadjuvant immunotherapy; The methylation immune index (Meti) is calculated using the following formula.
[0028] Meti = ∑PIIP / ∑NI PIIP represents the number of peaks that overlap with the promoters of immune genes. NI represents the number of immune genes.
[0029] S7. Using the Logistic Regression method, construct a multi-omics hybrid model to obtain the model risk score.
[0030] For the training set samples, the difference in methylation risk scores between time points T1 and T0 (∆T1-T0 Mets) and the ctDNA status at time point T1 are used as inputs. A multi-omics hybrid model is constructed using the Logistic Regression method to obtain the model risk score, as detailed in the formula below. The model risk score is used to predict patient efficacy, with the Youden index used as the threshold for the risk score; high scores are considered non-pCR, and low scores are considered pCR.
[0031]
[0032] In the formula, α and β are the coefficients of ctDNA state and ∆T1-T0 Mets in the mixture model, respectively, and b is the intercept of the mixture model. α, β, and b are equal to 47.8, 48.7, and -24.5, respectively.
[0033] like Figure 2 As shown, the results of the multiple methylation qPCR method verification are as follows: We selected 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, and 1 gene from 13 DMRs, resulting in 8190 possible combinations. The classification performance of non-pCR and pCR models was calculated using different gene combinations. The classification performance gradually decreased with fewer gene combinations. Finally, we selected 6 gene combinations to maintain an AUC of 1. Primers and probes were designed for these 6 DMR regions and the internal reference gene ACTB, with the sequences as follows:
[0034] The above primer combinations were used to verify the plasma scores of 15 cases of non-pCR and 12 cases of pCR in esophageal cancer.
[0035] The experimental procedure is as follows: cfDNA was extracted using a plasma DNA extraction kit. cfDNA transformation was performed using the BS transformation kit. The transformed DNA product was divided into two equal parts as templates, and multiplex qPCR was performed using ARL8B, SHOX2, RASSF1A, PTGER4 (tube 1) and EMX1, ZNF497, ACTB (tube 2), respectively. Subtract the ACTB CT value from the CT values of ARL8B, SHOX2, RASSF1A, PTGER4, EMX1, and ZNF497 respectively to obtain the normalized CT value of each gene. The default CT value for genes with no detected signal is 45. Dataset splitting: 70% of the samples from non-pCR and pCR were randomly selected as the training set, and the remaining 30% were selected as the test set.
[0036] Model Building and Optimization: To establish a prognostic model for esophageal cancer, the sklearn package in Python 3 was used. Model building and parameter optimization were performed based on the training set. The LogisticRegression model parameters included: regularization coefficient (C), penalty term, and loss function (solver). This study used a grid search method to fine-tune the parameters, and the final parameter settings were: (C=1.0, penalty='l2', solver=sag). The model formula is: y = 0.8206-0.1756*F1+0.1135*F2+0.4061*F3-0.0951*F4-0.0837*F5-1.3159*F6. (F1: ARL8B, F2: SHOX2, F3: RASSF1A, F4: PTGER4, F5: EMX1, F6: ZNF497) Threshold determination: The optimal threshold is determined by using the model prediction of cancer probability values of training set samples and the gold standard results input into the ROC curve in step 2. The probability value with the largest Youden index is selected as the probability value of the model. When the predicted probability value is 0.36, the Youden index is the highest.
[0037] Test results:
[0038] The results show that the classification model performs excellently, with an accuracy of 88%.
[0039] The above provides a detailed description of a method for monitoring the efficacy of esophageal squamous cell carcinoma based on multi-omics dynamic ctDNA, as provided in this application. The specific embodiments are described only to aid in understanding the method and its core concepts. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the scope of protection of the claims.
Claims
1. A method for monitoring the efficacy of esophageal squamous cell carcinoma treatment based on multi-omics dynamic ctDNA, characterized in that, Including the following steps: S1. Analyze ctDNA targeted sequencing data and cfMeDIP-seq data; Identification of S2 and WGBS differentially methylated regions; S3 and TCGA 450K methylation chip differential methylation probe identification; S4, cfMeDIP-seq differential methylation region identification; S5, methylation risk score; S6. Define the methylation immune index as an evaluation indicator to assess the prognosis of neoadjuvant immunotherapy; S7. Using the Logistic Regression method, construct a multi-omics hybrid model to obtain the model risk score; Analysis of ctDNA targeted sequencing data in step S1: The raw sequencing fastq data was filtered using Fastp software to remove adapter sequences, poly-N sequences, and low-quality reads, resulting in clean reads of ctDNA targeted sequencing data. The Burrows-Wheeler Aligner was used to align clean reads of ctDNA targeted sequencing data to the EnsemblGRCh37 / hg19 reference genome; PCR replicates were processed using Picard to obtain uniquely aligned reads; Local alignment and base recalibration were performed using the Genome Analysis Toolkit; GATK MuTect2, VarDict, and VarScan were used to identify single nucleotide variants, small insertions, and deletions. ANNOVAR was used for variant annotation based on multiple databases, including population frequency databases, disease or phenotype databases, and variant prediction tools. cfMeDIP-seq data analysis: FastQC was used to evaluate the quality of the raw cfMeDIP-seq FastQ data; Fastp software was used to remove connector sequences, poly-N sequences, and low-quality reads to obtain clean reads of cfMeDIP-seq data. The Burrows-Wheeler Aligner was used to align clean reads of the cfMeDIP-seq data to the EnsemblGRCh37 / hg19 reference genome; Use Picard MarkDuplicates to remove PCR duplicate reads; The peak region for each sample is obtained using the callpeak function in the MACS2 software. Identification of WGBS differentially methylated regions: Raw data of whole-genome bisulfite sequencing of ESCC tumors and matched healthy tissues were downloaded from the GEO database; Fastp software was used to remove adapter sequences, poly-N sequences, and low-quality reads to obtain clean reads of whole-genome bisulfite sequencing data; The clean reads of the whole-genome bisulfite sequencing data were aligned to the EnsemblGRCh37 / hg19 reference genome using Bismark software, and PCR duplicate reads were removed. The reference genome was divided into 300bp non-overlapping windows. The methylation level of the WGBS data in each window was defined as the total number of methylated reads in that window divided by the sum of methylated and non-methylated reads. After calculating the methylation level for each window, the windows are filtered, requiring that at least 70% of the samples have a methylation level higher than 0.7 or lower than 0.
25. For the filtered window, the differentially methylated regions were identified using ROTS software as WGBS-DMRs; Differential methylation probe identification using TCGA 450K methylation chip: Download the infinite human methylation 450 Bead Chips data for esophageal squamous cell carcinoma from the TCGA data portal, and download the relevant clinical information; Microarray data of a 450K human whole blood cohort containing 656 individuals were downloaded from the Gene Expression Omnibus; ChAMP software was used to perform differential methylation analysis between TCGA tumors and adjacent normal samples or blood-derived normal samples, resulting in 135,653 differential methylation probes. Identification of differentially methylated regions using cfMeDIP-seq: Before performing DMRs, the reference genome of the Ensembl GRCh37 / hg19 version was split into 300bp non-overlapping windows; For each sample, the read count for each bin is calculated using featureCounts; The bins are filtered as follows: only bins with more than 5 reads in at least 80% of the samples are retained; Use the total number of reads as the library size to calculate the RPKM for each bin; The ROTS method was used to identify DMRs with a threshold pvalue < 0.05, and the DMR count data was obtained. Further filtering was performed on all DMRs.
2. The method for monitoring the efficacy of esophageal squamous cell carcinoma treatment based on multi-omics dynamic ctDNA according to claim 1, characterized in that, Detailed process for methylation risk scoring: Calculate the number of RPKMs in the DMRs region of the cfMeDIP-seq obtained in the above steps; Using a random forest model, the importance of these DMRs to the prognostic effect was evaluated in the training set queue at time T0, and the top 60% were selected as candidate DMR regions. LASSO regression models were constructed based on candidate DMR regions in the training set queue at time T0, and methylation risk scores were obtained.
3. The method for monitoring the efficacy of esophageal squamous cell carcinoma treatment based on multi-omics dynamic ctDNA according to claim 2, characterized in that, Formula for calculating the methylation immune index: Meti = ∑PIIP / ∑NI , PIIP represents the number of peaks that overlap with the promoter of immune genes, and NI represents the number of immune genes.
4. The method for monitoring the efficacy of esophageal squamous cell carcinoma treatment based on multi-omics dynamic ctDNA according to claim 3, characterized in that, For the training set samples, the difference in methylation risk scores at time points Tx and T0 and the ctDNA status are used as inputs. The model risk score is used to predict the efficacy of the treatment for patients, and the Youden index is used as the threshold for the risk score. The risk scoring formula is as follows: In the formula, α and β are the coefficients of ctDNA state and ∆Tx-T0 Mets in the mixture model, respectively, and b is the intercept of the mixture model. α, β and b are equal to 47.8, 48.7 and -24.5, respectively.