Solid tumor MRD detection method and detection system based on ctDNA and computer readable medium
By using MRD qualitative algorithm and Monte-Carlo sampling method in MRD detection, the p-value of ctDNA positive was calculated, and the problem of background noise interference in plasma free DNA detection was solved, and the accuracy of the detection was improved.
Patent Information
- Application Number
- CN202510102066.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Plasma free DNA has background noise interference in MRD detection, resulting in detection errors.
The MRD qualitative algorithm was used to estimate the background error rate of single-base mutations, simulate the mutation readings, and calculate the p-value of ctDNA positive by Monte-Carlo sampling to improve the accuracy of the detection.
It improves the accuracy of solid tumor ctDNA MRD detection and reduces misjudgment caused by background noise.
Smart Images

Figure CN119943142A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a ctDNA-based solid tumor MRD detection method, a detection system and a computer-readable medium, and belongs to the technical field of tumor gene detection. Background Art
[0002] MRD, often translated as Minimal Residual Disease, mainly refers to cancer-derived cells or abnormal molecules that cannot be found by traditional imaging or laboratory methods but can be found by liquid biopsy after radical treatment of tumor patients. MRD in solid tumors usually refers to molecular residual lesions. Plasma ctDNA is the most mature medium for evaluating MRD, and its detection has been written into multiple guidelines / consensus. At present, the application value of MRD technology in tumors mainly has the following three points: Evaluate prognosis and select treatment: MRD detection based on ctDNA can stratify the risk of recurrence and evaluate the prognosis of patients. This can assist the clinician in identifying potential treatment exemptions and beneficiaries, and match precise treatment plans.
[0003] Monitor recurrence and intervene in time: MRD can detect disease recurrence earlier than imaging and traditional clinical methods, creating valuable opportunities for patients to intervene in time.
[0004] Evaluate efficacy and adjust regimen: ctDNA-based MRD monitoring can predict and evaluate treatment efficacy at an early stage, and provide a reference for drug discontinuation observation, regimen adjustment, and continued treatment.
[0005] MRD detection strategies can be divided into two categories based on whether the tissue is tested at baseline: tumor-informed strategies and tumor-unaware strategies. Compared with tumor-unaware strategies, tumor-informed strategies can reduce the risk of false positives caused by technical and biological background errors during ctDNA evaluation, and detect fewer ctDNA mutation reads, thus achieving better sensitivity.
[0006] Based on the tumor-informed strategy, the choice of subsequent MRD monitoring panel can be divided into: personalized panel and fixed panel (fixed panel is also called group customized panel). The personalized panel first tests the tumor tissue sample to understand the patient-specific genomic variation map, and then detects certain mutation sites in the tumor tissue through ctDNA. The representative company of personalized panel testing is Natera's Signatera technology. First, the patient's tumor tissue sample is tested through whole-exome WES, and 16 sites are selected for subsequent MRD monitoring. The fixed panel does not rely on the patient's specific mutations, but selects the same type of monitoring panel for a certain patient group, such as lung cancer, colorectal cancer, etc., which can not only detect existing mutations in the tissue, but also detect new mutations.
[0007] The following problems exist in the existing MRD detection technology: MRD status needs to be evaluated based on ctDNA mutations, but not all mutations detected will be considered MRD positive. At the technical implementation level, single-strand errors or random errors may occur during library enrichment and PCR. At this time, if the sequencing depth is simply increased without background purification, it will only amplify the interference with MRD interpretation; at the same time, clonal hematopoietic mutations will also bring biological background noise. Misjudging the above mutations as tumor-specific will inevitably strongly interfere with the judgment of the detection signal, thereby affecting the specificity of MRD detection, and ultimately causing huge deviations in clinical practice. Summary of the invention
[0008] The technical problem to be solved by the present invention is: when plasma free DNA is tested for MRD, the detection error interference caused by background noise. The present invention proposes a method for qualitatively testing the baseline mutation sites of this tissue in plasma using an MRD qualitative algorithm during dynamic plasma testing. This method estimates the background error rate of single-base mutations, simulates the readings based on this error rate, and finally compares them with the actual observed readings. If the simulated readings are equal to or greater than the actual observed readings, the p-value of ctDNA positivity is calculated by the number of simulated experiments. The p-value calculation method is based on the Monte-Carlo sampling method, which calculates the p-value by simulating a certain number of experiments (such as 10,000 times) and counting the number of simulated readings greater than or equal to the actual observed readings.
[0009] A solid tumor ctDNA MRD detection method, which is used for non-therapeutic and diagnostic purposes, comprises the following steps: Step 1, performing high-throughput targeted sequencing on the subject's plasma sample to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations in his tissue sample and has tissue sample positive somatic mutation data; Step 2, for the sequencing data of the obtained plasma samples, calculating the background mutation frequency of each type of three-base mutation; Step 3, comparing the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the positive somatic mutations in the tissue sample; Step 4, based on the background mutation frequency of the three-base mutation type obtained in step 2, at the corresponding site of the positive somatic mutation in the tissue sample obtained in step 3, determine whether there is a mutation at the site in the plasma sample by random simulation; Step 5, repeat step 4 to simulate, and calculate the probability p of no mutation under the simulation conditions; In step 6, the following judgments are made: (1) If the plasma mutations found in the high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is greater than or equal to 3, the sample is judged to be MRD-positive; (2) If the plasma mutations found in the high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is less than 3, the sample is judged to be MRD-negative; (3) For other cases other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD-positive.
[0010] In step 2, the three-base mutation type refers to a combination of 12 basic single-base mutation forms A>T, A>G, A>C, C>A, C>T, C>G, G>A, G>C, G>T, T>A, T>C, T>G and one base upstream and downstream of the reference genome at that position.
[0011] The background mutation frequency of three-base mutation types is obtained through the following steps: stacking the BAM file, reference sequence file and tissue baseline positive mutation to generate pileup format data for the specified area, filtering out common SNP sites and blacklist sites, and then filtering out sites with too small sequencing depth or too large total frequency of non-reference bases; the mutation frequency of each type of three-base mutation type is counted as the background mutation rate.
[0012] The sequencing depth is too small when it is less than 3-6, and the total frequency of non-reference bases is too high when it is greater than 8-12%.
[0013] The number of simulations is 200-100000 times.
[0014] The threshold is 0.001-0.05.
[0015] The random simulation method refers to: simulating according to the binomial distribution to simulate whether there is a mutation at each site; if the total average mutation frequency of the simulation is greater than or equal to the actual observed average mutation frequency, and the total number of mutation site reads of the simulation is greater than or equal to the actual observed total number of mutation site reads, then it is assumed that the plasma sample has a mutation, otherwise it has no mutation.
[0016] In the simulation of binomial distribution, random simulation calculation is performed at the mutation site according to the sequencing depth in sequencing and the background mutation frequency of the corresponding three-base mutation type.
[0017] A solid tumor ctDNA MRD detection system, which is used for non-therapeutic and diagnostic purposes, comprising: A sequencing module, used for performing high-throughput targeted sequencing on a plasma sample of a subject to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations of his tissue sample and has tissue sample positive somatic mutation data; A background mutation frequency calculation module is used to calculate the background mutation frequencies of various three-base mutation types for the sequencing data of the obtained plasma samples; A mutation data analysis module, used to compare the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the positive somatic mutations in the tissue sample; A mutation simulation generation module is used to determine whether there is a mutation at the site in the plasma sample by random simulation at the corresponding site of the positive somatic mutation in the obtained tissue sample based on the background mutation frequency of the three-base mutation type; and repeat the simulation to calculate the probability p of no mutation under the simulation conditions; The MRD determination module is used to make the following determinations: (1) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is greater than or equal to 3, the sample is determined to be MRD positive; (2) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is less than 3, the sample is determined to be MRD negative; (3) For other situations other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD positive.
[0018] A computer-readable medium having recorded thereon a computer program capable of executing the above-mentioned solid tumor ctDNA MRD detection method.
[0019] The beneficial effects of this patent are: by calculating the background error rate of single-base mutations, simulating mutation readings based on the error rate, and comparing them with the actual observed mutation readings, the number of times the simulated mutation readings are greater than the actual observed mutation readings. The p-value of the sample ctDNA positive is calculated by multiple simulation results at the mutation site; the accuracy of solid tumor ctDNA MRD detection is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is the IGV visualization of patient samples with inconsistent MRD status before and after the improvement of this patented algorithm. DETAILED DESCRIPTION
[0021] In the detection method of this patent, the cancer type is first determined, and finally 14 cancer types (lung adenocarcinoma, colorectal cancer, breast cancer, gastric cancer, lung squamous cell carcinoma, liver cancer, cervical cancer, bile duct cancer, esophageal cancer, endometrial cancer, diffuse large B-cell lymphoma, prostate cancer, skin melanoma, and brain glioma) are determined by analyzing the samples. When performing sample analysis, the tumor sample data included are the mutation site enrichment regions of nearly 7,000 tumor WES samples obtained by actual testing by the patent applicant, and on this basis, the mutation hotspots covering the COSMIC database and the clinical driver mutation library are added to determine the site area for detection, and then the MRD probe is designed.
[0022] The specific sample information is shown in Table 1: Table 1 Sample type and quantity
[0023] After obtaining the above data, when performing site screening, common pseudomutations are first filtered through the healthy human data baseline database. Common pseudomutations here usually refer to those variations that appear in the gene sequence and do not actually lead to changes in protein function or the appearance of disease phenotypes. These pseudomutations include: 1) Pseudogenes: Pseudogenes are non-functional DNA fragments similar to functional gene sequences. They may be redundant copies of functional genes and are usually identified as pseudogenes due to frameshift mutations or premature termination codons. Pseudogenes may produce low levels of transcription to generate RNA, but usually cannot produce functional final protein products. 2) Single nucleotide polymorphisms (SNPs): Throughout the human genome, a common mutation mode is called single nucleotide polymorphism (SNP). SNP is a single base change in DNA, and these polymorphisms are often used as genomic landmarks or "markers". However, not all SNPs will lead to functional changes, and some may be neutral or pseudomutations. 3) Common disease-common variation hypothesis (CD-CV): This hypothesis holds that the susceptibility to common complex diseases is caused by common variations at certain sites in the population, especially in the coding or regulatory regions of genes. These common variants (MAF between 1% and 5%) are not necessarily pathogenic. This means that among these common variants, some may not cause disease, but are pseudomutations. 4) Common Disease-Rare Variant Hypothesis (CD-RV): This hypothesis holds that the genetic factors of complex diseases are mainly composed of many rare variants with low frequencies (generally MAF <1%) and high pathogenic risks. This means that when analyzing rare variants, some pseudomutations that do not cause disease may be encountered. These pseudomutations can be obtained through self-built databases obtained by the applicant through sample testing and analysis, or through public databases, such as HGMD, ClinVar, TCGA and other databases.
[0024] Furthermore, the mutation type is limited. The mutation type screening used in the method of this patent is: Missense mutation (missense_mutation), frameshift deletion mutation (frame_shift_del), nonsense mutation (nonsense_mutation), frameshift insertion mutation (frame_shift_ins), in-frame insertion (in_frame_ins), in-frame deletion (in_frame_del), transcription start site (translation_start_site), stop codon mutation (nonstop_mutation), silent mutation (Silent), splice site (Splice_Site).
[0025] Furthermore, the obtained sites need to be further screened to ensure that the variation at each site has appeared in at least a certain number of samples. In this embodiment, the variation sites are retained to occur in at least 20 samples. The purpose of this step is to exclude some mutations that may be caused by errors in the sequencing process to avoid false positives.
[0026] Furthermore, the regions of the mutations that meet the conditions and whose hotspot coordinates are within 10 bp are merged. After merging, if the merged region is less than 80 bp, it can be extended to the same length at both ends to further expand to a region of 80 bp (or 81 bp). If the length of the merged region is greater than 80 bp, the region will not be extended. Then, the performance of each interval of the site interval set obtained in the previous step in the pan-cancer data is counted; and the pan-cancer type and each cancer type count statistics are performed. The pan-cancer type is a mutation carried in any tumor sample, and is counted by sample (for example, sample A has 2 mutations in interval α, but the sample count is recorded as 1).
[0027] After tracing back the coordinates of the main contributing variant sites in each interval in the previous step and merging them with the COSMIC database (V97), further add microsatellite site regions and fusion regions. For the obtained intervals, if the interval is <=120bp, expand it to a 120bp interval; if the interval is already greater than 120bp, use 120bp as a sliding window and split it into 120bp regions in sequence. If the last region is less than 120, design from the end to the front. The microsatellite loci included here were screened in the following way: Screening of satellite sites: a. Whole exome sequencing (WES) data of healthy subjects and colorectal cancer MSI-positive samples (PCR method, obtained by testing 5 loci (NR-27, NR-24, NR-21, BAT-25 and BAT-26)) with qualified quality control (sequencing depth>150X), 100 samples each.
[0028] b. From the WES genome coverage area, single nucleotide repeat sequences with a length of 10 to 20 bp were screened as candidate satellite sites.
[0029] c. Preliminary screening of satellite loci: more than 80% of the samples must have a microsatellite locus coverage depth greater than 50% of the average depth of the samples themselves.
[0030] d. For the qualified satellite loci screened in the previous step, the coverage depth and the number of alleles of different lengths of each sample at each microsatellite locus were counted, and the frequency of each allele was calculated. The SVM algorithm was used as a feature for binary classification training, and the loci with an AUC value greater than 0.85 were selected as satellite loci with high MSI resolution.
[0031] The fusion genes selected here are ALK, ROS1, RET, NTRK, NRG, FGFR, MET, EGFR, HER2 and BRAF.
[0032] After screening the above sites and regions, a 120bp probe was designed. After ctDNA extraction, library construction, library enrichment, library quantification, cyclization and DNB preparation, high-throughput sequencing was performed on the DNBSEQ-T7 sequencer; after sequencing, the raw data was quality controlled and bioinformatics analyzed to obtain reliable biological information.
[0033] For mutations detected in tissue baseline samples, the MRD qualitative algorithm is used to perform qualitative testing of the tissue baseline mutation sites in plasma during dynamic plasma testing; The specific steps are as follows: MRD positivity was determined by examining whether the tissue baseline positive mutations and the level of new mutations in plasma samples were significantly higher than the background mutation rate.
[0034] The core idea is to estimate the background error rate of single-base mutations, simulate the readings based on this error rate, and finally compare them with the actual observed readings. If the simulated readings are equal to or greater than the actual observed readings, the p-value of ctDNA positivity is calculated by the number of simulated experiments. The p-value calculation method is based on the Monte-Carlo sampling method, which calculates the p-value by simulating a certain number of experiments (such as 10,000 times) and counting the number of simulated readings greater than or equal to the actual observed readings.
[0035] Specifically, the depth is sampled based on the background error rate, and then this sampling result is compared with the actual observed base readings to determine the results of the simulation experiment. Finally, through multiple simulation experiments, the probability of the simulated reading being equal to or greater than the actual observed reading is calculated as the p-value for ctDNA positivity.
[0036] The MRD qualitative algorithm consists of four steps: 1) Data preprocessing In this step, multiple input files including BAM files, reference sequence files, and lists of tissue baseline positive mutations and new mutations are first read. The samtools tool is used for pileup operation to generate pileup format data for the specified region. In this process, for each site in the read pileup, the previous and next bases and the reference genotype of the site are extracted to form a three-base reading context (tri-nucleotide context, TNC). Tri-base mutation refers to the combination of 12 basic single-base mutation forms A>T, A>G, A>C, C>A, C>T, C>G, G>A, G>C, G>T, T>A, T>C, T>G and one base upstream and downstream of the reference genome at the position. There are 192 types in total, which are used to accurately calculate the sample TNC background.
[0037] 2) Calculation of background error rate Background error (or background noise) refers to systematic errors in the sequencing process, which often lead to incorrect mutation calls in a specific sequence environment. The pileup operation in the BAM file is to stack the sequencing reads at each position for analysis, and evaluate and filter the sites with mutations. The calculation of background error rate mainly includes the following steps: a. The pileup data format describes all the aligned reads at a given position. For each row in the pileup, a series of filtering operations are performed using the three-base read context extracted previously.
[0038] b. This includes filtering out common population polymorphic sites and blacklisting regions and sites that are characterized by errors that occur under normal circumstances or variable sites that are known not to cause disease.
[0039] c. Filter the depth (sequencing depth or coverage, that is, the number of sequencing reads at a site). Sites with a depth less than 5 or a sum of base frequencies exceeding 10% will be filtered out. If the sequencing depth of a site is less than 5, the data at this site may be considered unreliable due to insufficient coverage. Too low coverage may result in the inability to accurately detect true variants because of possible random errors or insufficient sequencing. When the total frequency of non-reference bases at a site exceeds 10%, the high frequency of non-reference bases may indicate sequencing errors or sample contamination, which will also affect the accuracy of the results. Through such filtering, data sites that may cause inaccurate results due to sequencing quality issues can be removed, thereby improving the accuracy and reliability of subsequent analysis and ensuring the accuracy of the results. Too low coverage and too high non-reference base frequencies may cause inaccurate results.
[0040] For each target site, the mutation frequency is counted, and then each possible mutation (Ref, Alt) key-value pair is generated in turn, and the frequency of the key-value pair is recorded. The (Ref, Alt) and mutation frequency in the three-base context are put into a hash table.
[0041] Through this hash table, the mutation frequencies of all 12 single-base mutations in 16 contexts can be calculated. The frequency of each mutation type is equal to the total number of mutations of that type divided by the depth of the corresponding triplet type base pair, which can reflect the degree of error in a specific sequence environment. This ratio is the calculated background error rate.
[0042] 3) Estimation of ctDNA positive p value Monte Carlo simulation method was used to estimate the positive probability of ctDNA: a) When performing plasma MRD analysis, the subjects have previously undergone mutation testing of tissue samples. Here, we use the binomial distribution to simulate whether each tissue baseline positive mutation in the plasma sample will mutate. The probability of success (i.e., mutation) is the background mutation rate calculated above. The generated random numbers simulate the mutation distribution we expect based on the background mutation rate. Specifically, for M mutation sites (including each tissue baseline positive mutation), we use the background mutation rate p b (calculated previously) to simulate and determine whether each site will mutate. For each mutation site i, the sequencing depth of the mutation site is D i , where i=1,2,…M, becomes a random variable X i Indicates the number of reads with mutations at site i:
[0043] If site i mutates, X i > 0, otherwise X i <0.
[0044] b) If the number of simulated mutation site reads is greater than or equal to the number of actually observed mutation site reads, and the simulated average mutation frequency is greater than or equal to the actually observed average mutation frequency, then the plasma sample is considered to have mutations. Otherwise, it is considered to have no mutations. Specifically, calculate the total number of site reads S where mutations occur in the simulation:
[0045] Let S obs is the number of reads at the site where the mutation actually occurred, f obsis the average mutation frequency of the M mutation sites actually observed. For the simulated data, the average mutation frequency is:
[0046] We compare the simulation results with the observed data. For each simulation, if , then I = 1, otherwise I = 0. I is the result of the above comparison.
[0047] c) This process was repeated multiple times (the default is 10,000 times) in order to make the results more stable and to obtain a p-value. The p-value can be understood as the probability that no mutation occurs. In other words, the smaller the p-value, the greater the possibility of mutation. At this time, there is a greater degree of confidence that these mutations are not background noise but real mutation signals, indicating that ctDNA may exist in the plasma sample. Specifically, by repeating the simulation N times, we can estimate the p-value (that is, the probability that the observed mutation frequency and number of mutation sites exceed the actual observed value under the null hypothesis):
[0048] in is the comparison between the j-th simulation result and the observed data.
[0049] In this way, it is possible to assess whether the observed mutations are statistically significant and thus determine the probability of ctDNA positivity.
[0050] 4) MRD qualitative analysis The final steps for the MRD characterization of the sample are: a. Check whether the mutations in plasma contain the baseline positive mutations in tissues. If so, the sample is considered MRD positive (which means that when the patient's plasma MRD test is performed, the same mutations that have been found in the patient's tissues can still be found).
[0051] b. To ensure the effectiveness of the Monte Carlo simulation, if the sum of the support numbers of baseline positive mutations in all tissues is less than 3, the sample is considered MRD negative. That is, only the mutation sites in the corresponding tissues are considered. If the sum of the reads of mutations at all these sites in the plasma is less than 3, it is considered negative.
[0052] In other cases, MRD is qualitatively determined based on the p-value obtained in the above simulation. We use p<0.01 as the threshold for MRD positivity. The final result will generate a detailed mutation report, including information such as mutation site, mutation type, mutation frequency, background error rate, etc.
[0053] This algorithm provides an effective method for calculating ctDNA positivity, which draws on the Monte-Carlo method in statistics and combines the concept of background error rate in genomics. It can be effectively applied to ctDNA positivity analysis of clinical samples.
[0054] Example 1 Testing the degree of improvement of the MRD qualitative algorithm of the present invention on the sensitivity of MRD analysis In order to test the degree of improvement in the sensitivity of MRD analysis after algorithm upgrade, 78 samples of patients who had not been treated and pathologically diagnosed with advanced solid tumors were randomly screened (theoretically, the MRD status was positive), and the baseline tumor tissue samples, white blood cell control samples, and baseline plasma samples of each patient were sequenced using the MRD probe kit. The MRD status of plasma samples was evaluated using the pre-upgrade MRD analysis process (only MRD qualitative step a) and the post-upgrade MRD analysis process (including MRD qualitative steps a and b). The comparative analysis results are as follows. After adding the above-mentioned MRD qualitative step b, the number of MRD-positive samples increased by 3, and the sensitivity increased by 3.8%. The specific comparison results are shown in Table 2: Table 2 Comparison of MRD analysis sensitivity
[0055] One of the patient samples, such as Figure 1 As shown in the figure, tissue baseline variation was detected at multiple sites in plasma, but none of them met the single site positive judgment criteria. If only based on the judgment criteria of MRD qualitative step a, the sample MRD was judged to be negative; if MRD qualitative step b was added, the sample could be accurately judged to be MRD positive. It can be seen that adding MRD qualitative step b is very helpful to improve the sensitivity of MRD analysis.
[0056] Considering the main clinical application scenarios of MRD detection, including perioperative patients’ postoperative recurrence risk monitoring and local advanced inoperable patients’ tumor recurrence monitoring after radical chemoradiotherapy, 68 more samples of patients in stages I to III were added for sensitivity testing. The results of the comparative analysis before and after the MRD qualitative step upgrade are as follows. After adding the above-mentioned MRD qualitative step b, one MRD-positive patient in stage I and one MRD-positive patient in stage III were added, a total of 2 MRD-positive samples were added, and the sensitivity was increased by 2.9%. The specific comparison results are shown in Table 3: Table 3 Comparison of MRD analysis sensitivity
[0057] From the above data, it can be seen that whether it is an advanced tumor or a stage I~III tumor, adding the MRD qualitative step b can improve the overall MRD analysis sensitivity.
[0058] Example 2 Testing the Effect of the MRD Qualitative Algorithm of the Present Invention on the Specificity of MRD Analysis In order to test the impact of the MRD analysis algorithm upgrade on the analysis specificity, we selected 100 healthy human plasma samples and sequenced them using the MRD probe kit. Theoretically, there are no tumor somatic mutations in the plasma samples of healthy subjects. We mixed the plasma sequencing data of each healthy person with the mutation list of 78 patients with advanced tumors. 100 healthy people × 78 tumor mutation tables can generate 7,800 samples. The MRD qualitative step was upgraded before and after the upgrade analysis process to determine the MRD status. The comparative analysis results of the two analysis processes are shown in Table 4: Table 4 Comparison of MRD analysis sensitivity
[0059] After adding the Monte Carlo model for sampling analysis, the specificity of MRD analysis can be improved from 81.9% to 98.2%, an increase of 16.3%. This shows that adding the Monte Carlo model for sampling analysis can not only improve sensitivity, but also improve specificity.
Claims
1. A method for detecting ctDNA MRD in solid tumors, which is used for non-therapeutic and diagnostic purposes, comprising the following steps: Step 1, performing high-throughput targeted sequencing on the plasma sample of the subject to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations of his tissue sample and has tissue sample positive somatic mutation data; Step 2, for the sequencing data of the obtained plasma samples, calculating the background mutation frequency of each type of three-base mutation; Step 3, comparing the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the positive somatic mutations in the tissue sample; Step 4, based on the background mutation frequency of the three-base mutation type obtained in step 2, at the corresponding site of the positive somatic mutation in the tissue sample obtained in step 3, determine whether there is a mutation at the site in the plasma sample by random simulation; Step 5, repeat step 4 to simulate, and calculate the probability p of no mutation under the simulation conditions; In step 6, the following judgments are made: (1) If the plasma mutations found in the high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is greater than or equal to 3, the sample is judged to be MRD-positive; (2) If the plasma mutations found in the high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is less than 3, the sample is judged to be MRD-negative; (3) For other cases other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD-positive.
2. A solid tumor ctDNA MRD detection system, which is used for non-therapeutic and diagnostic purposes, characterized in that: include: A sequencing module, used for performing high-throughput targeted sequencing on a plasma sample of a subject to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations of his tissue sample and has tissue sample positive somatic mutation data; A background mutation frequency calculation module is used to calculate the background mutation frequencies of various three-base mutation types for the sequencing data of the obtained plasma samples; A mutation data analysis module, used to compare the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the positive somatic mutations in the tissue sample; A mutation simulation generation module is used to determine whether there is a mutation at the site in the plasma sample by random simulation at the corresponding site of the positive somatic mutation in the obtained tissue sample based on the background mutation frequency of the three-base mutation type; and repeat the simulation to calculate the probability p of no mutation under the simulation conditions; The MRD determination module is used to make the following determinations: (1) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is greater than or equal to 3, the sample is determined to be MRD positive; (2) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads of all mutations is less than 3, the sample is determined to be MRD negative; (3) For other situations other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD positive.
3. The solid tumor ctDNA MRD detection system according to claim 2, characterized in that: The three-base mutation type refers to the combination of 12 basic single-base mutation forms A>T, A>G, A>C, C>A, C>T, C>G, G>A, G>C, G>T, T>A, T>C, T>G and one base upstream and downstream of the reference genome at that position.
4. The solid tumor ctDNA MRD detection system according to claim 3, characterized in that: The background mutation frequency of three-base mutation types is obtained through the following steps: stacking the BAM file, reference sequence file and tissue baseline positive mutation to generate pileup format data for the specified area, filtering out common SNP sites and blacklist sites, and then filtering out sites with too small sequencing depth or too large total frequency of non-reference bases; the mutation frequency of each type of three-base mutation type is counted as the background mutation rate.
5. The solid tumor ctDNA MRD detection system according to claim 4, characterized in that: Too low a sequencing depth means less than 3-6, and too high a total frequency of non-reference bases means greater than 8-12%.
6. The solid tumor ctDNA MRD detection system according to claim 2, characterized in that: The number of simulations is 200-100000 times.
7. The solid tumor ctDNA MRD detection system according to claim 2, characterized in that: The threshold is 0.001-0.
05.
8. The solid tumor ctDNA MRD detection system according to claim 2, characterized in that: The random simulation method refers to: simulating according to binomial distribution to simulate whether there is a mutation at each site; If the simulated total average mutation frequency is greater than or equal to the actual observed average mutation frequency, and the simulated total mutation site reads are greater than or equal to the actual observed total mutation site reads, then the plasma sample is assumed to have a mutation, otherwise it is assumed to have no mutation.
9. The solid tumor ctDNA MRD detection system according to claim 8, characterized in that: In the simulation of binomial distribution, random simulation calculation is performed at the mutation site according to the sequencing depth in sequencing and the background mutation frequency of the corresponding three-base mutation type.
10. A computer readable medium, characterized in that It records a computer program for running the solid tumor ctDNA MRD detection method according to claim 1.
Citation Information
Patent Citations
Mononucleotide variation detecting method based on blood circulation tumor DNA, device and storage medium
CN110010197A
Identification and use of circulating nucleic acid tumor markers
CN113337604A
Device for detecting MRD marker based on linkage gene mutation
CN116064755A
Double background noise mutation removal method based on blood circulating tumor DNA
CN116356001A
Cancer recurrence risk assessment method and device, electronic equipment and storage medium
CN117672507A