A ctDNA-based solid tumor MRD detection method, detection system, and computer-readable medium

By calculating the p-value through the MRD qualitative algorithm and Monte-Carlo sampling method, the background noise interference in plasma free DNA detection was resolved, the accuracy and sensitivity of ctDNA MRD detection were improved, and the risk of false positives was reduced.

CN119943142BActive Publication Date: 2025-09-26GENESEEQ TECH INC +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510102066.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-09-26
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

When MRD detection is performed on plasma free DNA, background noise can cause detection errors and interference, affecting the judgment of detection signals and causing deviations in clinical practice.

Method used

The MRD qualitative algorithm was used to estimate the background error rate of single-base mutations, simulate mutation reads, and calculate the p-value of ctDNA positivity using the Monte-Carlo sampling method. MRD detection of plasma samples was performed in combination with high-throughput targeted sequencing and reference genome alignment.

Benefits of technology

It improves the accuracy of ctDNA MRD detection in solid tumors, reduces the risk of false positives, and improves the sensitivity and specificity of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943142B_ABST
    Figure CN119943142B_ABST
Patent Text Reader

Abstract

The present invention relates to a ctDNA-based solid tumor MRD detection method, detection system, and computer-readable medium, and belongs to the technical field of tumor gene detection. The present invention proposes a method for qualitatively characterizing plasma sample MRD based on tissue baseline mutation sites during dynamic plasma testing. This method calculates the background error rate of single-base mutations, simulates mutation readings based on this error rate, and compares them with the actual observed mutation readings. The number of times the simulated mutation readings are greater than the actual observed mutation readings is statistically analyzed. The p-value of the sample ctDNA positive is calculated based on multiple simulation results at the mutation site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a ctDNA-based solid tumor MRD detection method, detection system and computer-readable medium, belonging to the technical field of tumor gene detection. Background Art

[0002] MRD, often translated as minimal residual disease, primarily refers to cancer-derived cells or abnormal molecules that remain undetected by traditional imaging or laboratory methods but can be detected through liquid biopsies after radical treatment. MRD in solid tumors typically refers to molecular residual disease. Plasma ctDNA is the most mature medium for assessing MRD, and its detection has been included in multiple guidelines / consensus reports. Currently, the application value of MRD technology in oncology lies in the following three key areas:

[0003] Prognosis assessment and treatment selection: ctDNA-based MRD testing can stratify recurrence risk and assess patient prognosis. This can assist clinicians in identifying potential treatment exemptions and beneficiaries, and tailor treatment plans accordingly.

[0004] Monitor recurrence and intervene promptly: MRD can detect disease recurrence earlier than imaging and traditional clinical methods, creating valuable opportunities for patients to intervene in a timely manner.

[0005] Evaluate efficacy and adjust regimens: ctDNA-based MRD monitoring can predict and evaluate treatment efficacy early, providing a reference for discontinuation of treatment observation, regimen adjustment, and continued treatment.

[0006] MRD detection strategies can be categorized into two main categories based on whether baseline tissue testing is performed: tumor-informed and tumor-unaware strategies. Compared to tumor-unaware strategies, tumor-informed strategies can reduce the risk of false positives due to technical and biological errors during ctDNA assessment and detect ctDNA mutations with fewer reads, thereby achieving better sensitivity.

[0007] Based on the tumor-informed strategy, the choice of subsequent MRD monitoring panel can be divided into: personalized panel and fixed panel (fixed panel is also called group customized panel). Personalized panel means first testing the tumor tissue sample to understand the patient's specific genomic variation map, and then detecting certain mutation sites in the tumor tissue through ctDNA. The representative company of personalized panel testing is Natera's Signatera technology, which first detects the patient's tumor tissue sample through whole-exome WES, and screens 16 sites for subsequent MRD monitoring. Fixed panel does not rely on the patient's specific variation, but selects the same type of monitoring panel for a certain patient group, such as lung cancer, colorectal cancer, etc., which can not only detect existing mutations in the tissue, but also detect new mutations.

[0008] The following problems exist in existing MRD detection technologies: MRD status needs to be assessed based on ctDNA mutations, but not every detected mutation will be considered MRD-positive. At the technical implementation level, single-strand errors or random errors may be generated during library enrichment and PCR. Increasing sequencing depth without background purification will only amplify the interference with MRD interpretation. At the same time, clonal hematopoietic mutations will also introduce biological background noise. Misinterpreting these mutations as tumor-specific will inevitably strongly interfere with the judgment of the detection signal, thereby affecting the specificity of MRD detection and ultimately causing huge deviations in clinical practice. Summary of the Invention

[0009] The technical problem to be solved by the present invention is the interference caused by background noise when performing MRD detection on plasma free DNA. The present invention proposes a method for qualitatively testing the tissue baseline mutation sites in plasma using an MRD qualitative algorithm during dynamic plasma testing. This method estimates the background error rate of single-base mutations, simulates readings based on this error rate, and compares them with the actual observed readings. If the simulated reading is equal to or greater than the actual observed reading, the p-value for ctDNA positivity is calculated based on the number of simulated experiments. The p-value is calculated based on the Monte-Carlo sampling method, which simulates a certain number of experiments (e.g., 10,000 times) and counts the number of simulated readings that are greater than or equal to the actual observed readings, thereby calculating the p-value.

[0010] A method for detecting ctDNA MRD in solid tumors, which is used for non-therapeutic and diagnostic purposes, comprises the following steps:

[0011] Step 1: performing high-throughput targeted sequencing on a plasma sample of a subject to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations on his or her tissue sample and has tissue sample-positive somatic mutation data;

[0012] Step 2: Calculate the background mutation frequency of each three-base mutation type based on the sequencing data of the obtained plasma sample;

[0013] Step 3: aligning the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the tissue sample's positive somatic mutations;

[0014] Step 4: Based on the background mutation frequency of the three-base mutation type obtained in step 2, at the corresponding site of the tissue sample positive somatic mutation obtained in step 3, determine whether there is a mutation at the site in the plasma sample by random simulation;

[0015] Step 5: Repeat step 4 to simulate and calculate the probability p of no mutation under the simulation conditions;

[0016] In step 6, the following judgments are made: (1) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads for all mutations is greater than or equal to 3, the sample is judged to be MRD-positive; (2) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads for all mutations is less than 3, the sample is judged to be MRD-negative; (3) For other cases other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD-positive.

[0017] In step 2, the three-base mutation type refers to a combination of 12 basic single-base mutation forms A>T, A>G, A>C, C>A, C>T, C>G, G>A, G>C, G>T, T>A, T>C, T>G and one base upstream and downstream of the reference genome at that position.

[0018] The background mutation frequency of three-base mutation types is obtained by stacking the BAM file, reference sequence file, and tissue baseline positive mutations to generate pileup format data for the specified region. After filtering out common SNP sites and blacklist sites, sites with too low sequencing depth or too high total frequency of non-reference bases are filtered out. The mutation frequency of each type of three-base mutation type is calculated as the background mutation rate.

[0019] Too low a sequencing depth means less than 3-6, and too high a total frequency of non-reference bases means greater than 8-12%.

[0020] The number of simulations is 200-100000 times.

[0021] The threshold is 0.001-0.05.

[0022] The random simulation method is as follows: simulation is performed according to a binomial distribution to simulate whether there is a mutation at each site; if the total average mutation frequency of the simulation is greater than or equal to the actual observed average mutation frequency, and the total number of simulated mutation site reads is greater than or equal to the actual number of mutation site reads, then the plasma sample is assumed to have a mutation, otherwise it is assumed to have no mutation.

[0023] In the simulation of the binomial distribution, random simulation calculations are performed at the mutation site based on the sequencing depth in sequencing and the background mutation frequency of the corresponding three-base mutation type.

[0024] A solid tumor ctDNA MRD detection system, which is used for non-therapeutic and diagnostic purposes, comprising:

[0025] A sequencing module for performing high-throughput targeted sequencing on a plasma sample of a subject to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations on his or her tissue sample and has tissue sample-positive somatic mutation data;

[0026] A background mutation frequency calculation module is used to calculate the background mutation frequency of various three-base mutation types based on the sequencing data of the obtained plasma samples;

[0027] A mutation data analysis module is used to compare the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the positive somatic mutations in the tissue sample;

[0028] A mutation simulation generation module is used to determine whether a mutation exists at a site in the plasma sample based on the background mutation frequency of the obtained three-base mutation type, using a random simulation method at the corresponding site of the obtained positive somatic mutation in the tissue sample; and repeatedly simulate to calculate the probability p of no mutation under the simulation conditions;

[0029] The MRD determination module is used to make the following determinations: (1) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads for all mutations is greater than or equal to 3, the sample is determined to be MRD-positive; (2) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads for all mutations is less than 3, the sample is determined to be MRD-negative; (3) For other cases other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD-positive.

[0030] A computer-readable medium having recorded thereon a computer program capable of executing the above-mentioned solid tumor ctDNA MRD detection method.

[0031] The beneficial effects of this patent are: by calculating the background error rate of single-base mutations, simulating mutation reads based on this error rate, and comparing them with the actual observed mutation reads, the simulated mutation reads are statistically greater than the number of actually observed mutation reads. The p-value of the sample ctDNA positive is calculated by multiple simulation results at the mutation site, thereby improving the accuracy of solid tumor ctDNA MRD detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is an IGV visualization of patient samples with inconsistent MRD status before and after the improvement of this patented algorithm. DETAILED DESCRIPTION

[0033] In the detection method of this patent, the cancer type is first determined, and finally determined by analyzing 14 cancer types (lung adenocarcinoma, colorectal cancer, breast cancer, gastric cancer, lung squamous cell carcinoma, liver cancer, cervical cancer, bile duct cancer, esophageal cancer, endometrial cancer, diffuse large B-cell lymphoma, prostate cancer, skin melanoma, and brain glioma). When conducting sample analysis, the tumor sample data included are the mutation site enrichment regions of nearly 7,000 tumor WES samples obtained by actual testing by the patent applicant. On this basis, the mutation hotspots and clinical driver mutation library covering the COSMIC database are added to determine the site area for detection, and then the MRD probe is designed.

[0034] Specific sample information is shown in Table 1:

[0035] Table 1 Sample type and quantity

[0036]

[0037] After acquiring the aforementioned data, site screening begins by filtering for common pseudomutations using a baseline database of healthy human data. Common pseudomutations generally refer to variations in gene sequences that do not actually lead to altered protein function or disease phenotypes. These pseudomutations include: 1) Pseudogenes: Pseudogenes are nonfunctional DNA fragments with sequence similarities to functional genes. They may be redundant copies of functional genes and are typically identified as pseudogenes due to frameshift mutations or premature stop codons. Pseudogenes may produce low levels of RNA transcription but typically fail to produce a functional final protein product. 2) Single nucleotide polymorphisms (SNPs): A common mutation pattern throughout the human genome is called a single nucleotide polymorphism (SNP). SNPs are single base changes in DNA and often serve as genomic landmarks or "markers." However, not all SNPs lead to altered function; some may be neutral or pseudomutations. 3) Common disease-common variant hypothesis (CD-CV): This hypothesis posits that susceptibility to common complex diseases is caused by common variation at specific loci, particularly in coding or regulatory regions of genes. These common variants (with a MAF between 1% and 5%) are not necessarily pathogenic. This means that some of these common variants may not cause disease but may instead be pseudomutations. 4) Common Disease-Rare Variant Hypothesis (CD-RV): This hypothesis posits that the genetic basis of complex diseases is primarily composed of numerous rare variants with low frequencies (generally MAF <1%) and high pathogenic risk. This means that when analyzing rare variants, some pseudomutations that do not cause disease may be encountered. These pseudomutations can be obtained from a self-built database generated by the applicant through sample testing and analysis, or from public databases such as HGMD, ClinVar, and TCGA.

[0038] Furthermore, the mutation types are further limited. The mutation type screening used in the method of this patent is:

[0039] Missense mutation (missense_mutation), frameshift deletion mutation (frame_shift_del), nonsense mutation (nonsense_mutation), frameshift insertion mutation (frame_shift_ins), in-frame insertion (in_frame_ins), in-frame deletion (in_frame_del), transcription start site (translation_start_site), stop codon mutation (nonstop_mutation), silent mutation (Silent), splice site (Splice_Site).

[0040] Furthermore, the obtained sites need to be further screened to ensure that the variation at each site has appeared in at least a certain number of samples. In this embodiment, the variation sites are set to occur in at least 20 samples. The purpose of this step is to exclude some mutations that may be caused by errors in the sequencing process and avoid false positives.

[0041] Furthermore, for mutations that meet the conditions and whose hotspot coordinates are within 10 bp, regions are merged. After merging, if the merged region is less than 80 bp, it can be extended to the same length at both ends to further expand to a region of 80 bp (or 81 bp). If the length of the merged region is greater than 80 bp, the region is not extended.

[0042] The performance of each interval in the locus interval set obtained in the previous step in the pan-cancer data is then analyzed. Pan-cancer and individual cancer type counts are performed. A pan-cancer type is defined as any tumor sample carrying a mutation, and is counted per sample (e.g., if sample A has two mutations within interval α, the sample count is recorded as one).

[0043] After tracing back the coordinates of the main contributing variant sites in each interval in the previous step and merging them with the COSMIC database (V97), further microsatellite loci and fusion regions were added. For the obtained intervals, if the interval is <= 120 bp, it was expanded to a 120 bp interval; if the interval is larger than 120 bp, it was split into 120 bp regions in sequence using 120 bp as a sliding window. If the last region is less than 120, the design was carried out from the end to the front.

[0044] The microsatellite loci included here were screened in the following way:

[0045] Screening of satellite sites:

[0046] a. Whole-exome sequencing (WES) data from healthy controls and colorectal cancer MSI-positive samples (PCR-based, targeting five loci: NR-27, NR-24, NR-21, BAT-25, and BAT-26) were used, with 100 samples each, after passing quality control (sequencing depth >150X).

[0047] b. Screen single nucleotide repeat sequences with a length of 10 to 20 bp within the WES genome coverage region as candidate satellite loci.

[0048] c. Initial screening of satellite loci: screening ensures that more than 80% of the samples have a microsatellite locus coverage depth greater than 50% of the average depth of the samples themselves.

[0049] d. For the qualified satellite loci screened in the previous step, count the coverage depth and number of alleles of different lengths at each microsatellite locus for each sample, calculate the frequency of each allele, and use these as features to train a binary classification using the SVM algorithm. Select loci with an AUC greater than 0.85 as satellite loci with high MSI resolution.

[0050] The fusion genes selected here are ALK, ROS1, RET, NTRK, NRG, FGFR, MET, EGFR, HER2 and BRAF.

[0051] After screening the above sites and regions, a 120bp probe was designed. Following ctDNA extraction, library construction, library enrichment, library quantification, circularization, and DNB preparation, high-throughput sequencing was performed using the DNBSEQ-T7 sequencer. After sequencing, the raw data was subjected to quality control and bioinformatics analysis to obtain reliable biological information.

[0052] For mutations detected in tissue baseline samples, the MRD qualitative algorithm is used to perform qualitative testing of the tissue baseline mutation sites in plasma during dynamic plasma testing;

[0053] The specific steps are as follows:

[0054] MRD positivity was determined by examining whether the levels of tissue baseline positive mutations and new mutations in plasma samples were significantly higher than the background mutation rate.

[0055] The core idea is to estimate the background error rate for single-base mutations, simulate readouts based on this error rate, and compare them with the actual observed readouts. If the simulated readout is equal to or greater than the actual observed readout, the p-value for ctDNA positivity is calculated based on the number of simulated experiments. The p-value is calculated based on the Monte-Carlo sampling method, which simulates the experiment a certain number of times (e.g., 10,000 times) and counts the number of times the simulated readout is greater than or equal to the actual observed readout.

[0056] Specifically, the depth is sampled based on the background error rate, and this sampling result is then compared with the actual observed base readings to determine the results of the simulation experiment. Finally, through multiple simulation experiments, the probability of the simulated reading being equal to or greater than the actual observed reading is calculated as the p-value for ctDNA positivity.

[0057] The MRD qualitative algorithm consists of four steps:

[0058] 1) Data preprocessing

[0059] In this step, multiple input files are first read, including BAM files, reference sequence files, and lists of tissue baseline positive and emerging mutations. The samtools tool is used to perform a pileup operation, generating data in pileup format for the specified region. For each site in the read pileup, the preceding and following bases and the reference genotype for that site are extracted to form a tri-nucleotide context (TNC). Tri-nucleotide mutations are 192 combinations of the 12 basic single-base mutation forms (A>T, A>G, A>C, C>A, C>T, C>G, G>A, G>C, G>T, T>A, T>C, and T>G) with one base upstream and downstream of the reference genomic position. This is used to accurately calculate the sample TNC background.

[0060] 2) Calculation of background error rate

[0061] Background error (or background noise) refers to systematic errors in the sequencing process that often lead to incorrect mutation calls in a specific sequence context. The pileup operation in the BAM file stacks the sequencing reads at each position for analysis, evaluating and filtering sites with mutations. Calculating the background error rate mainly involves the following steps:

[0062] a. The pileup data format describes all aligned reads at a given position. For each row in the pileup, a series of filtering operations are performed using the three-base read context extracted previously.

[0063] b. This includes filtering out common population polymorphic sites and blacklisting regions and sites that are characterized by errors that occur under normal circumstances or variable sites that are known not to cause disease.

[0064] c. Filter the depth (sequencing depth or coverage, that is, the number of sequencing reads at a site). Sites with a depth less than 5 or a sum of base frequencies exceeding 10% will be filtered out. If the sequencing depth of a site is less than 5, the data for this site may be considered unreliable due to insufficient coverage. Too low coverage may make it impossible to accurately detect true variants because of random errors or insufficient sequencing. When the total frequency of non-reference bases at a site exceeds 10%, the high frequency of non-reference bases may indicate sequencing errors or sample contamination, which will also affect the accuracy of the results. Through such filtering, data sites that may cause inaccurate results due to sequencing quality issues can be removed, thereby improving the accuracy and reliability of subsequent analysis and ensuring the accuracy of the results. Too low coverage and too high non-reference base frequencies may cause inaccurate results.

[0065] For each target site, the mutation frequency is calculated. Then, a (Ref, Alt) key-value pair is generated for each possible mutation, and the frequency of each key-value pair is recorded. The (Ref, Alt) and mutation frequency in the context of three bases are placed in a hash table.

[0066] This hash table allows us to calculate the mutation frequencies of all 12 single-base mutations in 16 contexts. The frequency of each mutation type is equal to the total number of mutations of that type divided by the depth of the corresponding triplet-type base pair, reflecting the degree of error in a specific sequence context. This ratio is the calculated background error rate.

[0067] 3) Estimation of ctDNA positive p-value

[0068] Monte Carlo simulation method was used to estimate the positive probability of ctDNA:

[0069] a) When performing plasma MRD analysis, the subject has previously undergone mutation testing of tissue samples. Here, we use the binomial distribution to simulate whether each tissue baseline positive mutation in the plasma sample will mutate. The probability of success (i.e., mutation) is the background mutation rate calculated above. These generated random numbers simulate the mutation distribution we expect based on the background mutation rate. Specifically, for M mutation sites (including each tissue baseline positive mutation), we use the background mutation rate p b (calculated previously) to simulate and determine whether each site will mutate. For each mutation site i, the sequencing depth of the mutation site is D i , where i=1,2,...M, becomes a random variable X i Indicates the number of reads with mutations at site i:

[0070]

[0071] If site i mutates, X i > 0, otherwise X i <0.

[0072] b) If the number of simulated mutation site reads is greater than or equal to the number of actually observed mutation site reads, and the simulated average mutation frequency is greater than or equal to the actually observed average mutation frequency, the plasma sample is considered to have a mutation. Otherwise, it is considered to have no mutation. Specifically, calculate the total number of reads S for sites where mutations occur in the simulation:

[0073]

[0074] Let S obs is the number of reads at the site where the mutation actually occurred, f obs is the average mutation frequency of the M mutation sites actually observed. For the simulated data, the average mutation frequency is:

[0075]

[0076] We compare the simulation results with the observed data. For each simulation, if , then I = 1, otherwise I = 0. I is the result of the above comparison.

[0077] c) This process is repeated multiple times (the default is 10,000 times) to stabilize the results and generate a p-value. The p-value can be interpreted as the probability that no mutation has occurred. In other words, the smaller the p-value, the greater the likelihood of a mutation. At this point, there is greater confidence that these mutations are not background noise but rather true mutation signals, indicating the presence of ctDNA in the plasma sample. Specifically, by repeating the simulation N times, we can estimate the p-value (i.e., the probability, under the null hypothesis, that the observed mutation frequency and number of mutation sites exceed the actual observed values):

[0078]

[0079] in is the comparison between the j-th simulation result and the observed data.

[0080] In this way, it is possible to assess whether the observed mutations are statistically significant and thus determine the probability of ctDNA positivity.

[0081] 4) MRD qualitative analysis

[0082] The final steps for the qualitative analysis of MRD in samples are:

[0083] a. Check whether the plasma mutations contain the tissue baseline positive mutations. If so, the sample is considered MRD-positive (meaning that when the patient's plasma MRD test is performed, the same mutations that have been found in the patient's tissue can still be found).

[0084] b. To ensure the validity of the Monte Carlo simulation, a sample is considered MRD-negative if the sum of the support counts for baseline positive mutations across all tissues is less than 3. This means that only the mutation sites in the corresponding tissues are considered. If the sum of the number of reads for mutations at all these sites in plasma is less than 3, the sample is considered MRD-negative.

[0085] In other cases, MRD is qualitatively determined based on the p-values ​​obtained in the simulation above. We use p < 0.01 as the threshold for MRD positivity. The final result generates a detailed mutation report, including information such as mutation site, mutation type, mutation frequency, and background error rate.

[0086] This algorithm provides an effective method for calculating ctDNA positivity, which draws on the Monte-Carlo method in statistics and combines the concept of background error rate in genomics. It can be effectively applied to the ctDNA positivity analysis of clinical samples.

[0087] Example 1 Testing the degree of improvement of the sensitivity of MRD analysis by the MRD qualitative algorithm of the present invention

[0088] To test the extent to which the algorithm upgrade improves the sensitivity of MRD analysis, 78 untreated patient samples pathologically diagnosed with advanced solid tumors (theoretically, all had positive MRD status) were randomly screened. Each patient's baseline tumor tissue sample, white blood cell control sample, and baseline plasma sample were sequenced using an MRD probe kit. The MRD status of the plasma samples was assessed using the pre-upgrade MRD analysis process (only MRD qualitative step a) and the post-upgrade MRD analysis process (including MRD qualitative steps a and b). The comparative analysis results are as follows: after adding the above-mentioned MRD qualitative step b, the number of MRD-positive samples increased by 3, and the sensitivity increased by 3.8%. The specific comparison results are shown in Table 2:

[0089] Table 2 Comparison of MRD analysis sensitivity

[0090]

[0091] One of the patient samples, such as Figure 1As shown, when tissue baseline variation is detected at multiple sites in plasma, but none of them meet the criteria for single-site positivity, the sample would be MRD-negative based solely on the criteria of MRD qualitative step a. However, with the addition of MRD qualitative step b, the sample can be accurately determined to be MRD-positive. This demonstrates that the addition of MRD qualitative step b significantly improves the sensitivity of MRD analysis.

[0092] Considering that the primary clinical application scenarios of MRD testing also include perioperative recurrence risk monitoring in patients with postoperative disease and monitoring tumor recurrence after radical chemoradiotherapy in patients with locally advanced, inoperable disease, 68 additional samples from patients with stage I to III disease were tested for sensitivity. The results of the comparative analysis before and after the upgrade of the MRD qualitative step are as follows. After adding the aforementioned MRD qualitative step b, one additional MRD-positive case was found in a stage I patient and one additional MRD-positive case was found in a stage III patient, for a total of two additional MRD-positive samples, resulting in a 2.9% improvement in sensitivity. The specific comparative results are shown in Table 3:

[0093] Table 3 Comparison of MRD analysis sensitivity

[0094]

[0095] From the above data, it can be seen that adding the MRD qualitative step b can improve the overall MRD analysis sensitivity, whether it is an advanced tumor or a stage I to III tumor.

[0096] Example 2 Testing the Effect of the MRD Qualitative Algorithm of the Present Invention on the Specificity of MRD Analysis

[0097] To test the impact of the upgraded MRD analysis algorithm on analytical specificity, we selected 100 healthy human plasma samples for sequencing using an MRD probe kit. Plasma samples from healthy subjects are theoretically free of tumor somatic mutations. We combined the plasma sequencing data from each healthy individual with the mutation lists of 78 patients with advanced tumors. This combined data from 100 healthy individuals × 78 tumor mutation lists yielded 7,800 samples. MRD status was determined using the pre- and post-upgrade analysis processes for the MRD qualitative step. The comparative analysis results of the two analysis processes are shown in Table 4.

[0098] Table 4 Comparison of MRD analysis sensitivity

[0099]

[0100] Adding the Monte Carlo model to the sampling analysis improved the specificity of MRD analysis from 81.9% to 98.2%, a 16.3% increase. This shows that adding the Monte Carlo model to the sampling analysis step not only improves sensitivity but also specificity.

Claims

1. A method for detecting ctDNA MRD in solid tumors, which is used for non-therapeutic and diagnostic purposes, characterized in that: The steps include: Step 1: performing high-throughput targeted sequencing on a plasma sample of a subject to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations on his or her tissue sample and has tissue sample-positive somatic mutation data; Step 2: For the sequencing data of the obtained plasma samples, calculate the background mutation rate p of each type of three-base mutation b ; Step 3: aligning the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the tissue sample's positive somatic mutations; Step 4: Based on the background mutation rate of the three-base mutation type obtained in Step 2, at the corresponding site of the positive somatic mutation in the tissue sample obtained in Step 3, determine whether there is a mutation at the site in the plasma sample by random simulation; Step 5: Repeat step 4 to simulate and calculate the probability p of no mutation under the simulation conditions; Step 6: Make the following judgments: (1) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads for all mutations is greater than or equal to 3, the sample is judged to be MRD-positive; (2) If the plasma mutations found in high-throughput targeted sequencing include tissue mutations, and the number of reads for all mutations is less than 3, the sample is judged to be MRD-negative; (3) For other cases other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD-positive; In step 2, the background mutation rate of the three-base mutation type is obtained by stacking the BAM file, reference sequence file, and tissue baseline positive mutations to generate pileup format data for the specified region, filtering out common SNP sites and blacklist sites, and then filtering out sites with too low sequencing depth or too high total frequency of non-reference bases; The mutation frequency of each type of three-base mutation is calculated as the background mutation rate p b ; Too low sequencing depth means less than 6, and too high total frequency of non-reference bases means greater than 8%; In step 4, the random simulation method refers to: For M mutation sites, use the background mutation rate p b Simulate, for each mutation site i, the sequencing depth of the mutation site is D i , where i=1,2,...M, generates a random variable X i , X i Indicates the number of reads with mutation at site i, and defines X i ~Binomial(D i ,p b ), if site i mutates, X i >0, otherwise X i =0; Calculate the total number of reads S of sites where mutations occurred in the simulation, ,in is the indicator function, when X i >0, its value is 1. i =0, its value is 0; Let S obs is the number of reads at the site where the mutation actually occurred, f obs is the average mutation frequency of the M mutation sites actually observed. For the simulated data, the average mutation frequency is: ; For each simulation, if S ≥ S obs , and f≥f obs , then I=1, otherwise I=0, I is the result of a simulation comparison; In step 5, the probability p is obtained by the following formula: ; Where I(j) is the comparison between the j-th simulation result and the observed data, and N is the number of repeated simulations.

2. The method for detecting ctDNA MRD in solid tumors according to claim 1, wherein: The three-base mutation type refers to the combination of 12 basic single-base mutation forms A>T, A>G, A>C, C>A, C>T, C>G, G>A, G>C, G>T, T>A, T>C, T>G and one base upstream and downstream of the reference genome of the single-base mutation position.

3. The method for detecting ctDNA MRD in solid tumors according to claim 1, wherein: The number of simulations is 200-100000 times.

4. The method for detecting ctDNA MRD in solid tumors according to claim 1, wherein The threshold is 0.001-0.

05.

5. A solid tumor ctDNA MRD detection system for non-therapeutic and diagnostic purposes, characterized in that: include: A sequencing module for performing high-throughput targeted sequencing on a plasma sample of a subject to obtain sample sequencing data; the subject has undergone high-throughput targeted sequencing of somatic mutations on his or her tissue sample and has tissue sample-positive somatic mutation data; The background mutation rate calculation module is used to calculate the background mutation rate p of various three-base mutation types for the sequencing data of the obtained plasma samples. b ; A mutation data analysis module is used to compare the obtained sample sequencing data with the reference genome to obtain mutation data at the corresponding sites of the positive somatic mutations in the tissue sample; A mutation simulation generation module is used to determine whether a mutation exists at a site in the plasma sample based on the background mutation rate of the obtained three-base mutation type, using a random simulation method at the corresponding site of the positive somatic mutation in the obtained tissue sample; the simulation is repeated to calculate the probability p of no mutation under the simulation conditions; The MRD determination module is used to make the following determinations: (1) If the plasma mutations found in high-throughput targeted sequencing contain tissue mutations, and the number of reads for all mutations is greater than or equal to 3, the sample is determined to be MRD-positive; (2) If the plasma mutations found in high-throughput targeted sequencing contain tissue mutations, and the number of reads for all mutations is less than 3, the sample is determined to be MRD-negative; (3) For other cases other than (1) and (2), if p < threshold, the plasma sample is considered to be MRD-positive; The background mutation rate of the three-base mutation type was obtained by stacking the BAM file, reference sequence file, and tissue baseline positive mutations to generate pileup format data for the specified region. After filtering out common SNP sites and blacklist sites, sites with too low sequencing depth or too high total frequency of non-reference bases were filtered out. The mutation frequency of each type of three-base mutation is calculated as the background mutation rate p b ; Too low sequencing depth means less than 6, and too high total frequency of non-reference bases means greater than 8%; In the mutation simulation generation module, the random simulation method refers to: For M mutation sites, use the background mutation rate p b Simulate, for each mutation site i, the sequencing depth of the mutation site is D i , where i=1,2,...M, generates a random variable X i , X i Indicates the number of reads with mutation at site i, and defines X i ~Binomial(D i ,p b ), if site i mutates, X i >0, otherwise X i =0; Calculate the total number of reads S of sites where mutations occurred in the simulation, ,in is the indicator function, when X i >0, its value is 1. i =0, its value is 0; Let S obs is the number of reads at the site where the mutation actually occurred, f obs is the average mutation frequency of the M mutation sites actually observed. For the simulated data, the average mutation frequency is: ; For each simulation, if S ≥ S obs , and f≥f obs , then I=1, otherwise I=0, I is the result of a simulation comparison; In the MRD judgment module, the probability p is obtained by the following formula: ; Where I(j) is the comparison between the j-th simulation result and the observed data, and N is the number of repeated simulations.

6. The solid tumor ctDNA MRD detection system according to claim 5, characterized in that The three-base mutation type refers to the combination of 12 basic single-base mutation forms A>T, A>G, A>C, C>A, C>T, C>G, G>A, G>C, G>T, T>A, T>C, T>G and one base upstream and downstream of the reference genome of the single-base mutation position.

7. The solid tumor ctDNA MRD detection system according to claim 5, characterized in that The number of simulations is 200-100000 times.

8. The solid tumor ctDNA MRD detection system according to claim 5, characterized in that The threshold is 0.001-0.

05.

9. A computer-readable medium, characterized in that It describes a computer program for running the solid tumor ctDNA MRD detection method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Mononucleotide variation detecting method based on blood circulation tumor DNA, device and storage medium

    CN110010197A

  • Double background noise mutation removal method based on blood circulating tumor DNA

    CN116356001A