Micro residual focus monitoring method and system based on circulating tumor DNA

By employing a dual-panel collaborative detection architecture and UMI marker error correction technology, the insufficient sensitivity and specificity of UTUC MRD detection are addressed, achieving high coverage and low false-positive detection of UTUC, and providing a clinically valuable MRD monitoring solution.

CN121575110APending Publication Date: 2026-02-27MAIYUE BIOTECHNOLOGY (SUZHOU) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202512002485.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing MRD detection methods lack sufficient sensitivity and specificity for upper urothelial carcinoma (UTUC), especially in the Chinese population where they lack personalized design. Traditional fixed panels cannot cover important gene loci, resulting in high false positive rates and test results that do not have clinical guidance value.

Method used

Employing a dual-panel collaborative detection architecture, combining UTUC-specific fixed panel design, vacuum concentration process optimization, and UMI labeling and error correction technology, this study monitors minimal residual disease lesions of circulating tumor DNA through a combination of personalized and fixed panels. This includes the collaborative analysis of personalized panel target sets and leukocyte genomic DNA, eliminating clonal hematopoietic background noise.

Benefits of technology

It improves the sensitivity and specificity of MRD testing, reduces the risk of false positives, provides test results with greater clinical guidance value, and ensures high coverage and cost-effectiveness for UTUC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121575110A_ABST
    Figure CN121575110A_ABST
Patent Text Reader

Abstract

The invention discloses a high-specificity minimal residual disease (MRD) monitoring method and system based on circulating tumor DNA (ctDNA). The method comprises the following steps: receiving tumor tissue sequencing data of an UTUC patient, and generating a double-Panel target list containing personalized and fixed Panel; respectively extracting plasma cfDNA and leukocyte gDNA; performing vacuum concentration, hybrid capture and sequencing on the cfDNA library by using the double Panel lists, and performing deep sequencing on the leukocyte gDNA; constructing an individualized clonal hematopoietic mutation filtering database; actively filtering and rejecting clonal hematopoietic background mutation by utilizing a filtering database; and calculating an MRD load score based on the filtered tumor-derived mutation and outputting a report. The system comprises corresponding modules which are used for automatically executing the process. According to the invention, through cooperation of four major technologies of double-Panel design, process optimization, UMI error correction and active clonal hematopoietic filtration, MRD monitoring with extremely high sensitivity and specificity on UTUC is realized, false positive is significantly reduced, and the kit has drug resistance early warning potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular detection technology, specifically relating to a method and system for monitoring minimal residual disease lesions based on circulating tumor DNA. Background Technology

[0002] Minimal residual disease (MRD) refers to a small number of tumor cells remaining after cancer treatment and is a major risk factor for tumor recurrence. Currently, MRD monitoring is often performed through imaging or traditional molecular testing, but these methods have low sensitivity and poor specificity, easily leading to missed detections or misdiagnosis.

[0003] Liquid biopsy technology based on circulating tumor DNA (ctDNA) offers advantages such as being non-invasive and highly sensitive, providing a new solution for MRD monitoring. However, due to the extremely low concentration and severe fragmentation of ctDNA in blood, as well as the significant differences in mutation profiles among individuals, existing technologies face major challenges: traditional fixed-gene panels cannot cover patient-specific mutations, resulting in limited sensitivity; background noise such as sequencing errors and clonal hematopoiesis can easily lead to false positives; single detection strategies are difficult to adapt to the heterogeneity of different tumor types; existing MRD monitoring protocols lack optimized designs specifically for UTUC, especially failing to incorporate the mutational characteristics of the Chinese UTUC population, resulting in insufficient detection sensitivity and clinical relevance for this specific cancer type.

[0004] In existing technologies, MRD detection mostly employs fixed gene panels or tumor informed consent analysis strategies, but lacks a systematic solution that combines personalized testing with standardized analysis. Therefore, there is an urgent need in this field for an MRD monitoring protocol specifically designed for UTUC that balances high sensitivity, high specificity, and cost-effectiveness. Furthermore, genomic studies of UTUC in the Chinese population have shown that, in addition to FGFR3, genes such as HRAS, TP53, and KDM6A also exhibit characteristic high-frequency mutations. However, existing general or fixed panels based on other cancer types often fail to adequately cover these clinically significant gene loci in the Chinese UTUC population. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for monitoring minimal residual disease (MRD) based on circulating tumor DNA. Through a dual-panel collaborative detection architecture, combined with UTUC-specific fixed panel design, vacuum concentration process optimization, and UMI labeling and error correction technology, it achieves highly sensitive and specific MRD monitoring for upper urothelial carcinoma.

[0006] To achieve the above objectives, on the one hand, the present invention provides a method for monitoring minimal residual disease based on circulating tumor DNA, specifically for monitoring upper urothelial carcinoma (UTUC), comprising the following steps:

[0007] Receive whole-exome sequencing data of primary tumor tissue from UTUC patients;

[0008] Based on the whole exome sequencing data, the processor automatically screens out multiple high-frequency somatic mutation sites to form a personalized panel target set.

[0009] The personalized panel target set is combined with a predefined fixed panel target set for UTUC to generate a dual-panel target list.

[0010] Cell-free plasma DNA (cfDNA) and leukocyte genomic DNA (gDNA) were extracted from peripheral blood samples of the same subject.

[0011] Sequencing libraries were constructed based on extracted cell-free circulating DNA (cfDNA) from plasma, and sequencing library solutions were obtained.

[0012] The constructed sequencing library solution was then concentrated under vacuum.

[0013] The hybridization reaction solution containing capture probes based on the dual-panel target list is mixed with the concentrated sequencing library to perform a hybridization capture reaction;

[0014] The products after hybridization capture were sequenced to obtain plasma cfDNA sequencing data.

[0015] The leukocyte genomic DNA (gDNA) was sequenced using a dual-panel target list to obtain leukocyte control sequencing data;

[0016] Bioinformatics analysis was performed on the leukocyte control sequencing data to screen out somatic mutations with allele frequencies not lower than a predetermined threshold, and a personalized clonal hematopoietic mutation filtering database was constructed.

[0017] Bioinformatics analysis was performed on the plasma cfDNA sequencing data, including error correction and mutation detection based on unique molecular identifiers;

[0018] When analyzing the plasma cfDNA sequencing data, all detected mutations are compared with the filtering database, and mutations present in the filtering database are removed.

[0019] The MRD burden score was calculated based on the median allele frequency of all selected and removed mutations.

[0020] The output includes an MRD load report containing the MRD load score.

[0021] On the other hand, the present invention provides a minimal residual disease (MRD) monitoring system based on circulating tumor DNA for monitoring upper urothelial carcinoma (UTUC), comprising:

[0022] The sample and data input interface is used to receive whole-exome sequencing data of primary tumor tissue from UTUC patients, as well as peripheral blood samples from the same subject.

[0023] The sample processing and nucleic acid extraction module is connected to the sample and data input interface and is used to extract plasma circulating cell-free DNA (cfDNA) and leukocyte genomic DNA (gDNA) from peripheral blood samples, respectively.

[0024] The personalized panel customization module is connected to the sample and data input interface and is configured to automatically screen high-frequency somatic mutation sites based on the received tumor tissue sequencing data to form a personalized panel target set, which is then combined with a predefined fixed panel target set for UTUC to generate a dual-panel target list.

[0025] The library construction and processing unit, connected to the sample processing and nucleic acid extraction unit and the personalized panel customization unit, is configured as follows:

[0026] Sequencing libraries were constructed based on the circulating cell-free DNA (cfDNA) from plasma and the genomic DNA (gDNA) from leukocytes, respectively. The sequencing library solution constructed based on circulating cell-free DNA (cfDNA) from plasma was concentrated under vacuum. The hybridization reaction solution containing capture probes based on the dual-panel target list was mixed with the concentrated sequencing library to perform a hybridization capture reaction.

[0027] The sequencing unit, connected to the library construction and processing unit, is configured to sequence the products after hybridization capture to obtain plasma cfDNA sequencing data; and to sequence the leukocyte genomic DNA (gDNA) through the dual-panel target list to obtain leukocyte control sequencing data.

[0028] A bioinformatics integration analysis engine, connected to the sequencing unit, is configured to execute:

[0029] Bioinformatics analysis was performed on the leukocyte control sequencing data to screen out somatic mutations with allele frequencies not lower than a predetermined threshold, and a personalized clonal hematopoietic mutation filtering database was constructed. Bioinformatics analysis was performed on the plasma cfDNA sequencing data, including error correction based on unique molecular identifiers and mutation detection. All detected plasma cfDNA mutations were compared with the filtering database, and mutations present in the filtering database were removed. Based on the median allele frequencies of all the removed mutations, the MRD burden score was calculated.

[0030] A report generator, connected to the bioinformatics integration analysis engine, is used to output an MRD load report containing the MRD load score.

[0031] The beneficial effects of this invention are as follows: the dual-panel collaborative design of "personalized panel and UTUC-specific fixed panel" not only ensures high sensitivity in tracking patient-specific mutations, but also covers high-frequency hotspots and important drug resistance sites in the population through the fixed panel, thus avoiding missed detections.

[0032] This invention elevates paired leukocyte sequencing from a traditional auxiliary verification role to an indispensable pre-filtering step in the MRD detection process. This proactive exclusion method fundamentally improves the specificity of MRD detection within the low allele frequency range, reduces the risk of false positives due to clonal evolution of the hematopoietic system, and makes the test results more clinically valuable. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a flowchart of the method of the present invention.

[0035] Figure 2 This is a block diagram of the system architecture of the present invention.

[0036] Figure 3 Data was filtered for individualized clonal hematopoietic mutations in patient Chen.

[0037] Figure 4 Data was filtered for individualized clonal hematopoietic mutations in patient Zhao. Detailed Implementation

[0038] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0039] like Figure 1 As shown, this invention provides a method for monitoring minimal residual disease based on circulating tumor DNA, specifically for monitoring upper urothelial carcinoma (UTUC), comprising the following steps:

[0040] S1: Receive whole-exome sequencing data of primary tumor tissue from UTUC patients;

[0041] S2: Based on the whole-exome sequencing data, the processor automatically screens out multiple high-frequency somatic mutation sites to form a personalized panel target set. The screening criteria for these high-frequency somatic mutation sites can be as follows: based on the whole-exome sequencing data, using significance analysis tools such as MutSigCV (q value < 0.1), and selecting the top 40 somatic single nucleotide variants (SNVs) and small insertions / deletions (Indels) in terms of allele frequency. This criterion ensures that the selected personalized panel targets have high tumor abundance and representativeness.

[0042] S3: Combine the personalized panel target set with a predefined fixed panel target set for UTUC to generate a dual-panel target list.

[0043] The predefined fixed panel target set for UTUC was designed based on a meta-analysis of whole-exome data from a UTUC patient cohort (n=550). This set covers at least the following core genes associated with the occurrence, development, prognosis, and targeted therapy resistance of UTUC: FGFR3 (including known resistance sites such as p.V555M and p.Y373C), HRAS (common activating mutation sites), TP53 (inactivation mutation hotspots), and KDM6A (high-frequency inactivation mutations). The fixed panel contains 50 specific sites, ensuring effective monitoring of key molecular events related to UTUC even with unknown tumor mutation profiles.

[0044] S4: Collect circulating cell-free DNA (cfDNA) and leukocyte genomic DNA (gDNA) from peripheral blood samples of the same subject.

[0045] The extraction of leukocyte genomic DNA can be performed using the conventional magnetic bead method in the art. For example, using a commercially available blood genomic DNA extraction kit, the simplified steps include: mixing the blood sample with lysis buffer and incubating to release nucleic acids, then adding magnetic beads and binding buffer to capture DNA, washing repeatedly to remove impurities, and finally eluting to obtain purified leukocyte genomic DNA.

[0046] S5 constructs a sequencing library based on extracted plasma circulating cell-free DNA (cfDNA) to obtain a sequencing library solution;

[0047] S6: Concentrate the constructed sequencing library solution under vacuum;

[0048] S7: Mix the hybridization reaction solution containing capture probes based on the dual-panel target list with the concentrated sequencing library to perform a hybridization capture reaction;

[0049] S8: Sequencing the products after hybridization capture to obtain plasma cfDNA sequencing data; sequencing the leukocyte genomic DNA (gDNA) using a dual-panel target list to obtain leukocyte control sequencing data;

[0050] S9: Perform bioinformatics analysis on the white blood cell control sequencing data to screen out somatic mutations with allele frequencies not lower than a predetermined threshold and construct an individualized clonal hematopoietic mutation filtering database; perform bioinformatics analysis on the plasma cfDNA sequencing data, including error correction processing and mutation detection based on unique molecular identifiers;

[0051] Using the same dual-panel target list as the cfDNA assay, ultra-high-depth sequencing (UHMR) of the leukocyte gDNA library was performed at a depth of at least 100,000×. The resulting data were analyzed to identify all somatic mutations with a VAF of at least 0.1%. These mutations were considered to be patient-specific clonal hematopoietic backgrounds and were entered into a personalized clonal hematopoietic mutation filtering database. This database served as a pre-filter for subsequent analyses.

[0052] For example, ultra-high-depth sequencing (UHMS) of leukocyte gDNA at a depth of at least 100,000× was performed on the two subjects (patient Zhao and patient Chen). After analysis, all somatic mutations with a VAF of at least 0.1% were screened out, and individualized clonal hematopoietic mutation background databases were constructed for each. The results are shown in Table 1 (see Appendix). Figure 3 The individualized clonal hematopoietic mutation filtering data of patient Chen shown in Table 2 and Table 2 (see Appendix) Figure 4 The individualized clonal hematopoietic mutation filtering data of patient Zhao is shown in the figure.

[0053] As shown in Tables 1 and 2, there are significant differences in the gene spectrum and abundance of clonal hematopoietic mutations carried by different individuals, highlighting the limitations of using universal thresholds or population databases for filtering. This invention, by establishing a dedicated filtering database for each individual, can accurately identify and remove their unique background noise. For example, patient Zhao's background mutations include sites such as DUSP5 p.R53Q (0.10%) and TWIST1 p.R44C (0.19%), while patient Chen carries distinctly different mutations such as ZNF777 p.P619S (0.17%) and POLRMT p.R831H (0.17%). In subsequent cfDNA analyses targeting their respective individuals, these sites will be effectively filtered out, thereby avoiding false positives due to individual differences and significantly improving the specificity of tumor-derived mutation detection.

[0054] S10: When analyzing the plasma cfDNA sequencing data, all detected mutations are compared with the filtering database, and mutations present in the filtering database are removed.

[0055] For cfDNA sequencing data, UMI error correction is first performed. All mutations initially detected after error correction are immediately compared with the above-mentioned filtering database. Any mutation that is a perfect match (same genomic location and base change) found in the database is identified as a non-tumor-derived clonal hematopoietic mutation and removed from the results list.

[0056] S11: Calculate the MRD burden score based on the median allele frequency of all selected and removed mutations;

[0057] S12: Output an MRD load report containing the MRD load score.

[0058] The capture probe is selected from:

[0059] (1) The length is preferably 80-120 bp, and more preferably 100-120 bp. This is because this length can better balance hybridization kinetics and specificity. If it is too long, it is easy to introduce non-specific binding, and if it is too short, the hybridization efficiency is insufficient. Within this length range, the probe can ensure specific binding to the target sequence and has good hybridization efficiency. For example, probes with lengths of 100 bp, 110 bp or 120 bp can be designed.

[0060] (2) The GC content should be controlled between 40% and 60%, preferably between 45% and 55%. This range can prevent the probe secondary structure from becoming too complex due to excessively high GC content, and can also prevent the hybridization stability from being affected by excessively low GC content. In practice, the GC content of the probe sequence can be calculated and optimized by software.

[0061] (3) The sequence similarity with non-target regions of the normal human genome should not exceed 85%, preferably not exceeding 80%. Homology comparison should be performed using a BLAST tool with default parameter settings to ensure probe specificity. For example, probe sequences with more than 85% similarity to other regions of the human genome can be excluded.

[0062] The vacuum concentration is performed at 45-50℃ and 1000-2000 rpm for 20-40 minutes until the solution is dry or nearly dry. The specific steps are as follows: Take 500-1000 ng of the constructed sequencing library and add it to a PCR tube; place the PCR tube in a vacuum concentration centrifuge and open the cap; set the temperature to 45-50℃, the rotation speed to 1000-2000 rpm, and the time to 20-40 minutes; observe the solution state until it is completely dry or nearly dry; immediately proceed with the subsequent hybridization reaction, avoiding prolonged storage. This significantly reduces the loss of cfDNA fragments and increases the concentration of reactants captured in subsequent hybridization, thereby improving the capture efficiency for low-frequency mutations. The capture probes include fixed panel probes and personalized panel probes, with a molar ratio of 1:1 to 5:1. The probe mixing ratio is accurately determined using qPCR quantification to ensure capture uniformity.

[0063] The hybridization reaction solution contains the following components: 1%-4% dextran sulfate, preferably 2-3%; 1-8% formamide, preferably 4-6%; and 0.2-2M sodium chloride, preferably 0.5-1M. The hybridization reaction solution is added to a dried library, vortexed for 30 seconds to ensure the library at the bottom of the tube dissolves, briefly centrifuged, and transferred to a PCR tube. The PCR instrument parameters are set as follows: place the hybridization reaction solution on the PCR instrument and run the program at 85 degrees Celsius for 5 minutes (95 degrees Celsius with a hot cap); maintain at 60 degrees Celsius for 12-18 hours. Subsequent library capture, amplification, and purification are then performed.

[0064] The hybridization capture reaction is carried out at 65°C for 12-18 hours, preferably 14-16 hours. Temperature is kept stable during the hybridization process to avoid fluctuations that could affect hybridization efficiency.

[0065] The hybridization capture reaction is followed by a washing step, the specific operation of which is as follows:

[0066] Use a washing buffer preheated to 60-65°C, which contains:

[0067] 0.1-1X SSPE, preferably 0.5X SSPE; where “0.1-1X” indicates the working concentration multiple, and SSPE is a saline buffer.

[0068] 0.005-0.2% N-lauroyl sarcosine sodium, preferably 0.01-0.1%;

[0069] The washing process includes:

[0070] First wash: Use 150 μL of preheated wash buffer and incubate at 60°C for 10 minutes;

[0071] Second wash: Use 150 μL of preheated wash buffer and incubate at 60°C for 10 minutes;

[0072] Third wash: Use 150 μL of preheated wash buffer and incubate at 60°C for 10 minutes.

[0073] The error correction process based on unique molecular identifiers includes the following steps:

[0074] (1) Cluster the sequencing reads according to their unique molecular identifier sequences, allowing for one base mismatch. Use an algorithm that combines exact matching and fuzzy matching to group reads with the same UMI sequence into the same cluster;

[0075] (2) Generate a consensus sequence for each unique molecular identifier cluster, requiring that the number of reads within the cluster is ≥3 and the majority base consistency is ≥80%. Clusters with fewer than 3 reads are discarded; clusters with base consistency below 80% are subject to further quality control.

[0076] (3) Screen for mutations with a mutation allele frequency ≥0.04% that exist in at least two independent and unique molecular identifier clusters. This dual verification mechanism effectively eliminates sequencing errors and false positives.

[0077] The bioinformatics analysis and screening of the fixed panel target set for UTUC are as follows:

[0078] Bioinformatics analysis was used to screen for genes with somatic mutations in this population, resulting in a broad initial candidate gene set. This set included genes such as TTC40, MUC19, KMT2D, TP53, TTN, GPR98, SYNE1, DYNC1H1, AHNAK2, DNAH2, MUC16, DNAH8, KMT2C, TERT, CSMD3, ARID1A, FGFR3, TRRAP, DNHD1, FAT3, LRP1B, RB1, RYR1, HSPG2, HMCN1, USH2A, NEB, PIK3CA, ZFHX4, PLEC, ABCA2, HERC1, CREBBP, HYDIN, RYR2, LRP2, CSMD1, LAMA5, KDM6A, TSC2, DNAH11, SPRED2, CELSR3, DST, and SPTBN5. Genes including ACACB, MACF1, HERC2, OBSCN, and MUC5B.

[0079] The fixed panel target set for UTUC includes multiple gene loci selected from the TERT, FGFR3, TP53, KDM6A, PIK3CA, HRAS, KMT2D, KMT2C, ARID1A, CREBBP, BLM, BRAF, BRCA2, EGFR, ERBB3, and FLNC genes.

[0080] The fixed panel target set consists of sites in the TERT, FGFR3, TP53, KDM6A, PIK3CA, HRAS, KMT2D, KMT2C, ARID1A, CREBBP, BLM, BRAF, BRCA2, EGFR, ERBB3, and FLNC genes.

[0081] The initial candidate gene set was evaluated in multiple dimensions, and the evaluation indicators included:

[0082] Incidence rate in the Chinese UTUC population;

[0083] Whether it is a known driver gene (such as TERT, FGFR3, TP53), a gene related to targeted therapy / immunotherapy (such as FGFR3, PIK3CA), or a gene associated with poor prognosis;

[0084] Whether it focuses on key oncogenic pathways (such as chromatin remodeling, RTK / RAS pathway, etc.).

[0085] Gene size and sequencing coverage difficulty;

[0086] The balance between the number of probes required to incorporate genes into a panel and the increased detection sensitivity.

[0087] Based on the comprehensive evaluation above, we selected a subset of genes that achieved the best balance between population coverage (>95%), clinical actionability (including driver and resistance sites), and testing cost. The final fixed panel target set for UTUC includes the following genes: TERT, FGFR3, TP53, KDM6A, PIK3CA, HRAS, KMT2D, KMT2C, ARID1A, CREBBP, BLM, BRAF, BRCA2, EGFR, ERBB3, and FLNC. This 16-gene combination was retrospectively validated in our validation cohort (N=84), achieving over 95% mutation coverage while keeping the panel size within a reasonable range for cost-effective testing.

[0088] From this fixed panel, we further identified core genes with extremely high characteristic and clinical value for UTUC in China, including TERT (promoter region), FGFR3 (containing drug resistance sites such as V555M and Y373C) and KDM6A, and used them as key targets for monitoring and early warning.

[0089] like Figure 2 As shown, the present invention proposes a monitoring system for implementing the above method, comprising:

[0090] The sample and data input interface is used to receive whole-exome sequencing data of primary tumor tissue from UTUC patients, as well as peripheral blood samples from the same subject.

[0091] The sample processing and nucleic acid extraction module is connected to the sample and data input interface and is used to extract plasma circulating cell-free DNA (cfDNA) and leukocyte genomic DNA (gDNA) from peripheral blood samples, respectively.

[0092] The personalized panel customization module is connected to the sample and data input interface and is configured to automatically screen high-frequency somatic mutation sites based on the received tumor tissue sequencing data to form a personalized panel target set, which is then combined with a predefined fixed panel target set for UTUC to generate a dual-panel target list.

[0093] The library construction and processing unit, connected to the sample processing and nucleic acid extraction unit and the personalized panel customization unit, is configured as follows:

[0094] (a) Sequencing libraries were constructed based on the circulating cell-free DNA (cfDNA) in plasma and the genomic DNA (gDNA) in leukocytes, respectively;

[0095] (b) Vacuum concentration of sequencing library solutions constructed based on circulating cell-free DNA (cfDNA) in plasma;

[0096] (c) Mix the hybridization reaction solution containing capture probes based on the dual-panel target list with the concentrated sequencing library to perform a hybridization capture reaction;

[0097] The sequencing unit, connected to the library construction and processing unit, is configured to sequence the products after hybridization capture to obtain plasma cfDNA sequencing data; and to sequence the leukocyte genomic DNA (gDNA) through the dual-panel target list to obtain leukocyte control sequencing data.

[0098] A bioinformatics integration analysis engine, connected to the sequencing unit, is configured to execute:

[0099] (a) Perform bioinformatics analysis on the leukocyte control sequencing data to screen out somatic mutations with allele frequencies not lower than a predetermined threshold and construct an individualized clonal hematopoietic mutation filtering database;

[0100] (b) The plasma cfDNA sequencing data were subjected to bioinformatics analysis, including error correction and mutation detection based on unique molecular identifiers;

[0101] (c) Compare all detected plasma cfDNA mutations with the filtering database and remove mutations present in the filtering database;

[0102] (d) Calculate the MRD burden score based on the median allele frequency of all selected and removed mutations;

[0103] A report generator, connected to the bioinformatics integration analysis engine, is used to output an MRD load report containing the MRD load score.

[0104] The monitoring system's workflow is as follows:

[0105] (1) The user submits the whole exome sequencing (WES) data file of the primary tumor tissue of the UTUC patient through the sample and data input interface and loads the peripheral blood sample from the same subject into the system.

[0106] (2) The sample processing and nucleic acid extraction unit automatically performs the following operations: First, it centrifuges the peripheral blood sample to obtain the plasma layer and the leukocyte layer; then, it extracts plasma circulating cell-free DNA (cfDNA) from the plasma and leukocyte genomic DNA (gDNA) from the leukocytes.

[0107] (3) The personalized panel customization unit automatically analyzes the received tumor tissue WES data and selects multiple high-frequency somatic mutation sites based on a preset algorithm (such as sorting by mutation frequency) to form a personalized panel target set; then, the set is combined with the fixed panel target set for UTUC pre-stored in the system to generate a comprehensive dual-panel target list.

[0108] (4) The library construction and processing unit receives the extracted cfDNA and automatically performs end repair, tailing, adapter ligation and pre-amplification steps to complete the construction of the sequencing library.

[0109] (5) The vacuum concentration subunit within the library construction and processing unit performs vacuum concentration on the constructed sequencing library solution at 45°C and 1500 rpm for 30 minutes until the solution is nearly dry. Subsequently, its hybridization capture subunit adds the hybridization reaction solution containing capture probes synthesized based on the dual-panel target list to the concentrated library and performs a hybridization reaction at 65°C for 16 hours.

[0110] (6) The sequencing unit (such as a high-throughput sequencer) sequences the cfDNA library products after hybridization capture to obtain plasma cfDNA sequencing data. At the same time, using the same dual-panel target list, the independently constructed leukocyte gDNA library is sequenced to obtain leukocyte control sequencing data.

[0111] (7) The bioinformatics integration and analysis engine processes two data streams in parallel:

[0112] First, analyze the leukocyte control sequencing data to screen out somatic mutations with allele frequencies not lower than a predetermined threshold (e.g., 0.1%), and construct an individualized clonal hematopoietic mutation filtering database.

[0113] Simultaneously, an error correction algorithm based on unique molecular identifiers (UMI) is applied to the plasma cfDNA sequencing data to perform high-confidence mutation detection.

[0114] Then, all mutations detected in cfDNA are automatically compared with the above-mentioned filtering database, and clonal hematopoietic-related mutations that match in the database are removed.

[0115] (8) The bioinformatics integration analysis engine calculates the median allele frequency of all tumor-derived mutations that have been screened and removed, and uses it as the MRD burden score.

[0116] (9) The report generator receives the MRD burden score, the list of detected mutations, and other related information output by the bioinformatics integration analysis engine, and automatically integrates them to generate a structured clinical monitoring report. This report not only includes the MRD burden score, but also outputs corresponding drug mutation warning prompts based on specific drug resistance sites (such as FGFR3 mutations) detected in the fixed panel, providing a reference for clinical treatment decisions.

[0117] In this invention, when analyzing plasma cfDNA sequencing data, all initially detected variants are compared with a filtering database. Any mutations present in the database are considered background signals of non-tumor origin and are removed in subsequent MRD analysis. Only those mutations not detected in leukocytes and meeting other quality criteria are retained as true tumor-derived molecular residual lesion markers for quantitative and dynamic tracking.

[0118] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0119] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for monitoring minimal residual disease based on circulating tumor DNA, characterized in that, The following steps are included in the monitoring of upper urothelial carcinoma (UTUC): Receive whole-exome sequencing data of primary tumor tissue from UTUC patients; Based on the whole exome sequencing data, the processor automatically screens out multiple high-frequency somatic mutation sites to form a personalized panel target set. The personalized panel target set is combined with a predefined fixed panel target set for UTUC to generate a dual-panel target list. Cell-free plasma DNA (cfDNA) and leukocyte genomic DNA (gDNA) were extracted from peripheral blood samples of the same subject. Sequencing libraries were constructed based on extracted cell-free circulating DNA (cfDNA) from plasma, and sequencing library solutions were obtained. The constructed sequencing library solution was then concentrated under vacuum. The hybridization reaction solution containing capture probes based on the dual-panel target list is mixed with the concentrated sequencing library to perform a hybridization capture reaction; The products after hybridization capture were sequenced to obtain plasma cfDNA sequencing data. The leukocyte genomic DNA (gDNA) was sequenced using a dual-panel target list to obtain leukocyte control sequencing data; Bioinformatics analysis was performed on the leukocyte control sequencing data to screen out somatic mutations with allele frequencies not lower than a predetermined threshold, and a personalized clonal hematopoietic mutation filtering database was constructed. Bioinformatics analysis was performed on the plasma cfDNA sequencing data, including error correction and mutation detection based on unique molecular identifiers; When analyzing the plasma cfDNA sequencing data, all detected mutations are compared with the filtering database, and mutations present in the filtering database are removed. The MRD burden score was calculated based on the median allele frequency of all selected and removed mutations. The output includes an MRD load report containing the MRD load score.

2. The method for monitoring minimal residual disease based on circulating tumor DNA as described in claim 1, characterized in that, The capture probe has a length of 80-120 bp, a GC content of 40%-60%, and a sequence identity of no more than 85% with the non-target region of the normal human genome.

3. The method for monitoring minimal residual disease based on circulating tumor DNA as described in claim 1, characterized in that, The vacuum concentration is carried out at 45-50℃ and 1000-2000 rpm for 20-40 minutes until the solution is dry or nearly dry.

4. The method for monitoring minimal residual disease based on circulating tumor DNA as described in claim 1, characterized in that, The hybridization reaction solution comprises: Dextran sulfated at a concentration of 1%-4%; Formamide at a concentration of 1-8%; Sodium chloride with a concentration of 0.2-2M; The capture probe includes a fixed panel probe and a personalized panel probe, with a molar concentration ratio of 1:1 to 5:

1.

5. The method for monitoring minimal residual disease based on circulating tumor DNA as described in claim 4, characterized in that, The hybridization capture reaction was carried out at 65°C for 12-18 hours.

6. The method for monitoring minimal residual disease based on circulating tumor DNA as described in claim 1, characterized in that, The hybridization capture reaction is followed by a washing step, which uses a washing buffer preheated to 60-65°C containing 0.1-1X SSPE and 0.005-0.2% N-lauroyl sarcosine sodium.

7. The method for monitoring minimal residual disease based on circulating tumor DNA as described in claim 1, characterized in that, The error correction process based on unique molecular identifiers includes: Sequencing reads are clustered according to their unique molecular identifier sequences, allowing for one base mismatch. Generate a consensus sequence for each unique molecular identifier cluster, requiring that the number of reads within the cluster is ≥3 and the majority base agreement is ≥80%; Mutations with a mutation allele frequency ≥ 0.02% and existing in at least two independent and unique molecular identifier clusters are selected.

8. A monitoring system for implementing the method according to any one of claims 1-7, characterized in that, include: The sample and data input interface is used to receive whole-exome sequencing data of primary tumor tissue from UTUC patients, as well as peripheral blood samples from the same subject. The sample processing and nucleic acid extraction module is connected to the sample and data input interface and is used to extract plasma circulating cell-free DNA (cfDNA) and leukocyte genomic DNA (gDNA) from peripheral blood samples, respectively. The personalized panel customization module is connected to the sample and data input interface and is configured to automatically screen high-frequency somatic mutation sites based on the received tumor tissue sequencing data to form a personalized panel target set, which is then combined with a predefined fixed panel target set for UTUC to generate a dual-panel target list. The library construction and processing unit, connected to the sample processing and nucleic acid extraction unit and the personalized panel customization unit, is configured to construct sequencing libraries based on the plasma circulating cell-free DNA (cfDNA) and the leukocyte genomic DNA (gDNA), respectively. The sequencing library solution constructed based on circulating cell-free DNA (cfDNA) in plasma was concentrated under vacuum; the hybridization reaction solution containing capture probes based on the dual-panel target list was mixed with the concentrated sequencing library to perform a hybridization capture reaction; The sequencing unit, connected to the library construction and processing unit, is configured to sequence the products after hybridization capture to obtain plasma cfDNA sequencing data; and to sequence the leukocyte genomic DNA (gDNA) through the dual-panel target list to obtain leukocyte control sequencing data. A bioinformatics integration analysis engine, connected to the sequencing unit, is configured to perform bioinformatics analysis on the leukocyte control sequencing data, screen for somatic mutations with allele frequencies not lower than a predetermined threshold, and construct an individualized clonal hematopoietic mutation filtering database; perform bioinformatics analysis on the plasma cfDNA sequencing data, including error correction based on unique molecular identifiers and mutation detection; compare all detected plasma cfDNA mutations with the filtering database and remove mutations present in the filtering database; and calculate the MRD burden score based on the median allele frequencies of all removed mutations. A report generator, connected to the bioinformatics integration analysis engine, is used to output an MRD load report containing the MRD load score.

9. The monitoring system as described in claim 8, characterized in that, The library construction and processing unit includes: The vacuum concentration unit is configured to concentrate the sequencing library solution under vacuum at 45-50℃ and 1000-2000 rpm for 20-40 minutes. The hybridization capture unit, connected to the vacuum concentration unit, is configured to mix the hybridization reaction solution containing the capture probe with the concentrated sequencing library and perform a hybridization reaction at 65°C for 12-18 hours.

Citation Information

Cited By

  • Data processing device computer program product for ctDNA variation detection and application

    CN122090959A