Design method of a probe for enhancing detection of molecular residual lesions and application thereof
Patent Information
- Application Number
- CN202510963633.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-07-11
AI Technical Summary
不过,传统的多基因探针集设计时,通常只覆盖关键的驱动基因以及与癌症发生发展密切相关的数百个基因,由于肿瘤存在异质性,难以保证MRD检测的灵敏度
[0112] The probe set design method disclosed herein ensures that the selected regions contribute to mutation detection to the greatest extent possible, eliminating redundancy and retaining the essential.
Smart Images

Figure CN121075423B_ABST
Abstract
Description
Technical Field
[0001] This disclosure pertains to the field of biotechnology, and in particular to the design method and application of a probe for enhancing the detection of molecular residual lesions. Background Technology
[0002] Cancer, a disease that seriously threatens human life and health, has become one of the leading causes of death due to its high incidence, high mortality, high recurrence rate, and the trend of younger age of onset. Choosing the most suitable treatment plan for cancer patients is a serious challenge facing clinical practice today. With the rapid development of tumor genome sequencing technology, it has become possible to develop personalized treatment plans for cancer patients using gene testing technology. Before cancer treatment, testing for point mutations, copy number variations, or gene rearrangements can provide prospective predictions of the efficacy of targeted and immunotherapies, offering effective references for doctors' clinical medication use and thus significantly improving the precision of treatment.
[0003] Besides targeted therapy, research on detecting potential lesions that cannot be detected by imaging (including positron emission tomography) or traditional laboratory methods by examining tumor-derived molecular abnormalities such as circulating tumor DNA (ctDNA) is also gaining momentum. MRD (molecular residual disease) detection can help doctors identify patients at high risk of recurrence earlier, providing a basis for subsequent clinical decisions. Depending on whether tumor tissue mutation information is referenced, MRD monitoring can be divided into tumor-informed and tumor-agnostic / The strategy. The Tumor-informed strategy requires reference to mutation information detected in tumor tissue, while the Tumor-agnostic / The strategy does not rely on the patient's tumor mutation information. However, peripheral blood contains not only tumor-derived ctDNA but also biological noise such as clonal hematopoiesis, which affects tumor-agnostic / The specificity and sensitivity of analytical methods were compared. In contrast, the tumor-informed strategy, which selects sites based on the patient's own tumor mutation profile and customizes personalized probes for MRD monitoring, effectively reduces false positive results caused by background noise, thereby improving detection performance. A retrospective study showed that the tumor-informed strategy significantly outperformed the tumor-agnostic / tumor-informed strategy in MRD detection. Strategy. Furthermore, a prospective study on non-small cell lung cancer, through head-to-head comparisons, also reached similar conclusions. Currently, most MRD detection products that have received FDA Breakthrough Device Designation employ a tumor-informed analysis strategy. The "Consensus on Molecular Residual Lesion Detection in Solid Tumors" recommends that, based on existing technical and clinical evidence, tumor-informed analysis methods should be given priority for MRD detection in solid tumors.
[0004] When employing a tumor-informed analysis strategy, tumor tissue detection can be based on either Western blotting (WES) or multi-gene probe sets. Whole-exome sequencing (WES) is known to provide a comprehensive picture of a patient's mutations, but it suffers from low sequencing depth and high cost. Compared to WES, multi-gene probe sets typically cover hundreds of genes closely related to cancer development and progression, offering a simpler analysis process and greater clinical applicability and feasibility. Studies in urothelial carcinoma have shown that personalized MRD strategies based on multi-gene probe sets demonstrate similar performance to WES-based personalized MRD in identifying MRD-positive patients and assessing disease-free survival (DFS) risk. Studies in non-small cell lung cancer and breast cancer have also demonstrated that multi-gene probe sets covering key driver genes can effectively support the screening of MRD-related tumor tissue mutation sites, thereby accurately predicting the patient's recurrence risk. However, traditional multi-gene probe set designs typically only cover key driver genes and hundreds of genes closely related to cancer development and progression. Due to tumor heterogeneity, it is difficult to guarantee the sensitivity of MRD detection. Summary of the Invention
[0005] To address at least one of the above-mentioned problems, this disclosure provides a method for designing probe sets. Probe sets prepared using the method provided in this disclosure can overcome inter-individual heterogeneity of tumors and improve detection sensitivity.
[0006] According to one aspect of this disclosure, a method for designing a probe set is provided, including the step of obtaining a set of markers, comprising:
[0007] S1, select one or more cancer-related genes to obtain set 1;
[0008] S2, based on the mutations in cancer patient samples, screen out one or more regions in the exons with high detection rates in cancer patients, and merge them with set 1 to obtain set 2;
[0009] S3, based on the mutations in the cancer patient samples, screen out one or more exon fragments that can improve the cancer patient coverage rate and the number of mutations covered by a single cancer patient, merge them with the set 2 to remove duplicates, and obtain the target region as a set of markers.
[0010] In some embodiments, the method further includes step S4, preparing a probe set comprising multiple oligonucleotides, the nucleic acid sequences of which are capable of hybridizing with the biomarker set obtained in step S3.
[0011] In some implementations, the cancer-related genes include genes associated with cancer prediction, prognosis, diagnostic value, medication, cancer treatment targets or development, and cancer immunotherapy.
[0012] In some embodiments, the set 1 in step S1 is obtained by screening from databases, literature reports, target genes of drugs approved by the FDA or NMPA, and guideline-recommended genes.
[0013] In some implementations, the database includes one or more of the following: CKB, oncoKB, clinicaltrial, civic, My Cancer Genome database, etc.
[0014] In some implementations, the guidelines include one or more of the following: CSCO guidelines, NCCN guidelines, ASCO guidelines, and ESMO guidelines.
[0015] In some embodiments, step S1 includes: obtaining a set 1 of gene exons that have cancer prediction, prognosis, diagnostic value, are associated with cancer treatment targets or development, cancer immunotherapy, and cancer medication.
[0016] In some embodiments, the sample in step S2 or S3 includes a tissue sample, preferably a cancer tissue sample, such as tissue, paraffin sections, and their processed forms.
[0017] In some embodiments, step S2, which involves screening for one or more regions with high detection rates of cancer patients in exons, includes the following steps:
[0018] T1) The detection rate of exon regions is calculated according to Formula 1:
[0019]
[0020] The exon region refers to the region captured and sequenced by the whole-exome capture probe within each exon; the number of patients covered by the exon region refers to the number of cancer patients with at least a predetermined number of mutations detected within each exon region; the exon region length is the number of base pairs (bp) contained in the exon region; the total number of patients is the total number of cancer patients; and the coefficient is 10. -7 ~1, preferably 10 -7 10 -6 10 -510 -4 10 -3 10 -2 10 -1 Or 1, more preferably 10 -5 ;
[0021] T2) Select exon regions whose detection rate is higher than the detection rate cutoff value, merge and remove duplicates to obtain one or more regions in the exon with high detection rate of cancer patients.
[0022] In some implementations, the preset quantity is 1 to 5, for example, 1, 2, 3, 4 or 5.
[0023] In some implementations, if the exon region contains an exon terminus (5' or 3'), splice site mutations are included when counting the number of patients covered.
[0024] In some specific implementations, the cutoff value is obtained as follows:
[0025] Different cutoff values are used to screen out exon regions whose detection rate is higher than the cutoff value. The total number of cancer patients covered by all screened exon regions and the total length of all screened exon regions are counted. The total detection rate of all screened exon regions is calculated as follows: Total detection rate of all screened exon regions = Total number of cancer patients covered by all screened exon regions / (Total length of all screened exon regions × coefficient × Total number of patients), where the total number of patients is the total number of cancer patients. Different cutoff values are plotted on the x-axis, and the total detection rate of the screened regions corresponding to different cutoff values is plotted on the y-axis to obtain a series of scatter points on the coordinate axis. The scatter points are connected sequentially with straight lines to obtain a line graph. The minimum cutoff value corresponding to the inflection point where the linear relationship between the cutoff value and the total detection rate changes significantly is the cutoff value in step T2).
[0026] Those skilled in the art will understand that capture sequencing can be performed using any conventional whole-exome capture probe.
[0027] In some implementations, the preset percentage is 1%.
[0028] In some implementations, the preset multiple is 1.1 to 2 times, including 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 or 2 times, preferably 1.5 times.
[0029] In some implementations, when the cancer is pan-cancer, regions with high detection rates for each type of single cancer patient or some multi-cancer patients are first obtained in the exons, and then the regions are merged and deduplicated to obtain exon regions with high detection rates in pan-cancer patients.
[0030] In some implementations, the one or more exon fragments in step S3 that can improve cancer patient coverage and mutation coverage per individual cancer patient are obtained through screening via the following steps:
[0031] P1) The regions captured and sequenced by the whole exome capture probe in cancer patient samples were divided into several mutation clusters;
[0032] P2) Determine the key mutational features of each mutation cluster to assess the target cancer patient coverage of each mutation cluster, the extent of increase in mutation coverage per patient for each mutation cluster, the extent of increase in mutation coverage per sample master clone for each mutation cluster, and the extent of increase in non-synonymous mutation coverage for each mutation cluster.
[0033] P3) Based on the key mutation features, the mutation clusters are screened to obtain one or more exon fragments that can improve the coverage of cancer patients and the mutation coverage of a single cancer patient.
[0034] In some implementations, when the cancer is pan-cancer, exon fragments that improve the coverage of single cancer patients or partial multi-cancer patients and the mutation coverage of individual cancer patients are first obtained according to steps P1)-P3), and then merged and deduplicated to obtain one or more exon fragments that improve the coverage of pan-cancer patients and the mutation coverage of individual cancer patients.
[0035] In some embodiments, the mutation clusters in step P1) are divided according to the following steps: the start position of the first mutation to the end position of the last mutation within the region captured by the whole exome capture probe sequencing in the cancer patient sample is taken as a mutation region; in each mutation region, mutations of every X bp are aggregated into a mutation cluster in the direction from 5' to 3', and those less than X bp are taken as a separate mutation cluster; and the start position of the first mutation to the end position of the last mutation in each mutation cluster is divided into a segment as a mutation cluster.
[0036] In some implementations, the length distribution of the mutation clusters includes less than or equal to X bp.
[0037] In some embodiments, X = 20 to 100, for example 20, 30, 40, 50, 60, 70, 80, 90 or 100, preferably 40.
[0038] In some implementations, the key mutation features in step P2) include one or more of the following: mutation cluster sample recurrence index RI, sample mutation detection enhancement index RII, sample master clone mutation detection enhancement index RIIc, and sample non-synonymous mutation detection enhancement index RII. f .
[0039] In some implementations, the RI is calculated using Formula 2:
[0040]
[0041] Where, n p This represents the number of cancer patients covered by each mutation cluster, where n represents the total number of said cancer patients.
[0042] In some implementations, k = 1 to 10 5 For example, 1, 10, 10 2 10 3 10 4 10 5 10 preferred 3 .
[0043] In some implementations, the RII is calculated using Formula 3:
[0044]
[0045] Where, N addi This represents the increase in the number of covered mutations in the i-th patient relative to set 2, where i is 1 to n, based on set 2 and each of the mutation clusters is added. p integers, l add This represents the increase in probe region length relative to set 2 when this mutation cluster is added.
[0046] In some implementations, l add For integers ≥ 0, when l add When the value is 0, the RII is not calculated, indicating that the inclusion of this area is of no value.
[0047] In some implementations, the N addi The result is obtained through calculation using Formula 4:
[0048]
[0049] Among them, the number of mutations covered by set 2 in patient i is N cover This indicates that when each mutation cluster is supplemented, the number of mutations covered by the increased patient coverage is N. only The preset design aims to achieve N mutation coverage per sample in the targeted capture region. g .
[0050] In some implementations, the N g =1 to 10, for example, N g =1, 2, 3, 4, 5, 6, 7, 8, 9, 10; preferred N g =4.
[0051] In some implementations, when set 2 does not cover the exon containing the mutation cluster, or when set 2 covers the exon containing the mutation cluster but the distance between set 2 and the mutation cluster is greater than Ybp, where Y = 40–100, then the increased probe region length l relative to set 2 is... add Calculate according to formula 5:
[0052] l add =ceil(L / X)*X+Y Formula 5;
[0053] Where the length L is the number of bases in the mutant cluster.
[0054] In some implementations, when set 2 covers the exon containing the mutation cluster, and there exists a set 2 coverage region that is less than or equal to Ybp from the mutation cluster, the increased probe region length l add Calculate according to Formula 6:
[0055]
[0056] Before the mutation clusters were added, set 2 had a total of N regions, and the length of region e was L. e After merging the regions in set 2 that are at a distance of Ybp or less from the mutation cluster into one region, there are a total of N1 regions. After merging, the length of region j is L. j .
[0057] In some implementations, the RIIc is calculated using Formula 7:
[0058]
[0059] Where N add-clonali This was done to improve the coverage of the master clone mutation level in the sample.
[0060] In some implementations, the N add-clonali Calculated according to Formula 8:
[0061]
[0062] Where, N cover-clonal This represents the number of primary clonal mutations covered by set 2 for each sample. When a mutation cluster is added, the number of primary clonal mutations covered by this sample is increased by N. only-clonal .
[0063] In some implementations, the RII f It is obtained through calculation using formula 9:
[0064]
[0065] Where N add-nbi To increase the number of non-synonymous mutations detected in the sample.
[0066] In some implementations, the N add-nbi Calculated according to Formula 10:
[0067]
[0068] Where N is used cover-nb This represents the number of non-synonymous mutations covered by set 2. When adding mutation clusters, the increased number of non-synonymous mutations covered by the samples is N. only-nb .
[0069] In some implementations, the screening in step P3) includes:
[0070] 1) Based on the mutation cluster sample recurrence index RI, sample mutation detection enhancement index RII, sample master clone mutation detection enhancement index RIIc, and sample non-synonymous mutation detection enhancement index RII f The mutation clusters are sorted from highest to lowest.
[0071] 2) Based on the mutation detection enhancement index (RII) of the sample, the sample recurrence index (RI) of the mutation cluster, the mutation detection enhancement index (RIIc) of the main clone of the sample, and the non-synonymous mutation detection enhancement index (RII) of the sample. f The mutation clusters are sorted in descending order of priority, and the first-ranked mutation cluster is selected and transferred to a specific region.
[0072] 3) Determine whether the area meets the requirements. If it does, proceed to step 4). If it does not, repeat steps 1) to 3) until the requirements are met.
[0073] 4) The defined regions that meet the requirements are identified as one or more exon segments that can improve cancer patient coverage and mutation coverage per individual cancer patient.
[0074] In some implementations, the requirement includes 90% to 99% of patients covering ≥3 to 5 mutations. For example: 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% of patients covering ≥3, 4, or 5 mutations.
[0075] In some implementations, the requirement includes 95% to 99% of patients covering ≥4 mutations.
[0076] In some implementations, the cancer patient samples in steps S2 and S3 may be the same or different.
[0077] In some implementations, the mutations in the cancer patient sample are detected or obtained from a database.
[0078] In some implementations, the detection includes whole-genome sequencing, targeted sequencing, PCR technology, in situ hybridization technology, microarray technology, etc.
[0079] Those skilled in the art will understand that probes can be designed to capture the markers using known probe design methods, including but not limited to tiling, exhaustive, shingled, or bi-stranded probes based on positive and negative chain designs.
[0080] In some embodiments, the probe design method includes placing the first probe in a region along a 5' to 3' direction, extending upstream by M nt from the starting boundary of the chromosome region where the marker is located, and then placing the next probe sequentially at intervals of M nt. The last probe covering the region extends downstream beyond the terminal boundary by a distance ≥ M nt. Each probe is 100 to 150 nt in length, for example, 120 nt, where M = 1 to 60 nt.
[0081] In other embodiments, multiplex PCR primers are designed to amplify the marker.
[0082] In some implementations, the cancer is selected from any of the following: the cancer is pan-cancer or a single cancer, wherein the pan-cancer includes at least two types of cancer.
[0083] In some specific embodiments, the cancers include, but are not limited to, any one or more of the following: lung cancer, breast cancer, colorectal cancer, bladder urothelial carcinoma, nasopharyngeal carcinoma, mature T-cell and NK-cell tumors, biliary tract tumors, bile duct cancer, sellar region tumors, non-Hodgkin's lymphoma, paraganglioma, peritoneal cancer, liver cancer, hepatobiliary duct cancer, hepatocellular carcinoma, liver space-occupying lesions, anal cancer, testicular cancer, cervical cancer, cervical adenocarcinoma, bone cancer, melanoma, ampullary cancer, and basal cell carcinoma. Familial adenomatous polyposis (FAP / AFAP), parathyroid carcinoma, thyroid cancer, mesothelioma, plasma cell myeloma, glioma, extranodal nasal NK / T-cell lymphoma, colorectal cancer, appendix cancer, lacrimal gland tumor, lymphoma, intracranial space-occupying lesion, intracranial tumor, ovarian cancer, ovarian serous cystadenocarcinoma, ovarian clear cell carcinoma, ovarian space-occupying lesion, follicular lymphoma, choroid plexus tumor, diffuse large B-cell lymphoma, urothelial carcinoma, embryonal tumor, skin cancer (non-melanotic). Pigmentoma), prostate cancer, gestational trophoblastic disease, breast cancer, breast sarcoma, chondrosarcoma, soft tissue sarcoma, neuroendocrine tumors, nerve sheath tumors, neuroepithelial tumors, neurofibromas, nephroblastoma, papillary renal cell carcinoma, renal pigment cell carcinoma, adrenal cortical tumors, clear cell renal cell carcinoma, renal cell carcinoma, germ cell tumors, esophageal cancer, esophageal and gastric cancer, pheochromocytoma, fallopian tube cancer, pineal region tumors, medulloblastoma, mantle cell lymphoma, head Neck cancer, head and neck squamous cell carcinoma, vulvar cancer, gastric cancer, gastrointestinal stromal tumor, tumor of the stomach or gastroesophageal junction, gastric adenocarcinoma, salivary gland cancer, small bowel cancer, sex cord-stromal cell tumor, thymic tumor, pancreatic cancer, pancreatic mass, hereditary leiomyomatosis and renal cell carcinoma syndrome (HLRCC), vaginal cancer, penile cancer, central nervous system tumors, cervical cancer, endometrial cancer, uterine sarcoma, histiocytic, dendritic cell tumors, leukemia, lymphoma and multiple myeloma.
[0084] According to another aspect of this disclosure, this disclosure provides a probe set obtained by the design method described herein.
[0085] According to another aspect of this disclosure, this disclosure provides a design apparatus for a probe set, the apparatus including a module for obtaining a set of markers, comprising:
[0086] The first screening module is used to screen one or more genes related to cancer to obtain set 1;
[0087] The second screening module is used to screen out one or more regions in the exons of cancer patients with high detection rates based on mutations in cancer patient samples, and merge them with set 1 to obtain set 2 after deduplication;
[0088] The third screening module is used to screen one or more exon fragments that can improve the coverage of cancer patients and the number of mutations in a single cancer patient based on the mutations in the cancer patient samples. These fragments are then merged with the set 2 and deduplicated to obtain the biomarker set.
[0089] In some embodiments, the device further includes a probe design module for designing and preparing a probe set comprising multiple oligonucleotides whose nucleic acid sequences are capable of hybridizing with the biomarker set.
[0090] In some implementations, the second screening module includes the following steps:
[0091] T1) The detection rate of exon regions is calculated according to Formula 1:
[0092]
[0093] The exon region refers to the region captured and sequenced by the whole-exome capture probe within each exon; the number of patients covered by the exon region refers to the number of cancer patients for whom at least 1, 2, 3, 4, or 5 mutations were detected in each exon region; the exon region length is the number of base pairs (bp) contained in the exon region; the total number of patients is the total number of cancer patients; and the coefficient is 10. -7 ~1, preferably 10 -7 10 -6 10 -5 10 -4 10 -3 10 -2 10 -1 Or 1, more preferably 10 -5 ;
[0094] T2) Select the exon regions whose detection rate is higher than the detection rate cutoff value, merge and remove duplicates to obtain the regions in the exon with high detection rate of cancer patients.
[0095] In some implementations, the third screening module includes the following steps:
[0096] P1) The regions captured and sequenced by the whole exome capture probe in cancer patient samples were divided into several mutation clusters;
[0097] P2) Determine the key mutational features of each mutation cluster to assess the target cancer patient coverage of each mutation cluster, the extent to which each mutation cluster increases the mutation coverage of a single patient, the extent to which each mutation cluster increases the main clonal mutation coverage of the sample, and the extent to which the mutation cluster increases non-synonymous mutation coverage.
[0098] P3) Based on the key mutation features, the mutation clusters are screened to obtain exon fragments that can improve the coverage of cancer patients and the mutation coverage of a single cancer patient.
[0099] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions that, when executed by a processor, enable the execution of the design method.
[0100] According to another aspect of this disclosure, an electronic device is provided, comprising: the proposed computer-readable storage medium; and a processor capable of executing computer instructions stored in the computer-readable storage medium.
[0101] According to another aspect of this disclosure, this disclosure provides a detection method comprising the steps of: using the probe set to capture nucleic acids in a sample collected from a subject and sequencing them to detect tumor-derived nucleic acid mutations.
[0102] In some implementations, the sample includes a natural sample or a standard.
[0103] In some implementations, the sample includes a cell sample or a cell-free biological sample.
[0104] In some implementations, the sample includes a tissue sample or a body fluid sample.
[0105] In some embodiments, the body fluid samples include: whole blood, serum, plasma, tears, aqueous humor, vitreous humor, saliva, sputum, nasopharyngeal swabs, oral swabs, milk, bronchoalveolar lavage fluid, pleural effusion, peritoneal fluid, lumbar vertebral or ventricular CSF, lymph, amniotic fluid, peritoneal fluid, prostatic fluid, semen, urine, feces, pus, vaginal discharge, tumor cells, and their processed forms.
[0106] In some embodiments, the tissue sample includes: tissue, paraffin sections, and their processed forms.
[0107] In some embodiments, the tumor-derived nucleic acid mutations include one or more of the following: nucleic acid mutations originating from minimal residual disease of tumors, nucleic acid mutations related to tumor medication, nucleic acid mutations related to early tumor screening, nucleic acid mutations related to tumor metastasis or recurrence monitoring, nucleic acid mutations related to tumor prognosis, and nucleic acid mutations related to tumor auxiliary diagnosis.
[0108] According to another aspect of this disclosure, the present disclosure provides the use of the design method in the preparation of a kit comprising the probe set.
[0109] In some implementations, the kit is used to detect tumor-derived nucleic acid mutations.
[0110] In some embodiments, the tumor-derived nucleic acid mutations include one or more of the following: nucleic acid mutations originating from minimal residual disease of tumors, nucleic acid mutations related to tumor medication, nucleic acid mutations related to early tumor screening, nucleic acid mutations related to tumor metastasis or recurrence monitoring, nucleic acid mutations related to tumor prognosis, and nucleic acid mutations related to tumor auxiliary diagnosis.
[0111] Beneficial effects:
[0112] The probe set design method disclosed herein ensures that the selected regions contribute to mutation detection to the greatest extent possible, eliminating redundancy and retaining the essential.
[0113] The probe set design method disclosed herein provides a probe set with broad coverage, meeting the needs of drug target detection, increasing the number of mutations, and effectively improving the number of covered mutation sites, thereby enhancing detection sensitivity. This is beneficial for improving detection performance in various scenarios such as MRD, diagnosis, medication, prognosis, and treatment. Detection using the probe set prepared by the method disclosed herein not only provides information on drug targets but also offers longer lead times in MRD monitoring, allowing physicians more time for clinical decision-making. It has broad clinical application prospects. Attached Figure Description
[0114] Figure 1 A flowchart of the marker screening described in Example 1 is shown.
[0115] Figure 2 An example diagram of TCGA breast cancer exon screening thresholds is shown.
[0116] Figure 3 A comparison chart of mutation densities is shown.
[0117] Figure 4 The number of mutations detected in patients in the clinical cohort is shown.
[0118] Figure 5 The results of the survival curves in the lung cancer clinical cohort are shown. Detailed Implementation
[0119] This disclosure provides a probe set design method that can not only take into account both cancer drug use and MRD monitoring, but also greatly overcome the inter-individual heterogeneity of tumors, improve the detection rate of patients and the number of mutations detected in patient tissues, thereby ensuring the sensitivity of subsequent MRD monitoring.
[0120] In some embodiments, the design method is based on screening gene ranges with predictive, prognostic, diagnostic, and immunological value. After steps of screening for regions with high detection rates using probe sets / WES and improving patient coverage, target regions are obtained as biomarkers for both cancer treatment and MRD monitoring. Further design of probe sets to capture these biomarkers is then performed. In some embodiments, when screening regions to improve patient coverage, evaluation is conducted based on key indices from four aspects: the target cancer patient coverage rate for each mutation cluster, the increase in mutation coverage per patient for each mutation cluster, the increase in the number of mutations in the main clonal mutations of the sample for each mutation cluster, and the increase in the number of non-synonymous mutations for each mutation cluster.
[0121] The use of the design method provided in this disclosure to prepare probe sets has the following advantages:
[0122] 1. The probe set has a higher overall mutation density value than WES, and the probe set has a stronger mutation capture capability.
[0123] 2. Increase the detection rate of patients and the number of mutation sites detected in patients, especially master clonal mutations and non-synonymous mutations.
[0124] 3. Effectively improves detection sensitivity and specificity, significantly extends lead time, and enhances the value of MRD monitoring.
[0125] definition
[0126] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly used in the field to which this disclosure pertains. For the purposes of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural forms, and vice versa.
[0127] In this disclosure, MRD can be an abbreviation for three terms: molecular residual disease, measurable residual disease, and minimal residual disease. MRD reflects the residual status of tumor lesions. After treatment, a small number of tumor cells may remain in the body of cancer patients. These tumor cells may be so few as to not cause any symptoms or signs and are usually undetectable by traditional methods such as cytological microscopy or serological tests. Detection requires highly sensitive modern cutting-edge technologies such as flow cytometry, PCR, and NGS. MRD refers to the small number of tumor cells that cannot be detected by these standard cytological analyses. If a patient has a positive MRD, it means that the patient has a higher risk of recurrence or a poorer prognosis.
[0128] In this disclosure, the term "driver mutation" refers to a mutation that provides a selective growth advantage in tumor cells. Driver mutations are causally involved in cancer formation, giving cancer cells a growth advantage, and are positively selected from the cancer-generating tissue microenvironment. Driver mutations are not essential for the maintenance of the final stage of cancer (although often they are), but they must be selected at some point in the cancer-generating cell line. In some embodiments, nucleic acid sequences are sequenced to detect nucleic acid variants, mutations, or variations. Methods for detecting sequence variants are known in the art, and sequence variants can be detected by any sequencing method known in the art.
[0129] In this disclosure, the term "lowest limit of detection (LoD)" refers to the lowest mutation frequency at which a detection sensitivity of ≥95% is achieved. In this disclosure, "sensitivity" refers to the probability of detecting a mutation or classifying a sample as MRD-positive in a given mutation / sample frequency.
[0130] In this disclosure, the term "sequencing data" refers to any sequence information about a nucleic acid molecule known to a person skilled in the art. Sequencing data may include information about DNA or RNA sequences that must be converted into nucleic acid sequences, modified nucleic acids, single-stranded or double-stranded sequences, or alternatively, amino acid sequences. Sequencing data may additionally include information about sequencing equipment, acquisition date, read length, sequencing orientation, source of the sequenced entity, adjacent sequences or reads, the presence of duplications, or any other suitable parameters known to a person skilled in the art. Sequencing data may be presented in any suitable format, file, encoding, or document known to a person skilled in the art.
[0131] In this disclosure, the term "tumor" refers to a mass or growth that is defined by itself as an abnormal new growth of cells that typically grow faster than normal cells and will continue to grow without treatment, sometimes causing damage to adjacent structures. Tumors can vary greatly in size. Tumors can be solid or fluid-filled. A tumor can refer to a benign (non-malignant, usually harmless) or malignant (capable of metastasis) growth. Some tumors may contain benign neoplastic cells (e.g., carcinoma in situ) as well as malignant cancer cells (e.g., adenocarcinoma). It should be understood that this includes growths located in multiple locations throughout the body. Therefore, for the purposes of this disclosure, tumors include primary tumors, lymph nodes, lymphatic tissue, and metastatic tumors.
[0132] In this disclosure, non-limiting examples of the cancers include lung cancer, breast cancer, colorectal cancer, bladder urothelial carcinoma, nasopharyngeal carcinoma, mature T-cell and NK-cell tumors, biliary tract tumors, bile duct cancer, sellar region tumors, non-Hodgkin's lymphoma, paraganglioma, peritoneal cancer, liver cancer, hepatobiliary duct cancer, hepatocellular carcinoma, liver space-occupying lesions, anal cancer, testicular cancer, cervical cancer, cervical adenocarcinoma, bone cancer, melanoma, ampullary cancer, basal cell carcinoma, and familial adenomatous polyposis (FAP). (AP / AFAP), parathyroid carcinoma, thyroid carcinoma, mesothelioma, plasma cell myeloma, glioma, extranodal nasal NK / T-cell lymphoma, colorectal cancer, appendix cancer, lacrimal gland tumor, lymphoma, intracranial space-occupying lesion, intracranial tumor, ovarian cancer, ovarian serous cystadenocarcinoma, ovarian clear cell carcinoma, ovarian space-occupying lesion, follicular lymphoma, choroid plexus tumor, diffuse large B-cell lymphoma, urothelial carcinoma, embryonal tumor, skin cancer (non-melanoma), prostate cancer, gestational trophoblastic disease, breast cancer, breast sarcoma, chondrosarcoma, soft tissue sarcoma, neuroendocrine tumors, nerve sheath tumors, neuroepithelial tumors, neurofibroma, nephroblastoma, renal papillary cell carcinoma, renal pigment cell carcinoma, adrenal cortical tumors, renal clear cell carcinoma, renal cell carcinoma, germ cell tumors, esophageal cancer, esophageal and gastric cancer, pheochromocytoma, fallopian tube cancer, pineal region tumors, medulloblastoma, mantle cell lymphoma, head and neck cancer. Squamous cell carcinoma of the head and neck, vulvar cancer, gastric cancer, gastrointestinal stromal tumor, tumor of the stomach or gastroesophageal junction, gastric adenocarcinoma, salivary gland cancer, small bowel cancer, sex cord-stromal cell tumor, thymic tumor, pancreatic cancer, pancreatic mass, hereditary leiomyomatosis and renal cell carcinoma syndrome (HLRCC), vaginal cancer, penile cancer, central nervous system tumors, cervical cancer, endometrial cancer, uterine sarcoma, histiocytic and dendritic cell tumors, leukemia, lymphoma and multiple myeloma.
[0133] In some implementations, the sequencing technologies mentioned include, but are not limited to, Illumina, BGI Genomics, and Geneplus.
[0134] In this disclosure, the term "pan-cancer" or "pan-cancer species" encompasses a wide range of cancer types. In some embodiments, the pan-cancer includes at least one of pan-solid tumors, hematologic malignancies, rare cancers, and childhood cancers. In some embodiments, pan-cancer research aims to reveal commonalities and differences among different cancers, advancing precision medicine and new drug development for cancer. By integrating data from multiple cancers, pan-cancer research provides a more comprehensive perspective for cancer prevention, diagnosis, and treatment.
[0135] The term "lung cancer" as used herein is used in the broadest sense to encompass all cancers that originate in the lungs. In some implementations, lung cancer may include both non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC). Non-small cell lung cancer may include lung adenocarcinoma, lung adenocarcinoma, and large cell carcinoma. The term "lung cancer" includes both locally acquired and metastatic lung cancer. The term "lung cancer" is classified based on tumor characteristics, lung function, and patient condition as follows: Stage I: Tumor confined to the lungs, without lymph node metastasis. Stage II: Large tumor or invasion of adjacent tissues, with possible local lymph node metastasis. Stage III: Tumor invading the mediastinum, chest wall, etc., with extensive lymph node metastasis. Stage IV: Distant metastasis, such as to the liver, brain, bone, heart, esophagus, mediastinum, trachea, recurrent laryngeal nerve, carina, vertebral body, or independent tumor nodules in different ipsilateral lobes of the lungs.
[0136] In this disclosure, the term "exon" is a segment of DNA in a gene that carries genetic information and encodes the amino acid sequence of a protein. After transcription, exons in the initial transcript are preserved, while introns are removed, and finally, the exons are joined together to form mature mRNA.
[0137] In this disclosure, the term "subject" means any animal, mammal, or human. A subject has, may have, or is suspected of having one or more diseases. A subject may have cancer, may exhibit cancer-related symptoms, may not exhibit cancer-related symptoms, or may not have been diagnosed with cancer. In some embodiments, the subject is a human being.
[0138] In this disclosure, the term "TCGA database" refers to the Cancer Genome Atlas Program (TCGA). It currently includes data from multiple cancers, encompassing genomic, transcriptomic, epigenetic, proteomic, and other omics data, as well as clinical sample information.
[0139] In this disclosure, the term "master clonal mutation" refers to a mutation present in all tumor cells, typically occurring in the early stages of tumor development. Conversely, "subclonal mutation" refers to a mutation present only in a subset of tumor cells, typically occurring in the later stages of tumor development, reflecting tumor evolution.
[0140] In this disclosure, the term "merging and deduplication" includes: retaining two or more regions that do not overlap with reference genome coordinates; for two or more regions that overlap with reference genome coordinates, or two or more adjacent regions (e.g., 0 bp apart), merging the upstream start base position of the upstream region to the downstream end base position of the downstream region into one region; and for two or more completely repetitive regions that are identical in length and start / end sites relative to the reference genome, retaining only one of the regions.
[0141] In this disclosure, the term "lead time" refers to the interval between the point at which a disease is detected through screening or early detection and the point at which the disease naturally progresses to the point where clinical symptoms appear or can be detected by routine diagnosis. A longer lead time means the disease may be detected earlier, potentially leading to better treatment opportunities and prognosis. In some embodiments, the screening includes imaging examinations. In some embodiments, the genetic testing includes detection of gene mutations using techniques such as sequencing or PCR. In some embodiments, the sequencing includes probe set capture sequencing, whole-genome sequencing, or multiplex PCR sequencing.
[0142] Unless the context clearly indicates otherwise, the terms “a” and “an” as used herein include plural references.
[0143] The term "about" as used herein is as understood by one of ordinary skill in the art and varies within a certain range depending on the context in which it is used. If one of ordinary skill in the art is unfamiliar with the use of this term in the context in which it is used, "about" will mean a particular value plus or minus 10%.
[0144] The symbols “×” or “*” used in this article can be used interchangeably to represent the mathematical multiplication sign. The symbols “ / ” or “—” used in this article can be used interchangeably to represent the mathematical division sign.
[0145] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this disclosure in any way. The actual scope of protection of this disclosure is set forth in the claims. In the following description, descriptions of well-known structures and techniques are omitted to avoid unnecessarily obscuring the concepts of this disclosure. Such structures and techniques have also been described in many publications. Unless otherwise specified, the devices, instruments, reagents, and / or kits used in the following embodiments are commercially available or obtained through conventional methods known to those skilled in the art.
[0146] Example
[0147] Example 1. Screening of biomarkers
[0148] The design of this detection chip needs to consider both: ① cancer drug detection, comprehensively encompassing guideline-level drug targets and pan-cancer hotspot mutations, including point mutations, small fragment insertions / deletions, gene rearrangements, and copy number variations; ② ensuring the sensitivity of subsequent MRD monitoring, with the tissue probe set design guaranteeing that most samples will detect at least four mutations. The biomarker screening flowchart is as follows: Figure 1 As shown.
[0149] 1. Information was crawled from the CKB, oncoKB, and civic databases to obtain genes with predictive, prognostic, and diagnostic value; genes related to tumor treatment targets or development and immunotherapy were extracted from clinical trial databases and literature reports; and gene exon regions were screened as set 1, based on target genes of drugs approved by the National Medical Products Administration (NMPA) or FDA, CSCO guidelines, or NCCN guidelines.
[0150] 2. Further screening of high-detection-rate regions in exons was conducted using whole-exome sequencing (WES) datasets.
[0151] 1) Obtained exon sequencing data from 8505 pan-cancer tissue samples, including WES sequencing data from 1301 clinical pan-cancer tissue samples (including 97 lung cancer patients, 69 colorectal cancer patients, 28 breast cancer patients, and 1107 other pan-cancer patients) (using the Core Exome Panel chip). This dataset contains WES sequencing data from v3.0 (supplier: Berklee, chip size: 33.9M) and 7204 pan-cancer tissue samples (including 1024 lung cancers, 561 colorectal cancers, 893 breast cancers, and 4726 other pan-cancer types) downloaded from the TCGA database. Among these, 5833 other pan-cancer samples include bladder urothelial carcinoma, biliary tract tumors, bile duct cancers, sellar region tumors, liver cancer, hepatobiliary duct cancer, hepatocellular carcinoma, testicular cancer, cervical cancer, bone cancer, melanoma, ampullary carcinoma, parathyroid carcinoma, thyroid cancer, mesothelioma, plasma cell myeloma, glioma, intracranial lesions, intracranial tumors, ovarian cancer, ovarian serous cystadenocarcinoma, ovarian lesions, choroid plexus tumors, urothelial carcinoma, and embryonic lesions. Fetal tumor samples, prostate cancer samples, chondrosarcoma samples, soft tissue sarcoma samples, neuroendocrine tumor samples, nerve sheath tumor samples, neuroepithelial tumor samples, neurofibroma samples, nephroblastoma samples, renal papillary cell carcinoma samples, renal pigment cell carcinoma samples, adrenal cortical tumor samples, renal clear cell carcinoma samples, renal cell carcinoma samples, germ cell tumor samples, esophageal cancer samples, fallopian tube cancer samples, head and neck cancer samples, head and neck squamous cell carcinoma samples, gastric cancer samples, gastrointestinal stromal tumor samples, gastric or gastroesophageal junction tumor samples, salivary gland cancer samples, sex cord-stromal cell tumor samples, thymic tumor samples, pancreatic cancer samples, pancreatic mass samples, hereditary leiomyomatosis and renal cell carcinoma syndrome (HLRCC) samples, central nervous system tumor samples, cervical cancer samples, endometrial cancer samples, uterine sarcoma samples.
[0152] Exon sequencing data from 8505 pan-cancer tissue samples underwent quality control and mutation calling. The sample sequencing data was aligned with the reference genome GRCH37 to obtain alignment data. After sequence alignment, BAM files were obtained, and GATK tools were used for mutation detection to obtain VCF files containing mutation information. Subsequently, the mutations obtained in the above detection process were annotated using BedAnno (version 1.20) software to obtain annotation files, thereby identifying point mutations in the samples and inserting missing mutation information.
[0153] Download the clinical information of the samples and the somatic mutation information from the WES results from the TCGA database. The somatic mutation information file format of the WES results is MAF (Mutation Annotation Format). The downloaded mutation results are filtered to remove mutations in the Variant_Classification column that are RNA, 3'Flank, 5'Flank, and intergenic regions, and to remove mutations in the Variant_Classification column that are Introns and are ≥50bp away from an exon.
[0154] 2) WES gene exon set: Based on the WES sequencing results obtained in step 1), the coordinate information of all exon regions is extracted from the annotation file according to the transcripts through public databases such as NCBI.
[0155] 3) Calculate the detection rate for each exon region:
[0156] Taking breast cancer as an example, the number of bases in the region captured by the whole exome capture probe (WES probe) in each exon is taken as the length of that exon region; and the number of breast cancer patients in the region captured by the WES probe in each exon is taken as the number of patients covered by each exon region, and the total number of patients is the total number of breast cancer patients.
[0157] Based on the mutation data from the clinical / TCGA sample set in step 2), the detection rate of the region captured by the WES probe in each exon (referred to as the exon region detection rate) is calculated as shown in Formula 1:
[0158]
[0159] Note: When counting the number of patients covered by an exon region, if the exon region contains exon ends (5' or 3'), splice site mutations are included in the count. The unit for exon region length in Formula 1 is bp, and the coefficient is 10. -5 .
[0160] For three types of patients—lung cancer, colorectal cancer, and other pan-cancer types—the detection rate of exon regions was statistically analyzed according to the above steps and Formula 1.
[0161] 4) Determine the cutoff value for the target cancer type using the following steps:
[0162] Taking breast cancer as an example, the detection rate of each exon region is calculated using Formula 1. When different cutoff values are used, exon regions with detection rates higher than the cutoff values are selected. The total number of breast cancer patients covered by all selected exon regions with detection rates higher than the cutoff values is counted, along with the total length of all selected exon regions with detection rates higher than the cutoff values. The overall detection rate of all selected exon regions is calculated as follows: Overall detection rate of all selected exon regions = Total number of breast cancer patients covered by all selected exon regions / (Total length of all selected exon regions × coefficient × Total number of breast cancer patients). Using different cutoff values as the x-axis and the overall detection rate of the selected regions as the y-axis, a series of scatter points are obtained on the coordinate system. Connecting these scatter points with straight lines yields a line graph (see...). Figure 2 The minimum cutoff value corresponding to the inflection point where the linear relationship between the cutoff value and the overall detection rate changes significantly is taken as the detection rate cutoff value for breast cancer (e.g., ...). Figure 2 (The location indicated by the middle arrow).
[0163] For patients with lung cancer, colorectal cancer, and other pan-cancer types, the detection rate cutoff value was determined according to the above method.
[0164] 5) For the datasets of lung cancer, colorectal cancer, breast cancer, and other pan-cancer types, select exon regions with detection rates higher than their respective cutoff values as high detection rate regions for lung cancer, colorectal cancer, breast cancer, and other pan-cancer types in the exons. Then, merge and remove duplicates of these high detection rate regions for lung cancer, colorectal cancer, breast cancer, and other pan-cancer types in the exons to obtain high detection rate regions for pan-cancer types in the exons. Finally, merge these high detection rate regions for pan-cancer types in the exons with set 1 and remove duplicates to obtain set 2.
[0165] 3. In addition to areas with high detection rates, additional areas that improve patient coverage and mutation coverage per patient were selected and added to set 2.
[0166] The following procedures were performed using tissue samples from the 8505 pan-cancer patients used in step 2:
[0167] The sample sets were divided into four categories: lung cancer, breast cancer, colorectal cancer, and other pan-cancers. For each sample set, the start position of the first mutation within each exon region and the end position of the last mutation were considered as a mutation region. Within each mutation region, mutations within a 40bp range (from 5' to 3') were clustered into a mutation cluster. The start position of the first mutation within each cluster and the end position of the last mutation were then divided into segments, also considered as a mutation cluster. If there was only one mutation within a 40bp range, the start and end positions of the mutation cluster were the coordinates of that mutation site. Therefore, the length of each mutation cluster was less than or equal to 40bp. Four key features were extracted from the mutation clusters, and these clusters were then screened based on these features. Regions that improved patient coverage and the number of mutations per patient were added to set 2, resulting in a biomarker set that considered both cancer medication and MRD monitoring.
[0168] The four key features mentioned in 3.1 include:
[0169] ① Candidate region sample recurrence index (RI)
[0170] The mutation cluster sample reproducibility index (RI) was used to assess the coverage of each mutation cluster in the four patient classes. For each mutation cluster, the RI was calculated as shown in Equation 2:
[0171]
[0172] Where X = 40, n p Let represent the number of samples covered by each mutation cluster in each sample set, and n represent the total number of samples in each sample set, where k = 10. 3 .
[0173] ② Sample mutation detection enhancement index (RII)
[0174] The Sample Mutation Detection Enhancement Index (RII) is used to assess the increase in mutation coverage for a single patient relative to set 2, for each mutation cluster, to determine the potential value of that region in improving sample mutation coverage. For each mutation cluster, the RII is calculated as shown in Equation 3:
[0175]
[0176] Where, N addi This represents the increase in the number of mutations covered by the i-th sample relative to set 2, where i is from 1 to n, based on set 2 and each of the mutation clusters is added. p integers, l add This represents the increase in probe region length relative to set 2 after adding this mutation cluster; l add For integers ≥ 0, when ladd When the value is 0, the RII is not calculated, indicating that the inclusion of this area is of no value.
[0177] N in Formula 3 addi The value is determined by formula 4:
[0178]
[0179] Among them, the number of mutations in the i-th patient covered by region 2 is used with N cover This indicates that when each mutation cluster is added, the increased number of mutations covered by the sample is N. only The preset design targets the desired mutation coverage for each sample, aiming to achieve Ng, where N is the number of mutations per sample. g =4.
[0180] The l mentioned in Formula 3 add The value is obtained by calculating using formula 5 or 6:
[0181] If set 2 does not cover the exon containing the mutation cluster, or if set 2 covers the exon containing the mutation cluster but the distance between set 2 and the mutation cluster is greater than Ybp, where Y = 80, then the increased probe region length l relative to set 2 is... add Calculate according to formula 5:
[0182] l add =ceil(L / X)*X+Y Formula 5,
[0183] Where X = 40, and length L is the number of bases in the mutant cluster.
[0184] Alternatively, if set 2 covers the exon containing the mutation cluster, and there exists a region covered by set 2 that is less than or equal to Ybp from the mutation cluster, then l is calculated according to formula 6. add :
[0185]
[0186] Where X = 40, Y = 80, before the mutation clusters were added, set 2 had a total of N regions, and the length of the e-th region was L. e Let e be an integer from 1 to N. After merging mutation clusters whose distance from a region in set 2 to a region is less than or equal to Ybp with the regions in set 2 into one region, there are a total of N1 regions. After merging, the length of the j-th region is L. j j = integers from 1 to N1.
[0187] ③ Sample master clone mutation detection enhancement index RIIc:
[0188] Sample Master Clonal Mutation Detection Enhancement Index: Measures the number of newly detected samples for each mutation cluster in response to the master clonal mutation, and is used to assess the extent of the increase in the number of samples for the master clonal mutation in that region.
[0189] For each mutation cluster, the sample master clone mutation detection enhancement index RIIc:
[0190]
[0191] Where N add-clonali To improve the coverage of the master clone mutation level in this sample, and calculated according to Formula 8, l add As shown in step ②, perform the calculation according to formula 5 or 6.
[0192]
[0193] Where, N cover-clonal This represents the number of primary clonal mutations covered by region 2 for each sample. When a mutation cluster is added, the number of primary clonal mutations covered by this sample is increased by N. only-clonal The software used to detect mutations in the master clone was PyCloneVI. g =4.
[0194] ④ Detection Enhancement Index (RII) for Non-Synonymous Mutations (Mutations other than Synonymous Mutations) in Samples f :
[0195] Detection Enhancement Index for Sample Nonsynonymous Mutations (RII) f Assess the extent to which mutation clusters increase the number of non-synonymous mutations, and calculate according to Formula 9:
[0196]
[0197] N in formula 9 add-nbi To increase the number of non-synonymous mutations detected in a sample, it is calculated according to formula 10, l add As shown in step ②, perform the calculation according to formula 5 or 6.
[0198]
[0199] In Formula 10, N is used. cover-nb This represents the number of non-synonymous mutations covered by set 2. When adding mutation clusters, the increased number of non-synonymous mutations covered by the samples is N. only-nb N g =4.
[0200] 3.2 The screening of mutation clusters based on the four key characteristics includes the following steps:
[0201] 1) For each type of sample set, divide all mutation clusters and apply the following criteria: mutation cluster sample recurrence index (RI), sample mutation detection enhancement index (RII), sample master clone mutation detection enhancement index (RIIc), and sample non-synonymous mutation detection enhancement index (RII). f Sort them from highest to lowest.
[0202] 2) Based on the mutation detection enhancement index (RII) of the sample, the sample recurrence index (RI) of the mutation cluster, the mutation detection enhancement index (RIIc) of the main clone of the sample, and the non-synonymous mutation detection enhancement index (RII) of the sample. f The mutations are sorted by priority from high to low, and the mutation cluster with the highest priority is selected and transferred to the designated region.
[0203] 3) Determine if the identified region meets the requirements. If it does, proceed to step 4). If not, repeat steps 1) through 3) until the requirements are met. The requirements are: Lung cancer, colorectal cancer, and breast cancer: 99% of patients cover ≥4 mutations; other pan-cancers: 95% of patients cover ≥4 mutations.
[0204] 4) Obtain the defined regions that meet the needs for lung cancer, colorectal cancer, breast cancer and other pan-cancers respectively, merge and remove duplicates to obtain the defined regions that improve patient coverage and mutation coverage per patient, and then merge and remove duplicates with set 2 to obtain a pan-cancer biomarker set that takes into account both cancer drug use and MRD monitoring.
[0205] Example 2. Comparison of mutation density between multi-gene probe sets and whole-exome sequencing
[0206] The effectiveness of targeted sequencing panel design can be quantitatively evaluated by mutation density, defined as the number of variant sites that can be detected per unit length of probe (see Equation 11). This metric directly measures the variant detection efficiency of different targeted capture strategies; a higher mutation density indicates a stronger ability of the strategy to capture the target variant and higher detection efficiency.
[0207] A 1× planar probe design was performed on the pan-cancer biomarker obtained in Example 1, with a probe length of 120 nt and a total size of 3.3M. Based on the designed probe set, the chromosomal regions covered by the probe set were extracted from the exon sequencing data of 1024 lung cancer samples (derived from 7204 pan-cancer tissue samples downloaded from the TCGA public database mentioned in step 2 of Example 1) for targeted sequencing simulation, and the mutation density was compared with the exon sequencing results. The mutation density calculation formula is shown in Formula 11, where d represents the mutation density, and n... m It is the number of mutations detected within the target capture area, l r This is the length of the target capture region. We calculated the variation density of each target capture region on all chromosomes, including the 22 autosomes and the sex chromosome (X), and visualized the results. Figure 3Statistical analysis showed that the probe set had an average mutation density of 0.080 and a median of 0.18, representing increases of 5.67-fold and 1.57-fold respectively compared to the WES average and median density. This overall mutation density distribution indicates that the probe set has a stronger mutation-capturing ability. This advantage was consistently observed on all autosomes and the X chromosome. To some extent, this phenomenon indirectly suggests that probe sets can more specifically capture mutations, resulting in higher efficiency and lower cost.
[0208]
[0209] Example 3. Assessing the achievement of the number of covered mutations in patients
[0210] 1) Using whole-exome sequencing data from tissue samples of 1301 clinical pan-cancer patients in step 2 of Example 1, the number of detectable mutations in the regions covered by the biomarkers screened in Example 1 was evaluated. The 1301 patients were categorized into lung cancer, breast cancer, colorectal cancer, and other pan-cancer types. The total number of mutation sites covering different proportions of patients in different cancer groups was statistically analyzed. Table 1 shows the statistical results, indicating that the biomarkers screened in Example 1 covered the following regions:
[0211] As shown in Table 1, the proportion of patients with more than 4 mutations in the biomarker coverage area screened in Example 1 is as follows: 93.6% for lung cancer, 95.7% for colorectal cancer, 100% for breast cancer, and 94.3% for other pan-cancer types.
[0212] Table 1. Sample reach rate of different mutation numbers in the WES clinical cohort.
[0213] lung cancer 93.6% Colorectal cancer 95.7% Breast cancer 100.0% Other pan-cancer 94.3%
[0214] The 1301 patients included 97 with lung cancer, 69 with colorectal cancer, 28 with breast cancer, and 1107 with other pan-cancer types. These other pan-cancer types included patients with biliary tract tumors, sellar region tumors, liver cancer, hepatobiliary duct cancer, testicular cancer, bone cancer, melanoma, ampullary cancer, parathyroid cancer, thyroid cancer, mesothelioma, plasma cell myeloma, glioma, intracranial space-occupying lesions, intracranial tumors, ovarian cancer, choroid plexus tumors, urothelial carcinoma, embryonal tumors, prostate cancer, chondrosarcoma, soft tissue sarcoma, neuroendocrine tumors, and nerve sheath tumors. Patients with neuroepithelial tumors, neurofibromas, nephroblastomas, adrenocortical tumors, renal cell carcinomas, germ cell tumors, esophageal cancers, fallopian tube cancers, head and neck cancers, gastrointestinal stromal tumors, tumors of the stomach or gastroesophageal junction, salivary gland cancers, sex cord-stromal cell tumors, thymic tumors, pancreatic cancers, pancreatic lesions, hereditary leiomyomatosis and renal cell carcinoma syndrome (HLRCC), central nervous system tumors, cervical cancers, endometrial cancers, and uterine sarcomas.
[0215] 2) Using samples from the TCGA public database, the number of detectable mutations in patients within the biomarker coverage areas selected in Example 1 was evaluated. A total of 7204 samples downloaded from the TCGA database described in step 2 of Example 1 were used, and statistical analysis was performed separately for four cancer types: lung cancer, colorectal cancer, breast cancer, and other pan-cancers. The results are shown in Table 2. Over 98% of patients had more than four mutations, with a median of approximately 10 mutations.
[0216] The statistical results of the proportion of patients whose biomarker coverage area reached 4 or more mutations in Example 1 are shown in Table 2. The proportions were 99.8% for lung cancer, 100% for colorectal cancer, 99.8% for breast cancer, and 98.1% for other cancers.
[0217] Table 2. Sample reach rate of TCGA cohort with different numbers of mutations
[0218]
[0219]
[0220] The 7204 patients included 1024 with lung cancer, 561 with colorectal cancer, 893 with breast cancer, and 4726 with other pan-cancer types. These other pan-cancer types included urothelial carcinoma of the bladder, bile duct carcinoma, hepatocellular carcinoma, cervical cancer, ovarian serous cystadenocarcinoma, melanoma, prostate cancer, papillary renal cell carcinoma, renal pigment cell carcinoma, clear cell renal carcinoma, esophageal cancer, head and neck squamous cell carcinoma, gastric cancer, pancreatic cancer, and endometrial cancer.
[0221] 3) The probe set designed in Example 2 was used to perform capture sequencing on an independent clinical cohort (200 pan-cancer patients, different from the samples in Example 1). The results showed that the median number of detectable mutations in the samples was 13. The number of detected mutations in the 200 patients is as follows: Figure 4 The results of a real-world clinical cohort analysis of 200 cases showed that 96.5% (193 / 200) of the patients had more than four mutations.
[0222] The 200 pan-cancer patients included those with non-small cell lung cancer, small intestinal cancer, lung cancer, tumors of the stomach or gastroesophageal junction, endometrial cancer, colorectal cancer, breast cancer, melanoma, biliary tract tumors, lung lesions, ovarian cancer, glioma, and neuroendocrine tumors.
[0223] Example 4. Validation results of MRD monitoring in a real clinical cohort of lung cancer patients indicating recurrence.
[0224] The efficacy of the probes designed in Example 1 was evaluated using cancer tissue samples and pre-recurrence plasma samples from a real clinical cohort of 87 patients with stage IB to IIIA lung cancer. Capture sequencing was performed to detect recurrence in patients. If cancer tissue mutations were detected in the plasma samples of the same patients, it indicated a potential risk of recurrence. The results are shown in Table 3, and the survival curves are as follows: Figure 5 As shown.
[0225] Survival curve results showed that, compared with ctDNA-negative patients, postoperative ctDNA-positive patients had a significantly higher risk of recurrence and a significantly shorter median recurrence-free survival (RFS) (HR: 16.8; 95% CI: 6.8–41.7; p<0.001). The time difference compared to imaging findings was 193 days.
[0226] Table 3 Detection Sensitivity and Specificity
[0227] Sensitivity 87.0%(20 / 23) Specificity 96.9%(62 / 64) Positive predictive value (PPV) 90.9%(20 / 22) Negative predictive value (NPV) 95.4%(62 / 65)
[0228] in:
[0229] Sensitivity refers to the proportion of people who test positive among all actual cancer patients;
[0230] Specificity refers to the proportion of people who test negative among all people who do not actually have cancer;
[0231] Positive predictive value (PPV) refers to the proportion of people who actually have cancer out of all those who test positive.
[0232] Negative predictive value (NPV) refers to the proportion of people who test negative but are actually not cancerous.
[0233] Example 5. Probe Design Method
[0234] The probe design for the marker set obtained in Example 1 was as follows: Following a 5' to 3' orientation, the first probe was placed 40 nt upstream of the chromosomal region containing the marker, followed by subsequent probes spaced 40 nt apart. The last probe covering the region extended downstream beyond the terminal boundary by at least 40 nt. Each probe was 120 nt long. This high-density shingled probe design ensures at least 2× probe coverage at each location within the target region. The following tests were performed:
[0235] 1) The high-density shingled probe design combination and the 1×tiled probe combination described in Example 2 were tested using the same samples to compare their effectiveness in achieving the same sequencing data volume. Ten pan-cancer samples (7 lung cancer, 1 breast cancer, 1 thyroid cancer, and 1 gastric cancer) were used for evaluation. The results are shown in Table 4. The ratio of the effective depth of the high-density shingled probe to the effective depth of the 1×tiled probe can reach more than 1.5 times. Similarly, to achieve the same effective depth, only 50-60% of the data volume required for the 1×tiled probe design is needed, effectively saving detection costs. Effective depth refers to the sequencing depth after data quality control and filtering. Effective depth can be used for variant detection, expression analysis, or other downstream analyses. The specific calculation method is: Effective depth = Number of reads after data quality control and filtering × Read length / Length of the target region covered by the probe.
[0236] Table 4. Effective depth performance of different probe design methods under the same data volume.
[0237]
[0238] 2) Hotspot analysis performance
[0239] The detection sensitivity of hotspot variants using the high-density shingled probe was evaluated by testing one multimutation standard (GW-OGTM006, tumor SNV 5% gDNA standard II, Jingliang).
[0240] Using the Kebai NA12878 gDNA standard, the standard was diluted to SNV mutation frequencies of 2% and 1%, with 5 replicates designed for each frequency. The results are shown in Table 5. With a low input of 30ng DNA, the high-density shingled probe can ensure 100% (17 / 17) detection of hotspot SNV and indel variants above 2%, and 90% (63 / 70) detection of variants above 1%; the detection sensitivity of variants is high.
[0241] Table 5. Detection of Hotspot Variations
[0242]
[0243]
[0244] The technical solutions disclosed herein are not limited to the specific embodiments described above. Any technical modifications made based on the technical solutions disclosed herein shall fall within the protection scope of this disclosure.
Claims
1. A method for designing a probe set, comprising the step of obtaining a set of markers, wherein: S1, screen for one or more genes associated with cancer to obtain set 1; S2, based on the mutations in the cancer patient samples, screen out one or more regions in the exons with a high detection rate of cancer patients, and merge them with set 1 to obtain set 2; S3, based on the mutations in the cancer patient samples, screen out one or more exon fragments that can improve the cancer patient coverage rate and the number of mutations covered by a single cancer patient, merge them with the set 2 to remove duplicates, and obtain the target region as a set of markers; Specifically, step S2, which involves screening out one or more regions in the exons with a high detection rate for cancer patients, includes the following steps: T1) The detection rate of exon regions is calculated according to Formula 1: Exon region detection rate = Formula 1; The exon region is the region captured and sequenced by the whole exon capture probe in each exon; the number of patients covered by the exon region is the number of cancer patients in which at least a predetermined number of mutations were detected in each exon region; the length of the exon region is the number of base pairs contained in the exon region; and the total number of patients is the total number of cancer patients. T2) Select the exon regions whose detection rate is higher than the detection rate cutoff value, merge and remove duplicates to obtain one or more regions in the exon with high detection rate in cancer patients; In step S3, the one or more exon fragments that can improve cancer patient coverage and mutation coverage of a single cancer patient are obtained through screening via the following steps: P1) The regions captured and sequenced by whole-exome capture probes in cancer patient samples were divided into several mutation clusters; P2) Determine the key mutational features of each mutation cluster to assess the target cancer patient coverage of each mutation cluster, the extent to which each mutation cluster increases the mutation coverage of a single patient, the extent to which each mutation cluster increases the main clonal mutation coverage of the sample, and the extent to which the mutation cluster increases non-synonymous mutation coverage. P3) Based on the key mutation features, the mutation clusters are screened to obtain one or more exon fragments that can improve the coverage of cancer patients and the mutation coverage of a single cancer patient.
2. The design method according to claim 1, further comprising step S4, preparing a probe set, the probe set comprising a plurality of oligonucleotides, wherein the nucleic acid sequences of the oligonucleotides are capable of hybridizing with the biomarker set obtained in step S3.
3. According to the design method of claim 1, the one or more cancer-related genes in step S1 include genes related to cancer prediction, prognosis, diagnostic value, medication, cancer treatment targets or occurrence and development, and cancer immunotherapy.
4. The design method according to claim 1, wherein the sample in step S2 or S3 includes a tissue sample.
5. The design method according to claim 1, wherein the sample in step S2 or S3 is a cancer tissue sample.
6. The design method according to claim 1, wherein the cutoff value in step T2) is obtained by the following method: By selecting different cutoff values, exon regions with detection rates higher than the cutoff values are identified. The total number of cancer patients covered by all selected exon regions and the total length of all selected exon regions are calculated. The overall detection rate of all selected exon regions is then calculated as follows: Overall detection rate of all selected exon regions = Total number of cancer patients covered by all selected exon regions / (Total length of all selected exon regions × coefficient × Total number of patients). Given the total number of cancer patients, we plot different cutoff values on the x-axis and the total detection rate of the screening area corresponding to different cutoff values on the y-axis, obtaining a series of scatter points on the coordinates. We then connect the scatter points sequentially with straight lines to obtain a line graph. The minimum cutoff value corresponding to the inflection point where the linear relationship between the cutoff value and the total detection rate changes significantly is the cutoff value mentioned in step T2).
7. According to the design method of claim 1, the mutation clusters in step P1) are divided according to the following steps: the start position of the first mutation to the end position of the last mutation within the region captured and sequenced by the whole exome capture probe in the cancer patient sample is taken as a mutation region; within each mutation region, mutations of every X bp are aggregated into a mutation cluster in the direction from 5' to 3', and mutations less than X bp are taken as a separate mutation cluster; and the start position of the first mutation to the end position of the last mutation in each mutation cluster is divided into a segment as a mutation cluster. =40.
8. The design method according to claim 7, wherein the key mutation features in step P2) include one or more of the following: mutation cluster sample recurrence index RI, sample mutation detection enhancement index RII, sample master clone mutation detection enhancement index RIIc, and sample non-synonymous mutation detection enhancement index. .
9. The method according to claim 8, wherein the RI is calculated using formula 2: Official 2; in, n p This indicates the number of cancer patients covered by each mutation cluster. n This represents the total number of patients with the aforementioned cancer. =10 3 .
10. The method according to claim 8, wherein the RII is calculated using formula 3: Official 3; in, This indicates the level of increase in the number of covered mutations for the i-th patient relative to set 2, based on set 2 and for each of the mutation clusters added. i From 1 to n p integers, This represents the increase in probe region length relative to set 2 after adding this mutation cluster; l add For integers ≥ 0, when l add When the value is 0, the RII is not calculated, indicating that the inclusion of this area is of no value.
11. The method according to claim 8, wherein the RIIc is calculated using formula 7: Official 7; in This was done to improve the coverage of the master clone mutation level in the sample.
12. The method according to claim 8, wherein... It is obtained through calculation using formula 9: Official 9; in To increase the number of non-synonymous mutations detected in the sample.
13. According to the design method of claim 9, the desired mutation coverage of the i-th sample in the preset targeted capture region is: N g .
14. The method according to claim 10, wherein N in formula 3 addi The result is obtained through calculation using Formula 4: Official 4; in, No. i Example: The number of mutations covered by set 2 in the patient and used... This indicates that when each mutation cluster is supplemented, the number of mutations covered by the increased patient coverage is [number missing]. =4.
15. The method according to claim 11, wherein in formula 7 Calculated according to Formula 8: Official 8; in, This represents the number of primary clonal mutations covered by set 2 for each sample; when a mutation cluster is added, the number of primary clonal mutations covered by that sample is increased. ; =4.
16. The method according to claim 12, wherein in formula 9 Calculated according to Formula 10: Official 10; Among them, the use of This represents the number of non-synonymous mutations covered by set 2; when adding mutation clusters, the increased number of non-synonymous mutations covered by the samples is... ; =4.
17. The design method according to claim 7, wherein when set 2 does not cover the exon containing the mutation cluster, or set 2 covers the exon containing the mutation cluster, but the distance between set 2 and the mutation cluster is greater than Y bp, then l add Calculate according to formula 5: Official 5; Where length The number of bases in the mutant cluster. =40, =80; or, If set 2 covers the exon containing the mutation cluster, and there exists a region covered by set 2 that is less than or equal to Ybp from the mutation cluster, then l add Calculate according to Formula 6: Official 6; in, Before the mutation clusters were added, set 2 had a total of N regions. e Length is L e After merging the regions in set 2 that are less than or equal to Y bp from the mutation cluster into one region, there are a total of N1 regions. After merging, the regions... j Length is L j , = 40, =80.
18. The design method according to claim 1, wherein the screening in step P3) includes: 1) Based on the mutation cluster sample recurrence index RI, sample mutation detection enhancement index RII, sample master clone mutation detection enhancement index RIIc, and sample non-synonymous mutation detection enhancement index The mutation clusters are sorted from highest to lowest. 2) Based on the mutation detection enhancement index (RII) of the sample, the sample recurrence index (RI) of the mutation cluster, the mutation detection enhancement index (RIIc) of the main clone of the sample, and the non-synonymous mutation detection enhancement index of the sample. The mutation clusters are sorted in descending order of priority, and the mutation cluster with the highest priority is selected and transferred into a specific region. 3) Determine whether the area meets the requirements. If it does, proceed to step 4). If it does not, repeat steps 1) to 3) until the requirements are met. 4) The defined regions that meet the requirements are identified as one or more exon segments that can improve cancer patient coverage and mutation coverage per individual cancer patient; The requirement includes 90% to 99% of patients having ≥3 to 5 mutations.
19. The design method according to claim 1, wherein the cancer is pan-cancer or a single cancer type, and when the cancer is pan-cancer, In step S2, exon regions with high detection rates in single-cancer patients or some multi-cancer patients are first obtained, and then combined and deduplicated to obtain exon regions with high detection rates in pan-cancer patients; and / or, In step S3, exon fragments that improve the coverage rate of single cancer patients or some multi-cancer patients and the mutation coverage of a single patient are first obtained according to steps P1)-P3), and then the exon fragments that improve the coverage rate of pan-cancer patients and the mutation coverage of a single cancer patient are obtained by merging and deduplication.
20. The design method according to claim 2, wherein the method for designing the probe in step S4 includes a tiling method, an exhaustive method, a shingled method, or a design of a double-stranded probe based on positive and negative chains.
21. The design method according to claim 2, wherein the probe design method comprises: Following the direction from 5' to 3', extend upstream by M nt from the starting boundary of the exon fragment where the marker is located to place the first probe in the region, and then arrange the next probe in sequence at intervals of M nt. The last probe covering the region extends downstream beyond the end boundary by a distance ≥ M nt. The length of each probe is 100~150nt, and M=1~60.
22. The design method according to claim 21, where M is 40.
23. The probe set prepared by the design method according to any one of claims 1 to 22.
24. A design apparatus for a probe set, characterized in that, The apparatus includes a module for obtaining a set of markers, comprising: The first screening module is used for one or more cancer-related genes to obtain set 1; The second screening module is used to screen out one or more regions in the exons of cancer patients with high detection rates based on mutations in cancer patient samples, and merge them with set 1 to obtain set 2 after deduplication; The third screening module is used to screen out one or more exon fragments that can improve the coverage of cancer patients and the number of mutations in a single cancer patient based on mutations in cancer patient samples, and merge them with set 2 to remove duplicates, thereby obtaining a set of biomarkers. The second screening module, which filters out one or more regions in the exons with a high detection rate for cancer patients, includes the following steps: T1) The detection rate of exon regions is calculated according to Formula 1: Exon region detection rate = Formula 1; The exon region is the region captured and sequenced by the whole exon capture probe in each exon; the number of patients covered by the exon region is the number of cancer patients in which at least a predetermined number of mutations were detected in each exon region; the length of the exon region is the number of base pairs contained in the exon region; and the total number of patients is the total number of cancer patients. T2) Select exon regions whose detection rate is higher than the detection rate cutoff value, merge and remove duplicates to obtain one or more regions in the exon with high detection rates in cancer patients. The third screening module includes the following steps: P1) The regions captured and sequenced by whole-exome capture probes in cancer patient samples were divided into several mutation clusters; P2) Determine the key mutational features of each mutation cluster to assess the target cancer patient coverage of each mutation cluster, the extent to which each mutation cluster increases the mutation coverage of a single patient, the extent to which each mutation cluster increases the main clonal mutation coverage of the sample, and the extent to which the mutation cluster increases non-synonymous mutation coverage. P3) Based on the key mutation features, the mutation clusters are screened to obtain one or more exon fragments that can improve the coverage of cancer patients and the mutation coverage of a single cancer patient.
25. The apparatus of claim 24, further comprising a probe design module for designing and preparing a probe set, the probe set comprising a plurality of oligonucleotides, the nucleic acid sequences of the oligonucleotides being capable of hybridizing with the biomarker set.
26. A computer-readable storage medium storing computer instructions that, when executed by a processor, enable the execution of the design method according to any one of claims 1 to 22.
27. An electronic device comprising: The proposed computer-readable storage medium; And a processor capable of executing computer instructions stored in the computer-readable storage medium of claim 26.
28. A detection method comprising the steps of: capturing nucleic acids in a sample collected from a subject using the probe set of claim 23 and sequencing them to detect tumor-derived nucleic acid mutations.
29. The detection method according to claim 28, wherein the sample comprises a natural sample or a standard.
30. The detection method according to claim 28, wherein the sample comprises a cell sample or a cell-free biological sample.
31. The detection method according to claim 28, wherein the sample includes a tissue sample or a body fluid sample.
32. The detection method according to claim 31, wherein the body fluid sample comprises: Whole blood, serum, plasma, tears, aqueous humor, vitreous fluid, saliva, sputum, nasopharyngeal swabs, oral swabs, milk, bronchoalveolar lavage fluid, pleural and peritoneal fluid, cerebrospinal fluid, lumbar or ventricular CSF, lymph, amniotic fluid, peritoneal fluid, prostatic fluid, semen, urine, feces, pus, vaginal discharge, tumor cells, and their processed products.
33. The detection method according to claim 31, wherein the tissue sample comprises: Tissues, paraffin sections, and their processed products.
34. Use of the design method according to any one of claims 1 to 22 in the preparation of a reagent kit.
35. The use according to claim 34, wherein the kit is used to detect nucleic acid mutations of tumor origin.
36. The use according to claim 35, wherein the tumor-derived nucleic acid mutation includes one or more of the following: nucleic acid mutations derived from minimal residual disease of tumors, nucleic acid mutations related to tumor medication, nucleic acid mutations related to early tumor screening, nucleic acid mutations related to tumor metastasis or recurrence monitoring, nucleic acid mutations related to tumor prognosis, and nucleic acid mutations related to tumor auxiliary diagnosis.
Citation Information
Patent Citations
Screening method of gene region set for detecting small residual focus of ovarian cancer, gene region set and detection system thereof
CN120220803A