A probe kit for detecting MRD in non-small cell lung cancer and its application

By using a specific probe set for non-small cell lung cancer to cover high-evidence-level drug-induced mutations and high-frequency mutations, this technology solves the problem that existing technologies cannot effectively monitor new drug-resistant mutations caused by tumor clonal evolution, and achieves efficient detection of MRD in non-small cell lung cancer.

CN118932064BActive Publication Date: 2025-11-14THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411173634.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-11-14
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

Existing personalized MRD probes cannot effectively monitor new drug resistance mutations caused by tumor clonal evolution, and cannot meet the high-efficiency detection needs of minimal residual disease in non-small cell lung cancer.

Method used

This invention provides a probe set specific to non-small cell lung cancer, covering high-evidence-level drug-use mutations and high-frequency mutations, including specific probes for genes such as EGFR, KRAS, ERBB2, BRAF, MET, ALK, TP53, PIK3CA, CTNNB1, SMAD4, CDKN2A, and NFE2L2, for detecting tumor evolution and new drug resistance mutations. The probe set can be used alone or in combination with other probe sets.

Benefits of technology

It improves capture efficiency and detection sensitivity, stabilizes the experimental system, and can accurately monitor tumor evolution and new drug resistance mutations, meeting clinical testing requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005010764510000071
    Figure BDA0005010764510000071
  • Figure BDA0005010764510000081
    Figure BDA0005010764510000081
  • Figure BDA0005010764510000082
    Figure BDA0005010764510000082
Patent Text Reader

Abstract

This disclosure provides a probe set for detecting MRD in non-small cell lung cancer and its application. The probe set obtained using the screening method provided in this disclosure incorporates drug-related mutations with high evidence levels and high-frequency mutations in non-small cell lung cancer, and can be used for MRD monitoring in non-small cell lung cancer. The probe set described in this disclosure can monitor tumor evolution and newly emerging drug-resistant mutations, can overcome the spatiotemporal heterogeneity of tumors to a certain extent, and can also improve capture efficiency, stabilize the experimental system, and meet clinical testing requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of nucleic acid sequencing. Specifically, this disclosure relates to a probe set for detecting MRD in non-small cell lung cancer and its application. Background Technology

[0002] Minimal residual disease (MRD) refers to the presence of residual tumor cells or tiny lesions in cancer patients who have achieved radiographic complete remission after radical treatment, but which are undetectable by imaging methods. This represents a latent stage of cancer progression. The number of residual tumor cells may be small and may not cause any signs or symptoms initially, but they can lead to future tumor progression, recurrence, or metastasis. Circulating tumor DNA (ctDNA) sequencing can detect these molecular abnormalities and is used for prognostic assessment, recurrence monitoring, efficacy evaluation, and adjuvant therapy decisions, providing important references for clinical treatment decisions.

[0003] MRD status based on ctDNA testing is closely related to the recurrence risk of non-small cell lung cancer (NSCLC). For early-stage lung cancer patients undergoing radical surgery, ctDNA positivity can predict recurrence earlier than traditional imaging. NSCLC patients with MRD positivity have a higher risk of recurrence, and MRD-positive patients can benefit from adjuvant therapy. In recent years, some domestic and international cancer diagnosis and treatment guidelines and expert consensus have gradually adopted ctDNA as a prognostic biomarker. my country's "Expert Consensus on Molecular Residual Lesions in Non-Small Cell Lung Cancer" recommends that "after radical resection of early-stage NSCLC patients, MRD positivity indicates a high risk of recurrence and requires close follow-up management; MRD testing is recommended every 3–6 months."

[0004] Existing personalized MRD probes cannot effectively monitor new drug resistance mutations that arise from tumor clonal evolution caused by selective pressure from drug therapy and other factors. How to effectively monitor new drug resistance mutations caused by clonal evolution is a challenge facing MRD detection. Summary of the Invention

[0005] To address at least one of the above-mentioned problems, this disclosure provides a gene biomarker, probe set, screening method, detection method, and its application for detecting minimal residual disease (MRD) in non-small cell lung cancer. The probe set provided in this disclosure exhibits excellent capture efficiency and depth coefficient, and good probe uniformity. Using the probe set provided in this disclosure, it is possible to detect tumor evolution and novel drug resistance mutations caused by clonal evolution in MRD, and it can also improve capture efficiency and stabilize the experimental system.

[0006] According to one aspect of this disclosure, a genetic biomarker for detecting minimal residual disease in non-small cell lung cancer is provided, said biomarker comprising any one or more of the following genes: EGFR, KRAS, ERBB2, BRAF, MET, ALK, TP53, PIK3CA, CTNNB1, SMAD4, CDKN2A, or NFE2L2.

[0007] In some embodiments, the genetic marker includes any one or more of the following regions: exons 18, 19, 20, and 21 of EGFR; exons 2, 3, and 4 of KRAS; exon 20 of ERBB2; exons 11 and 15 of BRAF; intron 13 and exon 14 of MET; exon 23 of ALK; exons 4-10 of TP53; exons 10 and 21 of PIK3CA; exon 3 of CTNNB1; exon 9 of SMAD4; exons 1 and 2 of CDKN2A; and exon 2 of NFE2L2.

[0008] In some embodiments, the probe set is used to detect the genetic markers described in this disclosure.

[0009] In some embodiments, the probe set has a nucleotide sequence as shown in any one or more of SEQ ID NO:1 to 57, or a nucleotide sequence having at least 85% sequence identity with it.

[0010] In some embodiments, the probe set includes any one or more of the following combinations:

[0011] Probes for detecting the ALK gene: having a nucleotide sequence as shown in SEQ ID NO:1 and / or 2, or a nucleotide sequence having at least 85% sequence identity with it; and / or

[0012] Probes for detecting the NFE2L2 gene: having nucleotide sequences as shown in any one or more of SEQ ID NO:3-5, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0013] Probes for detecting the CTNNB1 gene: having a nucleotide sequence as shown in SEQ ID NO:6 and / or 7, or a nucleotide sequence having at least 85% sequence identity with it; and / or

[0014] Probes for detecting the PIK3CA gene: having nucleotide sequences as shown in any one or more of SEQ ID NO: 8–12, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0015] Probes for detecting the EGFR gene: having nucleotide sequences as shown in any one or more of SEQ ID NO:13-20, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0016] Probes for detecting the MET gene: having nucleotide sequences as shown in any one or more of SEQ ID NO:21-23, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0017] Probes for detecting the BRAF gene: having nucleotide sequences as shown in any one or more of SEQ ID NO:24-27, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0018] Probes for detecting the CDKN2A gene: having nucleotide sequences as shown in any one or more of SEQ ID NO:28–32, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0019] Probes for detecting the KRAS gene: having nucleotide sequences as shown in any one or more of SEQ ID NO:33-38, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0020] Probes for detecting the TP53 gene: having nucleotide sequences as shown in any one or more of SEQ ID NO:39-53, or nucleotide sequences having at least 85% sequence identity with them; and / or

[0021] Probes for detecting the ERBB2 gene: having a nucleotide sequence as shown in SEQ ID NO:54 and / or 55, or a nucleotide sequence having at least 85% sequence identity with it; and / or

[0022] Probes for detecting the SMAD4 gene: having a nucleotide sequence as shown in SEQ ID NO:56 and / or 57, or a nucleotide sequence having at least 85% sequence identity with it.

[0023] According to another aspect of this disclosure, a kit for detecting minimal residual disease in non-small cell lung cancer is provided, the kit comprising the probe set.

[0024] According to another aspect of this disclosure, a method for detecting minimal residual disease in non-small cell lung cancer is provided using the probe set and / or the kit described in this disclosure, the method comprising using the probe set and / or the kit to detect minimal residual disease in non-small cell lung cancer.

[0025] According to another aspect of this disclosure, the use of the gene markers and / or probe sets described herein in the preparation of a kit for detecting minimal residual disease in non-small cell lung cancer is provided.

[0026] In some implementations, the kit is used for the detection of cell, tissue, or body fluid samples.

[0027] In some embodiments, the cell sample includes a lung cancer cell suspension sample.

[0028] In some embodiments, the bodily fluid samples include: saliva, whole blood, serum, plasma, milk, urine, lumbar or ventricular CSF, bile, lymph, prostatic fluid, semen, sputum, feces, tears, tumor cells, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural effusion, peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, vaginal discharge, and their processed forms.

[0029] In some embodiments, the tissue sample includes lung cancer tissue, paraffin sections, and their processed forms.

[0030] Beneficial effects:

[0031] The probe set provided in this disclosure includes drug-related mutations with high evidence levels and high-frequency mutations in non-small cell lung cancer (NSCLC). It can be used alone or in combination with other probe sets for NSCLC MRD monitoring, exhibiting good sensitivity and specificity. The NSCLC cancer-specific probe set described in this disclosure can monitor tumor evolution and newly emerging drug-resistant mutations, overcoming the spatiotemporal heterogeneity of tumors to some extent, while also improving capture efficiency and stabilizing the experimental system.

[0032] Clinical samples have verified that the probe set of this invention has excellent capture efficiency and depth coefficient, and good probe uniformity, which can meet the requirements of clinical testing. Attached Figure Description

[0033] Figure 1 The capture efficiency test results of the non-small cell lung cancer-specific probe group are shown.

[0034] Figure 2 The results of the non-small cell lung cancer specific probe group ≥0.2 times the mean depth percentage test are shown.

[0035] Figure 3 The results of the non-small cell lung cancer specific probe group ≥0.5 times the mean depth percentage test are shown.

[0036] Figure 4 The results of depth coefficient testing for non-small cell lung cancer specific probe groups are shown.

[0037] Figure 5The results show the coverage of tissue mutations in non-small cell lung cancer (NSCLC) patients by the specific probe group.

[0038] Figure 6 The results show the test results of the accuracy of MRD detection in non-small cell lung cancer when using a combination of non-small cell lung cancer-specific probe groups and patient-specific probe groups. Detailed Implementation

[0039] This disclosure provides a personalized MRD probe set specific to non-small cell lung cancer (NSCLC), covering drug-significant mutations and high-frequency NSCLC-specific mutations screened based on sequencing data from 12,854 NSCLC cases reported in TCGA, COSMIC, and literature, for NSCLC MRD monitoring. This probe set covers 12 NSCLC-related drug- or high-frequency genes: EGFR, KRAS, ERBB2, BRAF, MET, BRAF, ALK, TP53, PIK3CA, CTNNB1, CDKN2A, and NFE2L2, as shown in Table 1. The probe sequences are shown in Table 2. Each probe is 120 bp in length, and the 5' end of the probe is biotin-labeled.

[0040] The probe set described in this disclosure can achieve good coverage and capture of the target region, obtain sufficient read information on the target region, thereby playing a role in monitoring tumor evolution and new drug resistance mutations, while also stabilizing the experimental system and improving capture efficiency.

[0041] Those skilled in the art will understand that the probe sets disclosed herein can be used alone or in combination with any other conventional probe sets.

[0042] In some embodiments, the probe set of this disclosure can be used in combination with a patient-specific probe set for detecting MRD, such as the personalized probe set for MRD detection described in Chinese Patent Application No. 202310890251.0.

[0043] The probe set provided in this disclosure can accurately detect mutations related to minimal residual disease (MRD) in non-small cell lung cancer, exhibiting good sensitivity and specificity, and has promising application prospects in MRD monitoring of non-small cell lung cancer.

[0044] definition

[0045] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly used in the field to which this invention pertains. For the purposes of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural forms, and vice versa.

[0046] Unless the context clearly indicates otherwise, the terms “a” and “an” as used herein include plural references.

[0047] The term "about" as used herein is as understood by one of ordinary skill in the art and varies within a certain range depending on the context in which it is used. If one of ordinary skill in the art is unfamiliar with the use of this term in the context in which it is used, "about" will mean a particular value plus or minus 10%.

[0048] In this application, the term "lung cancer" has its general meaning in the art as a disease involving uncontrolled cell growth in lung tissue, which in some cases leads to metastasis. Most primary lung cancers are lung cancers originating from epithelial cells. In some embodiments, lung cancer can be stratified into any of the aforementioned stages (e.g., latent, stage 0, stage IA, stage IB, stage IIA, stage IIB, stage IIIA, stage IIIB, or stage IV). The main types of lung cancer are small cell lung cancer (SCLC) and non-small cell lung cancer (NSCLC). In one specific embodiment, the subject has non-small cell lung cancer. As used herein, the term "non-small cell lung cancer," also known as non-small cell lung carcinoma (NSCLC), refers to epithelial lung cancer other than small cell lung cancer (SCLC). There are three main subtypes: adenocarcinoma, squamous cell carcinoma of the lung, and large cell lung cancer. Other less common types of non-small cell lung cancer include pleomorphic, carcinoid, salivary gland carcinoma, and unclassified carcinoma.

[0049] In this application, the term "MRD" can be an abbreviation for three terms: molecular residual disease, measurable residual disease, and minimal residual disease. MRD reflects the residual status of tumor lesions. After treatment, a small number of tumor cells may remain in the body of the tumor patient. These tumor cells may be so few as to not cause any symptoms or signs and are usually undetectable by traditional methods such as cytological microscopy or serological tests. Detection requires highly sensitive modern cutting-edge technologies such as flow cytometry, PCR, and NGS.

[0050] In this application, the term "cfDNA" refers to circulating free DNA, including DNA molecules that are naturally present in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). Although cfDNA is originally present in one or more cells of a large, complex biological organism (e.g., a mammal), cfDNA undergoes release from the cells into fluids present in the organism, and can therefore be obtained by obtaining a sample of the fluid without requiring an in vitro cell lysis step.

[0051] In this application, the term "ctDNA" refers to circulating tumor DNA, which is a DNA fragment derived from apoptosis, necrosis, or secretion of tumor cells. It contains the same gene variations and epigenetic modifications as tumor tissue DNA, such as point mutations, gene rearrangements, fusions, copy number variations, and methylation modifications. ctDNA detection can be applied to various aspects of early cancer screening, diagnosis and staging, guiding targeted drug therapy, efficacy evaluation, and recurrence monitoring. Combining the information on tumor-specific gene variations and methylation carried by ctDNA helps improve the sensitivity and specificity of detection, enabling earlier detection of cancer and playing a significant role in early cancer screening.

[0052] In this application, the term "gene marker" refers to a molecular indicator that has specific biological characteristics, biochemical features, or aspects that can be used to determine the presence or absence of a particular disease or condition and / or the severity of a particular disease or condition.

[0053] In this application, the term "marker" refers to a parameter associated with one or more biomolecules (i.e., "genetic marker"), such as naturally or artificially synthesized nucleic acids.

[0054] In this application, the term "sequence identity" refers to the "sequence identity percentage" or "sequence identity percentage" between two polynucleotides, which is the number of identical matching positions shared by sequences within a comparison window, taking into account additions or deletions (i.e., vacancies) that must be introduced for optimal alignment of the two sequences. A matching position is any location where the same nucleotide is present in both the target and reference sequences. Since vacancies are not nucleotides, vacancies present in the target sequence are not counted. Similarly, since nucleotides from the target sequence are counted but nucleotides from the reference sequence are not, vacancies present in the reference sequence are not counted. At least 60% sequence identity includes at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the total length of the sequence having sequence identity.

[0055] The percentage of sequence identity can be calculated as follows: determine the number of positions in both sequences where the same amino acid residue or nucleic acid base appears (the number of matching positions), divide the number of matching positions by the total number of positions in the comparison window, and multiply the result by 100 to obtain the percentage of sequence identity. Sequence comparison and determination of the percentage of sequence identity between two sequences can be accomplished using software that is readily available online and downloadable. Suitable software programs are available from various sources for protein and nucleotide sequence alignment. A suitable program for determining the percentage of sequence identity is bl2seq, which is part of the BLAST program suite available from the National Center for Biotechnology Information (NCBI) website (blast.ncbi.nlm.nih.gov). Bl2seq uses either the BLASTN or BLASTP algorithm for comparing two sequences. BLASTN is used to compare nucleic acid sequences, while BLASTP is used to compare amino acid sequences. Other suitable programs are, for example, Needle, Stretcher, Water, or Matcher, which are part of the EMBOSS suite of bioinformatics programs and are also available from the European Institute of Bioinformatics (EBI) at www.ebi.ac.uk / Tools / psa.

[0056] In this application, the term "CSCO Guidelines" refers to the clinical practice guidelines for various malignant tumors published by the Chinese Society of Clinical Oncology.

[0057] In this application, the term "TCGA database" refers to the Cancer Genome Atlas Program (TCGA). Currently, it contains data from 20,000 patients across 33 cancers. This includes data from various omics disciplines such as genomics, transcriptomics, epigenetics, and proteomics, as well as clinical sample information.

[0058] In this application, the term "Cosmic Database" refers to the Cancer Somatic Mutation Catalogue, a comprehensive database that records in detail driver genes associated with human cancer.

[0059] As used in this article, the term “sequencing” refers to the process of determining the sequence (e.g., the identity and order of monomeric units) of a biomolecule, such as a nucleic acid, like DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, double-strand sequencing, cyclic sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturing temperature co-amplification PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, synthetic sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, DNA nanosphere sequencing (DNBSEQ), complex probe anchored polymerization sequencing (cPAS), and combinations thereof. In some implementations, sequencing can be performed using a gene analyzer, such as those commercially available from Illumina, Inc., Pacific Biosciences, Inc., Applied Biosystems / Thermo Fisher Scientific, or BGI Genomics Co., Ltd. Examples include BGI's DNBseq sequencing platforms such as BGISEQ-500, BGISEQ-50, MGISEQ-2000, MGISEQ-200, DNBSEQ-T7, DNBSEQ-G99, and DNBSEQ-T20X2, or Illumina's HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, and NovaSeq6000.

[0060] As used herein, the terms "computer-readable medium" (e.g., data storage, data storage, etc.) or "computer-readable storage medium" refer to any medium that participates in providing instructions to a processor for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, but are not limited to, optical discs, solid-state drives, and magnetic disks, such as storage devices. Examples of volatile media include, but are not limited to, dynamic memory, such as RAM.

[0061] Common forms of computer-readable media include, for example, floppy disks, floppy disks, hard disks, magnetic tapes or any other magnetic media, CD-ROMs, any other optical media, punched cards, paper tapes, any other physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cassette tapes, or any other tangible media from which a computer can read.

[0062] In addition to computer-readable media, data may be provided as signals on a transmission medium included in a communication device or system to provide one or more sequences of instructions to a processor of a computer system for execution. For example, a communication device may include a transceiver having signals indicating instructions and data. The instructions and data are configured to cause one or more processors to perform the functions outlined in this disclosure. Representative examples of data communication transmission connections may include, for example, telephone modem connections, wide area networks (WANs), local area networks (LANs), infrared data connections, NFC connections, etc.

[0063] The following embodiments and accompanying drawings are provided to aid in understanding the present invention. However, it should be understood that these embodiments and drawings are for illustrative purposes only and do not constitute any limitation. The actual scope of protection of the present invention is set forth in the claims. It should be noted that various modifications and improvements made by those skilled in the art based on this inventive concept are all within the scope of protection of the present invention. In the following description, descriptions of well-known structures and techniques are omitted to avoid unnecessarily obscuring the concepts of this disclosure. Such structures and techniques have also been described in many publications. Unless otherwise specified, the equipment, instruments, reagents, and / or kits used in the following embodiments are commercially available or obtained through conventional methods known to those skilled in the art.

[0064] Example

[0065] Example 1: Screening of targets and probe combinations for non-small cell lung cancer

[0066] To screen for drug targets and high-frequency mutation sites in non-small cell lung cancer (NSCLC), we analyzed and calculated drug targets and high-frequency mutation regions related to NSCLC based on the database and sequencing data of 12,854 NSCLC patients, as shown in Table 1.

[0067] Table 1. Combination of cancer markers for non-small cell lung cancer detection

[0068]

[0069]

[0070] Based on the mutations and high-frequency mutation regions of non-small cell lung cancer, a cancer-specific probe set was designed, and the probe sequences are shown in Table 2.

[0071] Table 2. Non-small cell lung cancer specific MRD probe group

[0072]

[0073]

[0074]

[0075]

[0076] Example 2: Evaluation of the capture efficiency and uniformity of non-small cell lung cancer specific probes on clinical samples

[0077] The performance of non-small cell lung cancer-specific probes was tested using cfDNA samples from clinical patients with non-small cell lung cancer to evaluate the probe's capture efficiency and uniformity.

[0078] (1) Experimental procedure:

[0079] ①cfDNA extraction: Whole blood samples were centrifuged at 1,600g and 16,000g in two steps to separate plasma and remove cell debris from the plasma. Then, MagMAX was used to extract the plasma. TM Cell-free DNA isolation kit (Thermo Fisher) was used for plasma cfDNA extraction using magnetic beads.

[0080] ② Library Construction: The cfDNA was end-repaired and an "A" was added using a customized human molecular residual lesion (MRD) detection kit (GenePlus). Then, the DNA underwent adapter ligation, purification, pre-capture PCR (Non-C-PCR), and further purification to obtain a pre-capture intermediate library. Samples with acceptable intermediate library concentrations were then subjected to subsequent hybridization and elution.

[0081] ③ Hybridization Capture: Using a customized human molecular residual lesion (MRD) detection kit (GenePlus), libraries that passed concentration quality control were subjected to pooling, evaporation, hybridization with mixed probes, elution, PCR of the elution products, and purification to obtain a hybridized general library. The general library was then sequenced after passing concentration and fragment distribution quality control.

[0082] ④ Sequencing and FASTQ data output: Paired-end (PE100) sequencing was performed using the Gene+Seq2000 sequencer, and the data was split into fastq files after being downloaded from the sequencer.

[0083] ⑤ Data alignment and BAM file generation: Before data alignment, Realseq2 software (version: 1.1.6) was used to: (1) remove UMIs from the ends of reads and save them in the read name; (2) filter low-quality reads. The obtained fastq was aligned to the human reference genome (version: hs37d5) using BWA software (version: 0.7.15-r1140) to generate the initial alignment result BAM file. Then, Realseq2 software (version: 201808) was used to perform clustering and error correction on the PCR repeat reads in the initial alignment result file with the help of UMIs. The indel regions within a 50bp range extending from both ends of the detection chip were re-aligned using common indel mutations from the Thousand Talents Database and the dbSNP (version: 138) database. The base quality values ​​within a 50bp range extending from both ends of the detection chip were re-corrected using information from the Thousand Talents Database, the dbSNP (version: 138) database, and the COSMIC database.

[0084] ⑥ Sample quality control: (1) Sample pairing error: The bioinformatics process judges the sample pairing status by calculating the consistency between homozygous sites in the control sample within a 50bp range extending from both ends of the chip interval and the tumor sample. (2) Sample contamination: The Calculate Contamination module in GATK (version: 4.1.4) software is used to combine the bam file information of the control and tumor samples. The reads information of the supporting reference bases in the homozygous sites in the test samples are read and counted to evaluate the cross-contamination of the samples.

[0085] ⑦ Mutation calling: This product detects single nucleotide variants (SNVs) and insertion / deletion mutations (Indels) within a 50bp range extending from both ends of the chip capture region, as well as SVs. The mutations obtained from the above detection process are annotated using the following databases: (1) Gene Annotation Database (version: NCBI release 104); (2) dbSNP database (version: 147); (3) tgp database (version: phase3); (4) COSMIC database (version: v80); (5) ExAC database (version: 0.3.1); (6) clinvar database (version: 20200701). The mutations obtained from the above steps are filtered, and reliable mutations are retained.

[0086] In this embodiment, cfDNA samples from 10 clinical patients with non-small cell lung cancer were used to test probe capture efficiency, uniformity, and probe depth coefficient. The probes used in step ③, hybridization capture, were either a personalized probe set for detecting non-small cell lung cancer (the nucleotide sequences of the personalized probes are shown in Table 3 as SEQ ID NO: 58-79, hereinafter referred to as the non-small cell lung cancer personalized probe set) used alone, or a non-small cell lung cancer specific probe set provided in this application (the nucleotide sequences of the probes are shown in Table 2 as SEQ ID NO: 1-57, hereinafter referred to as the non-small cell lung cancer specific probe set) used alone, or a combination of both.

[0087] Table 3. Nucleotide sequences of personalized probe sets for detecting non-small cell lung cancer

[0088]

[0089]

[0090] (2) Probe capture efficiency

[0091] The probe capture efficiency test results for clinical samples are as follows: Figure 1 As shown in the figure. The results indicated that the capture efficiency of the personalized probe group for non-small cell lung cancer (NSCLC) ranged from 24.5% to 37.0%, with a median capture efficiency of 32.8%. The capture efficiency of the specific probe group for NSCLC ranged from 48.7% to 57.6%, with a median capture efficiency of 52.5%. The capture efficiency of the combined use of the personalized probe group and the specific probe group for NSCLC ranged from 50.8% to 58.7%, with a median capture efficiency of 53.8%. These results demonstrate that the capture efficiency of the specific probe group for NSCLC is significantly increased compared to using the personalized probe group alone, and that the combined use of the personalized probe group and the specific probe group for NSCLC can significantly improve the capture efficiency.

[0092] (3) Probe uniformity

[0093] The probe homogeneity test results of the non-small cell lung cancer specific probe group alone are as follows: Figure 2 and Figure 3 As shown. The results are as follows. Figure 2 As shown, the percentage of probes with a depth ≥0.2 times the average depth was 100% in all 10 cfDNA samples. Figure 3 As shown, among the 10 cfDNA samples, the percentage of probes with a depth ≥0.5 times the average depth was 98.7% at the minimum, 99.7% at the maximum, and 99.3% at the median. These results indicate that the non-small cell lung cancer-specific probe group exhibits excellent homogeneity and meets the needs of clinical testing.

[0094] (4) Probe depth coefficient

[0095] The depth coefficient results of the non-small cell lung cancer specific probe group are as follows: Figure 4 As shown in the figure. The test results show that the median value of the probe depth coefficient of the 10 samples is between 1.00 and 1.04, indicating that the non-small cell lung cancer specific probe group has a relatively consistent depth and the depth coefficient is stable in different samples.

[0096] Example 3: The accuracy of using a non-small cell lung cancer specific probe set to detect the MRD status of non-small cell lung cancer.

[0097] Nineteen clinical non-small cell lung cancer (NSCLC) patients were included to confirm the accuracy of NSCLC-specific probes in detecting MRD status. First, a comprehensive genomic analysis was performed on FFPE samples from the 19 patients to identify the mutation profiles for each patient. Then, based on the mutation profiles, 2-20 mutations were selected to design patient-specific personalized probes. The sequences of these patient-specific personalized probes are shown in Table 4. These probes were used in conjunction with the NSCLC-specific probe group for postoperative blood MRD detection. Patients were followed up every 3-6 months postoperatively, and 20 mL of peripheral blood was collected for MRD detection. A total of 40 cfDNA samples from the 19 clinical NSCLC patients were analyzed.

[0098] Table 4. Personalized probes customized for 19 clinical non-small cell lung cancer patients

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107] Note: If the mutations detected in the tissue are covered by the non-small cell lung cancer-specific probe set, no personalized probes will be customized for MRD monitoring.

[0108] First, the coverage of mutations detected in tissues from 19 patients by the non-small cell lung cancer-specific probe group was evaluated. Results are as follows: Figure 5As shown, 100% of the tissue samples were covered by the non-small cell lung cancer-specific probe group with at least one mutation, and the median number of mutations covered per sample was 2, indicating that the non-small cell lung cancer-specific probe group has good coverage for non-small cell lung cancer patients.

[0109] Then, the accuracy of combining the cancer-specific probe group and the personalized probe group in detecting MRD status in non-small cell lung cancer was evaluated. Imaging confirmed that 9 of the 19 patients had recurrence, and 10 were relapse-free. MRD results are as follows: Figure 6 As shown, among the 9 recurrent patients, 8 were identified as MRD-positive, with a sensitivity of 89% (8 / 9), and all 10 non-recurrent patients were identified as MRD-negative, with a specificity of 100% (10 / 10). This indicates that the non-small cell lung cancer (NSCLC) tumor-specific probe group provided by this invention has high accuracy for NSCLC MRD monitoring. Furthermore, among the 8 recurrent and MRD-positive patients, mutations detected in 9 plasma samples from 6 patients were covered by the NSCLC tumor-specific probe group. The detected mutations are shown in Table 5, with a patient coverage rate of 75% (6 / 8), indicating that the NSCLC tumor-specific probe group has high coverage for NSCLC MRD-positive patients.

[0110] Table 5. Coverage results of the 19 non-small cell lung cancer specific probe group for MRD-positive patients.

[0111]

[0112] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.

Claims

1. A probe set for detecting minimal residual disease in non-small cell lung cancer, characterized in that, The probe set is used to detect gene markers of minimal residual disease in non-small cell lung cancer, and the gene markers are composed of the following genes: EGFR, KRAS, ERBB2, BRAF, MET, ALK, TP53, PIK3CA, CTNNB1, SMAD4, CDKN2A and NFE2L2. The genetic markers cover the following regions: exons 18, 19, 20, and 21 of EGFR; exons 2, 3, and 4 of KRAS; exon 20 of ERBB2; exons 11 and 15 of BRAF; introns 13 and 14 of MET; exon 23 of ALK; exons 4-10 of TP53; exons 10 and 21 of PIK3CA; exon 3 of CTNNB1; exon 9 of SMAD4; exons 1 and 2 of CDKN2A; and exon 2 of NFE2L2. The probe set contains nucleotide sequences as shown in SEQ ID NO:1 to 57.

2. A kit for detecting minimal residual disease in non-small cell lung cancer, characterized in that, The kit includes the probe set as described in claim 1.

3. The use of the probe set according to claim 1 in the preparation of a kit for detecting minimal residual disease in non-small cell lung cancer.

4. The application according to claim 3, characterized in that, The kit is used for the detection of cell, tissue, or body fluid samples.

5. The application according to claim 4, characterized in that, The body fluid samples include: saliva, whole blood, serum, plasma, breast milk, urine, lumbar or ventricular CSF, bile, lymph, prostatic fluid, semen, sputum, feces, tears, tumor cells, bronchoalveolar lavage fluid, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural effusion, peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, vaginal discharge, and their processed forms; and / or, The tissue samples included lung cancer tissue, paraffin sections, and their processed forms.

Citation Information

Patent Citations

  • A design method for a personalized probe set for MRD detection and its application

    CN117144002B

  • Application of gene marker in prognosis evaluation of non-small cell lung cancer, detection device and computer readable medium

    CN116287233A