Biomarker, kit and application for detecting microsatellite instability of cancer
By using biomarker compositions and high-resolution melting curve methods, the problems of low sensitivity and insufficient applicability in the detection of cancer microsatellite instability in existing technologies have been solved, enabling efficient and accurate detection of endometrial cancer, gastric cancer, and colorectal cancer.
Patent Information
- Application Number
- CN202511134214.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing technologies for detecting cancer microsatellite instability (MSI) suffer from low sensitivity, complex operation, high cost, and limited applicability, especially in endometrial cancer, gastric cancer, and colorectal cancer, where efficient and accurate detection is difficult.
A biomarker composition, including a specificly located single nucleotide tandem repeat sequence marker, combined with a high-resolution melting curve method, was used to design specific primers and probes to achieve efficient detection of cancer microsatellite instability.
It achieves highly sensitive, specific, and low-cost detection of MSI status in endometrial cancer, gastric cancer, and colorectal cancer, with a 100% concordance rate with the IHC method.
Smart Images

Figure CN120775981B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of microsatellite instability detection technology, specifically relating to biomarkers, reagent kits, and applications for detecting cancer microsatellite instability. Background Technology
[0002] Microsatellites, also known as short tandem repeats (STRs), are DNA sequences typically composed of short, repeating units of 1 to 6 bases, usually repeated 10 to 60 times, and distributed throughout the human genome. Because of their structure as multiple base repetitions, these sequences are prone to errors during DNA replication. Genes in the human body with mismatch repair (MMR) functions, primarily MLH1, MSH2, MSH6, and PMS2, identify and repair errors that occur during replication to maintain the stability of microsatellite sequences. However, when mismatch repair is deficient (MMRd), errors generated during microsatellite replication cannot be corrected and accumulate, leading to changes in microsatellite sequence length or base composition. This phenomenon is known as microsatellite instability (MSI). The MSI status of tumors can be classified into three phenotypes: high-frequency MSI (MSI-H), low-frequency MSI (MSI-L), and microsatellite stable (MSS). Although MSI-L tumors differ from MSS tumors in terms of MSI level, their pathological features and most molecular characteristics are very consistent.
[0003] The phenomenon of MSI (Multiple Sign-Induced Infection) was first discovered in hereditary colorectal cancer, and subsequent studies have confirmed its presence in various cancers, including endometrial cancer, gastric cancer, urethral cancer, and ovarian cancer. Endometrial cancer exhibits the highest MSI frequency, at 21.9%; small bowel cancer is second, at 14.3%; colorectal cancer and gastric cancer have MSI frequencies of 10.2% and 8.5%, respectively. In other cancer types, the MSI frequency is relatively low; for example, ovarian cancer has an MSI rate of 2%–10%, and pancreatic cancer has an MSI rate of only 1%–2%.
[0004] MSI testing is an important initial screening method for Lynch syndrome, and MSI status is a key factor in predicting the prognosis and efficacy of adjuvant chemotherapy in colorectal cancer. It is also a predictive indicator of the efficacy of immunotherapy in advanced solid tumors (including colorectal cancer, gastric cancer, and endometrial cancer). Particularly in endometrial cancer, MSI status is an important marker for molecular subtyping and prognostic assessment. MSI has become a pan-cancer molecular target, and its importance is constantly increasing. Therefore, clinically, there is a need for a precise detection method applicable to these three cancer types with high MSI frequency to effectively identify MSI status.
[0005] Currently, the main clinical methods for detecting MSI are immunohistochemistry (IHC), PCR + capillary electrophoresis, NGS, and PCR + high-resolution melting curve analysis.
[0006] IHC (Inductively Coupled Chromatography) primarily determines MSI status by detecting the absence of MMR proteins and is currently the most widely used method in clinical practice. Specifically, the absence of expression of one or more of the four MMR proteins indicates deficient mismatch repair (dMMR), equivalent to MSI-H; while the complete expression of all four proteins indicates proficient mismatch repair (pMMR), equivalent to MSI-L / MSS. However, the interpretation of this method requires pathologists, and there is no unified authoritative standard in China. Procedures may vary between hospitals, and subjective judgments by different laboratory personnel may also lead to different results. Furthermore, the antibodies used by different manufacturers and their staining efficiency for different tissues and cell types vary; some antibodies may not effectively bind to the target protein, resulting in unreliable staining results or failure to detect low-expression proteins. Therefore, molecular-level detection methods are currently recommended in clinical practice for cross-verification with IHC results.
[0007] PCR-capillary electrophoresis combines multiplex PCR amplification with capillary electrophoresis to select specific microsatellite loci and simultaneously detect MSI status in both patient tumor tissue and paired normal tissue. The loci used in this method are primarily based on or slightly modified combinations of loci recommended by the National Cancer Institute (NCI). One typical combination includes two single nucleotides (BAT-25 and BAT-26) and three dinucleotides (D2S123, D5S346, and D17S250). The test compares the DNA of tumor and normal tissue from the same patient. If two or more of the five microsatellite sequences are mutated, the tumor is classified as high microsatellite instability (MSI-H); if only one sequence is mutated, it is low microsatellite instability (MSI-L); and if none of the sequences are mutated, it is microsatellite stable (MSS). Another typical combination consists of five single nucleotide repeat sequences (BAT-25, BAT-26, NR-21, NR-22, and NR-24). This method is the "gold standard" for MSI molecular-level detection, but it first requires multiplex PCR amplification followed by capillary electrophoresis of the products, making the entire process cumbersome and taking over 5 hours. Even with a uracil glycosylase (UDG) anti-contamination system, the electrophoresis of PCR products during detection still carries a relatively high risk of product contamination. Furthermore, because this method relies on capillary electrophoresis, its detection sensitivity is relatively low, requiring a tumor DNA content greater than 20% or even 30%. Finally, this method requires not only tumor tissue but also normal tissue for control testing, increasing the consumption of clinical samples and raising testing costs.
[0008] NGS (Non-Genome Sequencing) employs whole-genome sequencing, whole-exome sequencing, or targeted sequencing technologies to detect samples, covering anywhere from several to hundreds of thousands of microsatellite loci, providing a more comprehensive assessment of MSI status. However, the unique structure of these repetitive sequences makes NGS sequence capture and sequencing more challenging than with conventional sequences. This results in greater complexity in library preparation, sequencing parameter tuning, and bioinformatics analysis, making the overall process the most complex and time-consuming, typically requiring several days or even weeks. Furthermore, NGS often detects dozens or even hundreds of loci and is not specifically designed for MSI detection. It is typically used in conjunction with tumor somatic mutation detection or tumor mutational burden (TMB) testing, including MSI loci and providing MSI status information. While this adds more genetic information, it unnecessarily increases costs and time compared to MSI detection itself, hindering the large-scale clinical adoption of MSI testing. Finally, there is currently a lack of unified standards for interpreting the results of NGS detection of MSI. Different sequencing platforms and the algorithms on which the panels are based have different MSI positive thresholds, which increases the difficulty of using this method in clinical practice and is not conducive to large-scale promotion.
[0009] The PCR + high-resolution melting curve method combines multiplex PCR and high-resolution melting curve analysis. The PCR amplification products are analyzed directly using high-resolution melting curves without opening the container. Due to the difference in microsatellite and MSS sequence lengths between MSI-H and MSS, the Tm values of the product melting curves will differ. The MSI status is determined based on the difference in Tm values. This method eliminates the need for capillary electrophoresis, simplifying the procedure. Furthermore, by specifically selecting MSI sites, normal tissue controls are not required, allowing detection only in tumor tissues, reducing costs and making it more suitable for clinical use. The key to this method is the selection of highly sensitive and specific MSI sites and combinations. Currently, common site combinations are mostly selected for colorectal cancer and are more suitable for this type of cancer. These sites and combinations are not very sensitive for other cancer types, especially endometrial cancer, and have lower detection sensitivity compared to IHC, potentially missing a large number of samples. For example, Dedeurwaerdere F, Claes KB, Van Dorpe J, Rottiers I, Van der Meulen J, Breyne J, Swaerts K, Martens G. Comparison of microsatellite instability detection by immunohistochemistry and molecular techniques in colorectal and endometrial cancer. Sci Rep. 2021 Jun 18;11(1):12880. doi: 10.1038 / s41598-021-91974-x. PMID: 34145315; PMCID: PMC8213758 (Reference 1) describes a performance comparison of a combination of MSI sites based on melting curve analysis in colorectal and endometrial cancer. This combination of MSI sites includes 7 genes: DIDO1、 ACVR2A, MRE11, BTBD7, SULF2, SEC31A and RYR3 Its concordance rate with IHC in detecting colorectal cancer samples is 100%, but its concordance rate with IHC in detecting endometrial cancer samples is only 79%. Furthermore, there is limited data on gastric cancer detection, making it difficult to accurately assess its performance and thus it cannot yet meet clinical needs.
[0010] Given the shortcomings of the existing detection methods, and considering the actual clinical testing needs and the frequency of MSI in different cancers, there is an urgent need to develop a method that is highly sensitive, easy to operate, low in cost, and widely applicable, so as to achieve efficient detection of MSI status in cancers including endometrial cancer, gastric cancer, and colorectal cancer. Summary of the Invention
[0011] Based on this, the purpose of the present invention is to provide a biomarker for detecting cancer microsatellite instability. The biomarker is a composition that can simultaneously achieve efficient detection of MSI status in endometrial cancer, gastric cancer, and colorectal cancer, with the detection results showing a concordance rate of up to 100% with the IHC method.
[0012] A first aspect of the present invention is to provide a biomarker composition comprising the following third, seventh, eighth, eleventh, thirteenth, fourteenth, and sixteenth biomarkers:
[0013] The third biomarker is located in the FOXP2 gene, starting at chr7:114690176, and contains 11 consecutive A single nucleotide tandem repeat sequences.
[0014] The seventh biomarker is located in the HAL gene, starting at chr12:95985886, and contains 11 consecutive T single nucleotide tandem repeat sequences.
[0015] The eighth biomarker is located in the TSHZ2 gene, starting at chr20:53494836, and contains 11 consecutive A single nucleotide tandem repeat sequences.
[0016] The eleventh biomarker is located in the LRBA gene, starting at chr4:150914351, and contains 11 consecutive A single nucleotide tandem repeat sequences.
[0017] The thirteenth biomarker is located in the DEFB105A gene, starting at chr8:7822206, and contains a single nucleotide tandem repeat sequence of 9 consecutive A's.
[0018] The fourteenth biomarker is located in the JPT2 gene, starting at chr16:1701589, and contains 11 consecutive T single nucleotide tandem repeat sequences.
[0019] The sixteenth biomarker is located in the PTPRF gene, starting at chr1:43622447, and contains 11 consecutive T single nucleotide tandem repeat sequences.
[0020] In some embodiments, the biomarker further includes at least one of the following second, fourth, and ninth biomarkers:
[0021] The second biomarker is located in the ZBTB37 gene, starting at chr1:173899196, and contains 11 consecutive T single nucleotide tandem repeat sequences.
[0022] The fourth biomarker is located in the NBAS gene, starting at chr2:15396386, and contains 11 consecutive A single nucleotide tandem repeat sequences.
[0023] The ninth biomarker is located in the LRRC8D gene, starting at chr1:89935847, and contains 11 consecutive T single nucleotide tandem repeat sequences.
[0024] In some of these embodiments, the biomarker composition is as follows (1), (2), or (3):
[0025] (1) The third, fourth, seventh, eighth, eleventh, thirteenth, fourteenth and sixteenth markers;
[0026] (2) The second, third, seventh, eighth, eleventh, thirteenth, fourteenth and sixteenth markers;
[0027] (3) The third, seventh, eighth, ninth, eleventh, thirteenth, fourteenth and sixteenth markers.
[0028] In some preferred embodiments, the biomarker composition is (1).
[0029] A second aspect of the present invention is to provide the use of the biomarker composition described above in the preparation of a reagent for detecting cancer microsatellite instability.
[0030] A third aspect of the invention is to provide the use of the biomarker composition described above in the preparation of reagents for detecting cancer.
[0031] A fourth aspect of the present invention is to provide a kit for detecting cancer or cancer microsatellite instability, comprising reagents for detecting mutations in the biomarker composition as described above.
[0032] In some implementations, a high-resolution melting curve method is used, and the kit includes amplification primers and probes in groups C, G, H, K, M, N, and P:
[0033] Group C: SEQ ID NO: 29 and SEQ ID NO: 30, and SEQ ID NO: 75, which are for the third marker;
[0034] Group G: SEQ ID NO: 37 and SEQ ID NO: 38, and SEQ ID NO: 79 for the seventh marker;
[0035] Group H: SEQ ID NO: 39 and SEQ ID NO: 40, and SEQ ID NO: 80 for the eighth marker;
[0036] Group K: SEQ ID NO: 45 and SEQ ID NO: 46, and SEQ ID NO: 83 for the eleventh marker;
[0037] Group M: SEQ ID NO: 49 and SEQ ID NO: 50, and SEQ ID NO: 85, which are for the thirteenth marker;
[0038] Group N: SEQ ID NO: 51 and SEQ ID NO: 52, and SEQ ID NO: 86, for the fourteenth marker;
[0039] Group P: SEQ ID NO: 55 and SEQ ID NO: 56, and SEQ ID NO: 88 for the sixteenth marker.
[0040] In some embodiments, the kit further includes at least one of the following groups: Group B, Group D, and Group I amplification primers and probes:
[0041] Group B: SEQ ID NO: 27 and SEQ ID NO: 28, and SEQ ID NO: 74, for the second marker;
[0042] Group D: SEQ ID NO: 31 and SEQ ID NO: 32, and SEQ ID NO: 76 for the fourth marker;
[0043] Group I: SEQ ID NO: 41 and SEQ ID NO: 42, and SEQ ID NO: 81, for the ninth marker.
[0044] In some embodiments, the kit includes amplification primers and probes in groups C, D, G, H, K, M, N, and P.
[0045] In some embodiments, the kit includes amplification primers and probes in groups B, C, G, H, K, M, N, and P.
[0046] In some embodiments, the kit includes amplification primers and probes in groups C, G, H, I, K, M, N, and P.
[0047] In some of these embodiments, the cancer includes at least one of endometrial cancer, gastric cancer, and colorectal cancer.
[0048] A fifth aspect of the present invention is to provide a method for detecting microsatellite instability for non-diagnostic purposes, comprising the following steps: obtaining DNA from a sample to be tested, and using the kit described above to detect mutations in the biomarker composition described above.
[0049] In some of these embodiments, the mutation status of the biomarker composition (1) is detected using a high-resolution melting curve method. During the detection, the amplification primers and probes of the biomarker are grouped as follows: M group and P group are in the same reaction tube, C group and D group are in the same reaction tube, G group and K group are in the same reaction tube, and H group and N group are in the same reaction tube.
[0050] This invention, through extensive research and continuous optimization of analytical methods and combination strategies, has identified a biomarker composition suitable for simultaneously detecting microsatellite instability in endometrial cancer, gastric cancer, and colorectal cancer. This biomarker composition can simultaneously detect microsatellite instability in these three cancers with as few as eight biomarkers, exhibiting high sensitivity and specificity, with a 100% concordance rate with IHC methods. The number of biomarkers in this invention's biomarker composition is significantly lower than that of existing compositions achieving equivalent detection results, effectively reducing detection costs and improving efficiency, thus possessing significant clinical application value.
[0051] This invention also provides a kit for detecting mutations in the biomarker composition. The kit includes amplification primers and probes for the biomarker composition, which can effectively distinguish between wild-type and mutant types of each biomarker in the composition based on Tm values. It exhibits high sensitivity and specificity, accurately detecting microsatellite instability in samples. Furthermore, the detection operation is simple, and the result interpretation is more reasonable and intuitive. Attached Figure Description
[0052] Figure 1 This is the amplification curve of the reference sample corresponding to the detection limit in tube 1 of grouping method 1 in Example 7.
[0053] Figure 2 This is the amplification curve of the reference sample corresponding to the detection limit in tube 2 of grouping method 1 in Example 7.
[0054] Figure 3 This is the amplification curve of the reference sample corresponding to the detection limit in tube 3 of grouping method 1 in Example 7.
[0055] Figure 4 This is the amplification curve of the reference sample corresponding to the detection limit in tube 4 of grouping method 1 in Example 7.
[0056] Figure 5This is the amplification curve of the reference sample corresponding to the detection limit in tube 1 of grouping method 2 in Example 7.
[0057] Figure 6 This is the amplification curve of the reference sample corresponding to the detection limit in tube 2 of grouping method 2 in Example 7.
[0058] Figure 7 This is the amplification curve of the reference sample corresponding to the detection limit in tube 3 of grouping method 2 in Example 7.
[0059] Figure 8 This is the amplification curve of the reference sample corresponding to the detection limit in tube 4 of grouping method 2 in Example 7. Detailed Implementation
[0060] To facilitate understanding of the present invention, a more complete description will be provided below. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the present invention.
[0061] Unless otherwise specified, experimental methods in the following examples were performed under standard conditions, such as those described in the fourth edition of *Molecular Cloning: A Laboratory Manual*, edited by Green and Sambrook, published in 2013, or according to the manufacturer's recommendations. All commonly used chemical reagents used in the examples are commercially available products.
[0062] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. The term "and / or" as used in this invention includes any and all combinations of one or more of the associated listed items.
[0063] The present invention will be further described in detail below with reference to specific embodiments.
[0064] Example 1: NGS-based site screening
[0065] First, a site panel for NGS sequencing was constructed, incorporating data from various databases, including the TCGA database (TCGA-UCEC, TCGA-STAD, TCGA-COAD, and TCGA-READ), the GEO database, and the cBioPortal database, as well as literature sources, to collect high-frequency MSI sites. Analysis revealed very little MSI information in these databases, especially for the Chinese population. To better reflect the actual MSI detection situation in the Chinese population and to cover as many sites as possible to avoid omissions, whole-exome sequencing was performed on 16 tumor tissue samples to analyze high-frequency MSI sites. Combining the high-frequency MSI sites obtained from the literature and those from sequencing real Chinese population samples, an NGS site panel covering 3237 MSI sites was constructed. Then, samples from three cancer types—endometrial cancer, gastric cancer, and colorectal cancer—were collected, and a customized panel was used to simultaneously perform targeted sequencing on tumor tissue and paired normal tissue. The sequencing results were analyzed using bioinformatics software, combined with IHC results from clinical samples, to select corresponding high-frequency sites for each of the three cancer types for further PCR screening. The specific loci are shown in Table 1 below, mapped to the GRCh38 / hg38 human reference genome.
[0066] Table 1
[0067]
[0068]
[0069]
[0070]
[0071]
[0072] Twelve gene loci, including UBE2Z, ZBTB37, FOXP2, NBAS, ZBTB16, SEPTIN7, HAL, TSHZ2, LRRC8D, PBRM1, LRBA, and NRCAM, were identified from endometrial cancer samples; six gene loci, including DEFB105A, JPT2, PI15, PTPRF, ACVR2A, and NBAS, were identified from gastric cancer samples; and eight gene loci, including ACVR2A, CENPQ, SLC22A9, LRIG2, DIDO1, MRE11, TGFBR2, and PSIP1, were identified from colorectal cancer samples.
[0073] Example 2 Construction of a single-site PCR detection system
[0074] This invention uses PCR combined with high-resolution melting curve analysis to detect the MSI sites obtained in Example 1, and determines the MSI status based on the difference in melting curve Tm values. During detection, DNA fragment products containing the MSI sites are amplified by PCR, and the mutation status of the MSI sites is analyzed using the melting curve Tm values. In this embodiment, primers and probes were designed for the DNA mutation sequences containing the MSI sites shown in Table 2 below; the underlined areas indicate the locations of the MSI sites.
[0075] Table 2
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085] Upstream and downstream primers were designed for the mutant sequences in Table 2. The Tm values of the primers were designed to be between 58℃ and 61℃, and the 3' ends of the primers should be free of hairpins or dimers, with the amplified product length controlled between 100-500 bp. PCR amplification employed an asymmetric amplification method, where one primer was a "restriction primer" used in smaller quantities, and the other was an "excess primer" used in larger quantities. The purpose of using asymmetric amplification was to amplify more single-stranded template for specific probe binding, facilitating subsequent melting curve analysis. After screening, the upstream and downstream primer sequences for each marker are shown in Table 3 below.
[0086] Table 3
[0087]
[0088]
[0089]
[0090]
[0091]
[0092] Detection probes were designed for the MSI biomarkers screened by NGS in Example 1. The probes were designed for the mutant sequences described in Table 2, specifically using sequences missing one base as templates. This ensures a 100% match between the probe and the mutant template, but not a 100% match with the wild-type template, resulting in different Tm values for the probe-template binding products. Specifically, the probe's Tm value should be between 40℃ and 70℃. Furthermore, to prevent degradation by the proofreading activity of DNA amplification enzymes at the 3'-5' ends, the probes can be protected with thiolated bases at the 3' end, with 1-4 thiolated bases. The probes can be molecular beacon probes or TaqMan probes. Both ends of the probe require modification with reporter and quencher groups. Reporter group modifications include FAM, VIC, ROX, Cy5, RED, Cy3, etc., while quencher groups include BHQ1, BHQ2, BHQ3, MGB, and Dabcyl, etc. After screening, the probe design sequences for each MSI biomarker are shown in Table 4 below. "The mark refers to thiomodification."
[0093] Table 4
[0094]
[0095] A single-site detection system was constructed using the primer and probe combination designed above, and a reference standard was set up to evaluate the performance of the detection system. PCR reaction solution was prepared using primers, probes, buffer, and DNA polymerase. The amount of restriction primer was 0.16 μM, the amount of excess primer was 0.64 μM, the amount of probe was 0.08 μM, the amount of magnesium ions was 2 mM, the amount of dNTPs was 0.15 mM, the amount of DNA polymerase was 1.4 U, and the amount of UNG enzyme was 0.2 U. The buffer volume was 5 μl, which was supplemented with nuclease-free water to a final volume of 23 μl. Finally, 2 μl of the reference sample was added to prepare a 25 μl complete PCR system for instrumental analysis. The detection process primarily used a real-time PCR instrument with high-resolution melting curve functionality. Denaturation and annealing extension temperatures were set, and the amplification cycle count was set to 60 cycles. The melting curve collection temperature range was set to 30-70℃. The detection process involved first performing PCR amplification to enrich the single-stranded product, then detecting the binding of the single-stranded template and probe during the temperature rise process. Changes in fluorescence signal intensity were observed, and the derivative of the fluorescence change was used to obtain the melting curve peak diagram for interpretation. The instrumental procedure is shown in Table 5 below.
[0096] Table 5
[0097]
[0098] Criteria for judging the results of unit spot system tests: For negative reference samples corresponding to unit spot system tests, the result peak should be single-peaked, with a peak only appearing at low Tm values, indicating good specificity of the system at that site; for positive reference samples corresponding to unit spot system tests, the result peak should be double-peaked or single-peaked, with a peak at high Tm values, indicating good accuracy of the system at that site; for limit of detection reference samples corresponding to unit spot system tests, the result peak should be double-peaked, with two Tm values, one low Tm value and one high Tm value, indicating good limit of detection of the system at that site. If the test results of the above three reference samples are all consistent, it indicates that the performance of the unit spot system meets the requirements for subsequent testing.
[0099] Example 3: Single Cancer Site Detection and Screening
[0100] Following the PCR detection system construction method described in Example 2 above, a detection system with 24 biomarker loci was constructed. A total of 238 samples were collected, including 62 cases of endometrial cancer, 88 cases of gastric cancer, and 88 cases of colorectal cancer, as the test set. Biomarkers for the corresponding cancer types were used for detection. Specifically, 62 endometrial cancer samples were tested using a 12-gene locus system including UBE2Z, ZBTB37, FOXP2, NBAS, ZBTB16, SEPTIN7, HAL, TSHZ2, LRRC8D, PBRM1, LRBA, and NRCAM; 88 gastric cancer samples were tested using a 6-gene locus system including DEFB105A, JPT2, PI15, PTPRF, ACVR2A, and NBAS; and 88 colorectal cancer samples were tested using a 8-gene locus system including ACVR2A, CENPQ, SLC22A9, LRIG2, DIDO1, MRE11, TGFBR2, and PSIP1. The IHC results were used as the gold standard to evaluate the detection performance of each locus for a single cancer type. Based on a review of the literature, IHC, a protein-level detection method, tends to detect more positive results than PCR, a molecular-level detection method. Typically, IHC may show a positive result while PCR combined with capillary electrophoresis may show a negative result. To improve consistency with IHC, IHC is chosen as the gold standard for comparative validation.
[0101] Endometrial cancer samples were tested using a single site of endometrial cancer, and the results were compared with those obtained by IHC. Detailed results for each site are shown in Table 6 below.
[0102] Table 6
[0103]
[0104] Twelve sites were selected for combined detection of endometrial cancer samples. The traditional method of identifying two or more sites as positive was used to determine the MSI-H method. The results were compared with those of the IHC method, as shown in Table 7 below.
[0105] Table 7
[0106]
[0107] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.
[0108] Gastric cancer samples were tested using a single site of gastric cancer, and the results were compared with those obtained by IHC. The specific results for each site are shown in Table 8 below.
[0109] Table 8
[0110]
[0111] Six gastric cancer sites were selected for combined detection of gastric cancer samples. The traditional method of identifying two or more sites as positive was used to determine the MSI-H method. The results were compared with those of the IHC method, as shown in Table 9 below.
[0112] Table 9
[0113]
[0114] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.
[0115] The results of single-site colorectal cancer detection in colorectal cancer samples were compared with those of IHC, and the results are shown in Table 10 below.
[0116] Table 10
[0117]
[0118] Eight sites were selected for colorectal cancer testing. The traditional method of MSI-H was used, where two or more sites were positive. The results were compared with those of the IHC method, as shown in Table 11 below.
[0119] Table 11
[0120]
[0121] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.
[0122] As can be seen, the combination of high-frequency sites selected by NGS for the three cancer types showed 100% consistency with the IHC method. To cover all three cancer types, all the aforementioned sites could be included. However, this would result in too many sites, leading to a large workload and sample consumption. Furthermore, too many sites can reduce specificity. Therefore, further screening of these sites is needed to obtain a combination with optimal performance and an appropriate number of sites.
[0123] Example 4: Cross-cancer site detection and screening
[0124] Further filtering of loci is needed. The above loci should be used to cross-test other cancer types to see if any loci with the same effect can be eliminated. Specifically, a 12-loci system (UBE2Z, ZBTB37, FOXP2, NBAS, ZBTB16, SEPTIN7, HAL, TSHZ2, LRRC8D, PBRM1, LRBA, and NRCAM) was used to detect the above-mentioned gastric and colorectal cancer samples, respectively; a 6-loci system (DEFB105A, JPT2, PI15, PTPRF, ACVR2A, and NBAS) was used to detect the above-mentioned endometrial and colorectal cancer samples, respectively; and an 8-loci system (ACVR2A, CENPQ, SLC22A9, LRIG2, DIDO1, MRE11, TGFBR2, and PSIP1) was used to detect the above-mentioned endometrial and gastric cancer samples, respectively.
[0125] The results of gastric cancer samples were compared with those obtained by IHC using a single site of endometrial cancer detection, and the results are shown in Table 12 below.
[0126] Table 12
[0127]
[0128] Twelve sites of endometrial cancer were selected for testing gastric cancer samples. The traditional method of identifying two or more sites as positive was used to determine the MSI-H method. The results were compared with those of the IHC method, as shown in Table 13 below.
[0129] Table 13
[0130]
[0131] Positive concordance rate: 95.35%; negative concordance rate: 88.89%; overall concordance rate: 92.05%.
[0132] It can be seen that the 12-site combination for endometrial cancer has excellent sensitivity for detecting gastric cancer samples, but it cannot completely replace gastric cancer sites. Furthermore, the number of sites is too large, requiring improvement in specificity; some sites need to be removed.
[0133] The results of single-site detection of endometrial cancer in colorectal cancer samples were compared with those of IHC, and the results are shown in Table 14 below.
[0134] Table 14
[0135]
[0136] Twelve sites were selected from endometrial cancer for combined detection of colorectal cancer samples. The traditional method of identifying two or more sites as positive was used to determine the MSI-H method. The results were compared with those of the IHC method, as shown in Table 15 below.
[0137] Table 15
[0138]
[0139] Positive compliance rate: 100%; Negative compliance rate: 92.11%; Overall compliance rate: 96.59%.
[0140] As can be seen, the 12 loci for endometrial cancer detection have very high sensitivity for colorectal cancer, even exceeding the sensitivity of loci selected from colorectal cancer itself. The combination of the 12 loci achieves 100% consistency with the IHC method, completely replacing the role of colorectal cancer loci in colorectal cancer detection. This is somewhat unexpected, because theoretically, colorectal cancer and gastric cancer are more strongly associated, while the association with endometrial cancer is relatively weak. Here, the endometrial cancer loci demonstrate good performance in colorectal cancer, which is very helpful in reducing the number of loci. On the other hand, due to the excessive number of loci, specificity needs improvement, and some loci need to be removed.
[0141] Based on the results of the above 12-locus combination for endometrial cancer detection in three cancer types, certain genes at specific loci were removed. Specifically, the gene ZBTB16, with a sensitivity of less than 70% for detecting endometrial cancer samples, was removed; the gene UBE2Z, with a sensitivity of less than 75% for detecting gastric cancer, was removed; and the gene SEPTIN7, with a specificity of less than 90% for detecting colorectal cancer, was removed.
[0142] The results of gastric cancer site detection in endometrial cancer samples were compared with those of IHC, and the results are shown in Table 16 below.
[0143] Table 16
[0144]
[0145] Six gastric cancer sites were selected for testing in combination to detect endometrial cancer samples. The traditional method of identifying two or more sites as positive was used to determine the MSI-H method. The results were compared with those of the IHC method, as shown in Table 17 below.
[0146] Table 17
[0147]
[0148] Positive concordance rate: 86.67%; negative concordance rate: 100%; overall concordance rate: 90.32%.
[0149] It can be seen that the sensitivity of the 6-site combination for gastric cancer in detecting endometrial cancer samples is only 86.67%, and it cannot replace the endometrial cancer site.
[0150] The results of single-site gastric cancer detection in colorectal cancer samples were compared with those obtained by IHC, and the results are shown in Table 18 below.
[0151] Table 18
[0152]
[0153] Six gastric cancer sites were selected for combined detection of colorectal cancer samples. The traditional method of identifying two or more sites as positive was used to determine the MSI-H method. The results were compared with those of the IHC method, as shown in Table 19 below.
[0154] Table 19
[0155]
[0156] Positive compliance rate: 96.00%; Negative compliance rate: 100%; Overall compliance rate: 97.73%.
[0157] It can be seen that the combination of 6 gastric cancer sites has excellent sensitivity for detecting colorectal cancer samples, reaching 96%, but it still cannot completely replace colorectal cancer sites.
[0158] Based on the above-mentioned 6-locus combination for gastric cancer detection results across three cancer types, some gene loci were removed, primarily those with relatively low efficacy in endometrial cancer. Specifically, genes ACVR2A and PI15, with a sensitivity of less than 40% in detecting endometrial cancer samples, were removed.
[0159] The results of single-site colorectal cancer detection in endometrial cancer samples were compared with those of IHC, and the results are shown in Table 20 below.
[0160] Table 20
[0161]
[0162] Eight sites of colorectal cancer were selected for testing of endometrial cancer samples. The traditional method of MSI-H was used, where two or more sites were positive. The results were compared with those of IHC, as shown in Table 21 below.
[0163] Table 21
[0164]
[0165] Positive concordance rate: 75.56%; negative concordance rate: 100%; overall concordance rate: 82.26%.
[0166] It can be seen that the 8-site combination for colorectal cancer has relatively poor sensitivity for detecting endometrial cancer samples, reaching only 75.56%, which is consistent with the results reported in Reference 1. Colorectal cancer sites have too little effect on endometrial cancer samples to replace them.
[0167] The results of gastric cancer samples were compared with those obtained by IHC using single-site colorectal cancer detection, and the results are shown in Table 22 below.
[0168] Table 22
[0169]
[0170] Eight sites of colorectal cancer were selected for testing gastric cancer samples. The traditional method of MSI-H, which considers two or more sites as positive, was first used. The results were compared with those of IHC, and the results are shown in Table 23 below.
[0171] Table 23
[0172]
[0173] Positive concordance rate: 95.35%; negative concordance rate: 100%; overall concordance rate: 97.73%.
[0174] It can be seen that the combination of 8 sites for colorectal cancer has excellent sensitivity for detecting gastric cancer samples, reaching 95.35%, but it cannot completely replace gastric cancer sites.
[0175] Based on the above-mentioned 8-locus combination for colorectal cancer detection results across three cancer types, some gene loci were removed, primarily those with lower efficacy in endometrial cancer. Specifically, genes ACVR2A, CENPQ, SLC22A9, and TGFBR2, which had a sensitivity of less than 40% in detecting endometrial cancer samples, were removed.
[0176] Example 5: Combined Screening
[0177] Based on the detection results of Examples 3 and 4 above, the biomarker loci for the three cancer types showed that the endometrial cancer loci contributed significantly, but some gastric cancer loci were still needed to improve the detection sensitivity for gastric cancer. Colorectal cancer loci performed well in detecting colorectal and gastric cancer, but performed poorly in endometrial cancer, resulting in limited overall effectiveness. Some loci could be removed, relying on the endometrial cancer loci to ensure the detection performance for colorectal cancer. After removal, the remaining loci genes were: ZBTB37, FOXP2, NBAS, HAL, TSHZ2, LRRC8D, PBRM1, LRBA, NRCAM, DEFB105A, JPT2, PTPRF, LRIG2, DIDO1, MRE11, and PSIP1, a total of 16 biomarker genes.
[0178] The 16 filtered sites will be further screened and determined for combinations. First, a random combination method will be used to construct site combinations and determine the results for analysis of the test set samples in Example 3. Specifically, 4, 5, 6, 7, etc., sites will be randomly selected from the 16 sites, increasing sequentially. A sample with 2 or more positive sites will be determined as MSI-H. The results will be compared with those from the IHC method to evaluate the consistency, sensitivity, and specificity of each combination, and the combination with the best performance and the fewest sites will be selected.
[0179] Four markers were randomly selected from the 16 markers to form a combination, resulting in 1820 combinations. Among them, three combinations had the highest consistency, reaching 98.32%. The specific results are shown in Table 24 below.
[0180] Table 24
[0181]
[0182] Five markers were randomly selected from the 16 markers to form a combination, resulting in 4368 combinations. Among them, five combinations had the highest consistency, reaching 99.16%. The specific results are shown in Table 25 below.
[0183] Table 25
[0184]
[0185] Six markers were randomly selected from the 16 markers to form a combination, resulting in 8008 combinations. Among them, two combinations had the highest consistency, reaching 99.58%. The specific results are shown in Table 26 below.
[0186] Table 26
[0187]
[0188] As can be seen, when selecting 6 biomarkers, the sensitivity of the combination already reaches 100%, but the specificity does not. Increasing the number of biomarkers further will not improve sensitivity; in fact, specificity may decrease. A careful analysis of the results of the 6-biomarker combination revealed that a gastric cancer sample had 2 positive biomarkers, classifying it as MSI-H, but the IHC result was pMMR, resulting in a specificity below 100%. Therefore, for gastric cancer, the positive criterion should be adjusted to 3 or more biomarkers for a positive result, thus classifying the sample as MSI-H. Similarly, in colorectal cancer sample testing, because the sensitivity of endometrial cancer biomarkers is very high, samples with only 2 positive biomarkers are almost nonexistent; nearly all samples have 4 or more positive biomarkers. Therefore, for colorectal cancer, the positive criterion can also be adjusted to 3 or more positive biomarkers for a positive result, thus classifying the sample as MSI-H.
[0189] Therefore, some adjustments were made to the result interpretation. Specifically, the sample was classified as MSI-H based on 3 or more positive sites. This was compared with the results of the IHC method to re-evaluate the consistency, sensitivity, and specificity of each random combination of the 16 sites, and to select the combination with the best performance and the fewest sites.
[0190] Six markers were randomly selected from the 16 markers to form a combination, resulting in 8008 possible combinations. One combination showed the highest consistency, reaching 98.32%. The specific results are shown in Table 27 below.
[0191] Table 27
[0192]
[0193] Seven markers were randomly selected from the 16 markers to form a combination, resulting in 11,440 combinations. Among them, six combinations had the highest consistency, reaching 98.74%. The specific results are shown in Table 28 below.
[0194] Table 28
[0195]
[0196] Eight markers were randomly selected from the 16 markers to form a combination, resulting in 12,870 combinations. Among them, two combinations had the highest consistency, reaching 99.58%. The specific results are shown in Table 29 below.
[0197] Table 29
[0198]
[0199] Nine markers were randomly selected from the 16 markers to form a combination, resulting in 11,440 combinations. Among them, 11 combinations showed the highest consistency, reaching 99.58%, which was not higher than the combination of eight loci. The specific results are shown in Table 30 below.
[0200] Table 30
[0201]
[0202] Ten markers were randomly selected from 16 markers to form a combination, resulting in 8008 possible combinations. Among these, three combinations showed the highest consistency, reaching 100%. This indicates that selecting combinations of 10 loci yielded the highest performance. The specific results are shown in Figure 31.
[0203] Table 31
[0204]
[0205] A traditional randomized combination method was used to construct the locus combinations, and a sample was classified as MSI-H if three or more loci were positive. Selecting 10 loci ensured 100% consistency, sensitivity, and specificity. Analysis of the test results revealed that the increased number of loci was necessary to guarantee that endometrial cancer samples would show three or more positive loci.
[0206] Ten sites are a lot; while the detection performance is excellent, the cost is relatively high. We would like to further reduce the number of sites to achieve the same performance.
[0207] The above method uses traditional random combination to determine the results. Next, we want to use machine learning to build a new model to see if the selected locus combinations will cause differences. Specifically, using the IHC results as the gold standard and the detection results of 16 loci as data, we will build a LASSO regression model. The random combination does not consider the contribution of each locus; if the cumulative positive result is greater than a set threshold, it is judged as MSI-H. The LASSO regression model collects the detection results of each locus and, based on the gold standard results, judges the importance of each locus, selecting and retaining the loci with the greatest impact on the results. We choose the minimum value for the λ parameter and use 10-fold cross-validation. The final model selects 9 loci: ZBTB37, FOXP2, NBAS, HAL, TSHZ2, PBRM1, LRBA, JPT2, and PTPRF.
[0208] The LASSO regression model showed 100% consistency with the IHC method in detecting the sample results. The specific data is shown in Table 32 below.
[0209] Table 32
[0210]
[0211] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.
[0212] The LASSO regression model can achieve optimal consistency even with the number of loci reduced to nine. However, LASSO regression requires calculating the probability of each locus test result by inputting it into the constructed equation, assessing the probability of occurrence, and then confirming whether it is positive or negative. This process is still more complex than directly calculating the number of positive loci. Therefore, we aim to further reduce the number of loci, lower testing costs, and choose a simpler method for result interpretation.
[0213] In the previous analysis of randomization methods, it was determined that for gastric and colorectal cancer samples, a positive result at 3 or more loci was sufficient to classify the sample as MSI-H. For endometrial cancer samples, if this method were also used—meaning all three cancer types are classified as MSI-H based on 3 positive loci—10 loci would be needed to achieve optimal consistency. However, for endometrial cancer, it was found that some endometrial samples can be classified as MSI-H based on 2 positive loci, and the total number of positive loci in MSI-H samples is significantly lower than that in colorectal and gastric cancer. This means that if endometrial cancer samples are also classified as MSI-H based on 3 positive loci, even more loci would need to be included. If a 2-positive locus classification method is used for endometrial cancer samples, the required total number of loci can be reduced, but the specificity of endometrial cancer testing must also be maintained at 100%.
[0214] Therefore, the result interpretation method was further modified. For endometrial cancer samples, a positive result at two or more sites indicates MSI-H; for gastric and colorectal cancer samples, a positive result at three or more sites indicates MSI-H. Two interpretation methods were adopted based on the tumor type. Due to the modified result interpretation method, greater attention needs to be paid to the detection specificity of endometrial cancer samples.
[0215] Furthermore, analysis of the optimal combinations of sites in the above embodiments revealed that LRIG2, DIDO1, MRE11, and PSIP1 sites contributed relatively little to the improvement of combined performance, and other sites could completely replace them. Additionally, the PBRM1 site had relatively poor overall specificity. These five sites were eliminated to further narrow down the selected sites, choosing the most important ones. Specifically, a selection was made from 11 sites: ZBTB37, FOXP2, NBAS, HAL, TSHZ2, LRRC8D, LRBA, NRCAM, DEFB105A, JPT2, and PTPRF.
[0216] Specifically, a random combination method is used to further filter the above 11 sites. For endometrial cancer samples, if two or more sites are positive, the sample is determined to be MSI-H. The results are compared with those of the IHC method to calculate the performance of each combination and select the optimal combination.
[0217] For endometrial cancer samples, 5 biomarkers were randomly selected from 11 biomarkers to form a combination, resulting in 462 combinations. Among them, 9 combinations showed the highest consistency, reaching 98.39%. The specific results are shown in Table 33 below.
[0218] Table 33
[0219]
[0220] For endometrial cancer samples, 6 biomarkers were randomly selected from 11 biomarkers to form a combination, resulting in 462 combinations. Among them, 3 combinations showed the highest consistency, reaching 100%. The specific results are shown in Table 34 below.
[0221] Table 34
[0222]
[0223] For gastric and colorectal cancer samples, samples with 3 or more positive sites are identified as MSI-H. The results are compared with those of IHC, the performance of each combination is calculated, and the optimal combination is selected.
[0224] For gastric cancer and colorectal cancer samples, 6 biomarkers were randomly selected from 11 biomarkers to form a combination, resulting in a total of 462 combinations. Among them, 20 combinations had the highest consistency, reaching 100%. The specific results are shown in Table 35 below.
[0225] Table 35
[0226]
[0227]
[0228] For endometrial cancer samples, a positive result of 2 or more positivity points is required, and selecting 6 loci achieves optimal performance, with 3 possible combinations. For gastric and colorectal cancer, a positive result of 3 or more positivity points also requires 6 loci to achieve optimal performance, with 20 possible combinations. Now, we need to combine the loci from these two assessment methods, that is, the 3 combinations for endometrial cancer, and then merge them sequentially with the 20 combinations for gastric and colorectal cancer. The resulting combinations will contain loci that meet the requirements for both endometrial and gastric / colorectal cancer. A total of 60 (3) loci can be obtained. 20) New combinations are proposed. These new combinations may contain duplicate sites, which need to be deduplicated. This will satisfy the criteria for positivity (2 or more positive results) in endometrial cancer, as well as positivity (3 or more positive results) in gastric and endometrial cancer, achieving optimal performance. The combination with the fewest duplicate sites after deduplication should be selected as the optimal combination. There are a total of 3 such combinations, as detailed in Table 36 below.
[0229] Table 36
[0230]
[0231] These three combinations require the fewest loci, only 8, and for the detection results of three cancer types—endometrial cancer, gastric cancer, and colorectal cancer—the consistency with the IHC method can reach 100%.
[0232] Example 6: Combination Verification
[0233] In Examples 3-5 above, 238 samples (62 cases of endometrial cancer, 88 cases of gastric cancer, and 88 cases of colorectal cancer) were used as the test set to screen loci and determine the optimal combination and result determination method. Three combinations were obtained: combination a: ZBTB37, FOXP2, HAL, TSHZ2, LRBA, DEFB105A, JPT2, and PTPRF (8 loci); combination b: FOXP2, HAL, TSHZ2, LRRC8D, LRBA, DEFB105A, JPT2, and PTPRF (8 loci); and combination c: FOXP2, NBAS, HAL, TSHZ2, LRBA, DEFB105A, JPT2, and PTPRF (8 loci). In this example, 37 new cases of endometrial cancer, 64 cases of gastric cancer, and 51 cases of colorectal cancer were collected, totaling 152 samples, as the validation set to verify the performance of combinations a, b, and c obtained from the new determination method. Among them, the positive concordance rate is the sensitivity, the negative concordance rate is the specificity, and the overall concordance rate is the consistency.
[0234] Choose combination a. For endometrial cancer samples, if two or more sites are positive, it is determined to be MSI-H. For gastric cancer and colorectal cancer samples, if three or more sites are positive, it is determined to be MSI-H. The results are compared with the IHC method detection results, as shown in Table 37 below.
[0235] Table 37
[0236]
[0237] Positive concordance rate: 98.81%; negative concordance rate: 100%; overall concordance rate: 99.34%.
[0238] Selecting combination b, for endometrial cancer samples, two or more positive sites are selected to determine MSI-H; for gastric cancer and colorectal cancer samples, three or more positive sites are selected to determine MSI-H. The results are compared with those of the IHC method, as shown in Table 38 below.
[0239] Table 38
[0240]
[0241] Positive concordance rate: 97.62%; negative concordance rate: 100%; overall concordance rate: 98.68%.
[0242] Selecting combination c, for endometrial cancer samples, two or more sites being positive is defined as MSI-H; for gastric cancer and colorectal cancer samples, three or more sites being positive is defined as MSI-H. The results are compared with those of the IHC method, as shown in Table 39 below.
[0243] Table 39
[0244]
[0245] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.
[0246] The validation set results show that combination c has the highest consistency and best performance. Similarly, the validation set results show that there are indeed two positive dMMR samples in endometrial cancer, two positive pMMR samples in gastric cancer, and a relatively large number of positive sites in colorectal cancer MSI-H samples. Therefore, it is more reasonable to use two different result interpretation methods for the three cancer types.
[0247] Based on the comprehensive evaluation of various locus combinations using the test and validation sets, the optimal combination of eight loci—FOXP2, NBAS, HAL, TSHZ2, LRBA, DEFB105A, JPT2, and PTPRF—was selected. This combination requires the fewest detection loci and has a more reasonable result interpretation method, achieving 100% consistency with IHC detection. Through thorough validation with 390 samples, it can meet clinical needs for MSI detection in endometrial cancer, gastric cancer, and colorectal cancer.
[0248] Example 7 Construction of MSI detection kit
[0249] Using the optimal site combination described above (combination c), an MSI detection kit was constructed with a total of 8 sites, which can be divided into 4 reaction systems. Each reaction system detects 2 sites, which further reduces the amount of sample used during detection and simplifies the operation process and time. Specifically, the probes of the site detection system are modified with FAM and VIC fluorescent groups, and the detection probe of the internal reference gene is modified with Cy5 fluorescent group (the internal reference gene is used to evaluate the quality of the test sample and control the total amount of the test sample. Poor sample quality or too little or too much sample may affect the detection results. Therefore, the internal reference gene is used to monitor the sample quality and ensure the reliability of the detection results). In this embodiment, GAPDH is selected as the internal reference gene, and its amplification products and primer and probe sequences are shown in Table 40 below.
[0250] Table 40
[0251]
[0252] Such a reaction system contains three fluorescent probes, the specific forms of which are shown in Table 41 below.
[0253] Table 41
[0254]
[0255] In addition, reference standards were prepared for each of the eight sites to be detected, for performance evaluation of the kit. An MSI-H reference standard containing 5% tumor DNA was selected as the detection limit reference standard, an MSI-H reference standard containing 20% tumor DNA was selected as the positive reference standard, and a reference standard containing 20 ng of wild-type DNA was selected as the negative reference standard.
[0256] In a triple fluorescence reaction system, the reactions of each fluorescence channel competitively consume materials within the system, and the resulting products affect subsequent reactions. Therefore, each fluorescence channel system influences the performance of the others. Crucially, detection systems at two sites within the same reaction system can affect each other's performance, reducing site sensitivity compared to a single fluorescence channel. Therefore, to ensure good detection sensitivity at each of the eight sites, site grouping and tuning are necessary. Based on the detection performance of the single system at 8 sites, various combinations of sites were tested. By adjusting the amount of primers, probes, and magnesium ions used for each reaction tube at two sites, a combination with excellent overall performance was obtained, as follows: For the FAM channel site, the amount of restriction primer was 0.12 μM, the amount of excess primer was 0.48 μM, and the amount of probe was 0.2 μM; for the VIC channel site, the amount of restriction primer was 0.16 μM, the amount of excess primer was 0.64 μM, and the amount of probe was 0.12 μM; for the internal control, the amount of upstream primer was 0.08 μM, the amount of downstream primer was 0.08 μM, and the amount of probe was 0.04 μM; the amount of magnesium ions was 2 mM, the amount of dNTPs was 0.15 mM, the amount of DNA polymerase was 1.4 U, and the amount of UNG enzyme was 0.2 U. The buffer volume was 5 μl, which was supplemented with nuclease-free water to a final volume of 23 μl. Finally, 2 μl of the reference sample was added to prepare a complete PCR system of 25 μl for instrumental analysis. The PCR reaction procedure was the same as in Example 2.
[0257] This combination can accurately detect 5% of tumor DNA at 8 sites in 4 reaction tubes. The detection results of two typical site grouping methods are shown below. Grouping method 1 is detailed in Table 42, and grouping method 2 is detailed in Table 43.
[0258] Table 42
[0259]
[0260] Table 43
[0261]
[0262] The amplification curves of the four reaction tubes corresponding to the detection limit reference in grouping method 1 are shown below. Figures 1-4 As shown, where, Figure 1 The amplification curve of the reference sample corresponding to the detection limit is shown for tube #1. Figure 2 The amplification curve of the reference sample corresponding to the detection limit is shown for tube #2. Figure 3 This is the amplification curve of the reference sample corresponding to the detection limit for tube #3. Figure 4 The amplification curve for the reference sample corresponding to the detection limit of tube #4 is shown below; the amplification curves for the reference samples corresponding to the detection limits of the four reaction tubes in grouping method 2 are shown below. Figures 5-8 As shown, where, Figure 5 The amplification curve of the reference sample corresponding to the detection limit is shown for tube #1. Figure 6 The amplification curve of the reference sample corresponding to the detection limit is shown for tube #2. Figure 7 This is the amplification curve of the reference sample corresponding to the detection limit for tube #3. Figure 8 This is the amplification curve of the reference sample corresponding to the detection limit for tube #4. (Comparison) Figures 1-4 and Figures 5-8 As can be seen, the melting peak diagram of the four reaction tubes corresponding to the detection limit reference in grouping method 2 is better, and the effective peaks can be identified more clearly. Therefore, grouping method 2 was ultimately chosen as the site grouping for the kit.
[0263] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. Use of a reagent for detecting a biomarker composition in the manufacture of a kit for detecting microsatellite instability in cancer, characterized in that, The biomarker composition is as follows (1), (2) or (3): (1) the third marker, the fourth marker, the seventh marker, the eighth marker, the eleventh marker, the thirteenth marker, the fourteenth marker and the sixteenth marker; (2) the second marker, the third marker, the seventh marker, the eighth marker, the eleventh marker, the thirteenth marker, the fourteenth marker and the sixteenth marker; (3) the third marker, the seventh marker, the eighth marker, the ninth marker, the eleventh marker, the thirteenth marker, the fourteenth marker and the sixteenth marker; The second marker is located in the ZBTB37 gene, the starting position is chrl:173899196, and contains a single nucleotide tandem repeat sequence of 11 consecutive T; The third marker is located in the FOXP2 gene, the starting position is chr7:114690176, and contains a single nucleotide tandem repeat sequence of 11 consecutive A; The fourth marker is located in the NBAS gene, the starting position is chr2:15396386, and contains a single nucleotide tandem repeat sequence of 11 consecutive A; The seventh marker is located in the HAL gene, the starting position is chr12:95985886, and contains a single nucleotide tandem repeat sequence of 11 consecutive T; The eighth marker is located in the TSHZ2 gene, the starting position is chr20:53494836, and contains a single nucleotide tandem repeat sequence of 11 consecutive A; The ninth marker is located in the LRRC8D gene, the starting position is chrl:89935847, and contains a single nucleotide tandem repeat sequence of 11 consecutive T; The eleventh marker is located in the LRBA gene, the starting position is chr4:150914351, and contains a single nucleotide tandem repeat sequence of 11 consecutive A; The thirteenth marker is located in the DEFB105A gene, the starting position is chr8:7822206, and contains a single nucleotide tandem repeat sequence of 9 consecutive A; The fourteenth marker is located in the JPT2 gene, the starting position is chr16:1701589, and contains a single nucleotide tandem repeat sequence of 11 consecutive T; The sixteenth marker is located in the PTPRF gene, the starting position is chrl:43622447, and contains a single nucleotide tandem repeat sequence of 11 consecutive T; The reference genome of the marker is GRCh38 / hg38; The cancer is at least one of endometrial cancer, gastric cancer and colorectal cancer.
2. Use of reagents for detecting a biomarker composition in the manufacture of a kit for detecting cancer, characterized in that, The biomarker composition is as follows (1), (2) or (3): (1) the third marker, the fourth marker, the seventh marker, the eighth marker, the eleventh marker, the thirteenth marker, the fourteenth marker and the sixteenth marker; (2) the second marker, the third marker, the seventh marker, the eighth marker, the eleventh marker, the thirteenth marker, the fourteenth marker and the sixteenth marker; (3) the third marker, the seventh marker, the eighth marker, the ninth marker, the eleventh marker, the thirteenth marker, the fourteenth marker, and the sixteenth marker; the second marker is located in the ZBTB37 gene, and has a start position of chrl:173899196 and contains a single nucleotide tandem repeat sequence of 11 consecutive Ts; the third marker is located in the FOXP2 gene, and has a start position of chr7:114690176 and contains a single nucleotide tandem repeat sequence of 11 consecutive As; the fourth marker is located in the NBAS gene, and has a start position of chr2:15396386 and contains a single nucleotide tandem repeat sequence of 11 consecutive As; the seventh marker is located in the HAL gene, and has a start position of chr12:95985886 and contains a single nucleotide tandem repeat sequence of 11 consecutive Ts; the eighth marker is located in the TSHZ2 gene, and has a start position of chr20:53494836 and contains a single nucleotide tandem repeat sequence of 11 consecutive As; the ninth marker is located in the LRRC8D gene, and has a start position of chrl:89935847 and contains a single nucleotide tandem repeat sequence of 11 consecutive Ts; the eleventh marker is located in the LRBA gene, and has a start position of chr4:150914351 and contains a single nucleotide tandem repeat sequence of 11 consecutive As; the thirteenth marker is located in the DEFB105A gene, and has a start position of chr8:7822206 and contains a single nucleotide tandem repeat sequence of 9 consecutive As; the fourteenth marker is located in the JPT2 gene, and has a start position of chr16:1701589 and contains a single nucleotide tandem repeat sequence of 11 consecutive Ts; the sixteenth marker is located in the PTPRF gene, and has a start position of chrl:43622447 and contains a single nucleotide tandem repeat sequence of 11 consecutive Ts; the reference genome of the marker is GRCh38 / hg38; the cancer is at least one of endometrial cancer, gastric cancer, and colorectal cancer.
3. A kit for detecting cancer or microsatellite instability of cancer, characterized by, The kit comprises the following amplification primers and probes for the biomarker composition (1), (2), or (3) by using the high-resolution melting curve method: the amplification primers and probes for the biomarker composition (1) comprise groups C, D, G, H, K, M, N, and P; the amplification primers and probes for the biomarker composition (2) comprise groups B, C, G, H, K, M, N, and P; the amplification primers and probes for the biomarker composition (3) comprise groups C, G, H, I, K, M, N, and P; group B: SEQ ID NO: 27 and SEQ ID NO: 28 for the second marker, and SEQ ID NO: 74; group C: SEQ ID NO: 29 and SEQ ID NO: 30 for the third marker, and SEQ ID NO: 75; Group D: SEQ ID NO: 31 and SEQ ID NO: 32 against the fourth marker, and SEQ ID NO: 76; Group G: SEQ ID NO: 37 and SEQ ID NO: 38 against the seventh marker, and SEQ ID NO: 79; Group H: SEQ ID NO: 39 and SEQ ID NO: 40 against the eighth marker, and SEQ ID NO: 80; Group I: SEQ ID NO: 41 and SEQ ID NO: 42 against the ninth marker, and SEQ ID NO: 81; Group K: SEQ ID NO: 45 and SEQ ID NO: 46 against the eleventh marker, and SEQ ID NO: 83; Group M: SEQ ID NO: 49 and SEQ ID NO: 50 against the thirteenth marker, and SEQ ID NO: 85; Group N: SEQ ID NO: 51 and SEQ ID NO: 52 against the fourteenth marker, and SEQ ID NO: 86; Group P: SEQ ID NO: 55 and SEQ ID NO: 56 against the sixteenth marker, and SEQ ID NO: 88; the cancer is at least one of endometrial cancer, gastric cancer, and colorectal cancer.
Citation Information
Patent Citations
Biomarker group for detecting microsatellite instability in cancer as well as detection kit and application of biomarker group
CN116479128A
Methods and compositions for the molecular diagnosis of microsatellite instability and treatments for cancer
US20230317206A1