Marker for detecting microsatellite instability of endometrial cancer, kit and application

By optimizing the microsatellite instability-related biomarker sites for endometrial cancer and employing the high-resolution melting curve method, a biomarker composition suitable for endometrial cancer was selected. This solved the problem of low sensitivity of existing detection methods in endometrial cancer and achieved efficient and low-cost detection results.

CN120924666APending Publication Date: 2025-11-11SOUTH CHINA UNIV OF TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511227753.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing microsatellite instability (MSI) detection methods in endometrial cancer suffer from low sensitivity and high cost, failing to meet clinical needs. In particular, MSI site combinations based on colorectal cancer screening have poor detection performance in endometrial cancer.

Method used

The screening method for microsatellite instability-related biomarker sites was optimized. Combined with the detection limit of the PCR system, a biomarker composition suitable for endometrial cancer was selected, including biomarkers such as HAL, NBAS, ZBTB37, TSHZ2 and FOXP2, and the high-resolution melting curve method was used for detection.

Benefits of technology

This method achieves highly sensitive and low-cost detection of microsatellite instability in endometrial cancer, with 100% consistency with the IHC method, reducing detection costs and improving detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120924666A_ABST
    Figure CN120924666A_ABST
Patent Text Reader

Abstract

The invention relates to a marker for detecting microsatellite instability of endometrial cancer, a kit and application. The marker composition comprises a first marker, a second marker, a third marker, a fourth marker and a fifth marker. The marker composition which is higher in sensitivity, better in specificity and stronger in compatibility to the endometrial cancer microsatellite instability is obtained through screening, when the marker composition is used for detecting the endometrial cancer microsatellite instability under the condition that only five markers (HAL, NBAS, ZBTB37, TSHZ2 and FOXP2) are contained, the coincidence rate of the detection result and an IHC method is as high as 100%, and the marker composition can be used for detecting the endometrial cancer microsatellite instability. According to the present invention, the endometrial cancer detection performance is far higher than the endometrial cancer detection performance of the similar high-resolution melting curve method site composition reported at present, such that the detection cost can be effectively reduced, the efficiency can be effectively improved, the detection lower limit can achieve the tumor DNA content of 5%, the result bimodal map distinguishing is significant, the result determination is convenient, and the overall detection performance is excellent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of microsatellite instability detection technology, specifically relating to biomarkers, reagent kits, and applications for detecting microsatellite instability in endometrial cancer. Background Technology

[0002] Endometrial cancer (EC) is one of the most common malignant tumors in gynecology worldwide, and its incidence has been rising steadily in recent years. It is currently second only to cervical cancer, accounting for approximately 20%–30% of all gynecological malignancies. Although over 90% of EC cases are sporadic, 5%–10% of patients have a hereditary component; this type of hereditary endometrial cancer is commonly referred to as Lynch syndrome (LS). The pathogenesis of LS is closely related to germline mutations in DNA mismatch repair (MMR) genes, primarily involving MLH1, MSH2, MSH6, and PMS2. These genes are mainly responsible for checking for and correcting errors during DNA replication. Mutations in MMR genes lead to the loss of MMR protein function, resulting in mismatch repair deficiency (MMRd). If DNA repair capacity declines, replication errors persist, and spontaneous somatic mutations increase significantly. Especially for microsatellite sequences, these short tandem repeats (STRs) are more prone to replication errors due to their sequence structure, which accumulate and ultimately lead to microsatellite instability (MSI) and increased tumor susceptibility. Based on the number of affected sites, they are classified as highly unstable (MSI-H), poorly unstable (MSI-L), and stable (MSS).

[0003] The MSI phenomenon was initially discovered in hereditary colorectal cancer, and subsequent studies have confirmed its presence in various cancers, including endometrial cancer, gastric cancer, urethral cancer, and ovarian cancer. Endometrial cancer exhibits the highest MSI frequency, at 21.9%. The 2021 Chinese Expert Consensus on Molecular Detection of Endometrial Cancer explicitly recommends MMR / MSI status testing for all newly diagnosed endometrial cancer patients. This testing has the following clinical value: 1) Diagnosis: dMMR or MSI-H can serve as diagnostic markers for endometrial cancer; 2) Screening: Test results can be used for Lynch syndrome screening; 3) Prognosis: Molecular subtyping is one of the prognostic targets; 4) Prediction of immunotherapy efficacy: Estimating the value of immune checkpoint inhibitors.

[0004] Currently, clinical methods for detecting MSI mainly include immunohistochemistry (IHC), PCR + capillary electrophoresis, NGS, and PCR + high-resolution melting curve analysis, each with its own advantages and disadvantages. IHC primarily determines MSI status by detecting the absence of MMR proteins and is currently the most widely used method in clinical practice. PCR + high-resolution melting curve analysis has seen increasing application in MSI detection in recent years. Its technical principle involves closed-tube multiplex PCR amplification of tumor tissue-specific microsatellite loci, followed by direct melting curve analysis of the amplified products. Due to the length polymorphism differences between the amplified products of MSI-H and MSS samples, their double-stranded DNA melting temperatures (Tm values) will exhibit characteristic shifts. MSI status can be determined by analyzing the Tm value change patterns. Compared to traditional capillary electrophoresis, this technique eliminates the need for electrophoresis; melting curve detection is performed directly after PCR, making the overall operation simpler and the detection time only 1 to 2 hours. Secondly, the PCR + high-resolution melting curve method offers significantly improved detection sensitivity compared to capillary electrophoresis, generally achieving a tumor tissue content of 10% or even lower. Furthermore, due to the use of shorter MSI sites, even minute base deletions or insertions can be distinguished by the melting curve. Further screening for more specific sites allows for accurate identification of MSI status using only tumor tissue, eliminating the need for normal tissue and reducing sample and testing costs. This technology offers clear advantages over IHC and capillary electrophoresis. However, the key to this technology lies in the selection of MSI sites and the construction of the detection system. While some literature and patents report combinations of MSI sites based on the melting curve method, most of these combinations are based on sites screened for colorectal cancer. While they perform well in colorectal cancer, their sensitivity drops significantly for endometrial cancer, failing to meet clinical needs. Reference 1 describes a combination of microsatellite instability detection (MSI) sites in colorectal and endometrial cancer, comparing its concordance with IHC in these two cancers. Specifically, the MSI site combination includes seven genes: DIDO1, ACVR2A, MRE11, BTBD7, SULF2, SEC31A, and RYR3. This combination achieved 100% concordance with IHC in colorectal cancer samples, but only 79% concordance in endometrial cancer samples.The study "Detection of microsatellite instability with Idylla MSI assay in colorectal and endometrial cancer" (Reference 2) also recorded similar results. The combination of the seven sites mentioned above showed a sensitivity of only 72.7% compared to the IHC method in detecting endometrial cancer samples. Furthermore, for EC samples, this combination required a tumor tissue proportion of over 30% to achieve good performance. Additionally, patent CN112708679B (Reference 3) also describes 22 MSI sites that performed excellently in colorectal cancer, and indicates that the mutant sequences were determined by NGS validation. The mutant sequences are either 1 bp repeats or possibly 2 bp repeats. However, further data on the performance of these sites in endometrial cancer is still lacking.

[0005] Given the shortcomings of the existing detection methods, there is a clinical need for a method that is highly sensitive, easy to operate, low in cost, and widely applicable, in order to achieve efficient detection of MSI status in endometrial cancer. Summary of the Invention

[0006] Based on this, the purpose of the present invention is to provide biomarkers, kits and applications for detecting microsatellite instability in endometrial cancer. The biomarkers are compositions that have the advantages of high sensitivity, high specificity and low cost when used to detect microsatellite instability in endometrial cancer.

[0007] A first aspect of the present invention is to provide a biomarker composition comprising a first biomarker, a second biomarker, a third biomarker, a fourth biomarker, and a fifth biomarker:

[0008] The first biomarker is located in the HAL gene, starting at chr12:95985886, and contains 11 consecutive T single nucleotide tandem repeat sequences.

[0009] The second biomarker is located in the NBAS gene, starting at chr2:15396386, and contains 11 consecutive A single nucleotide tandem repeats.

[0010] The third biomarker is located in the ZBTB37 gene, starting at chr1:173899196, and contains 11 consecutive T single nucleotide tandem repeat sequences.

[0011] The fourth biomarker is located in the TSHZ2 gene, starting at chr20:53494836, and contains 11 consecutive A single nucleotide tandem repeat sequences.

[0012] The fifth biomarker is located in the FOXP2 gene, starting at chr7:114690176, and contains 11 consecutive A single nucleotide tandem repeats.

[0013] In some embodiments, the biomarker composition further includes at least one of the following sixth, seventh, and eighth biomarkers;

[0014] The sixth biomarker is located in the LRRC8D gene, starting at chr1:89935847, and contains 11 consecutive T single nucleotide tandem repeat sequences.

[0015] The seventh biomarker is located in the PBRM1 gene, starting at chr3:52662295, and contains 11 consecutive A single nucleotide tandem repeat sequences.

[0016] The eighth biomarker is located in the LRBA gene, starting at chr4:150914351, and contains 11 consecutive A single nucleotide tandem repeats.

[0017] In some of these embodiments, the biomarker composition is (1), (2), or (3):

[0018] (1): First marker, second marker, third marker, fourth marker, fifth marker and sixth marker;

[0019] (2): First marker, second marker, third marker, fourth marker, fifth marker, sixth marker and seventh marker;

[0020] (3) First marker, second marker, third marker, fourth marker, fifth marker, sixth marker, seventh marker and eighth marker.

[0021] A second aspect of the present invention is the use of reagents for detecting the biomarker compositions described above in the preparation of a kit for detecting microsatellite instability in endometrial cancer.

[0022] A third aspect of the invention is the use of reagents for detecting the biomarker compositions described above in the preparation of endometrial cancer detection kits.

[0023] A fourth aspect of the present invention is to provide a kit for detecting endometrial cancer or endometrial cancer microsatellite instability, comprising reagents for detecting mutations in the biomarker composition as described above.

[0024] In some embodiments, the high-resolution melting curve method is used, and the kit includes amplification primers and probes in groups A, B, C, D, and E as follows:

[0025] Group A: Restriction primers as shown in SEQ ID NO: 20 and excess primers as shown in SEQ ID NO: 21, and probes as shown in SEQ ID NO: 58, targeting the first marker;

[0026] Group B: Restriction primers as shown in SEQ ID NO: 22 and excess primers as shown in SEQ ID NO: 23, and probes as shown in SEQ ID NO: 59, targeting the second biomarker;

[0027] Group C: Restriction primers as shown in SEQ ID NO: 24 and excess primers as shown in SEQ ID NO: 25, and probes as shown in SEQ ID NO: 60, targeting the third biomarker;

[0028] Group D: Restriction primers as shown in SEQ ID NO: 26 and excess primers as shown in SEQ ID NO: 27, and probes as shown in SEQ ID NO: 61, targeting the fourth biomarker;

[0029] Group E: Restriction primers as shown in SEQ ID NO: 28 and excess primers as shown in SEQ ID NO: 29 for the fifth marker, and probes as shown in SEQ ID NO: 62.

[0030] In some embodiments, the kit further includes at least one of the following groups: F, G, and H amplification primers and probes:

[0031] Group F: Restriction primers as shown in SEQ ID NO: 30 and excess primers as shown in SEQ ID NO: 31, and probes as shown in SEQ ID NO: 63, targeting the sixth biomarker;

[0032] Group G: Restriction primers as shown in SEQ ID NO: 32 and excess primers as shown in SEQ ID NO: 33 for the seventh marker, and probes as shown in SEQ ID NO: 64;

[0033] Group H: Restriction primers as shown in SEQ ID NO: 34 and excess primers as shown in SEQ ID NO: 35 for the eighth marker, and probes as shown in SEQ ID NO: 65.

[0034] In some of these embodiments, 1 to 4 bases from the 3' end of each probe are thiolated.

[0035] A fifth aspect of the present invention is to provide a method for detecting microsatellite instability in endometrial cancer for non-diagnostic purposes, comprising the following steps: obtaining DNA from a sample to be tested, and detecting mutations in the biomarker composition described above using a kit as described above.

[0036] In some of these embodiments, a high-resolution melting curve method is used, and the amount of each restriction primer in the PCR reaction system is 0.01 μM-0.2 μM;

[0037] The amount of each excess primer used in the PCR reaction system is 0.5 μM-1.5 μM;

[0038] The amount of each probe used in the PCR reaction system is 0.05 μM-0.25 μM.

[0039] In some preferred embodiments, the amount of each restriction primer used in the PCR reaction system is 0.1 μM-0.2 μM;

[0040] The amount of each excess primer used in the PCR reaction system is 0.5 μM-1 μM;

[0041] The amount of each probe used in the PCR reaction system is 0.05 μM-0.1 μM.

[0042] In some embodiments, the PCR reaction system further includes DNA polymerase, wherein the amount of DNA polymerase used is 0.5U-1.5U.

[0043] In some preferred embodiments, the amount of DNA polymerase used is 1U-1.5U.

[0044] This invention optimizes the screening method for microsatellite instability-related biomarkers in endometrial cancer. Instead of using conventional methods to select biomarkers by calculating the AUC of NGS data, it combines the detection limit of the biomarker PCR system, requiring that the biomarker has a variation rate of more than 3% in endometrial cancer MSI-H samples. This way, the biomarkers obtained by NGS screening perform better in subsequent PCR-based detection, thus enabling the selection of biomarkers that are more suitable for endometrial cancer. Further screening and combination of microsatellite instability-related biomarkers for endometrial cancer obtained from initial screening were conducted. The performance of biomarker sites with different mutation forms was compared, and the optimal mutation form (the mutation form with a 1bp deletion) was determined. Finally, a biomarker composition with higher sensitivity, better specificity, and stronger compatibility for endometrial cancer microsatellite instability was obtained. When the biomarker composition contained only 5 biomarkers (HAL, NBAS, ZBTB37, TSHZ2, and FOXP2) for detecting endometrial cancer microsatellite instability, the detection results had a concordance rate of up to 100% with the IHC method, which is far higher than the performance of similar PCR + high-resolution melting curve method biomarker compositions for endometrial cancer detection reported to date.

[0045] The marker composition of the present invention can achieve better detection results with a significantly smaller number of markers than other existing solutions, thereby effectively reducing detection costs and improving efficiency.

[0046] Furthermore, by optimizing the screening method for microsatellite instability-related biomarkers in endometrial cancer and the PCR amplification system, this invention achieves a detection limit of 5% for tumor DNA content, and the bimodal graph of the results is clearly distinguishable, facilitating result interpretation. Overall, the detection performance is excellent. Attached Figure Description

[0047] Figure 1 The image shows the detection results of one of the samples for which a perfectly matched probe system was designed to target the mutation type of the eleventh biomarker with a 2bp deletion.

[0048] Figure 2 The image shows the detection results of one sample for which a perfectly matched probe system was designed to target the mutation type of missing 1bp of the seventh biomarker.

[0049] Figure 3 This is a graph showing the results of the HAL gene system reaction tube detection of the reference sample in Example 6.

[0050] Figure 4 This is a graph showing the results of the NBAS gene system reaction tube detection of the reference sample in Example 6.

[0051] Figure 5 This is a graph showing the results of the ZBTB37 gene system reaction tube detection of the reference sample in Example 6.

[0052] Figure 6 This is a graph showing the results of the TSHZ2 gene system reaction tube detection of the reference sample in Example 6.

[0053] Figure 7 This is a graph showing the results of the FOXP2 gene system reaction tube detection of the reference sample in Example 6. Detailed Implementation

[0054] To facilitate understanding of the present invention, a more complete description will be provided below. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the present invention.

[0055] Unless otherwise specified, experimental methods in the following examples were performed under standard conditions, such as those described in the fourth edition of *Molecular Cloning: A Laboratory Manual*, edited by Green and Sambrook, published in 2013, or according to the manufacturer's recommendations. All commonly used chemical reagents used in the examples are commercially available products.

[0056] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. The term "and / or" as used in this invention includes any and all combinations of one or more of the associated listed items.

[0057] The present invention will be further described in detail below with reference to specific embodiments.

[0058] Example 1: NGS Sequencing and Initial Site Screening

[0059] We collected MSI locus information from the TCGA, GEO, and cBioPortal databases, as well as literature and patents, selecting high-frequency MSI loci. However, this information primarily focuses on colorectal cancer MSIs, with very little information on endometrial cancer MSIs, and even less on MSIs in the Chinese population, making it difficult to calculate the frequency of these loci in the Chinese population. Since gene mutation frequencies vary across different ethnic groups, we cannot directly use research loci from Caucasians or other non-Chinese populations; further validation is required. Therefore, to better reflect the actual MSI detection situation in the Chinese population and to cover as many excellent loci as possible to avoid omissions, we first used 16 endometrial cancer samples for whole-exome and non-coding region sequencing to analyze high-frequency MSI loci. Combining the high-frequency MSI loci obtained from the above data with those obtained from sequencing real Chinese population samples, we constructed an NGS locus panel covering 2465 MSI loci. Then, we collected endometrial cancer samples and performed targeted sequencing using a customized panel, with tumor tissue and paired normal tissue being sequenced simultaneously.

[0060] The data from the assay were processed using bioinformatics software to obtain the final labeled BAM file. Then, software such as MSISensor or MANTIS was used to obtain the mutation count of all MSI sites for each sample. Considering the differences between the positive / negative determination methods for NGS and PCR, it's important to understand that PCR requires a certain level of tumor tissue variation in the MSI-H sample for accurate detection. NGS, on the other hand, primarily determines positive / negative results by calculating and comparing statistically significant differences between the two methods. However, in practical PCR testing, very small differences cannot be accurately detected; a significant difference is required. To ensure that the sites selected by NGS better meet the needs of subsequent PCR applications, we did not simply calculate the AUC of each site based on NGS data and IHC results, then select high-AUC sites. Instead, we considered the detection limits achievable by our constructed PCR-high-resolution melting curve method system, requiring a mutation rate greater than 3% in the tumor sample from the NGS data for classification as MSI-H. We increased the difference, imposing higher sensitivity requirements on the sites, which better meets the sensitivity requirements of subsequent PCR testing. This prevents missed detections due to insufficient mutation differences. In this way, we screened out sites with relatively high sensitivity for subsequent PCR validation.

[0061] Furthermore, according to the records in references 1-3 above, the mutated sequence may be missing 1 bp or 2 bp; according to the MSI definition, it could also be inserted 1 bp. Of these three mutation types, the first two showed relatively good results in colorectal cancer, but their performance in endometrial cancer is unknown. Therefore, we analyzed the three mutation types—1 bp deletion, 2 bp deletion, and 1 bp insertion—at each site. Compared with the results of IHC detection, we selected sites with high consistency among the three mutation types for subsequent PCR testing. The specific sites are shown in Table 1 below, mapped to the GRCh38 / hg38 human reference genome:

[0062] Table 1

[0063]

[0064]

[0065] Among them, the nine sites from the first to the ninth marker showed high consistency in mutations with a 1bp deletion; the five sites from the tenth to the fourteenth marker showed high consistency in mutations with a 2bp deletion; and the five sites from the fifteenth to the nineteenth marker showed high consistency in mutations with a 1bp insertion.

[0066] Example 2 Construction of a single-site PCR detection system

[0067] This embodiment uses PCR combined with high-resolution melting curve analysis to detect the MSI sites obtained in Example 1. During detection, DNA fragments containing the MSI sites are amplified by PCR, and the mutation status of the MSI sites is analyzed using the melting curve Tm value. In this embodiment, primers and probes are designed for the DNA mutation sequences containing the MSI sites shown in Table 2 below. The underlined areas indicate the locations of the MSI sites. The mutation sequences correspond to the mutation types in Table 1 above, namely, sequences with a 1 bp deletion, a 2 bp deletion, and a 1 bp insertion.

[0068] Table 2

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077] For each MSI biomarker mutation sequence in Table 2, corresponding upstream and downstream primers were designed. Primer design must meet the following requirements: Tm value controlled between 58℃ and 61℃; the 3' end must not contain hairpin structures or secondary structures such as dimers; the length of the amplification product should be between 100bp and 500bp.

[0078] The PCR amplification process employs an asymmetric amplification method. Specifically, one primer serves as a "restriction primer" and is used in a relatively small amount; the other primer is used as an "excess primer" and is used in a relatively large amount. The purpose of this asymmetric amplification method is to amplify more single-stranded template so that it can specifically bind to the probe, thereby providing a stronger fluorescence signal for subsequent melting curve analysis. The primer sequences are shown in Table 3 below:

[0079] Table 3

[0080]

[0081]

[0082]

[0083]

[0084] Detection probes were designed for the MSI biomarkers screened by NGS in Example 1. These probes were specifically designed for the mutant sequences described in Table 2. Specifically, probes were designed based on three types of mutant sequence templates: 1 bp deletion, 2 bp deletion, and 1 bp insertion. The specific mutation type and site requirements were consistent with those described above. That is, probes were designed for the 9 sites (1st to 9th biomarkers) for mutant sequences with a 1 bp deletion; probes were designed for the 5 sites (10th to 14th biomarkers) for mutant sequences with a 2 bp deletion; and probes were designed for the 5 sites (15th to 19th biomarkers) for mutant sequences with a 1 bp insertion. The probes achieved 100% matching with the mutant template but not with the wild-type template. As a result, the Tm values ​​of the two probe-template binding products differed, and the subsequent melting curves were used to distinguish the MSI state based on the difference in Tm.

[0085] The Tm value of the probe should be controlled between 40℃ and 70℃. In addition, to prevent the probe from being degraded by the 3'-5' end corrective activity of DNA amplification enzymes, the 3' end of the probe can be protected by thiolation modification, with the number of thiolated bases ranging from 1 to 4.

[0086] The probe can be in the form of a molecular beacon probe or a Taqman probe. Both ends of the probe need to be modified with reporter and quencher groups. Reporter group modification options include FAM, VIC, ROX, Cy5, RED, Cy3, etc., while quencher group modification options include BHQ1, BHQ2, BHQ3, MGB, and Dabcyl, etc. The probe design sequences are shown in Table 4 below, where "..." "The label indicates thiomodification:"

[0087] Table 4

[0088]

[0089]

[0090] A single-site detection system was constructed using the primer and probe combination designed above, and its performance was evaluated by setting relevant reference standards. The PCR reaction solution was prepared from water, buffer, primers, probes, and DNA polymerase. Specifically, the amount of restriction primer was 0.16 μM, the amount of excess primer was 0.6 μM, the amount of probe was 0.08 μM, the amount of magnesium ions was 3 mM, the amount of dNTPs was 0.2 mM, the amount of DNA polymerase was 1.4 U, and the amount of UNG enzyme was 0.2 U. The buffer volume is 5 μl, which is supplemented to 23 μl with nuclease-free water. Finally, 2 μl of the test reference is added to the prepared reaction system to prepare a 25 μl complete PCR system for instrument detection.

[0091] The detection instrument requires a real-time quantitative PCR instrument with high-resolution melting curve function. The denaturation, annealing, and extension temperatures should be set separately, and the amplification cycle number should be set to 60 cycles. The melting curve collection temperature range should be set to 30-70℃. The detection process consists of two stages: first, PCR amplification is performed to enrich single-stranded products; then, the binding of the single-stranded template to the probe is detected during the temperature rise; finally, based on the collected fluorescence signal from the instrument, the corresponding melting curve peaks are plotted for result interpretation. The instrumentation procedure is shown in Table 5 below:

[0092] Table 5

[0093]

[0094] Criteria for judging the results of unit spot system tests: For negative reference samples corresponding to unit spot system tests, the result chromatogram should show a single peak, with a peak only appearing at low Tm values, indicating good specificity of the system. For positive reference samples corresponding to unit spot system tests, the result chromatogram should show a double peak or a single peak, with a peak at high Tm values, indicating good accuracy of the system. For limit of detection reference samples corresponding to unit spot system tests, the result chromatogram should show a double peak, with two Tm values, one low Tm value and one high Tm value, indicating good limit of detection of the system. If the test results of the above three reference samples all meet the criteria, it indicates that the performance of the unit spot system is excellent.

[0095] Example 3: Initial screening of PCR detection system sites

[0096] Following the PCR detection system construction method described in Example 2 above, a detection system with 19 loci was constructed. Sixty endometrial cancer samples were collected as a test set, and the existing IHC results were used as the gold standard to evaluate the detection performance at each locus. Based on the descriptions in references 1-3, IHC, a protein-level detection method, typically detects more positive results than PCR, a molecular-level method. Often, IHC may be positive while PCR + capillary electrophoresis is negative. Therefore, to improve consistency with IHC, we chose IHC as the gold standard for comparative validation. Detailed detection results for each locus are shown in Table 6 below:

[0097] Table 6

[0098]

[0099] The results of marker site detection show that the sensitivity of these sites is not very high, significantly lower than the detection performance of combined colorectal cancer sites reported in the literature. However, the effect of missing 1 bp sites is better than the performance of combined colorectal cancer sites in endometrial cancer reported in the literature. This indicates that sites suitable for colorectal cancer are not necessarily suitable for endometrial cancer, and the variation in MSI samples in endometrial cancer is indeed significantly smaller than that in colorectal cancer samples. Therefore, it is essential to set the requirement that the variation rate of tumor samples in NGS data be greater than 3% in Example 1 for site screening. Because the variation is small in endometrial cancer, high AUC sites identified by NGS data may not be able to distinguish small variations during PCR validation, thus showing low sensitivity. Therefore, if only high AUC sites are selected, low-sensitivity sites at the PCR level will be selected, and many truly excellent sites will be missed.

[0100] Further analysis showed that the overall sensitivity of 1bp deletion sites was higher than that of 2bp deletion sites and 1bp insertion sites, with the 1bp insertion site exhibiting the worst performance. This is presumably because the degree of base changes at the MSI site in endometrial cancer MSI-H samples is lower than that in colorectal cancer MSI-H samples, and the mutations are mainly 1bp deletions. Therefore, 2bp deletion and 1bp insertion sites have lower sensitivity than 1bp deletion sites.

[0101] Furthermore, if we design probes to test samples for markers with a 2bp deletion, the samples will contain both 2bp and 1bp deletion mutations, with some samples having a higher number of 1bp deletions. Since we are designing probes for the target mutant (i.e., perfectly matching the 2bp deletion sequence), neither the 1bp deletion sequence nor the wild-type sequence will perfectly match the probe. The Tm values ​​of the binding compounds of these two sequences and the probe will be lower than those of the perfectly matched binding compounds with the 2bp deletion. Moreover, the wild-type sequence differs from the probe by two bases, resulting in a larger mismatch and an even lower Tm value. For such samples, some samples will exhibit three peaks in their melting curve peak diagram: the lowest Tm value wild-type peak, the middle Tm value 1bp deletion mutation peak, and the highest Tm value target mutation peak with a 2bp deletion. The close proximity of the 1bp deletion mutation peak and the 2bp deletion target mutation peak will affect the peak height and Tm value of the target peak, reduce site sensitivity, and interfere with the interpretation of the graph results. Figure 1 The image shows the detection results of one sample of the eleventh marker.

[0102] If a probe is designed to test a sample containing a 1bp deletion marker, both 1bp and 2bp deletion mutations will be present, though fewer in number. Similarly, if the probe is designed for a 1bp deletion, both the 2bp deletion sequence and the wild-type sequence have a one-base mismatch with the probe. The Tm values ​​of the products bound to the probe from these two sequences are similar, but significantly different from the Tm values ​​of the perfectly matched products with the 1bp deletion. For samples with a high number of 2bp deletion mutations, there will not be three melting peaks, only two: a low Tm value peak from the wild-type and 2bp deletion mutations, and a high Tm value peak from the target 1bp deletion mutation. The 2bp deletion mutation peak largely overlaps with the wild-type peak, having minimal impact on the target peak, thus not reducing site sensitivity or affecting the interpretation of the graph results. Figure 2 The image shows the detection results for one sample of the seventh biomarker.

[0103] As we can see from the sensitivity results of the three mutation types detected above, the 2bp deletion type accounts for a relatively small proportion of endometrial cancer. Therefore, the 1bp deletion mutation site is the optimal choice in terms of sensitivity and the display of the result peak diagram. Thus, we selected the 1bp deletion mutation site for the next step of combination screening.

[0104] Example 4: Combination Construction and Optimal Combination Selection

[0105] Based on the detection results of Example 3 above, the sites with a 2bp deletion and a 1bp insertion were removed, and the retained sites were: HAL, NBAS, ZBTB37, TSHZ2, FOXP2, LRRC8D, PBRM1, LRBA and NRCAM, a total of 9 marker genes.

[0106] We will further screen and determine the combinations of the nine filtered loci, using the test set samples from Example 3. We first constructed the locus combinations and result determination methods using a random combination approach. Specifically, we randomly selected 4, 5, 6, etc., from the nine loci, increasing the number sequentially. Initially, we used the traditional method, where a sample was classified as MSI-H if two or more loci were positive, and compared the results with those from the IHC method to evaluate the consistency, sensitivity, and specificity of each combination, ultimately selecting the locus combination with the best performance and the fewest loci.

[0107] Four markers were randomly selected from the nine markers to form a combination, resulting in 126 possible combinations. Among these, three combinations showed the highest consistency, reaching 98.33%. The specific results are shown in Table 7 below.

[0108] Table 7

[0109]

[0110] Five markers were randomly selected from the nine markers to form a combination, resulting in 126 possible combinations. Among these, ten combinations showed the highest consistency, reaching 100%. One typical combination is shown in Table 8 below:

[0111] Table 8

[0112]

[0113] As can be seen, when there are 5 biomarkers, selecting specific combinations of sites can achieve 100% consistency. Further increasing the number of sites will not increase sensitivity, but will reduce the probability of false negatives, although specificity may decrease, and will also increase testing costs.

[0114] Six markers were randomly selected from nine markers to form a combination, resulting in 84 possible combinations. Among these, 33 combinations showed the highest consistency, reaching 100%. One typical combination is shown in Table 9 below:

[0115] Table 9

[0116]

[0117] All 84 combinations showed 100% specificity, and the specificity did not decrease with the increase in the number of loci. This combination adds LRRC8D to the above 5-locus combination of HAL, NBAS, ZBTB37, TSHZ2, and FOXP2, and can also be used as an alternative combination scheme.

[0118] Increasing the number of loci beyond six will not increase sensitivity, and specificity may even decrease. To ensure specificity and match the number of loci, the positive locus threshold can be increased. That is, if two or more positive loci indicate MSI-H, this could be changed to three or more positive loci indicating MSI-H.

[0119] We wanted to examine the performance and specific locus characteristics of the optimal combination under this new determination method. Therefore, we defined the sample as MSI-H if three or more loci were positive, and compared the results with those of the IHC method. We calculated the consistency, sensitivity, and specificity of each random combination of the nine loci, and screened the combination with the best performance and the fewest loci.

[0120] Six markers were randomly selected from the nine markers to form a combination, resulting in 84 possible combinations. One combination showed the highest consistency (98.33%), as detailed in Table 10 below.

[0121] Table 10

[0122]

[0123] Seven markers were randomly selected from the nine markers to form a combination, resulting in 36 possible combinations. One combination showed the highest consistency, reaching 100%. The specific results are shown in Table 11 below.

[0124] Table 11

[0125]

[0126] When selecting 7 sites, consistency can reach a maximum of 100%. This combination of 7 sites is based on the optimal combination of 6 sites (HAL, NBAS, ZBTB37, TSHZ2, FOXP2, LRRC8D) for 2 or more positive sites, with the addition of PBRM1, which achieves the optimal performance when 3 or more sites are positive as the threshold.

[0127] Adding more loci does not increase sensitivity or decrease specificity. This is because the positive threshold has already been raised, ensuring better specificity. Whether it's the threshold for 2 positives or 3 positives, increasing the number of loci does not affect specificity. We speculate there are two reasons for this: firstly, the nine loci we included have excellent specificity, and the combination of these loci, even with an increased number of loci, does not significantly impact specificity; secondly, the sequence variation in MSI-H samples from endometrial cancer is relatively small, meaning the sensitivity per locus is not very high, and loci need to work together to achieve better results. Comparing references 1-3, the sensitivity per locus for colorectal cancer screening in colorectal cancer is superior to that for endometrial cancer screening, further reflecting that screening for loci in endometrial cancer is more challenging than in colorectal cancer.

[0128] The random combination method for determining the result does not consider the importance of each locus in the combination; it simply determines the final MSI status based on the number of positive results. Since our final result is a binary classification of positive and negative, we can also determine the result by constructing a relevant mathematical model. Specifically, we use machine learning methods to build a new model to see if there are any differences in the selected locus combinations.

[0129] We used the IHC results as the gold standard and the detection results of the above 9 loci as data to construct a LASSO regression model. The model collects the detection results of each locus and judges the importance of each locus based on the gold standard results, selecting and retaining the loci with the greatest impact on the results, i.e., the most important loci, and assigning them a certain coefficient. We chose the λ parameter as the minimum value and used 10-fold cross-validation to construct the LASSO regression model. The final model selected 8 loci: HAL, NBAS, ZBTB37, TSHZ2, FOXP2, LRRC8D, PBRM1, and LRBA.

[0130] The LASSO regression model showed 100% consistency with the IHC method in detecting the sample results. Specific data are shown in Table 12 below.

[0131] Table 12

[0132]

[0133] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.

[0134] As can be seen, the loci selected by the LASSO regression model are based on the seven optimal loci combinations (HAL, NBAS, ZBTB37, TSHZ2, FOXP2, LRRC8D, PBRM1) for 3 or more positive cases, with the addition of LRBA.

[0135] In summary, we have obtained four site combinations, all of which show 100% consistency with the results obtained by IHC.

[0136] Combination a has 5 loci: HAL, NBAS, ZBTB37, TSHZ2, and FOXP2;

[0137] Combination b consists of 6 loci: HAL, NBAS, ZBTB37, TSHZ2, FOXP2, and LRRC8D;

[0138] Combination c has 7 loci: HAL, NBAS, ZBTB37, TSHZ2, FOXP2, LRRC8D, and PBRM1;

[0139] The d combination has 8 loci: HAL, NBAS, ZBTB37, TSHZ2, FOXP2, LRRC8D, PBRM1, and LRBA.

[0140] For combinations a and b, the determination method is that if two or more sites are positive, the sample is determined to be MSI-H; for combination c, the determination method is that if three or more sites are positive, the sample is determined to be MSI-H; and for combination d, the result is determined using the LASSO regression model.

[0141] Example 5: Performance of the validation set sample validation combinatorial test

[0142] We collected 56 new endometrial cancer samples as a validation set to verify the performance of the four combinations a, b, c, and d mentioned above. The positive concordance rate represents sensitivity, the negative concordance rate represents specificity, and the overall concordance rate represents consistency.

[0143] For combination a, selecting two or more positive sites is considered as MSI-H. The results are compared with those of the IHC method, as shown in Table 13 below:

[0144] Table 13

[0145]

[0146] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.

[0147] For combination b, selecting two or more positive sites is considered MSI-H. The results are compared with those of the IHC method, as shown in Table 14 below:

[0148] Table 14

[0149]

[0150] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.

[0151] Combination c, selecting 3 or more positive sites, is defined as MSI-H. The results are compared with those of IHC detection, as shown in Table 15 below:

[0152] Table 15

[0153]

[0154] Positive concordance rate: 96.55%; Negative concordance rate: 100%; Overall concordance rate: 98.21%.

[0155] Select combination d, use the LASSO regression model to determine the MSI status of the sample, and compare it with the detection results of the IHC method. The results are shown in Table 16 below:

[0156] Table 16

[0157]

[0158] Positive compliance rate: 100%; Negative compliance rate: 100%; Overall compliance rate: 100%.

[0159] The validation set results showed that combinations a, b, and d had the highest consistency and the best performance. Combination c had a false negative because one sample had only two positive sites, so its sensitivity did not reach 100%.

[0160] Combining the test and validation set data, the four combinations were compared. Combination a had the fewest loci and achieved the best consistency. Combination b added one locus without affecting specificity. Combination c had better specificity but also had the risk of false negatives due to the increased positive threshold. Combination d achieved the best detection performance in both the test and validation sets, but had the most loci. In addition, the LASSO regression model required calculating the probability value of each locus by inputting the detection result into the constructed equation, and then evaluating the probability of occurrence, i.e., whether the final MSI status was positive or negative. The whole process was more complicated than directly calculating the total number of positive loci.

[0161] In summary, we selected combination a, which offers the best overall performance: HAL, NBAS, ZBTB37, TSHZ2, and FOXP2. This combination achieves the highest consistency while using the fewest possible sites, is the simplest to operate, and minimizes costs, thus meeting the needs for widespread clinical testing of endometrial cancer MSI.

[0162] Example 6: Construction of a complete MSI detection kit

[0163] Using the optimal site combination a described above, an MSI detection kit was constructed. A dual-fluorescence channel detection system was employed, consisting of a site detection system and an internal control detection system. Specifically, the probe of the site detection system was modified with a FAM fluorescent group, and the detection probe of the internal control gene was modified with a Cy5 fluorescent group. The internal control gene was used to evaluate the quality of the test samples and control the total amount of sample to be tested. Poor sample quality or insufficient or excessive sample loading can affect the accuracy of the detection results; therefore, an internal control gene is used to monitor sample quality and ensure the reliability of the detection results. The internal control gene can be GAPDH or other human housekeeping genes. In this embodiment, GAPDH was selected as the internal control gene, and its amplification products, primers, and probe sequences are shown in Table 17 below.

[0164] Table 17

[0165]

[0166] This reaction system contains two fluorescent probes, and the specific system combination is shown in Table 18 below. The specific composition of the reaction system is as follows: the amount of restriction primer is 0.16 μM, the amount of excess primer is 0.6 μM, and the amount of probe is 0.08 μM; the amount of internal control upstream primer is 0.08 μM, the amount of internal control downstream primer is 0.08 μM, and the amount of internal control probe is 0.04 μM; the amount of magnesium ions is 3 mM, the amount of dNTPs is 0.2 mM, the amount of DNA polymerase is 1.4 U, the amount of UNG enzyme is 0.2 U, and 5 The buffer volume was 5 μl, which was then supplemented with nuclease-free water to a final volume of 23 μl. Finally, 2 μl of the reference sample was added to the prepared reaction system to create a 25 μl complete PCR system for instrumental analysis. The PCR reaction procedure was the same as in Example 2.

[0167] Table 18

[0168]

[0169] In addition, we prepared reference standards for each of the five detected sites for kit performance evaluation. We selected the MSI-H reference standard containing 5% tumor DNA as the detection limit reference standard, the MSI-H reference standard containing 20% ​​tumor DNA as the positive reference standard, and the reference standard containing 20 ng of wild-type DNA as the negative reference standard. The results of the five reaction tubes detecting the reference standards are shown in the figures below. Figures 3-7 As shown. Among them, Figure 3 The result diagram of the reference sample for HAL gene system detection; Figure 4 The result diagram of the NBAS genetic system testing reference sample; Figure 5 The results of the ZBTB37 gene system detection reference sample are shown in the figure. Figure 6 The result diagram of the reference sample for TSHZ2 gene system detection; Figure 7 The image shows the results of the FOXP2 gene system detection reference sample. As can be seen, the kit's five reaction systems accurately distinguish between MSI-H and MSS samples, with a detection limit reaching 5% of tumor DNA content. The bimodal graphs of the results are clearly distinct, facilitating result interpretation, and the overall performance is excellent.

[0170] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A biomarker composition, characterized in that, The biomarker composition includes the following first biomarker, second biomarker, third biomarker, fourth biomarker and fifth biomarker: The first biomarker is located in the HAL gene, starting at chr12:95985886, and contains 11 consecutive T single nucleotide tandem repeat sequences. The second biomarker is located in the NBAS gene, starting at chr2:15396386, and contains 11 consecutive A single nucleotide tandem repeats. The third biomarker is located in the ZBTB37 gene, starting at chr1:173899196, and contains 11 consecutive T single nucleotide tandem repeat sequences. The fourth biomarker is located in the TSHZ2 gene, starting at chr20:53494836, and contains 11 consecutive A single nucleotide tandem repeat sequences. The fifth biomarker is located in the FOXP2 gene, starting at chr7:114690176, and contains 11 consecutive A single nucleotide tandem repeats.

2. The biomarker composition according to claim 1, characterized in that, The biomarker composition further includes at least one of the following sixth, seventh, and eighth biomarkers; The sixth biomarker is located in the LRRC8D gene, starting at chr1:89935847, and contains 11 consecutive T single nucleotide tandem repeat sequences. The seventh biomarker is located in the PBRM1 gene, starting at chr3:52662295, and contains 11 consecutive A single nucleotide tandem repeat sequences. The eighth biomarker is located in the LRBA gene, starting at chr4:150914351, and contains 11 consecutive A single nucleotide tandem repeats.

3. The biomarker composition according to claim 2, characterized in that, The biomarker composition is (1), (2), or (3): (1): First marker, second marker, third marker, fourth marker, fifth marker and sixth marker; (2): First marker, second marker, third marker, fourth marker, fifth marker, sixth marker and seventh marker; (3) First marker, second marker, third marker, fourth marker, fifth marker, sixth marker, seventh marker and eighth marker.

4. The use of the reagent for detecting the biomarker composition as described in any one of claims 1 to 3 in the preparation of a microsatellite instability detection kit for endometrial cancer.

5. The use of the reagent for detecting the biomarker composition as described in any one of claims 1 to 3 in the preparation of an endometrial cancer detection kit.

6. A kit for detecting endometrial cancer, characterized in that, Including reagents for detecting mutations in the biomarker composition as described in any one of claims 1 to 3.

7. The kit according to claim 6, characterized in that, The kit employs a high-resolution melting curve method and includes the following groups of amplification primers and probes: Group A, Group B, Group C, Group D, and Group E. Group A: Restriction primers as shown in SEQ ID NO: 20 and excess primers as shown in SEQ ID NO: 21, and probes as shown in SEQ ID NO: 58, targeting the first marker; Group B: Restriction primers as shown in SEQ ID NO: 22 and excess primers as shown in SEQ ID NO: 23, and probes as shown in SEQ ID NO: 59, targeting the second biomarker; Group C: Restriction primers as shown in SEQ ID NO: 24 and excess primers as shown in SEQ ID NO: 25, and probes as shown in SEQ ID NO: 60, targeting the third biomarker; Group D: Restriction primers as shown in SEQ ID NO: 26 and excess primers as shown in SEQ ID NO: 27, and probes as shown in SEQ ID NO: 61, targeting the fourth biomarker; Group E: Restriction primers as shown in SEQ ID NO: 28 and excess primers as shown in SEQ ID NO: 29 for the fifth marker, and probes as shown in SEQ ID NO:

62.

8. The kit according to claim 7, characterized in that, The kit also includes at least one of the following groups of amplification primers and probes: Group F, Group G, and Group H: Group F: Restriction primers as shown in SEQ ID NO: 30 and excess primers as shown in SEQ ID NO: 31, and probes as shown in SEQ ID NO: 63, targeting the sixth biomarker; Group G: Restriction primers as shown in SEQ ID NO: 32 and excess primers as shown in SEQ ID NO: 33 for the seventh marker, and probes as shown in SEQ ID NO: 64; Group H: Restriction primers as shown in SEQ ID NO: 34 and excess primers as shown in SEQ ID NO: 35 for the eighth marker, and probes as shown in SEQ ID NO:

65.

9. The kit according to claim 7, characterized in that, Each probe has 1 to 4 bases from its 3' end that are thiolated.

10. A method for detecting microsatellite instability in endometrial cancer for non-diagnostic purposes, characterized in that, Includes the following steps: Obtain DNA from the sample to be tested, and use the kit as described in any one of claims 6 to 9 to detect the mutation status of the biomarker composition as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Biomarkers for detecting microsatellite instability in populations and their applications

    CN112708679B