Biomarkers for detecting human microsatellite instability and their applications

A panel of biomarkers for MSI detection using fluorescent PCR and specific genomic markers addresses the limitations of current methods, offering accurate, cost-effective, and rapid MSI analysis.

JP7772807B2Active Publication Date: 2025-11-18AMOY DIAGNOSTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023545850
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-27
Filing Date
2022-01-21
Publication Date
2025-11-18
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

Current methods for detecting microsatellite instability (MSI) are costly, time-consuming, require specialized equipment, and are prone to contamination, with limited versatility and sensitivity, making them unsuitable for widespread adoption.

Method used

A panel of biomarkers, including 22 homopolymeric repeat markers located at specific genomic positions, is used for MSI detection, combined with a fluorescent PCR platform and molecular beacon or light-off/light-on probes for rapid, accurate analysis.

Benefits of technology

The biomarker panel provides high detection accuracy, low cost, and fast turnaround time, reducing equipment costs and eliminating the need for complex software analysis, while maintaining high sensitivity and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007772807000042
    Figure 0007772807000042
  • Figure 0007772807000043
    Figure 0007772807000043
  • Figure 0007772807000044
    Figure 0007772807000044
Patent Text Reader

Abstract

The present invention provides a biomarker group for detecting human microsatellite instability, the biomarker group including at least two of markers 1 to 22, and use thereof. Msl, which detects human populations using the biomarker group, has high detection accuracy, a versatile detection platform, and low detection costs.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of microsatellite instability detection, and specifically to a group of biomarkers for detecting human microsatellite instability and their applications. [Background technology]

[0002] Microsatellites (MS), also known as short tandem repeats (STRs), are DNA sequences consisting of a few nucleotides (usually 1–6) repeated in tandem in the human genome. The repeat length ranges from 5–50, and approximately 500,000 such sequences are widely distributed throughout the human genome. When short tandem repeats are replicated, strand slippage due to highly repeated bases leads to replication errors at a much higher rate than in normal sequences. In healthy individuals, mismatches in short tandem repeats are promptly repaired by the DNA mismatch repair (MMR) mechanism, maintaining the length of the short tandem repeats. Abnormalities in DNA mismatch repair function result in uncorrected replication errors in short tandem repeats, leading to changes in repeat length or base composition, a condition known as microsatellite instability (MSI).

[0003] Currently, the gold standard for MSI detection is the detection of five microsatellite markers using PCR-capillary electrophoresis. This requires analysis of the amplified products using a first-generation sequencer, and was first proposed by the National Cancer Institute (NCI). This panel includes two single-nucleotide repeat markers and three dinucleotide repeat markers. Following a 2004 revision, the dinucleotide repeat markers were replaced with single-nucleotide repeat markers to improve detection sensitivity and specificity. When analyzing MSI detection results, if two or more markers cause microsatellite instability, the result is high microsatellite instability (MSI-H); if only one marker is detected, the result is low microsatellite instability (MSI-L); and if no marker indicates microsatellite instability, the result is microsatellite stable (MSS). This gold-standard detection method has the following obvious drawbacks that make it difficult to widely use MSI detection: 1. When detecting each sample, the sample to be detected and the control sample must be detected at the same time, which increases the detection cost and workload. 2. The detection process is lengthy, requiring PCR amplification, amplification product processing, first-generation sequencer fragment analysis detection, specialized software analysis, etc., making it difficult to complete MSI detection quickly. 3. The high cost of purchasing the equipment, as well as the reagents and consumables used in first-generation sequencer kits, increase the cost of MSI detection. 4. The assay platform is not very versatile, there are few products related to capillary electrophoresis in traditional assays, and the level of expertise in capillary electrophoresis platforms in each laboratory is low. 5. Because the concentration of the amplified product is high, the reaction tube for the amplified product must be opened many times during detection, which can easily lead to contamination.

[0004] Immunohistochemistry (IHC) can be used to detect MMR-associated proteins (MLH1, MSH2, MSH6, and PMS2). The corresponding proteins are fluorescently labeled and then microscopically examined for their expression in individual cells, allowing for further analysis of MMR dysfunction. If all four proteins are detectable in all cells, MMR expression is normal (pMMR), whereas if one or more proteins are absent in the cells, MMR expression is absent (dMMR). While IHC can be developed for basic detection, significant differences in the performance of different detection kits, relatively low detection sensitivity, and high physician demands for visual observation make it difficult to adopt as a standard approach for MSI detection.

[0005] NGS can be used to detect MSI markers by preparing a library of multiple markers in a sample, sequencing the library using a second-generation sequencer, and analyzing whether the repeat length of each marker has changed using a sensitive algorithm to further determine the MSI status. While MSI detection using NGS platforms has matured, its potential applications are limited by drawbacks such as the lack of continuous improvement in the fidelity of sequencing reagents, high detection costs, long detection cycles, and significant differences in detection algorithms between different detection kits.

[0006] All three of the above methods for detecting MSI have serious flaws, so there is an urgent need for detection methods with high detection accuracy, a common detection platform, low detection costs, and short detection cycles to confirm MSI status. Summary of the Invention [Problem to be solved by the invention]

[0007] The object of the present invention is to overcome the drawbacks of the prior art and to provide a group of biomarkers for detecting human microsatellite instability.

[0008] Another object of the present invention is to provide applications of the above-mentioned group of biomarkers for detecting human microsatellite instability. [Means for solving the problem]

[0009] The technical solutions of the present invention are as follows: A panel of biomarkers for detecting human microsatellite instability, comprising at least two of the following biomarkers:

[0010] a first marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human EIF4E3 gene, and starting at chr3:71739332; a second marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human UBAC2 gene, starting at chr13:99890849; a third marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human TAOK3 gene, and beginning at chr12:118675984; a fourth marker containing a homopolymeric repeat of 11 consecutive T bases and localizing to the human IFT140 gene, beginning at chr16:1612145; a fifth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human PRR5-ARHGAP8 gene, and beginning at chr22:45205158; a sixth marker containing a homopolymeric repeat sequence of 10 consecutive A bases and located in the human AVIL gene, beginning at chr12:58202496; a seventh marker containing a homopolymeric repeat sequence of eight consecutive A bases, located in the human ACVR2A gene, and beginning at chr2:148683685; an eighth marker containing a homopolymeric repeat sequence of 12 consecutive A bases and located in the human PPP1CC gene, beginning at chr12:111160513; a ninth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human RBM14-RBM4 gene, and beginning at chr11:66410771; a tenth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human SDHC gene, and beginning at chr1:161309335; an eleventh marker containing a homopolymeric repeat sequence of 12 consecutive A bases and located in the human PUM2 gene, starting at chr2:20526998; a twelfth marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human DEC1 gene, starting at chr9:118164375; a thirteenth marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human COL11A1 gene, and beginning at chr1:103468855; a fourteenth marker containing a homopolymeric repeat of 11 consecutive T bases and located in the human YTHDF3 gene, starting at chr8:64081981; a fifteenth marker containing a homopolymeric repeat sequence of 12 consecutive T bases and located in the human ACTL6A gene, starting at chr3:179291295; a sixteenth marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human MCM3AP gene, and beginning at chr21:47703455; a seventeenth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human SPDL1 gene, and starting at chr5:169020337; an 18th marker containing a homopolymeric repeat sequence of 12 consecutive T bases and located in the human SMARCA2 gene, starting at chr9:2083325; a 19th marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human PROSER1 gene, and beginning at chr13:39608335; a 20th marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human MZB1 gene, and beginning at chr5:138723143; a 21st marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human ANKIB1 gene, starting at chr7:92020452; A group of biomarkers for detecting human microsatellite instability, comprising at least two of the following: a 22nd marker that contains a homopolymer repeat of 10 consecutive A bases, is localized in the human BRD4 gene, and begins at chr19:15366000; and

[0011] In a preferred embodiment of the invention, the group of biomarkers consists of at least 2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10, or 11, or 12, or 13, or 14, or 15, or 16, or 17, or 18, or 19, or 20, or 21 of the biomarkers, or consists of all of the markers.

[0012] Other embodiments of the present invention are as follows. 1. A method for analyzing human microsatellite instability, comprising: a first marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human EIF4E3 gene, and starting at chr3:71739332; a second marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human UBAC2 gene, starting at chr13:99890849; a third marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human TAOK3 gene, and beginning at chr12:118675984; a fourth marker containing a homopolymeric repeat of 11 consecutive T bases and localizing to the human IFT140 gene, beginning at chr16:1612145; a fifth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human PRR5-ARHGAP8 gene, and beginning at chr22:45205158; a sixth marker containing a homopolymeric repeat sequence of 10 consecutive A bases and located in the human AVIL gene, beginning at chr12:58202496; a seventh marker containing a homopolymeric repeat sequence of eight consecutive A bases, located in the human ACVR2A gene, and beginning at chr2:148683685; an eighth marker containing a homopolymeric repeat sequence of 12 consecutive A bases and located in the human PPP1CC gene, beginning at chr12:111160513; a ninth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human RBM14-RBM4 gene, and beginning at chr11:66410771; a tenth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human SDHC gene, and beginning at chr1:161309335; an eleventh marker containing a homopolymeric repeat sequence of 12 consecutive A bases and located in the human PUM2 gene, starting at chr2:20526998; a twelfth marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human DEC1 gene, starting at chr9:118164375; a thirteenth marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human COL11A1 gene, and beginning at chr1:103468855; a fourteenth marker containing a homopolymeric repeat of 11 consecutive T bases and located in the human YTHDF3 gene, starting at chr8:64081981; a fifteenth marker containing a homopolymeric repeat sequence of 12 consecutive T bases and located in the human ACTL6A gene, starting at chr3:179291295; a sixteenth marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human MCM3AP gene, and beginning at chr21:47703455; a seventeenth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human SPDL1 gene, and starting at chr5:169020337; an eighth marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human SMARCA2 gene, starting at chr9:2083325; a 19th marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human PROSER1 gene, and beginning at chr13:39608335; a 20th marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human MZB1 gene, and beginning at chr5:138723143; a 21st marker containing a homopolymeric repeat sequence of 12 consecutive T bases and located in the human ANKIB1 gene, starting at chr7:92020452; A method for analyzing human microsatellite instability, comprising determining the number of overlapping bases in at least two homopolymer repeat sequences from the group consisting of: a 22nd marker containing a homopolymer repeat sequence of 10 consecutive A bases, located in the human BRD4 gene, and starting at chr19:15366000.

[0013] In a preferred embodiment of the invention, the number of bases in the homopolymer repeat sequence and its mutant forms for at least 2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10, or 11, or 12, or 13, or 14, or 15, or 16, or 17, or 18, or 19, or 20, or 21 of said biomarkers, or for all of said markers, is determined.

[0014] In a preferred embodiment of the invention, the method further comprises the step of amplifying the region of said homopolymer repeat sequence or a mutant form thereof.

[0015] More preferably, melting curve data is generated after the detection process is completed.

[0016] More preferably, the amplifying step comprises using molecular beacon probes.

[0017] More preferably, the amplification further comprises a limiting primer and an excess primer corresponding to each region of the homopolymer repeat sequence or a mutant form thereof, the limiting primer having the same orientation as the molecular beacon probe, the single-stranded template generated from the excess primer complementarily pairing with the molecular beacon probe, and the amount of excess primer is greater than the amount of limiting primer.

[0018] More preferably, the ratio of limiting primer to excess primer is 1:3 to 1:20.

[0019] More preferably, the amplification step uses a Light off / Light on probe.

[0020] More preferably, the amplification further comprises a limiting primer and an excess primer corresponding to each region of the homopolymer repeat sequence or a mutant form thereof, the limiting primer having the same orientation as the Light off / Light on probe, the single-stranded template generated from the excess primer complementarily pairing with the Light off / Light on probe, and the amount of excess primer is greater than the amount of limiting primer.

[0021] More preferably, the ratio of limiting primer to excess primer is 1:3 to 1:20.

[0022] More preferably, the number of cycles of the amplification is 40 to 70.

[0023] More preferably, the amplification is carried out using a high-fidelity DNA polymerase.

[0024] Yet another embodiment of the present invention is as follows. 1. A kit for analyzing microsatellite instability in a human biological sample, comprising: a first marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human EIF4E3 gene, and starting at chr3:71739332; a second marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human UBAC2 gene, starting at chr13:99890849; a third marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human TAOK3 gene, and beginning at chr12:118675984; A fourth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human IFT140 gene, begins at chr16:1612145. a fifth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human PRR5-ARHGAP8 gene, and beginning at chr22:45205158; a sixth marker containing a homopolymeric repeat sequence of 10 consecutive A bases and located in the human AVIL gene, beginning at chr12:58202496; a seventh marker containing a homopolymeric repeat sequence of eight consecutive A bases, located in the human ACVR2A gene, and beginning at chr2:148683685; an eighth marker containing a homopolymeric repeat sequence of 12 consecutive A bases and located in the human PPP1CC gene, beginning at chr12:111160513; a ninth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human RBM14-RBM4 gene, and beginning at chr11:66410771; a tenth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human SDHC gene, and beginning at chr1:161309335; an eleventh marker containing a homopolymeric repeat sequence of 12 consecutive A bases and located in the human PUM2 gene, starting at chr2:20526998; a twelfth marker containing 12 consecutive homopolymeric repeats of T and located in the human DEC1 gene, starting at chr9:118164375; a thirteenth marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human COL11A1 gene, and beginning at chr1:103468855; a fourteenth marker containing a homopolymeric repeat of 11 consecutive T bases and located in the human YTHDF3 gene, starting at chr8:64081981; a fifteenth marker containing a homopolymeric repeat sequence of 12 consecutive T bases and located in the human ACTL6A gene, starting at chr3:179291295; a sixteenth marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human MCM3AP gene, and beginning at chr21:47703455; a seventeenth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human SPDL1 gene, and starting at chr5:169020337; an 18th marker containing a homopolymeric repeat sequence of 12 consecutive T bases and located in the human SMARCA2 gene, starting at chr9:2083325; a 19th marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human PROSER1 gene, and beginning at chr13:39608335; a 20th marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human MZB1 gene, and beginning at chr5:138723143; The 21st marker contains a homopolymeric repeat sequence of 12 consecutive T bases and is located in the human ANKIB1 gene, starting at chr7:92020452. A kit for analyzing microsatellite instability in a human biological sample, comprising a primer pair and a probe for amplifying at least two markers from the group consisting of: a 22nd marker comprising a homopolymeric repeat sequence of 10 consecutive A bases, located in the human BRD4 gene, and starting at chr19:15366000; and

[0025] In a preferred embodiment of the present invention, the kit comprises primer pairs and probes for amplifying the number of bases of homopolymer repeat sequences of at least 2, or 3, or 4, or 5, or 6, or 7, or 8, or 9, or 10, or 11, or 12, or 13, or 14, or 15, or 16, or 17, or 18, or 19, or 20, or 21 of the biomarkers, or the number of bases of homopolymer repeat sequences of all of the biomarkers.

[0026] In a preferred embodiment of the invention, the probe is a molecular beacon probe.

[0027] More preferably, the primer pair consists of a limiting primer and an excess primer corresponding to each of the homopolymer repeat sequences, the limiting primer has the same orientation as the molecular beacon probe, the single-stranded template generated from the excess primer complementarily pairs with the molecular beacon probe, and the amount of the excess primer is greater than the amount of the limiting primer.

[0028] More preferably, the ratio of limiting primer to excess primer is 1:3 to 1:20.

[0029] In a preferred embodiment of the invention, the probes are a Light off / Light on probe pair.

[0030] More preferably, the primer pair consists of a limiting primer and an excess primer corresponding to each of the homopolymer repeat sequences, the limiting primer has the same orientation as the Light off / Light on probe, the single-stranded template generated from the excess primer complementarily pairs with the Light off / Light on probe, and the amount of excess primer is greater than the amount of limiting primer.

[0031] More preferably, the ratio of the Limiting Primer to the Excess Primer is 1:3 to 1:20.

[0032] In a preferred embodiment of the present invention, the number of cycles of amplification using the primer pair is 40 to 70 cycles.

[0033] In a preferred embodiment of the present invention, amplification with said primer pair uses a high-fidelity DNA polymerase. [Effects of the Invention]

[0034] The advantageous effects of the present invention are as follows: 1. The detection of human MSI according to the present invention has high detection accuracy, a versatile detection platform, and low detection costs. 2. The homopolymer repeat lengths detected by the present invention are concentrated in the range of 7 to 14 bp, and the markers can detect MSI status in humans. 3. The present invention can perform detection using a fluorescent PCR platform, and the equipment price is reduced from one million to about 150,000 compared with first-generation gene sequencing or second-generation gene sequencing. 4. The detection platform of the present invention reduces the coupling between instruments and reagents, increases the versatility of reagent kit detection, and the detection judgment does not require software algorithms but is directly determined based on the "presence / absence" of melting peaks, allowing rapid judgment by the naked eye. 5. The kit of the present invention can complete the detection of 24 to 48 samples in about 2 hours using a single fluorescent PCR device, which has reasonable throughput and fast detection time. 6. This invention uses a melting curve for detection, and fluorescently labeled probes are used for detection. The probes match the mutant sequences, and the mutant sequences are verified and determined by NGS. The mutant sequences consist of 1-bp repeated base changes, but 2-bp base changes are also possible. [Brief explanation of the drawings]

[0035] [Figure 1] 1 shows the source and screening process of sites relevant to MSI detection in Example 1 of the present invention. [Figure 2] 1 shows the distribution of mutations in markers of MSS (negative) and MSI (positive) samples in Example 1 of the present invention. [Figure 3] FIG. 10 is a line diagram showing the detection of different samples by a molecular beacon probe and a light-off / light-on probe in Example 2 of the present invention. [Figure 4] FIG. 10 is a line graph of the melting curves of clinical samples for marker detection in Example 4 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0036] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0037] Example 1 The biomarker group of the present invention is derived from sequencing data such as human WGS and WES, and MS markers are selected according to the scheme in FIG. 1.

[0038] MS markers were analyzed based on the difference in tandem repeat length between MSI-negative and MSI-positive samples. The length variations of MS markers in NGS sequencing data were analyzed using human genome databases such as NCBI, Ensembl, and UCSC as reference sequences. NGS sequencing data from human colorectal cancer samples was collected using WGS and WES to determine the detected length of each MS marker. The repeat length variations of the detected sequence were compared with the reference sequence to determine the variations in each tissue sample in colorectal cancer patients. Among these, the length of MS markers in MSI samples should be altered, with a higher incidence of mutations with decreased or increased length. The repeat length of MS markers in MSS samples should match the repeat length of the reference sequence.

[0039] By establishing a training set and a test set, the sensitivity and specificity of the selected sites in different grouped samples were verified, and the AUC of each MS marker was calculated based on the sensitivity and specificity data of the MS marker using SPASS software. The closer the AUC is to 1.0, the better the marker's ability to distinguish between MSI and MSS samples. The selected markers are listed in Table 1 below: Table 1: MS marker information selected after analysis of sequencing data JPEG0007772807000001.jpg126170

[0040] Clinical samples were detected using the "gold standard" PCR-capillary electrophoresis method, and the MSI status of each sample was determined as the standard result for NGS site screening. After analyzing the sequencing data, oligos were designed for selected markers. Clinical samples for which MSI status was determined using the gold standard were then compiled into libraries and subjected to NGS sequencing. The half-lengths of each marker in MSI and MSS samples were determined using bioassays. Using human genome databases such as NCBI, Ensembl, and UCSC as standard reference sequences, the detection sensitivity of each marker in MSI samples and the detection specificity of each marker in MSS samples were analyzed. The results of plotting the negative and positive results of the marker mutant sequences are shown in Figure 2. Software such as SPASS was used to analyze the detection results of each marker in MSI and MSS samples, demonstrating that more markers can effectively distinguish between MSI and MSS samples. The AUC information for each marker's detection in MSS and MSI samples is shown in Table 2 below. Table 2: AUC information for distinguishing MSI and MSS samples for markers screened by NGS JPEG0007772807000002.jpg127170

[0041] The analysis results showed that the 22 markers tested had excellent discrimination ability when detecting MSS and MSI samples, respectively, with AUC values ​​all greater than 0.95, and could specifically discriminate negative-positive samples.

[0042] <Example 2> In this example, each MS marker obtained in Example 1 is detected by the fluorescent PCR melting curve method, and analysis is performed using a fluorescent probe during melting curve detection. The sequences of the detected markers are indicated by underlined sequences, and the sequences of each marker are as follows: 1.EIF4E3(chr3:71739332,SEQ ID NO.01): JPEG0007772807000003.jpg221702.UBAC2(chr13:99890849,SEQ ID NO.02) JPEG0007772807000004.jpg221703.TAOK3(chr12:118675984,SEQ ID NO.03) JPEG0007772807000005.jpg221704.IFT140(chr16:1612145,SEQ ID NO.04) JPEG0007772807000006.jpg221705.PRR5-ARHGAP8(chr22:45205158,SEQ ID NO.05) JPEG0007772807000007.jpg221706.AVIL(chr12:58202496,SEQ ID NO.06) JPEG0007772807000008.jpg221707.ACVR2A(chr2:148683685,SEQ ID NO.07) JPEG0007772807000009.jpg221708.PPP1CC(chr12:111160513,SEQ ID NO.08) JPEG0007772807000010.jpg221709.RBM14-RBM4(chr11:66410771,SEQ ID NO.09) JPEG0007772807000011.jpg2217010.SDHC(chr1:161309335,SEQ ID NO.10) JPEG0007772807000012.jpg2217011.PUM2(chr2:20526998,SEQ ID NO.11) JPEG0007772807000013.jpg2217012.DEC1(chr9:118164375,SEQ ID NO.12) JPEG0007772807000014.jpg2217013.COL11A1(chr11:103468855,SEQ ID NO.13) JPEG0007772807000015.jpg2217014.YTHDF3(chr8:64081981,SEQ ID NO.14) JPEG0007772807000016.jpg2217015.ACTL6A(chr3:179291295,SEQ ID NO.15) JPEG0007772807000017.jpg2217016.MCM3AP(chr21:47703455,SEQ ID NO.16) JPEG0007772807000018.jpg2217017.SPDL1(chr5:169020337,SEQ ID NO.17) JPEG0007772807000019.jpg2217018.SMARCA2(chr9:2083325,SEQ ID NO.18) JPEG0007772807000020.jpg2217019.PROSER1(chr13:39608335,SEQ ID NO.19) JPEG0007772807000021.jpg2217020.MZB1(chr5:138723143,SEQ ID NO.20) JPEG0007772807000022.jpg2217021.ANKIB1(chr7:92020452,SEQ ID NO.21) JPEG0007772807000023.jpg2217022.BRD4(chr19:15366000,SEQ ID NO.22) JPEG0007772807000024.jpg22170

[0043] For the above markers screened by NGS, upstream and downstream primers should be designed to amplify the desired marker sequence and have an appropriate amplicon length. During amplification, the amounts of the upstream and downstream primers should be adjusted so that one primer sufficiently hybridizes with the probe to form a single-stranded amplified oligonucleotide. The primer with the lower dosage is designated the "limiting primer," while the other primer with the higher dosage is designated the "excess primer." The probe is oriented in the same direction as the limiting primer. The design sequences of the marker primers are shown in Table 3 below. Table 3: Marker primer information JPEG0007772807000025.jpg140170JPEG0007772807000026.jpg229170

[0044] A detection probe was designed for the marker screened by NGS in Example 1. The probe was positioned between the upstream and downstream primers, and the mutant sequence was designed to have a Tm value in the range of 30 to 55°C. The probe can be protected at the 5' end with a thiol modification to prevent degradation of the proofreading activity by 3' to 5'. The probe must also be modified with a quenching or blocking group at the 3' end to prevent extension by 5' to 3' polymerization activity. The designed probe may be a molecular beacon probe or a light-off / light-on probe, and single-stranded oligonucleotides produced by excess primers can complementarily pair with the probe. The probe must be modified with a fluorophore containing a modification such as FAM, VIC, HEX, ROX, CY3, or CY5; a quenching group containing a modification such as BHQ1, BHQ2, or BHQ3; and a blocking group containing a modification such as an amino group, a dideoxy group, or a C3 / C6 spacer. The designed sequences of the marker probes are shown in Table 4 below: Table 4: Marker probe information JPEG0007772807000027.jpg138170JPEG0007772807000028.jpg226170JPEG0007772807000029.jpg178170

[0045] A detection reaction mixture was prepared using limiting primers, excess primers, and a probe to amplify the MS marker. Detection was performed using a fluorescent PCR system. The annealing temperature was determined based on the primer Tm. The number of amplification cycles was set to 60, and the melting curve collection temperature range was set to 30–55°C. After amplification, melting curves of the probe and single-stranded oligonucleotide template were collected. The negative derivative of the temperature versus fluorescence was calculated to obtain a melting peak profile and interpret the results. The melting curve detection line is shown in Figure 3. The peak of the molecular beacon probe is upward-pointing. If the sample contains a pure mutant template, there is a melting peak at a high Tm. If the sample contains a pure wild-type template, there is a melting peak at a low Tm. If the sample contains both mutant and wild-type templates, there are melting peaks at both high and low Tm. When a light-off / light-on probe is used, the melting peak points downward, and if the sample is a pure mutant template, the melting peak is at a high Tm value; if the sample is a pure wild-type template, the melting peak is at a low Tm value; and if both mutant and wild-type templates are present in the sample, the melting peak is at both a high and a low Tm value.

[0046] Example 3 The kit of the present invention is an MSI detection kit consisting of the biomarkers in Example 2. Prior to marker selection, it is necessary to analyze the sensitivity and specificity of each marker in MSI-negative and MSI-positive samples. Each MS marker is detected using a molecular beacon probe or a light-off / light-on probe in MSI and MSS samples. After detection is complete, the detection sensitivity of each marker in MSI samples and the detection specificity in MSS samples are compiled and statistically analyzed. The detection sensitivity and specificity of each MS marker are shown in the table below: Table 5: MSI detection sensitivity and specificity of markers JPEG0007772807000030.jpg126170

[0047] As can be seen from the detection results, each marker in Example 2 has high detection performance when detecting both MSI and MSS samples, but combining markers is necessary to further ensure the detection performance of the kit. When detecting MSS samples, each marker ensured 100% specificity, and no nonspecific detection was observed. Different markers had different detection sensitivities for MSI samples, ranging from 89.02 to 96.34%, so combining different markers is necessary when using the kit.

[0048] Example 4 An MSI detection kit is constructed using a combination of different markers as described in Example 2. A reasonable amount of MS markers should be used when constructing the kit. Using too many MS markers in the kit increases the manufacturing cost of the kit, wastes sample DNA during detection, increases the detection workload, and reduces detection throughput. Using too few MS markers in the kit may result in risks of missed detections and false positives, potentially reducing the kit's detection performance. For each MS marker, a consistency analysis of the detection results is performed to determine whether there are any overlaps, missed detections, or other indications between the detection results of different MS markers.

[0049] The combinations of markers include four sets of two-marker combinations, three sets of four-marker combinations, two sets of eight-marker combinations, and one set of ten-marker combinations. By randomly selecting markers, two markers are combined to create four sets of two-marker combinations, by randomly selecting markers, four markers are combined to create three sets of four-marker combinations, by randomly selecting markers, eight markers are combined to create two sets of eight-marker combinations, and by randomly selecting markers, ten markers are combined to create one set of ten-marker combinations. Specific combination information is shown in the table below. Table 6: MSI detection sensitivity and specificity of marker combinations JPEG0007772807000031.jpg128170

[0050] After selecting MSI markers, the reaction tubes for the markers can be combined to detect different markers in the same reaction tube. An internal control gene can be added to the single-tube reaction mixture to control sample quality, detection reagents, and the detection process, improving the stability and reliability of the kit.

[0051] The 10 combinations in Table 6 were compared for consistency with the gold standard "PCR-capillary electrophoresis." Clinical samples were blindly processed, and the same clinical samples were separately detected using the 10 combinations in Table 6 and the "PCR-capillary electrophoresis" kit. The test results could be interpreted directly and quickly without the need for additional algorithms or graphic analysis software. After detection was completed, the detection consistency of the two reagents was analyzed by comparing them with the gold standard assay. An example of the detection line shape for a clinical sample is shown in Figure 4. The test results of the kit of the present invention and the gold standard test kit were analyzed for consistency. The sensitivity, specificity, positive predictor, negative predictor, overall accuracy, and Kappa value were summarized, and the consistency of the clinical sample tests is shown in Tables 7 to 16. Table 7: Comparison of agreement between the gold standard and Group 1 combinations JPEG0007772807000032.jpg49170Sensitivity of the present kit = 100% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 8: Comparison of agreement between the gold standard and Group 2 combinations JPEG0007772807000033.jpg49170Sensitivity of the present kit = 99.39% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 99.53% Total precision = 100% Kappa = 1.0 Table 9: Comparison of agreement between the gold standard and the Group 3 combination JPEG0007772807000034.jpg49170Sensitivity of the present kit = 100% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 10: Comparison of agreement between the gold standard and Group 4 combinations JPEG0007772807000035.jpg49170Sensitivity of the present kit = 100% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 11: Comparison of agreement between the gold standard and the Group 5 combination JPEG0007772807000036.jpg49170Sensitivity of the present kit = 100% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 12: Comparison of agreement between the gold standard and Group 6 combinations JPEG0007772807000037.jpg49170Sensitivity of the present kit = 100% Specificity of the present kit = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 13: Comparison of agreement between the gold standard and Group 7 combinations JPEG0007772807000038.jpg49170Sensitivity of the present kit = 100% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 14: Comparison of agreement between the gold standard and the Group 8 combination JPEG0007772807000039.jpg49170Sensitivity of the present kit = 100% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 15: Comparison of agreement between the gold standard and the Group 9 combination JPEG0007772807000040.jpg49170Sensitivity of the present kit = 100% Specificity of the present kit = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0 Table 16: Comparison of agreement between the gold standard and the group 10 combination JPEG0007772807000041.jpg49170Sensitivity of the present kit = 100% Specificity of the kit of the present invention = 100% Positive predictive value = 100% Negative predictive value = 100% Total precision = 100% Kappa = 1.0

[0052] As can be seen from the detection results, when the biomarkers in Example 2 were used, the detection performance was excellent even when different random combinations were performed, and the detection specificity was maintained at 100% for both random combinations, and the sensitivity was higher than 99%, showing excellent detection performance.

[0053] The above description is merely a preferred embodiment of the present invention, and therefore the scope of the present invention cannot be limited thereby. In other words, any equivalent changes and modifications based on the claims and the contents of the specification of the present invention fall within the comprehensive scope of the present invention. [Industrial Applicability]

[0054] The present invention discloses a group of biomarkers for detecting human microsatellite instability and applications thereof, the group of biomarkers including at least two of markers 1 to 22. Detection of human MSI according to the present invention is highly accurate, has a versatile detection platform, is inexpensive, and has industrial applicability.

Claims

1. 1. A method for analyzing human microsatellite instability in vitro, comprising: a first marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human EIF4E3 gene, and starting at chr3:71739332; a second marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human UBAC2 gene, starting at chr13:99890849; a fourth marker containing a homopolymeric repeat of 11 consecutive T bases and localizing to the human IFT140 gene, beginning at chr16:1612145; a fifth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human PRR5-ARHGAP8 gene, and beginning at chr22:45205158; a seventh marker containing a homopolymeric repeat sequence of eight consecutive A bases, located in the human ACVR2A gene, and beginning at chr2:148683685; a ninth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human RBM14-RBM4 gene, and beginning at chr11:66410771; A method for analyzing human microsatellite instability in vitro, comprising determining the number of bases in each homopolymer repeat sequence contained therein.

2. 2. The method of claim 1, further comprising the step of amplifying the region of said homopolymer repeat sequence or a mutant form thereof.

3. 3. The method of claim 2, wherein melting curve data is generated after the validation process is complete.

4. 3. The method of claim 2, wherein the amplification comprises using molecular beacon probes.

5. The method of claim 4, characterized in that the amplification further includes a limiting primer and an excess primer corresponding to each region of the homopolymer repeat sequence or its mutant form, the limiting primer has the same orientation as the molecular beacon probe, the single-stranded template generated from the excess primer complementarily pairs with the molecular beacon probe, and the amount of excess primer is greater than the amount of limiting primer.

6. The method of claim 5, wherein the ratio of limiting primer to excess primer is 1:3 to 1:

20.

7. 3. The method of claim 2, wherein the amplification comprises using a Light off / Light on probe.

8. The method of claim 7, characterized in that the amplification further comprises a limiting primer and an excess primer corresponding to each region of the homopolymer repeat sequence or its mutant form, the limiting primer having the same direction as the Light off / Light on probe, the single-stranded template generated from the excess primer complementarily pairing with the Light off / Light on probe, and the amount of the excess primer is greater than the amount of the limiting primer.

9. 9. The method of claim 8, wherein the ratio of limiting primer to excess primer is 1:3 to 1:

20.

10. The method according to claim 2, wherein the number of cycles of the amplification is 40 to 70.

11. The method of claim 2, wherein the amplification is carried out using a high-fidelity DNA polymerase.

12. 1. A kit for analyzing microsatellite instability in a human biological sample, comprising: a first marker containing a homopolymeric repeat sequence of 12 consecutive A bases, located in the human EIF4E3 gene, and starting at chr3:71739332; a second marker containing a homopolymeric repeat of 12 consecutive T bases and located in the human UBAC2 gene, starting at chr13:99890849; a fourth marker containing a homopolymeric repeat of 11 consecutive T bases and localizing to the human IFT140 gene, beginning at chr16:1612145; a fifth marker containing a homopolymeric repeat of 11 consecutive T bases, located in the human PRR5-ARHGAP8 gene, and beginning at chr22:45205158; a seventh marker containing a homopolymeric repeat sequence of eight consecutive A bases, located in the human ACVR2A gene, and beginning at chr2:148683685; a ninth marker containing a homopolymeric repeat sequence of 12 consecutive T bases, located in the human RBM14-RBM4 gene, and beginning at chr11:66410771; A kit for analyzing microsatellite instability in a human biological sample, comprising a primer pair and a probe for amplifying the number of bases in each homopolymer repeat sequence contained in the sample.

13. The kit of claim 12, wherein the probe is a molecular beacon probe.

14. The kit of claim 13, wherein the primer pair consists of a limiting primer and an excess primer corresponding to each of the homopolymer repeat sequences, the limiting primer has the same orientation as the molecular beacon probe, the single-stranded template generated from the excess primer complementarily pairs with the molecular beacon probe, and the amount of the excess primer is greater than the amount of the limiting primer.

15. The kit of claim 14, wherein the ratio of limiting primer to excess primer is 1:3 to 1:

20.

16. The kit of claim 12, wherein the probes are a Light off / Light on probe pair.

17. The kit of claim 16, wherein the primer pair consists of a limiting primer and an excess primer corresponding to each of the homopolymer repeat sequences, the limiting primer has the same direction as the Light off / Light on probe, the single-stranded template generated from the excess primer complementarily pairs with the Light off / Light on probe, and the amount of the excess primer is greater than the amount of the limiting primer.

18. 18. The kit of claim 17, wherein the ratio of the limiting primer to the excess primer is 1:3 to 1:

20.

19. The kit according to claim 12, wherein the number of cycles of amplification using the primer pair is 40 to 70 cycles.

20. The kit according to claim 12, characterized in that the amplification using the primer pair uses a high-fidelity DNA polymerase.

Citation Information

Patent Citations

  • Biomarker combination and kit for microsatellite instability detection and application thereof

    CN111471755A

  • Biomarker panel and methods for detecting microsatellite instability in cancers

    WO2019145306A1