Method for screening microsatellite sites and method for detecting microsatellite instability

By pre-setting unimodal and bimodal probability distribution models to screen microsatellite loci with lower polymorphism, and combining them with probe group capture for microsatellite instability detection, the problem of insufficient information utilization and stability in existing methods is solved, and efficient and robust single-sample detection is achieved.

CN116469464BActive Publication Date: 2026-05-013D BIOMEDICINE SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
3D BIOMEDICINE SCI & TECH CO LTD
Filing Date
2022-01-12
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for detecting microsatellite instability have shortcomings in terms of information utilization and robustness of judgment. Traditional PCR detection requires control samples, and NGS methods are relatively arbitrary in site selection, which limits detection performance and stability. Furthermore, existing NGS algorithms are difficult to accurately describe the read distribution of microsatellite sites.

Method used

The maximum likelihood estimate of microsatellite loci is calculated using pre-defined unimodal and bimodal probability distribution models. By screening out a set of microsatellite loci with lower polymorphism, these loci are captured and detected using probe sets. The likelihood ratio threshold is determined by combining the training set for further screening, thereby achieving robustness of single-sample detection.

Benefits of technology

It achieves efficient microsatellite instability detection without the need for control samples, and has good detection performance, with specificity and sensitivity reaching 100% and 98.37%, respectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116469464B_ABST
    Figure CN116469464B_ABST
Patent Text Reader

Abstract

The present application relates to a microsatellite site screening method capable of screening a microsatellite site with lower polymorphism. The present application also relates to a microsatellite instability detection method using a microsatellite site obtained by the microsatellite site screening method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene detection. More specifically, this invention relates to a method for screening conserved microsatellite loci and a method for detecting microsatellite instability using the microsatellite loci obtained therefrom. Background Technology

[0002] Microsatellites (MS) are sequences of repeating bases in the human genome, with repeating units ranging from 1 to 6 bp in length. Microsatellite instability (MSI) refers to changes in the number of repeats in a microsatellite, resulting in the appearance of new alleles. The direct cause of microsatellite instability is primarily the dysfunction of the intracellular DNA mismatch repair (MMR) mechanism. MMR dysfunction mainly has two causes: 1) mutations in mismatch repair genes (including MLH1, MSH2, MSH6, PMS2, and EPCAM), and 2) hypermethylation of the MLH1 promoter region. MSI is an important biomarker for tumor immunotherapy.

[0003] Stable, accurate, and efficient detection of microsatellite instability (MSI) has significant scientific and clinical implications. However, the manifestations of MS loci are relatively complex, and traditional detection methods based on PCR and capillary electrophoresis may lead to misdiagnosis due to insufficient information. Next-generation sequencing (NGS) technology provides more comprehensive information, making it possible to systematically and accurately determine the MS status, but existing MSI detection protocols still have room for improvement in terms of information utilization and robustness.

[0004] Currently, the main technologies for detecting MSI are: 1. MSI-PCR, which mainly utilizes five microsatellite loci provided by the Bethesda panel to compare tumor samples and control samples and observe whether there is a significant change in PCR amplification length. This method requires a control sample; otherwise, it cannot be detected. 2. Immunohistochemistry, which detects whether the expression of MLH1, MSH2, MSH6, and PMS2 proteins is abnormal. 3. NGS, which is based on the capture method to detect microsatellite instability. Most existing NGS methods only mention model establishment, but they are mostly based on a few loci for detection, or the selection of loci is relatively arbitrary, which cannot maximize the use of patient sample information and has great limitations in terms of detection performance and stability.

[0005] The existing technology has the following main drawbacks in clinical practice:

[0006] 1. MSI-PCR method. This method often requires manual interpretation of the PCR images to determine the presence of signals such as the generation of new peaks and peak drift, in order to identify unstable states. For situations where judgment is difficult, manual interpretation may be biased. Furthermore, this method requires control samples, and in practice, historically preserved samples are often unavailable for control, significantly limiting its application. On the other hand, the five MS loci in the Bethesda panel were initially developed for colorectal cancer, which may pose potential problems for pan-cancer applications.

[0007] 2. Immunohistochemistry. Because this method detects MMR-related proteins rather than MSI itself, and due to the complexity of sample processing and experiments involved, it is prone to false negatives and inconsistencies. The reported false negative rate is between 5% and 11%.

[0008] 3. NGS Method. Most existing NGS methods are limited to a small number of loci or lack detailed locus selection principles, resulting in somewhat arbitrary use of MS loci within the panel, which may interfere with the stability of the results. Furthermore, clinical historical samples often lack controls, necessitating single-sample detection capabilities. Typical MSI detection algorithms use historical reference sets for processing. When loci exhibit population polymorphism, a specific patient's MS status may differ significantly from the historical sample set, leading to misclassification as microsatellite instability (MSI or MSI-high). Only by using a large number of properly selected loci for cross-verification can this situation be effectively avoided, truly ensuring the robustness of single-sample detection. Additionally, there is considerable room for improvement in probe design and detection algorithms.

[0009] Furthermore, current NGS-based MSI detection algorithms all require explicit or implicit (e.g., hypothesis testing) use the read distribution information of various lengths at MS loci. Commonly used distribution functions include binomial distribution, hypergeometric distribution, and normal distribution. However, in reality, the read distribution of stable MS loci has the following characteristics: 1. The distribution is discrete; 2. The distribution peak is narrow, often concentrated at a specific length; 3. The distribution is non-centrosymmetric, and the probability of the actual read length measured at a specific locus being shorter than the expected length and longer than the expected length is not the same. Figure 1 These characteristics make it difficult for the standard distribution described above to accurately describe the read distribution of MS sites. Summary of the Invention

[0010] Therefore, the purpose of this invention is to provide a method for screening microsatellite loci that can screen out microsatellite loci with lower polymorphism (i.e., conservative).

[0011] Another object of the present invention is to provide a microsatellite instability detection method that uses microsatellite sites obtained by the microsatellite site screening method of the present invention.

[0012] In one aspect, the present invention relates to a method for screening microsatellite loci, the method comprising the following steps:

[0013] Obtain the first point set, which includes multiple captured stable microsatellite sites;

[0014] Based on multiple reads of each microsatellite site in the first site set, a unimodal maximum likelihood estimate of the microsatellite site is calculated using a preset unimodal probability distribution model, and a bimodal maximum likelihood estimate of the microsatellite site is calculated using a preset bimodal probability distribution model, wherein the preset bimodal probability distribution model considers more polymorphism than the preset unimodal probability distribution model.

[0015] For each microsatellite locus in the first locus set, obtain the likelihood ratio of the bimodal maximum likelihood estimate to the unimodal maximum likelihood estimate; and

[0016] Based on a preset judgment model including thresholds and the likelihood ratio of each microsatellite site in the first site set, the microsatellite sites in the first site set are screened to obtain a second microsatellite site set with lower polymorphism.

[0017] In one implementation, the microsatellite sites in the first site set are located in non-exon regions.

[0018] In one implementation, the preset decision model is obtained using a training set with positive and negative samples, and the preset decision model is:

[0019] Based on the likelihood ratios of the input microsatellite loci, plot the distribution of the likelihood ratios; and

[0020] The threshold is determined based on the intersection position of the main peak and the first distribution peak in the long tail of the distribution pattern.

[0021] In one implementation, the method further includes:

[0022] The preset determination model compares the likelihood ratio of each microsatellite site in the first site set with the threshold, and selects microsatellite sites with a likelihood ratio below the threshold as the second microsatellite site set.

[0023] In one implementation, calculating the unimodal maximum likelihood estimate of the microsatellite locus using a pre-defined unimodal probability distribution model includes:

[0024] For each of the plurality of reads, a unimodal maximum likelihood estimate is calculated for each read based on the stationary probabilities of the three events—missing, normal, and inserted—and their respective occurrence frequencies in each read; and

[0025] The joint maximum likelihood estimate of the multiple reads is obtained as the single-peak maximum likelihood estimate of the microsatellite locus.

[0026] In one implementation, calculating the bimodal maximum likelihood estimate of the microsatellite locus using a pre-defined bimodal probability distribution model includes:

[0027] For each of the plurality of read segments, a bimodal maximum likelihood estimate is calculated for each read segment based on the stationary probabilities of the three events—missing, normal, and inserted—and their respective occurrence frequencies in each read segment. The bimodal maximum likelihood estimate for each read segment is obtained based on a unimodal maximum likelihood estimate of the first peak and a unimodal maximum likelihood estimate of the second peak.

[0028] The joint maximum likelihood estimate of the multiple reads is obtained as the bimodal maximum likelihood estimate of the microsatellite locus.

[0029] In one implementation, the microsatellite site is 15 nucleotides or longer.

[0030] In one implementation, one or more sets of probes are used to capture microsatellite sites in the first site set:

[0031] Probe set 1, wherein the middle region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site;

[0032] Probe set 2, wherein the downstream region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site;

[0033] Probe set 3, wherein the upstream region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site; and

[0034] Probe set 4, wherein the ends of the probe sequences are located at the microsatellite sites, and optionally repeating bases are subtracted from the probe sequences corresponding to the microsatellite sites.

[0035] In one embodiment, the number of repeating bases is 0, 3, 5, 7, 9, or 11.

[0036] In one aspect, the present invention relates to a method for detecting microsatellite instability, the method comprising using a second set of loci obtained by a microsatellite loci screening method of the present invention.

[0037] According to the present invention, a constructed population dataset can be effectively used as a control without the need to provide control samples. When using the microsatellite locus set obtained through the microsatellite locus screening method of the present invention for microsatellite instability detection, it exhibits excellent detection performance, achieving a sensitivity of 98.37% with 100% specificity. Attached Figure Description

[0038] Figure 1 This is a schematic diagram showing the distribution of microsatellite loci reads. Here, "reads" represents the read segments.

[0039] Figure 2 This is a schematic diagram showing missing, normal, and insertion events in microsatellite site synthesis.

[0040] Figure 3 This shows a comparison between the distribution of site data (left column) and the maximum likelihood estimation (MLE) distribution (right column) for log-likelihood ratios around 10, 30, 50 and >100. Here, "reads" represents the reading segments.

[0041] Figure 4 Displays the distribution of likelihood ratios up to 100 and the filtering threshold division.

[0042] Figure 5 This is a schematic diagram of a microsatellite locus probe design.

[0043] Figure 6 This is a distribution diagram of the probe-captured signal. Detailed Implementation

[0044] The following description and examples illustrate embodiments of the present invention in detail. It should be understood that the present invention is not limited to the specific embodiments described herein and therefore can be modified. Those skilled in the art will recognize that many variations and modifications exist in the present invention, all of which are included within its scope.

[0045] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” used herein are also intended to include the plural forms. Furthermore, the open-ended expressions “comprising” and “including” are to be interpreted as potentially containing structural components or method steps not mentioned, but it should be noted that these open-ended expressions also cover situations where the invention consists only of the stated components and method steps (i.e., they cover the closed-ended expressions “consisting of…”).

[0046] As used throughout, a range is used as a shorthand to describe each and all values ​​within that range. Any value within a range, such as an integer value, a value incremented by one-tenth (when the range ends with one decimal place), or a value incremented by one-hundredth (when the range ends with two decimal places), can be chosen as the end of the range. For example, the range 0.1-10 is used to describe all values ​​within that range, such as 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8…9.5, 9.6, 9.7, 9.8, 9.9, and 10 (in increments of one-tenth), and includes all subranges, such as 0.1-1.0, 2.0-3.0, 4.0-5.0, 6.0-7.0, 8.0-9.0, etc.

[0047] All scientific and technical terms mentioned in this specification have the same meaning as commonly understood by those skilled in the art, and in case of conflict, the definitions in this specification shall prevail. To make the description of this invention easier to understand, some terms are explained below.

[0048] The term "next-generation sequencing" used in this article refers to a sequencing technology that uses the principle of "sequencing while synthesis," simultaneously performing parallel sequencing reactions on hundreds of thousands to millions of DNA molecules. The raw image data or electrochemical signals obtained are then analyzed through bioinformatics to ultimately obtain information such as the nucleic acid sequence or copy number of the sample. It is also known as high-throughput sequencing, deep sequencing, next-generation sequencing, and massively parallel sequencing (MPS). The basic procedure of next-generation sequencing involves randomly fragmenting the DNA to be tested into small fragments, constructing libraries through steps such as end repair, ligation of adapter sequences, and PCR, and finally sequencing using sequencers such as Illumina and Ion Torrent.

[0049] The term "capture sequencing" as used in this article refers to a technique that uses biotin-labeled DNA or RNA probes to capture target fragments in a DNA sample and then sequence them.

[0050] The term “microsatellite” as used in this article refers to a string of repeating bases in the human genome, with repeating units ranging in length from 1 to 6 bp.

[0051] The term "microsatellite instability" as used in this article refers to a change in the number of microsatellite repeats, resulting in the appearance of new alleles. Here, microsatellite instability and highly unstable microsatellites are used interchangeably.

[0052] The term “microsatellite stable (MSS)” used in this article refers to the opposite of microsatellite instability, meaning that the number of microsatellite repetitions does not result in changes that lead to the emergence of new alleles.

[0053] The term “missing” as used in this article refers to a read in MS site synthesis that skips the current template position and directly points to the next position.

[0054] As used in this article, the term "normal" or "normal synthesis" refers to the synthesis of a base at the current template position during MS site synthesis, pointing to the next position.

[0055] The term “insertion” as used in this article refers to the synthesis of a base at the current template position during MS site synthesis, while still pointing to the current position.

[0056] In one aspect, the present invention relates to a method for screening microsatellite loci, the method comprising the following steps:

[0057] Obtain the first point set, which includes multiple captured stable microsatellite sites;

[0058] Based on multiple reads of each microsatellite site in the first site set, a unimodal maximum likelihood estimate of the microsatellite site is calculated using a preset unimodal probability distribution model, and a bimodal maximum likelihood estimate of the microsatellite site is calculated using a preset bimodal probability distribution model, wherein the preset bimodal probability distribution model considers more polymorphism than the preset unimodal probability distribution model.

[0059] For each microsatellite locus in the first locus set, obtain the likelihood ratio of the bimodal maximum likelihood estimate to the unimodal maximum likelihood estimate; and

[0060] Based on a preset judgment model including thresholds and the likelihood ratio of each microsatellite site in the first site set, the microsatellite sites in the first site set are screened to obtain a second microsatellite site set with lower polymorphism.

[0061] In one implementation, the microsatellite sites in the first site set are located in non-exon regions.

[0062] In one implementation, the preset decision model is obtained using a training set with positive and negative samples, and the preset decision model is:

[0063] Based on the likelihood ratios of the input microsatellite loci, plot the distribution of the likelihood ratios; and

[0064] The threshold is determined based on the intersection position of the main peak and the first distribution peak in the long tail of the distribution pattern.

[0065] In one implementation, the method further includes:

[0066] The preset determination model compares the likelihood ratio of each microsatellite site in the first site set with the threshold, and selects microsatellite sites with a likelihood ratio below the threshold as the second microsatellite site set.

[0067] In one implementation, calculating the unimodal maximum likelihood estimate of the microsatellite locus using a pre-defined unimodal probability distribution model includes:

[0068] For each of the plurality of reads, a unimodal maximum likelihood estimate is calculated for each read based on the stationary probabilities of the three events—missing, normal, and inserted—and their respective occurrence frequencies in each read; and

[0069] The joint maximum likelihood estimate of the multiple reads is obtained as the single-peak maximum likelihood estimate of the microsatellite locus.

[0070] In one implementation, calculating the bimodal maximum likelihood estimate of the microsatellite locus using a pre-defined bimodal probability distribution model includes:

[0071] For each of the plurality of read segments, a bimodal maximum likelihood estimate is calculated for each read segment based on the stationary probabilities of the three events—missing, normal, and inserted—and their respective occurrence frequencies in each read segment. The bimodal maximum likelihood estimate for each read segment is obtained based on a unimodal maximum likelihood estimate of the first peak and a unimodal maximum likelihood estimate of the second peak.

[0072] The joint maximum likelihood estimate of the multiple reads is obtained as the bimodal maximum likelihood estimate of the microsatellite locus.

[0073] In one implementation, the microsatellite site is 15 nucleotides or longer.

[0074] In one implementation, one or more sets of probes are used to capture microsatellite sites in the first site set:

[0075] Probe set 1, wherein the middle region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site;

[0076] Probe set 2, wherein the downstream region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site;

[0077] Probe set 3, wherein the upstream region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site; and

[0078] Probe set 4, wherein the ends of the probe sequences are located at the microsatellite sites, and optionally repeating bases are subtracted from the probe sequences corresponding to the microsatellite sites.

[0079] In one embodiment, the number of repeating bases is 0, 3, 5, 7, 9, or 11.

[0080] In one aspect, the present invention relates to a method for detecting microsatellite instability, the method comprising using a second set of loci obtained by a microsatellite loci screening method of the present invention.

[0081] While various embodiments of the invention have been described above, it should be understood that they are provided by way of example only and not as limitations. Many changes to the disclosed embodiments may be made in accordance with the disclosure herein without departing from the spirit or scope of the invention. Therefore, the breadth and scope of the invention should not be limited by any of the embodiments described above.

[0082] All references mentioned herein are incorporated herein by reference. All publications and patent documents cited in this application are incorporated herein by reference for all purposes, and are cited as if they were individually cited. Example

[0083] Unless otherwise stated, all materials used in the embodiments herein are commercially available, and all specific experimental methods used to conduct the experiments are conventional experimental methods in the art or are performed according to the steps and conditions recommended by the manufacturer, and can be conventionally determined by those skilled in the art as needed.

[0084] Example 1. Wet test procedure

[0085] Step 1: Create the database.

[0086] 1. Reagents and consumables

[0087] a) KAPA Hyper Library preparation kit (KAPA, KK8504)

[0088] b) Duplex UMI Seq Adapters (IDT)

[0089] c) UDI Primer Mix (IDT)

[0090] d) Ampure XP Beads (BECKMAN COΜLTER,15672100)

[0091] e) 80% ethanol (prepare fresh before use)

[0092] f) Sterilization pipette tip (1000 / 200 / 100 / 10 μl)

[0093] g) Eppendorf low-adsorption EP tubes, PCR tubes

[0094] 2. Specific experimental steps

[0095] a) Preparation of experimental samples: DNA extraction was performed on paraffin-embedded (FFPE) samples using the ReliaPrep FFPE gDNA Miniprep System (Promega, Cat. NO: A2352) kit.

[0096] b) Add adapter: Input 200ng (fragmented) DNA, add Nuclease-free water to 25μl, vortex to mix and briefly centrifuge, perform the balance and A addition reaction (reaction system see Table 1, reaction operation see Table 2), centrifuge the reaction product of the previous step, and then perform the following reaction system (see Table 3 and Table 4).

[0097] Table 1

[0098] Reaction components Volume (μl) DNA 25 End Repair & A Tailing Buffer 3.5 End Repair&A tailing Enzyme Mix 1.5 total 30

[0099] Table 2

[0100] reaction temperature reaction time 20℃ 30min 65℃ 30min 4℃ ∞

[0101] Table 3

[0102] Reaction components Volume (μl) DNA after filling and adding A 30 UDI-Adapter (15uM) 5 Ligation buffer 15 DNA Ligase 5 total 55

[0103] Table 4

[0104] reaction temperature reaction time 20℃ 25min 4℃ ∞

[0105] c) Purification: Remove AMPure Beads from the 4°C freezer, vortex to mix, and incubate at room temperature for 30 minutes; transfer all the reaction product from the previous step to a 1.5 ml centrifuge tube, add 99 μl of vortexed AMPure Beads, vortex to mix, incubate at room temperature for 5 minutes, and then place on a magnetic rack; after the AMPure Beads are completely separated, remove the supernatant, keep the sample on the magnetic rack, add 200 μl of 80% ethanol, remove the supernatant after 1 minute, and then repeat the washing once more; air dry until no droplets remain; add 22 μl of Nuclease-free water, vortex to mix, incubate at room temperature for 5 minutes, and then place on a magnetic rack; after the AMPure Beads are completely separated, transfer 20 μl of supernatant to a new 0.2 ml PCR Tube.

[0106] d) PCR amplification: Prepare the following reaction mixture in 0.2 ml of PCR Tube (see Tables 5 and 6).

[0107] Table 5

[0108] Reaction components Volume (μl) Purified DNA 20 KAPA HiFi HotStart ReadyMix 25 P5 / P7 Primer Mix (10µM) 5 total 50

[0109] Table 6

[0110]

[0111] e) PCR product recovery: Transfer all the reaction products from the previous step to a 1.5 ml centrifuge tube, add 60 μl of vortexed AMPure Beads (1.2x), vortex to mix, incubate at room temperature for 5 min, then place on a magnetic rack. After the AMPure Beads are completely separated, discard the supernatant. Keep the sample on the magnetic rack, add 200 μl of 80% ethanol, discard the supernatant after 30 s, and repeat the washing once more. Add 31 μl of Nuclease-free water, vortex or adjust the 200 μl pipette to 25 μl, pipette 15 times to fully suspend the AMPure Beads, incubate at room temperature for 5 min, then place on a magnetic rack. After the AMPure Beads are completely separated, transfer 50 μl of supernatant to a 1.5 ml centrifuge tube. Label the sample name, library type, and library construction date. Proceed directly to the DNA library quality control procedure or temporarily store at 4℃.

[0112] Step 2 Capture:

[0113] 1. Experimental consumables and reagents

[0114] a) Human Cot-1 DNA (Invitrogen, 15279-011)

[0115] b) Blocking Oligo (IDT,1046637)

[0116] c) Probe library: working solution concentration is 1.5 μM (IDT)

[0117] d) Xgen CNV Backbone panel (IDT, 1080564 / 1080563)

[0118] e) Xgen hybridization buffer (IDT,1072281)

[0119] f) Ampure XP Beads, Dynabeads® M270 Streptavidin C1, 80% ethanol (prepare fresh for use), sterilization pipette tip (1000 / 200 / 10 μl).

[0120] 2. Specific experimental steps

[0121] a) Hybridization: In a 1.5 ml LoBind centrifuge tube, four DNA libraries were pooled together, with each library containing 500 ng of DNA. 5 μl of Human Cot-1 DNA and 2 μl of 1 / 2 Blocking Oligo TS MIX were added to each pool. After thorough mixing, the mixture was centrifuged to collect the mixture at the bottom of the tube. (The centrifuge tube was then sealed with sealing film, and five small holes were punched in the film.) The mixture was then vacuum-dried. 8.5 μl of xGen 2X Hybridization Buffer, 2.7 μl of xGen Hybridization Buffer Enhancer, and 1.8 μl of Nuclease-free water were added. After mixing by pipetting and centrifugation, the mixture was collected at the bottom of the tube and allowed to stand for 10 min to reconstitute. The 13 μl mixture obtained in the previous step was transferred to a 0.2 ml... Place the sample in a PCR tube and denature at 95°C for 10 min. Add 4 μl of probe and vortex for about 3 seconds to mix thoroughly (final volume 17 μl). Then incubate at 65°C overnight (PCR instrument hot cap temperature set to 75°C).

[0122] b) Enrichment: Prepare the buffer according to Table 7. Take 1x Wash Buffer 1 and all of 1x Stringent Wash Buffer at a rate of 100 μl / pool and incubate on a 65°C metal bath for at least 2 hours. Remove Dynabeads® M270 Streptavidin C1 from the 4°C freezer and allow it to equilibrate to room temperature for 30 min. Vortex the M270 until homogeneous, then take 100 μl / pool and place it on a magnetic rack. After the beads have completely separated, remove the supernatant. Add 1x BeadsWash Buffer at a rate of 200 μl / pool, vortex for 10 sec, and place it on a magnetic rack. After the beads have completely separated, remove the supernatant. Repeat the previous step once, then add 1x Beads Wash Buffer at a rate of 100 μl / pool. After vortexing the buffer, aliquot the mixture into 0.2ml PCR tubes. After 4 hours of hybridization, place the aliquoted M270 on a magnetic rack until the beads are completely separated, then remove the supernatant. Transfer 17μl of the hybridization mixture to the supernatant-free M270, pipette five times, then vortex for 3 seconds to fully suspend the beads. Return the tube to the PCR instrument and incubate for another 45 minutes, resuspending the beads every 12 minutes by vortexing for approximately 3 seconds. After incubation, add 100μl of preheated 1x Wash Buffer I, vortex for approximately 3 seconds, then transfer to a 1.5ml LoBind centrifuge tube. Vortex for approximately 3 seconds, then briefly centrifuge and immediately place the tube on a magnetic rack. After approximately 20 seconds, the beads are completely separated, and the supernatant is immediately removed. Add 200μl of preheated 1x Stringent Wash Buffer I. After 10 aspirations and 10 whistles, briefly centrifuge and immediately incubate on a 68°C metal bath for 5 min. Then place on a magnetic rack. After about 20 s, the beads will be completely separated. Immediately remove the supernatant. Repeat the previous step once. Add 200 μl of 1x Wash Buffer I, vortex for 2 min, briefly centrifuge, and then place on a magnetic rack. After the beads are completely separated, remove the supernatant. Add 200 μl of 1x Wash Buffer II, vortex for 1 min, briefly centrifuge, and then place on a magnetic rack. After the beads are completely separated, remove the supernatant. Add 200 μl of 1x Wash Buffer III, vortex for 30 s, briefly centrifuge, and then place on a magnetic rack. After the beads are completely separated, remove the supernatant.

[0123] Table 7

[0124] Concentrated buffer solution Concentrated buffer (μL) Nuclease-free water (μL) xGen 2X Bead Wash Buffer 250 250 xGen 10X Wash Buffer I* 30 270 xGen 10X Wash Buffer II 20 180 xGen 10X Wash Buffer III 20 180 xGen 10X Stringent Wash Buffer 40 360

[0125] c) PCR amplification: Resuspend the beads in 18 μl of Nuclease-free water, then transfer them to a 0.2 ml PCR tube and prepare the following reaction mix (see Tables 8 and 9).

[0126] Table 8

[0127] Reaction components Volume (μl) Heavy-suspension beads 20 KAPA HiFi HotStart ReadyMix 25 P7 primer 2.5 P5 primer 2.5 total 50

[0128] Table 9

[0129]

[0130] d) PCR product recovery: Remove AMPure Beads from the 4°C freezer and allow to equilibrate at room temperature for 30 min. Add 75 μl of AMPure Beads (1.5x) and mix thoroughly. After incubating at room temperature for 10 min, place the mixture on a magnetic rack. Once the Beads are completely separated, discard the supernatant. Keep the sample on the magnetic rack, add 400 μl of 80% ethanol, and discard the supernatant after 30 s. Repeat the washing process once more. Allow the sample to air dry until no droplets remain and the first crack appears on the Beads. Add 21.6 μl of Nuclease-free water, vortex or pipette 15 times to fully suspend the Beads, incubate at room temperature for 5 min, and then place the mixture on a magnetic rack. Once the Beads are completely separated, transfer 20 μl of the supernatant to a 1.5 ml LoBind centrifuge tube. Label the sample with the sample name, library type, and library construction date.

[0131] Step 3: Sequencing

[0132] 1. Sequencing platform: Novaseq 6000

[0133] 2. Sequencing length: PE100

[0134] Example 2. Probe Design Performance Demonstration

[0135] For each MS site, we designed 4 sets of probes, and the specific design scheme is as follows (see Figure 5 ):

[0136] a) The probe characteristics in probe group 1 are: a. The probe spans the MS site in the middle (the probe length is 120bp, that is, a 60bp region spans the MS site), b. For the MS sequence on the probe, we will subtract some repeating bases (subtracting lengths of 0, 3, 5, 7, 9, and 11 respectively).

[0137] b) The probe characteristics in probe group 2 are: a. The probe crosses the MS site in the middle and downstream (the probe length is 120bp, that is, the 80bp region crosses the MS site), b. For the MS sequence on the probe, we will remove some repeating bases (removing lengths of 0, 3, 5, 7, 9, and 11 respectively).

[0138] c) The probe characteristics in probe group 3 are: a. The probe crosses the MS site slightly upstream in the middle (the probe length is 120bp, that is, a 40bp region crosses the MS site), b. For the MS sequence on the probe, we will remove some repeating bases (removing lengths of 0, 3, 5, 7, 9, and 11 respectively).

[0139] d) The probe characteristics in probe group 4 are: a. The probe ends are located on the MS, b. For the MS sequence on the probe, we will subtract some repeating bases (subtracting lengths of 0, 3, 5, 7, 9, and 11 respectively).

[0140] In addition, when designing probe sequences, for MS sites with T and G bases on the sense strand, we use the negative strand sequence to design probes. This ensures that the designed probe sequences only contain A and C repeat sequences, which can effectively prevent hybridization between different probes due to complementary pairing of AT and CG.

[0141] To analyze probe capture performance, we used both the probes designed in this invention and a three-layer planar design probe near the site for library capture. The experimental method was similar to that described above. After obtaining tissue samples, DNA was first extracted using commercially available kits. For tissue-derived DNA, it was fragmented using sonication, DNA restriction enzyme digestion, or transposase technology to form double-stranded DNA fragments of approximately 200 bases in length. Then, libraries were constructed using corresponding kits for DNA from different sources. After library construction, libraries from one to four samples were mixed, and genomic fragments were captured using two different capture probes. The captured libraries were finally sequenced using a high-throughput sequencer.

[0142] This section uses sample tMSIbz-166 as an example to demonstrate the analysis results. Figure 6 As shown in the figure, each subplot represents one MS site. The main peak in each subplot represents the wild-type MS length, and the subpeaks on the left represent the MS positive signal. The figure shows that the designed probe has a significantly higher MS positive signal in the MS positive signal region than conventionally designed probes.

[0143] Example 3. Screening of MS sites

[0144] 3.1 Initial screening of MS loci

[0145] Based on whole exome sequencing (WES) data, a set of stable capture sites was selected, restricting these MS sites to non-exon regions. This is primarily because MSI in exon regions is more likely to cause phenotypic changes and cell death, thus exhibiting more mechanisms to maintain its conservation, leading to signal distortion and affecting the detection of positive signals. A total of 500 sites were selected here, constituting site set 1 (all sites in Appendix 1).

[0146] 3.2 Screening for MS site polymorphisms

[0147] MS locus polymorphism is relatively common in the population, which makes it difficult to distinguish between MS negative and positive results, and even poses a risk of introducing false positives. Therefore, selecting relatively conserved MS loci is particularly important. This step focuses on screening polymorphic MS loci. The general approach is as follows: Step 1, we assume that the locus has two states: 1) conserved mode, i.e., unimodal; 2) polymorphic mode, i.e., bimodal. Step 2, calculate the maximum likelihood value for the conserved mode and the polymorphic mode respectively, and calculate the ratio of their maximum likelihood values ​​(here, the polymorphic mode is the numerator, and the conserved mode likelihood value is the denominator). Theoretically, the larger the ratio, the greater the probability of the polymorphic mode. Step 3, plot the distribution of the likelihood ratio based on the initially selected 500 loci, determine the polymorphism judgment threshold, and finally determine the MS loci that need to be removed. The specific steps are as follows:

[0148] 3.2.1 Calculation of the single-peaked mode likelihood value of MS sites

[0149] a) First, a probabilistic model needs to be established to determine the changes in read length at MS sites. We assume that reads can be deleted or inserted during synthesis or sequencing. Although the occurrence of deletion, normal, and insertion events should be influenced by preceding events, there should be a stationary probability P = {p1, p2, p3} for the entire system, which can describe the proportion of the three events.

[0150] b) Subsequently, we establish the measured length X = {x1, x2, ..., x...} i ,…x n The relationship between x and the probabilities P of the three events, i Let be the measured length, and 'n' represent the number of observations at that site. For a template of length N, generate a template of length x. i The sequence of events can have more than one possible combination, but should satisfy the following relationship:

[0151] 1. The last event must be either missing or normal.

[0152] 2. If the last event is normal, then the normal event can take values ​​from 0 to min ((N-1), (xi -1), the sum of the number of missing events and the number of normal events equals N-1, and the sum of the number of normal events and the number of inserted events equals x. i -1.

[0153] 3. If the last event is missing, then the normal event can take values ​​from 0 to min ((N-1), x i The sum of the number of missing events and the number of normal events equals N-1, and the sum of the number of normal events and the number of inserted events equals x. i .

[0154] Based on the above rules, we can iterate through all event combinations that satisfy this condition. Then, we calculate the probability of each event combination occurring:

[0155] For the three events—missing, normal, and inserted—the probability P = {p1, p2, p3}. If the last event is normal, the occurrence counts of any event combination j are {j1, j2, j3}, and the probability of combination j occurring is j1 + j2 + j3. If the last event is missing, the occurrence counts of any event combination k are {k1, k2, k3}, and the probability of combination k occurring is k1 + k2 + k3. Therefore, the likelihood function for xi is:

[0156]

[0157] Here is the calculation θ The parameters are N and P = {p1, p2, p3}.

[0158] in Let p2 be the probability of a normal event occurring. Assuming the last synthesis is normal, the probability of event combination j is:

[0159]

[0160] in, Let p1 be the probability of the missing event occurring. Given that the last synthesis is missing, the probability of event combination k.

[0161]

[0162] Given only the probability of xi being recorded once, the joint probability L1 of that site is:

[0163]

[0164] It can be obtained through maximum likelihood estimation. θ The estimates of (i.e., N and P) and the corresponding maximum likelihood values. L1。

[0165] 3.2.2 Calculation of the bimodal mode likelihood value of MS loci

[0166] A bimodal pattern consists of two single peaks. θ 1 The parameters representing peak 1 (i.e., N1 and P1) θ 2 The parameters representing peak 2 (i.e., N2 and P2) are given. α represents the proportion of peak 1 in the doublet, and 1-α represents the proportion of peak 2. Then, for the observed value x... i The likelihood function has the following form:

[0167]

[0168] Where parameters Includes proportionality coefficient and the parameters of the two peaks and , For peak 1 and Correspondingly For peak 2 and .

[0169] The joint probability of this site L2 for:

[0170]

[0171] It can be obtained through maximum likelihood estimation. θ' The estimation of (i.e., N1, N2, P1, P2) and the corresponding L2 value.

[0172] 3.2.3 Calculation of maximum likelihood ratio and determination of threshold

[0173] Based on sections 3.2.1 and 3.2.2 above, the likelihood values ​​L1 and L2 for unimodal and bimodal patterns can be obtained, from which the maximum likelihood ratio (L2 / L1) can be calculated. A distribution plot of the likelihood ratio is shown below. Figure 4 The likelihood ratios of each site generally exhibit a log-normal distribution (referred to as the main peak), but in the tail region, other distributions are mixed, forming a long tail that can extend to the hundreds. The intersection value of the first distribution peak in the long tail with the main peak, 42, is taken as the filtering threshold. That is, a likelihood ratio ≥ 42 is considered a polymorphic site, and a likelihood ratio < 42 is considered a non-polymorphic site.

[0174] To examine the reasonableness of this threshold, four MS loci were randomly selected for polymorphism review. The likelihood ratios of these four loci were 10, 30, 50, and 120, respectively. Figure 3 It can be seen that the more the original data exhibits multiple peak separations, the higher the likelihood ratio value, and the distributions with likelihood ratios of 50 and 120 are clearly polymorphic. The observations are consistent with expectations.

[0175] Ultimately, considering that excessively short sites might lead to experimental bias, the length of the fitted MS sites was set to be greater than or equal to 15 nucleotides. A total of 100 sites were actually screened, which constitutes site set 2 (the first 100 sites in Appendix 1).

[0176] Example 4. Establishing a mathematical model to distinguish between MSS and MSI-H

[0177] 1. Preprocessing of sequencing data

[0178] The sequencing files (BCL format) were converted to sequence files (FASTQ format) using bcl2fastq v2.19.0 software. Then, the sequence files were quality controlled (QC) and filtered using fastp v0.20.0 software to remove low-quality sequences. The filtered sequences were then aligned to the reference genome hg19 using bwa v0.7.12 software to generate alignment files (BAM format). Finally, the alignment files were sorted and deduplicated using picard v1.8.0_221 software.

[0179] 2. The modeling steps are as follows:

[0180] 1) Select a PCR-validated positive set (30 positive samples) and a negative set (60 negative samples), and perform sequencing with an average sequencing depth of 2000X after deduplication as the target depth to ensure sufficient site coverage. For each site to be detected in MS site set 2, calculate the cumulative distribution of the repeat region length of that site;

[0181] 2) For each locus in the positive and negative sets, find the length of the repeating region corresponding to the maximum difference in the mean of the cumulative distribution of samples in the two datasets (define this length as the cutpoint, i.e., the threshold for determining whether a specific read supports MSI or MSS). The cumulative probability of the repeating region length corresponding to the cutpoint in the baseline distribution is the p-value of that locus. i .

[0182] 5) For the site to be detected, it is classified according to the relationship between the length of the repeat region of the read and the size of the cut point. Reads with a length less than or equal to the length of the cut point are defined as MSI reads, and reads with a length greater than the length of the cut point are defined as MSS reads.

[0183] 6) We assume that site (i) has N i For each read segment, the baseline probability at that location is p. i The dividing point is C. i Based on the segmentation point, we find that the number of MSI reads at this site is n. i If there are 1000 segments, then the number of MSS segments read is N. i -n i The following formula gives us the probability P under this condition:

[0184]

[0185] According to this formula, the cumulative probability P(X≥n) of detecting an MS read at this locus being greater than or equal to the current value ni can be calculated. i If the probability is less than 0.001, we consider this point to significantly support MSI-H.

[0186] 7) Finally, the significance status is calculated by combining 100 sites, that is, the ratio of the number of sites supporting MSI-H to the total number of sites; the cutoff value is obtained according to the MSI-score distribution of the samples in the training set (i.e. the aforementioned positive set and negative set). The cutoff value is 0.355. A value higher than 0.355 is MSI-H, and a value lower than 0.355 is MSS.

[0187] Example 5. Algorithm Performance Verification

[0188] The algorithm was validated using MSI-PCR as the gold standard. 153 FFPE samples from gastric cancer were selected, including 123 positive samples and 30 negative samples.

[0189] Final validation results: sensitivity was 98.37%, and specificity was 100% (Table 10). In the table, PCR refers to the MSI-PCR method, and NGS refers to the method according to the present invention. Furthermore, the detection results of the example samples are shown in Appendix 2 below.

[0190] Table 10

[0191]

[0192] Here, we also use the industry-standard MSI sensor as a comparison method. The MSI sensor's sensitivity is only 94.3%.

[0193] Appendix 1: The set of sites involved in this invention, wherein the first 100 are selected sites.

[0194] chromosome Starting position End position chr1 150595354 150595368 chr1 216373476 216373488 chr2 55559651 55559666 chr3 97605460 97605471 chr3 172052897 172052913 chr8 101540233 101540255 chr9 134744778 134744794 chr11 18596966 18596978 chr12 72067966 72067979 chr15 65312614 65312628 chr15 77642952 77642968 chr19 18583696 18583719 chr20 33033065 33033082 chr1 223156406 223156418 chr1 224641959 224641976 chr2 61577667 61577679 chr3 10128965 10128984 chr3 40352355 40352371 chr3 42671680 42671691 chr3 138216173 138216188 chr7 99697041 99697059 chr8 17422434 17422445 chr8 17503684 17503698 chr8 124232550 124232561 chr10 13726871 13726886 chr10 61956386 61956398 chr15 30010328 30010348 chr1 14109408 14109424 chr3 71739332 71739344 chr3 138193025 138193037 chr3 194896624 194896635 chr7 69583231 69583253 chr7 151433469 151433488 chr8 121013744 121013755 chr9 114341239 114341254 chr9 127287158 127287170 chr11 46569434 46569449 chr12 91372042 91372058 chr16 50609006 50609020 chr21 37734540 37734558 chr3 140678384 140678399 chr7 6230059 6230077 chr7 107167661 107167676 chr12 4874518 4874542 chr12 48360947 48360960 chr12 100635627 100635643 chr2 66664004 66664020 chr7 150327234 150327256 chr9 97555228 97555241 chr9 134014642 134014665 chr12 124242458 124242471 chr15 73428204 73428224 chr16 11862376 11862390 chr1 222904869 222904880 chr9 8341280 8341292 chr16 67805993 67806013 chr19 33464253 33464265 chr1 183481943 183481963 chr9 13206106 13206120 chr7 95775848 95775862 chr7 100319680 100319695 chr15 40005793 40005805 chr15 80460576 80460590 chr1 244803205 244803219 chr3 169525509 169525531 chr8 89128708 89128721 chr10 124810704 124810728 chr12 27841925 27841937 chr8 87751848 87751866 chr16 10783088 10783101 chr1 200080464 200080479 chr3 170801912 170801931 chr1 94964139 94964154 chr1 161721402 161721413 chr3 58368219 58368236 chr7 80456708 80456731 chr7 101747593 101747604 chr7 131163354 131163375 chr12 104461711 104461722 chr20 34099165 34099179 chr1 60338452 60338463 chr1 63998326 63998337 chr1 85120991 85121003 chr1 151733379 151733402 chr1 185269100 185269111 chr2 47702451 47702470 chr3 105238890 105238901 chr10 13538928 13538945 chr11 27414071 27414086 chr15 35174666 35174685 chr15 45725117 45725129 chr15 101550861 101550876 chr1 116940481 116940496 chr7 103202406 103202425 chr9 140637804 140637820 chr12 7294657 7294673 chr12 129566284 129566301 chr15 74513557 74513584 chr2 70456452 70456464 chr12 107237614 107237625 chr1 167871297 167871311 chr2 27597190 27597203 chr10 61840385 61840396 chr15 42503953 42503964 chr16 85105371 85105386 chr9 135773000 135773018 chr1 10177680 10177692 chr9 96259883 96259894 chr12 12247921 12247937 chr10 30333754 30333768 chr12 56966840 56966856 chr1 225586737 225586751 chr2 61009788 61009813 chr3 32571823 32571840 chr10 121675238 121675250 chr15 93007302 93007325 chr11 18111872 18111888 chr16 3808052 3808065 chr8 104075165 104075176 chr2 64210979 64210990 chr3 101370529 101370547 chr3 119154523 119154540 chr12 40258687 40258698 chr19 327260 327276 chr3 184642587 184642600 chr12 31633202 31633213 chr20 2552802 2552816 chr3 37550027 37550039 chr15 41991036 41991051 chr7 103206001 103206012 chr12 4645161 4645172 chr15 63063141 63063157 chr2 64112706 64112718 chr3 44815871 44815882 chr9 26906079 26906090 chr9 86354663 86354678 chr21 34903875 34903895 chr11 30900263 30900279 chr12 49397291 49397301 chr16 71808507 71808518 chr15 40942715 40942729 chr1 103468855 103468867 chr11 5687320 5687337 chr7 138417936 138417948 chr12 73012649 73012662 chr3 140675348 140675366 chr8 28651296 28651311 chr1 169586220 169586232 chr3 123086698 123086709 chr3 127820493 127820504 chr10 118891992 118892004 chr1 220369745 220369757 chr7 111474749 111474762 chr9 123886190 123886201 chr1 109525421 109525433 chr1 211280724 211280738 chr20 35445890 35445906 chr1 70694681 70694693 chr1 225707271 225707287 chr2 27466258 27466279 chr16 56532618 56532630 chr8 62550924 62550937 chr12 68690765 68690776 chr14 24837198 24837211 chr1 233167531 233167549 chr9 13193320 13193331 chr1 212142054 212142068 chr1 218610662 218610679 chr2 64139831 64139843 chr3 25792762 25792773 chr10 128780161 128780171 chr10 70181968 70181979 chr10 124810746 124810756 chr19 47153145 47153161 chr1 75042626 75042640 chr3 180320897 180320909 chr10 29169046 29169059 chr19 37721404 37721418 chr1 114940632 114940644 chr9 107556793 107556811 chr1 236590675 236590686 chr12 88926290 88926301 chr3 148583231 148583242 chr3 100570787 100570801 chr16 70595514 70595524 chr12 22063251 22063265 chr19 34832292 34832304 chr21 47720999 47721011 chr10 71664661 71664678 chr15 51768755 51768767 chr1 86590509 86590522 chr16 46652286 46652297 chr9 432140 432154 chr12 21327652 21327666 chr1 230798864 230798875 chr2 45806963 45806978 chr9 86293539 86293549 chr9 128064242 128064259 chr1 212905370 212905388 chr2 56611386 56611397 chr7 16725501 16725520 chr12 80648803 80648814 chr9 72000679 72000703 chr12 45771901 45771917 chr1 53267474 53267487 chr1 82452565 82452577 chr1 115226990 115227002 chr3 171065010 171065025 chr7 86522420 86522431 chr7 105122891 105122907 chr19 6908675 6908688 chr7 133933671 133933686 chr11 57582847 57582859 chr3 57632004 57632014 chr10 75576885 75576895 chr12 86407543 86407557 chr7 91714092 91714105 chr10 50031212 50031224 chr10 123686842 123686853 chr16 69965396 69965408 chr2 9552383 9552395 chr3 54420721 54420739 chr9 85926893 85926907 chr10 29162164 29162175 chr11 16117684 16117697 chr1 101194965 101194976 chr19 3543478 3543489 chr1 70742532 70742544 chr8 39114676 39114689 chr1 223984306 223984316 chr3 47860685 47860695 chr3 42668803 42668814 chr1 234565078 234565089 chr10 105990647 105990659 chr12 15776223 15776236 chr12 22215194 22215211 chr7 18668945 18668958 chr15 49935503 49935522 chr3 113442784 113442794 chr3 172536641 172536651 chr7 121699789 121699801 chr16 4927799 4927814 chr3 108081175 108081186 chr15 42714195 42714206 chr10 93752074 93752089 chr12 97303682 97303694 chr11 48142528 48142544 chr3 37360696 37360707 chr16 9010826 9010843 chr1 198711170 198711182 chr2 95816017 95816036 chr15 50581869 50581881 chr1 151134579 151134589 chr12 550911 550928 chr3 150276153 150276163 chr10 114293328 114293340 chr12 26628357 26628368 chr12 54920288 54920300 chr1 237619875 237619889 chr19 36564438 36564448 chr12 107002558 107002572 chr2 20130970 20130983 chr3 148858987 148859000 chr11 67209167 67209177 chr1 212969842 212969855 chr10 76349020 76349032 chr10 17204246 17204261 chr2 15319070 15319083 chr1 172512179 172512191 chr9 115060090 115060106 chr10 103771457 103771468 chr15 90984724 90984735 chr19 49815869 49815880 chr3 154002357 154002369 chr15 57545438 57545448 chr10 13152504 13152516 chr7 138453892 138453908 chr1 237819090 237819101 chr1 36298019 36298032 chr12 117402691 117402701 chr15 32393486 32393501 chr3 185654464 185654476 chr1 114186351 114186361 chr10 7270338 7270352 chr3 47604270 47604280 chr1 19476307 19476319 chr2 97357756 97357767 chr2 39964206 39964216 chr3 32746288 32746300 chr7 111926896 111926915 chr2 74076622 74076632 chr16 69492973 69492988 chr10 21997501 21997511 chr11 10820663 10820680 chr15 33923502 33923512 chr3 47857429 47857444 chr7 121942442 121942452 chr8 51314749 51314760 chr1 158637865 158637881 chr12 70329858 70329887 chr7 24329103 24329113 chr2 11913727 11913738 chr8 68334940 68334951 chr15 102185368 102185380 chr1 35826773 35826783 chr2 43971263 43971273 chr12 50754721 50754731 chr1 212194560 212194570 chr2 44934524 44934534 chr13 38171301 38171311 chr1 201773566 201773577 chr16 58585178 58585190 chr3 10429902 10429912 chr7 98648517 98648527 chrX 107827760 107827772 chr1 11026381 11026394 chr1 74818936 74818946 chr1 215914883 215914895 chr1 36019913 36019924 chr2 47641559 47641586 chr8 121293142 121293157 chr21 27855731 27855747 chr1 103356072 103356082 chr13 22140911 22140923 chr19 30927379 30927391 chr15 50873112 50873122 chr15 58891949 58891961 chr19 18857939 18857949 chr9 71691133 71691147 chr1 235993742 235993752 chr7 76959890 76959910 chr3 100378512 100378522 chr1 95354197 95354209 chr8 113323406 113323416 chr1 200143064 200143083 chr3 186295417 186295430 chr10 96025381 96025391 chr3 179333753 179333763 chr8 3443799 3443810 chr15 77695169 77695179 chr1 237038008 237038022 chr2 61319728 61319739 chr3 184552422 184552433 chr7 87323209 87323219 chr3 57439139 57439149 chr10 114886321 114886331 chr12 70990134 70990146 chr1 43664319 43664329 chr15 37328985 37328995 chr1 235301496 235301509 chr21 47821449 47821459 chr3 15839065 15839076 chr2 63182611 63182621 chr10 123996879 123996889 chr7 89933721 89933732 chr8 95160949 95160965 chr7 146805220 146805231 chr1 109349996 109350008 chr1 57349140 57349150 chr10 21958032 21958042 chr10 121678941 121678951 chr3 143551065 143551075 chr3 179448066 179448077 chr1 46165875 46165885 chr1 52912122 52912133 chr7 89982116 89982126 chr3 142259706 142259720 chr11 46638381 46638412 chr1 213341188 213341198 chr21 48064203 48064213 chr3 124998098 124998109 chr20 21143813 21143823 chr9 73458045 73458055 chr12 101689236 101689246 chr3 17418001 17418011 chr1 173797451 173797461 chr1 55573112 55573122 chr13 76134860 76134870 chr3 184100958 184100968 chr10 120810867 120810877 chr10 112044559 112044571 chr12 132496182 132496192 chr10 121335340 121335350 chr12 7170167 7170177 chr1 52897140 52897150 chr16 23117810 23117820 chr1 25571794 25571804 chr1 54502395 54502405 chr1 103463928 103463938 chr10 93788516 93788530 chr7 87005287 87005298 chr3 121151270 121151280 chr10 12708721 12708731 chr1 163038838 163038848 chr9 111903890 111903900 chr8 103846501 103846511 chr1 207075911 207075923 chr3 42260001 42260011 chr9 35342294 35342305 chr12 49581971 49581981 chr7 138333933 138333944 chr20 33587553 33587564 chr1 86241399 86241409 chr11 5013169 5013179 chr10 81053094 81053104 chr10 115925727 115925737 chr9 15727849 15727859 chr8 38959363 38959373 chr16 58313578 58313590 chr21 37626051 37626061 chr1 101440343 101440353 chr1 227171735 227171745 chr1 173545907 173545917 chr3 78689054 78689064 chr21 47647618 47647631 chr15 93492374 93492384 chr9 20948224 20948234 chr3 71021128 71021138 chr8 18413893 18413903 chr10 127789605 127789615 chr19 48533847 48533858 chr7 47546547 47546558 chr12 116549320 116549330 chr16 5517197 5517209 chr12 96408600 96408610 chr19 16472811 16472821 chr9 740769 740787 chr11 47310608 47310618 chr12 112117149 112117159 chr21 43549808 43549819 chr12 46758406 46758416 chr9 91970505 91970516 chr11 19167707 19167717 chr19 33289195 33289205 chr1 154426946 154426956 chr15 99439963 99439973 chr21 43974156 43974166 chr8 52284654 52284664 chr10 93695545 93695555 chr9 33338476 33338486 chr16 7759449 7759459 chr10 93748895 93748905 chr1 16200729 16200739 chr2 73635899 73635910 chr11 62488917 62488928 chr21 27937167 27937177 chr15 70992010 70992020 chr1 117633133 117633143 chr15 31251333 31251343 chr16 81969515 81969525 chr15 45898575 45898585 chr3 3194107 3194117 chr1 151665562 151665572 chr16 15892481 15892491 chr1 52927295 52927305 chr15 35664406 35664416 chr1 245977725 245977735 chr7 120764548 120764558 chr3 1269483 1269493 chr15 70987450 70987460 chr2 24930798 24930808 chr16 50267413 50267423 chr21 37591843 37591853 chr16 81062151 81062165 chr9 8484397 8484407 chr1 180803992 180804002 chr1 179961368 179961378 chr7 87329899 87329909 chr1 210414879 210414893 chr1 35486072 35486083 chr7 21747481 21747491 chr1 245222673 245222683 chr1 222753050 222753060 chr12 59307848 59307859 chr11 19197349 19197360 chr3 29739872 29739882 chr15 30033645 30033655 chr3 100992404 100992414 chr11 5664355 5664366 chr1 118495131 118495141 chr3 100583651 100583662 chr9 126202594 126202604 chr3 190345098 190345108 chr15 52656722 52656732 chr1 233225957 233225968 chr15 72260473 72260483 chr9 134168160 134168170 chr2 46583272 46583282 chr3 4878419 4878443 chr12 96927778 96927788 chr12 49581379 49581389 chr3 132196805 132196815 chr16 53493346 53493356 chr9 73164596 73164609 chr1 985443 985458 chr3 149686355 149686367 chr15 86278759 86278770 chr9 6254555 6254567 chr10 32742257 32742267 chr11 1481691 1481708

[0195] Appendix 2: Detection Results of Samples from Examples

[0196] Sample number Sample ID NGS-based MSI-score NGS result PCR result 1 tMSIbz-001 100 MSI-H MSI-H 3 tMSIbz-003 100 MSI-H MSI-H 4 tMSIbz-004 98 MSI-H MSI-H 5 tMSIbz-005 99 MSI-H MSI-H 7 tMSIbz-007 100 MSI-H MSI-H 8 tMSIbz-008 100 MSI-H MSI-H 9 tMSIbz-009 87 MSI-H MSI-H 10 tMSIbz-010 85 MSI-H MSI-H 11 tMSIbz-011 2 MSS MSS 12 tMSIbz-012 100 MSI-H MSI-H 13 tMSIbz-013 99 MSI-H MSI-H 14 tMSIbz-014 69 MSI-H MSI-H 15 tMSIbz-015 1 MSS MSS 16 tMSIbz-016 2 MSS MSI-L 17 tMSIbz-017 100 MSI-H MSI-H 18 tMSIbz-018 98 MSI-H MSI-H 19 tMSIbz-019 97 MSI-H MSI-H 20 tMSIbz-020 87 MSI-H MSI-H 21 tMSIbz-021 93 MSI-H MSI-H 23 tMSIbz-023 72 MSI-H MSI-H 24 tMSIbz-024 99 MSI-H MSI-H 25 tMSIbz-025 69 MSI-H MSI-H 26 tMSIbz-026 99 MSI-H MSI-H 27 tMSIbz-027 24 MSS MSI-H 28 tMSIbz-028 98 MSI-H MSI-H 29 tMSIbz-029 91 MSI-H MSI-H 30 tMSIbz-030 98 MSI-H MSI-H 31 tMSIbz-031 62 MSI-H MSI-H 32 tMSIbz-032 0 MSS MSS 33 tMSIbz-033 99 MSI-H MSI-H 34 tMSIbz-034 94 MSI-H MSI-H 35 tMSIbz-035 25 MSS MSI-L 36 tMSIbz-036 52 MSI-H MSI-H 37 tMSIbz-037 71 MSI-H MSI-H 38 tMSIbz-038 50 MSI-H MSI-H 39 tMSIbz-039 94 MSI-H MSI-H 40 tMSIbz-040 100 MSI-H MSI-H 41 tMSIbz-041 99 MSI-H MSI-H 42 tMSIbz-042 94 MSI-H MSI-H 43 tMSIbz-043 63 MSI-H MSI-H 45 tMSIbz-045 1 MSS MSS 46 tMSIbz-046 0 MSS MSS 47 tMSIbz-047 99 MSI-H MSI-H 48 tMSIbz-048 99 MSI-H MSI-H 49 tMSIbz-049 72 MSI-H MSI-H 50 tMSIbz-050 97 MSI-H MSI-H 51 tMSIbz-051 96 MSI-H MSI-H 53 tMSIbz-053 100 MSI-H MSI-H 54 tMSIbz-054 75 MSI-H MSI-H 55 tMSIbz-055 98 MSI-H MSI-H 56 tMSIbz-056 94 MSI-H MSI-H 57 tMSIbz-057 97 MSI-H MSI-H 58 tMSIbz-058 88 MSI-H MSI-H 60 tMSIbz-060 96 MSI-H MSI-H 61 tMSIbz-061 97 MSI-H MSI-H 62 tMSIbz-062 77 MSI-H MSI-H 63 tMSIbz-063 0 MSS MSS 64 tMSIbz-064 100 MSI-H MSI-H 65 tMSIbz-065 97 MSI-H MSI-H 66 tMSIbz-066 98 MSI-H MSI-H 67 tMSIbz-067 100 MSI-H MSI-H 68 tMSIbz-068 88 MSI-H MSI-H 69 tMSIbz-069 100 MSI-H MSI-H 70 tMSIbz-070 96 MSI-H MSI-H 71 tMSIbz-071 99 MSI-H MSI-H 72 tMSIbz-072 46 MSI-H MSI-H 73 tMSIbz-073 66 MSI-H MSI-H 74 tMSIbz-074 93 MSI-H MSI-H 75 tMSIbz-075 78 MSI-H MSI-H 76 tMSIbz-076 92 MSI-H MSI-H 77 tMSIbz-077 0 MSS MSS 78 tMSIbz-078 5 MSS MSS 79 tMSIbz-079 63 MSI-H MSI-H 80 tMSIbz-080 1 MSS MSS 81 tMSIbz-081 0 MSS MSS 83 tMSIbz-083 97 MSI-H MSI-H 84 tMSIbz-084 0 MSS MSS 86 tMSIbz-086 2 MSS MSS 87 tMSIbz-087 3 MSS MSS 88 tMSIbz-088 95 MSI-H MSI-H 89 tMSIbz-089 91 MSI-H MSI-H 90 tMSIbz-090 68 MSI-H MSI-H 91 tMSIbz-091 1 MSS MSS 92 tMSIbz-092 100 MSI-H MSI-H 93 tMSIbz-093 99 MSI-H MSI-H 94 tMSIbz-094 98 MSI-H MSI-H 95 tMSIbz-095 98 MSI-H MSI-H 96 tMSIbz-096 97 MSI-H MSI-H 97 tMSIbz-097 99 MSI-H MSI-H 98 tMSIbz-098 100 MSI-H MSI-H 99 tMSIbz-099 100 MSI-H MSI-H 100 tMSIbz-100 4 MSS MSS 101 tMSIbz-101 99 MSI-H MSI-H 102 tMSIbz-102 100 MSI-H MSI-H 103 tMSIbz-103 82 MSI-H MSI-H 104 tMSIbz-104 1 MSS MSS 105 tMSIbz-105 97 MSI-H MSI-H 107 tMSIbz-107 99 MSI-H MSI-H 108 tMSIbz-108 70 MSI-H MSI-H 110 tMSIbz-110 3 MSS MSS 113 tMSIbz-113 92 MSI-H MSI-H 114 tMSIbz-114 69 MSI-H MSI-H 115 tMSIbz-115 81 MSI-H MSI-H 116 tMSIbz-116 2 MSS MSS 117 tMSIbz-117 96 MSI-H MSI-H 118 tMSIbz-118 1 MSS MSS 120 tMSIbz-120 99 MSI-H MSI-H 121 tMSIbz-121 96 MSI-H MSI-H 122 tMSIbz-122 96 MSI-H MSI-H 123 tMSIbz-123 93 MSI-H MSI-H 124 tMSIbz-124 0 MSS MSS 125 tMSIbz-125 100 MSI-H MSI-H 127 tMSIbz-127 63 MSI-H MSI-H 128 tMSIbz-128 100 MSI-H MSI-H 129 tMSIbz-129 1 MSS MSS 130 tMSIbz-130 5 MSS MSS 131 tMSIbz-131 99 MSI-H MSI-H 132 tMSIbz-132 21 MSS MSI-H 133 tMSIbz-133 100 MSI-H MSI-H 134 tMSIbz-134 100 MSI-H MSI-H 135 tMSIbz-135 95 MSI-H MSI-H 136 tMSIbz-136 100 MSI-H MSI-H 138 tMSIbz-138 100 MSI-H MSI-H 139 tMSIbz-139 99 MSI-H MSI-H 140 tMSIbz-140 92 MSI-H MSI-H 141 tMSIbz-141 39 MSI-H MSI-H 142 tMSIbz-142 93 MSI-H MSI-H 144 tMSIbz-144 96 MSI-H MSI-H 145 tMSIbz-145 58 MSI-H MSI-H 146 tMSIbz-146 85 MSI-H MSI-H 147 tMSIbz-147 97 MSI-H MSI-H 148 tMSIbz-148 94 MSI-H MSI-H 149 tMSIbz-149 2 MSS MSS 151 tMSIbz-151 2 MSS MSS 152 tMSIbz-152 73 MSI-H MSI-H 153 tMSIbz-153 0 MSS MSS 154 tMSIbz-154 98 MSI-H MSI-H 155 tMSIbz-155 83 MSI-H MSI-H 156 tMSIbz-156 99 MSI-H MSI-H 158 tMSIbz-158 0 MSS MSS 159 tMSIbz-159 98 MSI-H MSI-H 160 tMSIbz-160 96 MSI-H MSI-H 161 tMSIbz-161 1 MSS MSS 162 tMSIbz-162 92 MSI-H MSI-H 163 tMSIbz-163 39 MSI-H MSI-H 164 tMSIbz-164 0 MSS MSS 165 tMSIbz-165 93 MSI-H MSI-H 166 tMSIbz-166 100 MSI-H MSI-H 167 tMSIbz-167 80 MSI-H MSI-H 168 tMSIbz-168 99 MSI-H MSI-H 169 tMSIbz-169 99 MSI-H MSI-H 170 tMSIbz-170 93 MSI-H MSI-H 171 tMSIbz-171 92 MSI-H MSI-H

Claims

1. A method for screening microsatellite loci, the method comprising the following steps: Obtain the first point set, which includes multiple captured stable microsatellite sites; Based on multiple reads of each microsatellite site in the first site set, a unimodal maximum likelihood estimate of the microsatellite site is calculated using a preset unimodal probability distribution model, and a bimodal maximum likelihood estimate of the microsatellite site is calculated using a preset bimodal probability distribution model, wherein the preset bimodal probability distribution model considers more polymorphism than the preset unimodal probability distribution model. For each microsatellite site in the first site set, the likelihood ratio of the bimodal maximum likelihood estimate to the unimodal maximum likelihood estimate is obtained; and Based on a preset judgment model including a threshold and the likelihood ratio of each microsatellite site in the first site set, the microsatellite sites in the first site set are screened to obtain a second microsatellite site set with lower polymorphism. The preset judgment model is obtained using a training set containing both positive and negative samples. The preset judgment model is as follows: Based on the likelihood ratios of the input microsatellite loci, plot the distribution of the likelihood ratios; and The threshold is determined based on the intersection position of the main peak and the first distribution peak in the long tail of the distribution pattern; Furthermore, the method also includes: The preset determination model compares the likelihood ratio of each microsatellite site in the first site set with the threshold, and selects microsatellite sites with a likelihood ratio below the threshold as the second microsatellite site set.

2. The method according to claim 1, wherein, The microsatellite loci in the first locus set are located in non-exon regions.

3. The method according to claim 1 or 2, wherein, The calculation of the unimodal maximum likelihood estimate of the microsatellite locus using a preset unimodal probability distribution model includes: For each of the plurality of reads, a unimodal maximum likelihood estimate is calculated for each read based on the stationary probabilities of the three events—missing, normal, and inserted—and their respective occurrence frequencies in each read; and The joint maximum likelihood estimate of the multiple reads is obtained as the single-peak maximum likelihood estimate of the microsatellite locus.

4. The method according to claim 1 or 2, wherein, The calculation of the bimodal maximum likelihood estimate of the microsatellite locus using a preset bimodal probability distribution model includes: For each of the plurality of read segments, a bimodal maximum likelihood estimate is calculated for each read segment based on the stationary probabilities of the three events—missing, normal, and inserted—and their respective occurrence frequencies in each read segment. The bimodal maximum likelihood estimate for each read segment is obtained based on a unimodal maximum likelihood estimate of the first peak and a unimodal maximum likelihood estimate of the second peak. The joint maximum likelihood estimate of the multiple reads is obtained as the bimodal maximum likelihood estimate of the microsatellite locus.

5. The method according to claim 1 or 2, wherein, The microsatellite locus is more than 15 nucleotides in length.

6. The method according to claim 1 or 2, wherein, Microsatellite loci in the first locus set are captured using one or more of the following sets of probes: Probe set 1, wherein the middle region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site; Probe set 2, wherein the downstream region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site; Probe set 3, wherein the upstream region of the probe sequence spans the microsatellite site, and optionally repeating bases are subtracted from the probe sequence corresponding to the microsatellite site; and Probe set 4, wherein the ends of the probe sequences are located at the microsatellite sites, and optionally repeating bases are subtracted from the probe sequences corresponding to the microsatellite sites.

7. The method according to claim 6, wherein, The number of repeating bases is 0, 3, 5, 7, 9, or 11.

8. A method for detecting microsatellite instability for non-diagnostic purposes, the method comprising using a second set of loci obtained by the screening method for microsatellite loci according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • BMSI (Microsatellite Instability) detection technology based on capture sequencing

    CN108913772A

  • Method for screening microsatellite loci and application thereof

    CN112662740A