Gene marker group and application thereof in microsatellite state determination

By employing genomic biomarker analysis and ensemble learning strategies, and utilizing multiple algorithms to detect microsatellite status, this approach addresses the issue of insufficient sensitivity in detecting low-tumor-content samples in existing technologies, achieving highly sensitive and accurate microsatellite status detection.

CN121629044APending Publication Date: 2026-03-10TIANJIN MEDICAL LAB BGI +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing MSI detection methods have poor sensitivity in samples with low tumor content and use a limited number of microsatellite loci, resulting in inaccurate detection.

Method used

A genome biomarker, including multiple MSI core loci, was used to detect microsatellite status using various algorithms. An ensemble learning strategy was employed to integrate the prediction results of multiple algorithms to determine the microsatellite status of the samples.

Benefits of technology

It improves the sensitivity and accuracy of microsatellite state detection in samples with low tumor proportions, and is suitable for detection in multiple modes, especially for test samples with low tumor proportions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121629044A_ABST
    Figure CN121629044A_ABST
Patent Text Reader

Abstract

The invention provides a gene marker group and application thereof in microsatellite state determination. The gene marker group comprises at least one of sites shown in a table 1. The gene marker group is suitable for high-sensitivity detection of a microsatellite state in a to-be-detected sample with a low tumor proportion in various modes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene detection, specifically to the genome of gene biomarkers and their application in microsatellite status determination. Background Technology

[0002] Microsatellite instability (MSI) is a molecular characteristic closely related to tumorigenesis and has significant clinical implications in the detection of various solid tumors. MSI primarily results from abnormalities in the mismatch repair (MMR) system during DNA replication. During DNA replication, DNA polymerase is prone to "slipping" when encountering short tandem repeat sequences, leading to the insertion or deletion of nucleotides at microsatellite sites. Normally, the mismatch repair system can recognize and repair this instability. However, if the MMR gene is mutated or abnormally hypermethylated, mismatch repair function is lost, and spontaneous high-frequency length variations in microsatellites cannot be repaired in time, thus forming MSI.

[0003] Currently, there are two main approaches to detecting MSI: one based on the cause and the other on the result. The cause-based approach typically uses immunohistochemistry (IHC) to detect MMR protein expression or high-throughput sequencing to directly detect MMR gene mutations. The result-based approach often uses PCR-capillary electrophoresis to amplify the DNA sequence of microsatellite loci and analyze the size of the amplified products to determine the MSI status, or high-throughput sequencing to detect whether the MSI loci are distributed similarly to controls.

[0004] Although existing MSI detection technologies have played an important role in tumor research and treatment, there are still some shortcomings, such as the limited number of microsatellite loci (MS) selected for determining MSI status and poor sensitivity of MSI detection in samples with low tumor content.

[0005] Therefore, existing MSI detection methods still need improvement. Summary of the Invention

[0006] This application aims to address at least one of the technical problems existing in the prior art. To this end, one objective of this application is to provide a means for accurately determining the microsatellite state of a low proportion of tumor tissue.

[0007] In a first aspect of this application, a set of genetic markers is proposed. According to embodiments of this application, the set of genetic markers includes at least one of the sites shown in Table 1:

[0008] Table 1

[0009]

[0010]

[0011] The aforementioned set of genetic biomarkers is suitable for highly sensitive detection of microsatellite status in test samples with low tumor proportions under various modes.

[0012] In a second aspect of this application, the application proposes the use of the aforementioned genomic biomarker set in microsatellite state detection.

[0013] The aforementioned set of genetic markers can be used for accurate detection of microsatellite status, and is particularly suitable for test samples with a low tumor proportion.

[0014] In a third aspect, this application proposes a method for determining the microsatellite state. According to an embodiment of this application, the method includes: acquiring sequencing data of a sample to be tested, the sequencing data including multiple MSI core sites; determining the MSI state information of the multiple MSI core sites using at least one algorithm; inputting the MSI state information obtained by the at least one algorithm into multiple trained prediction models to determine multiple MSI state prediction results of the at least one algorithm; and determining the microsatellite state of the sample to be tested based on the multiple MSI state prediction results of the at least one algorithm.

[0015] The aforementioned method provides a large number of sites (MSI core sites) that can be used to determine the microsatellite instability state. By using the ensemble approach, the MSI state of a single MSI core site is detected through multiple algorithms, and the MSI state prediction results obtained by each algorithm are integrated to obtain the microsatellite state result of the sample to be tested. This avoids tedious manual data processing and analysis, and significantly improves detection sensitivity and accuracy. It is especially suitable for the detection of samples with low tumor proportions under various modes.

[0016] In a fourth aspect, this application proposes a microsatellite state determination system. According to an embodiment of this application, the system includes: a sequencing data acquisition module for acquiring sequencing data of a sample to be tested, the sequencing data including multiple MSI core sites; an MSI state information determination module for determining the MSI state information of the multiple MSI core sites using at least one algorithm; an MSI state prediction result determination module for inputting the MSI state information obtained by the at least one algorithm into multiple trained prediction models to determine multiple MSI state prediction results of the at least one algorithm; and a microsatellite state judgment module for determining the microsatellite state of the sample to be tested based on the multiple MSI state prediction results of the at least one algorithm.

[0017] The aforementioned system can efficiently, sensitively, and accurately predict the microsatellite state of samples with a low tumor proportion. By utilizing an ensemble learning strategy, it not only enhances detection sensitivity and specificity but also improves the overall system's robustness and noise resistance, making it widely applicable and capable of effectively distinguishing between stable and unstable microsatellite states in complex and noisy data.

[0018] In a fifth aspect, this application provides a computer program product. According to an embodiment of this application, the computer program product includes: computer instructions; when some or all of the computer instructions are executed on a computer, the microsatellite state determination method as described in the third aspect of this application is performed.

[0019] In a sixth aspect, this application proposes a computing device. According to an embodiment of this application, the computing device includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the microsatellite state determination method as described in the third aspect of this application.

[0020] In a seventh aspect, this application provides a computer-readable storage medium. According to an embodiment of this application, the computer-readable storage medium stores computer instructions or programs that, when executed on a computer, cause the microsatellite state determination method as described in the third aspect of this application to be performed.

[0021] The aforementioned computer program products, computing devices, and computer-readable storage media automatically execute the microsatellite state determination method through computer instructions, achieving high efficiency and automation, and improving detection efficiency, accuracy, sensitivity, and specificity. Furthermore, the instruction-based nature of the method ensures high consistency, stability, and reliability across various experimental scenarios.

[0022] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the distribution of MSI sites under different purity dilution conditions according to a specific embodiment of this application;

[0025] Figure 2This is a schematic diagram of the distribution of microsatellite sites provided according to a specific embodiment of this application;

[0026] Figure 3 This is a schematic diagram of a microsatellite status determination system provided according to a specific embodiment of this application;

[0027] Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present application. Detailed Implementation

[0028] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0029] In this text, the singular forms “a,” “an,” etc., are used to include the plural referents (one or more). “A group” or “a plurality” refers to two or more.

[0030] In this document, “include” or “include” is an open-ended expression that includes the content or circumstances specified or exemplified thereafter, as well as any content or circumstances that are applicable or consistent with the stated circumstances but not specifically listed.

[0031] In this document, the term "sequencing data" refers to data obtained by sequencing the sample to be tested. Sequencing can be performed using a sequencing platform. This application does not impose any particular limitation on the sequencing platform; it can be any common second-generation sequencing platform or third-generation sequencing platform. This application preferably uses a second-generation sequencing platform. Sequencing platforms that can be used for the above sequencing in this application include, but are not limited to, GenoCare™ 1600 single-molecule sequencing platform, GenoLab™ M, FASTASeq300 and SURFSeq™ 5000 high-throughput sequencing platforms from GenoCare Biotechnology, HiSeq / Miseq / Nextseq / Novaseq sequencing platforms from Illumina, Ion Torrent platform from Thermo Fisher / Life Technologies, and BGISEQ and MGISEQ / DNBSEQ platforms from BGI Genomics.

[0032] In this paper, the term "MSI core site" refers to a specific site used to predict the state of microsatellites.

[0033] In this paper, the term "microsatellite (MS)" refers to short tandem repeat sequences in the genome, the number of which varies greatly among individuals.

[0034] In this paper, the term "microsatellite instability (MSI)" refers to the phenomenon in tumor cells where the length of microsatellites changes abnormally due to the insertion or deletion of repetitive sequences compared to normal cells.

[0035] In this document, the term "MSI state" refers to the presence of MSI or unstable microsatellite (sites), i.e., variations in the number of repetitive DNA nucleotide units in clonal or somatic microsatellites. In this application, MSI state includes MSS / MSS-L (microsatellite stable / microsatellite low instability) or MSI-H (microsatellite high instability). MSS / MSS-L indicates that the number of repetitive segments in microsatellite sites does not differ significantly between tumor and normal cells, while MSI-H indicates that the number of repetitive segments present in microsatellite sites differs significantly from the number of repetitive segments in normal cell DNA.

[0036] In this paper, the term "predictive model or machine learning model" refers to a computational model or algorithm that can automatically perform tasks such as prediction, classification, recognition, or decision-making by learning and analyzing input data. The learning process of this model is based on statistical principles and data pattern recognition, using a training dataset for parameter tuning and model optimization to improve its predictive or inference capabilities. Machine learning models can employ various algorithms and techniques, such as support vector machines, decision trees, random forests, and deep learning. These models can be trained and optimized through supervised learning, unsupervised learning, or reinforcement learning. In practical applications, machine learning models can be used in various fields, such as natural language processing, image recognition, pattern recognition, data mining, recommender systems, and predictive analytics. It has significant application potential in processing large-scale data, automated decision-making, and intelligent systems. The inventors of this application, by employing an ensemble strategy and machine learning methods, have obtained a large number of microsatellite loci for determining microsatellite states, enabling accurate and highly sensitive determination of the microsatellite states of samples in low-tumor-proportion samples.

[0037] The conventional "2B3D" algorithm includes two single nucleotide sites (BAT25, BAT26) and three dinucleotide sites (D2S123, D5S346, and D17S250). When determining the overall MSI status of a sample, 0 unstable sites are classified as MSS, 1 unstable site as MSI-L, and ≥2 unstable sites as MSI-H. In reality, the fewer sites detected, the more prone the interpretation of the sample's MSI status is to error. Although, due to cost considerations, authoritative domestic and international guidelines and expert consensus primarily recommend the 2B3D Panel and Promega Panel for MSI detection using PCR-capillary electrophoresis, the development of NGS technology presents limitations when only examining a small number of MSI sites. Similar limitations also exist in the Promega Panel.

[0038] Both PCR and NGS methods for detecting MSI are affected by the amount of tumor cells. For example... Figure 1 As shown, the inventors discovered in numerous experiments that the fragment distribution of a single site differs depending on the tumor purity. When the tumor purity is around 10%, the distribution trend of this microsatellite site in tumor and normal samples will be very similar, leading to poor sensitivity in MSI status detection.

[0039] To address the aforementioned problems, this application proposes a genome of genetic markers and its application in microsatellite state determination, a method for microsatellite state determination, a microsatellite state determination system, a computer program product, a computing device, and a computer storage medium. These are described in detail below:

[0040] Genome

[0041] On the one hand, this application proposes a set of genetic markers, which includes at least one of the sites shown in Table 1:

[0042] Table 1

[0043] ID CHR START END LEFT RIGHT MSI100 chr1 120053340 120053377 TTTTC GAGAC MSI207 chr1 180105259 180105301 GGAAA AAACT MSI283 chr2 39573062 39573089 GTCTC GAGTG MSI318 chr2 47641559 47641586 CAGGT GGGTT MSI365 chr2 51288511 51288553 GTGAT ATATT MSI376 chr2 95849361 95849384 TCCTA GTGAG MSI693 chr4 55598211 55598236 TTTGA GAGAA MSI801 chr5 112213678 112213718 TACTC TAAAT MSI1824 chr14 23652346 23652367 TTGCT GGCCA MSI1936 chr17 7572154 7572172 AAAAG GAAAA MSI1953 chr17 13171093 13171127 ACACG CCACC MSI1954 chr17 15359877 15359911 CCTTT GGCTC MSI1983 chr17 37152123 37152163 TGAAA TTGAA MSI2160 chr18 53161388 53161412 TATGG AGATT MSI2162 chr18 61873521 61873573 TATGC ACGAG MSI2165 chr18 72044533 72044569 AGGAC GCGCG

[0044] The aforementioned genomarker set can be used for accurate detection of microsatellite status.

[0045] In some examples of this application, the aforementioned genetic markers further include at least one of the sites shown in Table 2:

[0046] Table 2

[0047] ID CHR START END LEFT RIGHT MSI268 chr2 29449344 29449368 GCCTC GAGGA MSI395 chr2 111884996 111885015 CTTTC CCTCA MSI465 chr3 30691871 30691881 GAAGG GCCTG MSI682 chr4 25680309 25680328 TGTAA ACTGG MSI874 chr6 117642992 117643012 GGCAA GCAGA MSI1097 chr7 55273591 55273604 ATTTG GTATA MSI1153 chr7 116414203 116414214 GAAAT TTCTC MSI2166 chr18 72044569 72044581 TGTGT ATATA MSI2245 chr20 43955023 43955047 ATAGT TTAAT

[0048] To address the limitations of current methods for determining MSI status, which have limited available microsatellite loci and low sensitivity, the inventors have added the following loci to the existing microsatellite loci listed in Table 1 to enhance the accuracy and sensitivity of MSI status detection.

[0049] In some examples of this application, the aforementioned genetic markers further include at least one of the sites shown in Table 3:

[0050] Table 3

[0051]

[0052] .

[0053] Addressing the limitations of current methods for determining MSI status, which have limited available microsatellite loci and low sensitivity, the inventors have added the following loci to the existing microsatellite loci listed in Table 1 to enhance the accuracy and sensitivity of MSI status detection.

[0054] In some examples of this application, the aforementioned genetic markers further include at least one of the sites shown in Table 4:

[0055] Table 4

[0056] ID CHR START END LEFT RIGHT MSI359 chr2 48030639 48030647 AGATA TTCTT MSI728 chr5 79970914 79970922 GGGAC GGGCA MSI1546 chr11 125505377 125505386 GAAAG CATAC MSI1958 chr17 16029445 16029453 TCTTC CTGTT MSI2243 chr20 43954848 43954857 CTGCT CTATG .

[0057] Addressing the limitations of current methods for determining MSI status, which have limited available microsatellite loci and low sensitivity, the inventors have added the following loci to the existing microsatellite loci listed in Table 1 to enhance the accuracy and sensitivity of MSI status detection.

[0058] In some examples of this application, the aforementioned genetic markers further include at least one of the sites shown in Table 5:

[0059] Table 5

[0060]

[0061]

[0062] .

[0063] Addressing the limitations of current methods for determining MSI status, which have limited available microsatellite loci and low sensitivity, the inventors have added the following loci to the existing microsatellite loci listed in Table 1 to enhance the accuracy and sensitivity of MSI status detection.

[0064] It should be noted that, provided that experimental conditions permit, the gene marker loci shown in Table 1 and any combination of one or more gene marker loci in Tables 2-5 may be freely selected for the detection of MSI status.

[0065] The inventors discovered that detecting MSI status in samples with low tumor proportions based on the above-mentioned combination of gene marker sites has advantages such as high detection accuracy and high sensitivity.

[0066] Microsatellite status determination method

[0067] On the one hand, this application proposes a method for determining the state of a microsatellite, the method comprising:

[0068] 1) Obtain sequencing data of the sample to be tested, wherein the sequencing data includes multiple MSI core sites;

[0069] In some examples of this application, the sample to be tested may be derived from cell lines, biopsy, primary tissue, frozen tissue, formalin-fixed paraffin-embedded (FFPE) tissue, liquid biopsy, blood, serum, plasma, buffy coat, body fluid, visceral fluid, ascites, paracentesis, cerebrospinal fluid, saliva, urine, tears, semen, vaginal secretions, aspirate, lavage fluid, buccalswab, circulating tumor cells (CTC), cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), DNA, RNA, nucleic acids, purified nucleic acids, purified DNA, or purified RNA. In some preferred examples of this application, the aforementioned sample to be tested is selected from tissue samples or FFPE samples.

[0070] In practical MSI status detection, a test sample includes both healthy tissue and tumor tissue. The proportion of tumor tissue has a significant impact on MSI status detection. In some examples of this application, selecting a test sample with a tumor proportion of no less than 10% ensures the sensitivity and specificity of MSI status detection. The aforementioned tumor tissue can be selected from hereditary nonpolyposis colorectal cancer, colorectal cancer, gallbladder cancer, ovarian cancer, pancreatic cancer, prostate cancer, liver cancer, kidney cancer, lung cancer, gastric cancer, nasopharyngeal carcinoma, esophageal cancer, or endometrial cancer tissue.

[0071] In some examples of this application, candidate MSI sites can be obtained by searching for monobasic, dibasic, tribasic, tetrabasic, pentabasic, and hexabasic repeat sequences based on existing gene panels (such as BGI Genomics). In some specific examples of this application, sites that satisfy the condition of monobasic repeat count ≥ 5 and dibasic, tribasic, tetrabasic, pentabasic, and hexabasic repeat count ≥ 3 are selected as candidate MSI sites.

[0072] Candidate MSI sites obtained using the above methods may have reduced detection accuracy due to low coverage and conservative site performance. Therefore, in some examples of this application, core MSI sites can be obtained through quality control.

[0073] In some examples of this application, the quality control methods include at least one of the following:

[0074] a) MSI site sequencing depth quality control: Remove candidate MSI sites with sequencing depths lower than 20X (30X, 50X, etc. can be selected);

[0075] b) MSI site standard deviation quality control: Obtain the standard deviation information of MSI sites through algorithms (such as Manhattan distance algorithm) and remove MSI sites with a standard deviation of 0;

[0076] c) Quality control of unit point sensitivity in MSI sites: The status of MSI sites is detected by one or more algorithms (such as MSISensor algorithm, MSISensor-pro algorithm, etc.), and the sensitivity information of unit points in MSI sites is statistically analyzed. MSI sites with low sensitivity (such as sensitivity <0.5 or sensitivity of 0) are removed.

[0077] In other examples of this application, the aforementioned multiple MSI core sites are obtained by: acquiring multiple candidate MSI sites; determining the MSI state of the multiple candidate MSI sites based on at least one algorithm; and inputting the MSI state of the multiple candidate MSI sites into a trained machine learning model to determine the multiple MSI core sites.

[0078] After obtaining candidate MSI sites using the above method, one or more algorithms (such as MSISensor, MSISensor-pro, and Manhattan distance algorithm) are used to determine the state of each candidate MSI site. This state is then used as input to a trained machine learning model to determine the core MSI sites. For example, for any sample, the stable / unstable state of the i-th site is determined. If the site is unstable (MSI-H), its MSI state is marked as 1; if the site is stable (MSS / MSI-L), its MSI state is marked as 0. The 0 / 1 values ​​are then used as input to the trained machine learning model.

[0079] In some examples of this application, the machine learning models are selected from at least one of Lasso and random forest. The training methods used for these machine learning models are conventional in the art and will not be elaborated upon here.

[0080] The following example illustrates how the core sites of MSI are determined:

[0081] For example, using the Lasso machine learning model to screen MSI core sites includes:

[0082] The samples were divided into training and validation sets. Lasso dimensionality reduction was used to screen for core MSI loci, repeated 1000 times, and the model's AUC was recorded for each iteration. The more times an MSI locus was selected by the dimensionality reduction model in the 1000 random tests, the greater its role in MSI detection of the samples.

[0083] In a test of 193 samples (77 positive samples and 116 negative samples), the MSI site screening results obtained based on the MSISensor algorithm (parameters: single base repeat count ≥ 10; repeat unit length ≥ 2 and repeat count ≥ 5, totaling 60 candidate MSI sites) are as follows:

[0084] (1) In 1000 random tests, there were 209 AUCs = 100%, 641 AUCs ≥ 99%, 962 AUCs ≥ 98%, and 999 AUCs ≥ 97%; (2) After 1000 random tests, among all 60 candidate MSI sites examined, 14 MSI sites were screened ≥ 300 times (Table 6).

[0085] In a test of 193 samples (77 positive samples and 116 negative samples), the MSI site screening results obtained based on the MSISensor-pro algorithm (parameters: single base repeat count ≥ 8; repeat unit length ≥ 2 and repeat count ≥ 5, totaling 83 candidate MSI sites) are as follows:

[0086] (1) In 1000 random tests, there were 235 AUCs = 100%, 630 AUCs ≥ 99%, 965 AUCs ≥ 98%, and 1000 AUCs ≥ 97%; (2) After 1000 random tests, among all 83 candidate MSI loci examined, 14 MSI loci were screened ≥ 300 times (Table 6).

[0087] For example, using a random forest machine learning model to screen MSI core sites includes:

[0088] The samples were divided into training and validation sets (7:3 ratio), and a random forest model was used to screen MSI core sites. Based on the feature coefficients, the MSI core sites were determined.

[0089] In a test of 193 samples (77 positive samples and 116 negative samples), the MSI site screening results obtained based on the MSISensor algorithm (parameters: single base repeat count ≥10; repeat unit length ≥2 and repeat count ≥5, a total of 60 candidate MSI sites) are as follows: 16 sites have an eigenvalue coefficient >1 (Table 6).

[0090] In a test of 193 samples (77 positive samples and 116 negative samples), the MSI site screening results obtained based on the MSISensor-pro algorithm (parameters: single base repeat count ≥ 8; repeat unit length ≥ 2 and repeat count ≥ 5, a total of 83 candidate MSI sites) are as follows: 14 sites have an eigenvalue coefficient > 1 (Table 6).

[0091] Table 6

[0092]

[0093]

[0094] Those skilled in the art will understand that the number of sites obtained varies depending on the selection algorithm and algorithm parameters. The sites obtained above are merely an exemplary method shown in this application.

[0095] In some examples of this application, in addition to the core site screening methods described above, sites reported in literature or guidelines can also be obtained.

[0096] Among all the obtained MSI core loci, their importance is further ranked to suit different experimental platforms or detection methods. In some examples of this application, the aforementioned importance ranking methods include: adding quality control conditions, adjusting the parameters of the core loci screening method, etc.

[0097] In some examples of this application, multiple MSI core sites were obtained based on the above method, and their importance was ranked, including at least one of type I core sites, type II core sites, type III core sites, type IV core sites, and type V core sites.

[0098] Specifically, the aforementioned Type I core loci are shown in Table 1 above; the aforementioned Type II core loci are shown in Table 2 above; the aforementioned Type III core loci are shown in Table 3 above; the aforementioned Type IV core loci are shown in Table 4 above; and the aforementioned Type V core loci are shown in Table 5 above.

[0099] 2) Determine the MSI status information of the plurality of MSI core sites using at least one algorithm;

[0100] In some examples of this application, the aforementioned algorithms include at least one of MSISensor, MSISensor-pro, and the Manhattan distance algorithm. One or more of these algorithms are selected to determine the MSI state information of the multiple MSI core sites obtained in step 1), such as MSI-L / MSS or MSI-H.

[0101] The detection principle of the aforementioned MSISensor algorithm is based on examining whether the sample to be detected and the control sample are co-distributed at the same site using the chi-square distribution.

[0102] The aforementioned MSISensor-pro algorithm is based on examining whether the sample to be detected and the control sample are distributed at the same location using a multivariate distribution.

[0103] refer to Figure 2 The aforementioned Manhattan distance algorithm detects microsatellite instability by statistically analyzing the proportion of segments of different lengths at each microsatellite locus, i.e., the homogenized depth, and assigning an instability score (BGI-score) to each MSI locus based on the Manhattan distance. Specifically, the calculation of the Manhattan distance includes:

[0104]

[0105] Where d is the Manhattan distance, R T For the test sample to contain at least a portion of repeating unit length types corresponding to the read segment, R N T represents the length type of at least a portion of the corresponding read segment in the control sample. r N is the normalized read count of the r-th repeating unit length type in the sample to be tested. r The normalized read count is the number of the r-th repeating unit length type in the control sample.

[0106] The phrase "at least a portion of the repeating unit length types of the corresponding read segments exist in the test sample" means that in the test sample, there are certain read segments in which the repeating unit lengths belong to a specific type.

[0107] Specifically, suppose that in a certain sample to be tested, there are two repeating unit length types: 3 bases and 5 bases.

[0108] For 3-base repeat unit types, sequence fragments resembling "AAA" or "GAG" may be observed in reads. These reads have repeat units of the same length, i.e., 3 bases. By analyzing these reads, it can be determined that sequences with a 3-base repeat unit length are present in the sample.

[0109] For 5-base repeat unit types, sequence fragments such as "TTTTT" or "CGCGC" may be observed in reads. These reads have repeat units of the same length, i.e., 5 bases. By analyzing these reads, it can be determined that sequences with a 5-base repeat unit length are present in the sample.

[0110] Among them, the length type of at least a portion of the corresponding reading segments in the control sample is the same as explained above, and will not be repeated here.

[0111] The so-called R T ∪R N The union of all repeating unit types, or the union of some repeating unit types, including the test sample and the control sample.

[0112] The test samples are from individuals suspected of having tumors, while the control samples are from non-tumor tissues. In some examples, the relevant distribution information of the control samples may be obtained in related parallel experiments or through pre-sequencing.

[0113] In other examples of this application, the MSI status information of the core MSI sites can also be determined using parametric or non-parametric methods. The parametric methods include: comparing the mean or standard deviation of two distributions (normality), using the chi-square test for the proportions of different values ​​(normality), or employing gene enrichment techniques (drawing n reads from the population, where k reads are unstable and their probability conforms to a hypergeometric distribution). Non-parametric methods include: comparing the median, peak value, peak width, and number of peaks of two distributions, and calculating the sum of the probability density differences at various locations within the distribution function.

[0114] 3) Input the MSI state information obtained by the at least one algorithm into multiple trained prediction models respectively, and determine the multiple MSI state prediction results of the at least one algorithm;

[0115] In some examples of this application, inputting the MSI state information of N algorithms into M trained prediction models yields N×M MSI state prediction results for that algorithm. For instance, inputting the MSI state information obtained from two algorithms into 10 trained prediction models yields 20 MSI state prediction results for that algorithm.

[0116] 4) Based on the multiple MSI state prediction results of the at least one algorithm, determine the microsatellite state of the sample to be tested.

[0117] In some examples of this application, the aforementioned prediction models are selected from logistic regression, support vector machines, linear support vector machines, K-nearest neighbor algorithm, decision trees, or ensemble learning methods. The aforementioned ensemble learning methods may include Bagging, Boosting, Stacking, Voting, or Blending. The aforementioned prediction model training methods are conventional training methods in this field and will not be elaborated further here.

[0118] In some examples of this application, the inventors verified the sensitivity and specificity of 15 machine learning models and 3 algorithms in microsatellite state prediction.

[0119] The 15 machine learning models include: 5 simple machine learning models, namely: LogisticRegression, SVC, LinearSVC, KNeighborsClassifier, and DecisionTreeClassifier; and 10 ensemble machine learning models, namely: Integration, hard voting, soft voting, bagging_1, bagging_2, bagging_3, RandomForestClassifier, ExtraTreesClassifier, AdaBoostClassifier, and GradientBoostingClassifier.

[0120] The three algorithms are: MSISensor, MSISensor-pro, and Manhattan distance algorithm.

[0121] The inventors trained and evaluated various prediction models on 193 samples (77 positive; 116 negative) using existing model training methods (training set: validation set = 7:3). The evaluation results of microsatellite states using different algorithms and prediction models are shown in Table 7.

[0122] Table 7

[0123]

[0124]

[0125] Note: Algorithm 1 is MSISensor; Algorithm 2 is MSISensor-pro; Algorithm 3 is the Manhattan distance algorithm.

[0126] In some examples of this application, the microsatellite status of the aforementioned test sample is determined as follows:

[0127] The MSI state prediction result of the algorithm is determined based on the multiple MSI state prediction results; the microsatellite state of the sample to be tested is further determined based on the MSI state prediction result of the algorithm.

[0128] Using 3 algorithms and 10 prediction models as examples, the state judgment of microsatellites under test is explained:

[0129] In the first algorithm, the MSI-H prediction ratio exceeded half of the 10 prediction models, indicating that the microsatellite state predicted by this algorithm is MSI-H.

[0130] In the second algorithm, the proportion of MSI-H predictions in the 10 prediction models did not exceed half, so it is considered that the microsatellite state predicted by this algorithm is MSI-L / MSS.

[0131] In the third algorithm, more than half of the 10 prediction models predicted MSI-H, indicating that the microsatellite state predicted by this algorithm is MSI-H.

[0132] The three algorithms predict that more than half of the microsatellites in the test sample are in the MSI-H state, and therefore consider the MSI state of the test sample to be MSI-H.

[0133] It should be noted that the above-mentioned ratios can be set based on the sensitivity and specificity of the model or algorithm; this article is only providing an example.

[0134] The above method has the following advantages:

[0135] 1. By integrating multiple algorithms and machine learning models, the accuracy, sensitivity, and specificity of microsatellite state detection in samples with low tumor proportions were significantly improved;

[0136] 2. The choice of algorithm is flexible;

[0137] 3. It abandons the traditional NGS method of calculating MSI by "number of unstable sites among the sites under investigation / total number of sites under investigation". Instead, it adopts the idea of ​​"integration" of multiple machine learning models. This allows high-performance sites to be given higher weights during model training, enabling the model to perform better. Moreover, it can reduce the proportion of gray area samples in practical applications (the traditional method of calculating the proportion of unstable sites often leads to changes in the overall MSI state of the sample due to the difference in judgment of one or two candidate sites).

[0138] 4. This method ranks the obtained MSI core sites by importance, which can adapt to different experimental needs;

[0139] 5. The ensemble strategy employed in this method makes the final MSI state results of the samples more accurate. It reduces the deviation of the overall sample state caused by errors in the judgment of individual algorithms, individual machine learning models, or individual MSI candidate feature sites.

[0140] 6. This method is applicable not only to paired detection modes of tissue and normal control, but also to single-tissue detection modes of tumors;

[0141] 7. This method can also detect the MSI status of cfDNA samples.

[0142] Microsatellite status determination system

[0143] On the other hand, this application proposes a microsatellite state determination system. (Reference) Figure 3 The system includes: a sequencing data acquisition module 100, an MSI status information determination module 200, an MSI status prediction result determination module 300, and a microsatellite status judgment module 400.

[0144] Module 100 is used to acquire sequencing data of the sample to be tested, the sequencing data including multiple MSI core sites;

[0145] In some examples of this application, the proportion of tumors in the aforementioned test samples is not less than 10%. The aforementioned tumor types include: hereditary nonpolyposis colorectal cancer, colorectal cancer, gallbladder cancer, ovarian cancer, pancreatic cancer, prostate cancer, liver cancer, kidney cancer, lung cancer, stomach cancer, nasopharyngeal cancer, esophageal cancer, or endometrial cancer.

[0146] In some examples of this application, the aforementioned multiple MSI core sites are obtained by: acquiring multiple candidate MSI sites; determining the MSI state of the multiple candidate MSI sites based on at least one algorithm; and inputting the MSI state of the multiple candidate MSI sites into a trained machine learning model to determine the multiple MSI core sites.

[0147] In some examples of this application, the sequencing depth of the aforementioned MSI core sites is not less than 20X.

[0148] In some examples of this application, the aforementioned machine learning models are selected from at least one of Lasso and random forest.

[0149] In some examples of this application, the aforementioned algorithms include at least one of MSISensor, MSISensor-pro, and the Manhattan distance algorithm.

[0150] In some examples of this application, the aforementioned Manhattan distance is calculated as follows:

[0151]

[0152] Where d is the Manhattan distance, R T For the test sample to contain at least a portion of repeating unit length types corresponding to the read segment, R N T represents the length type of at least a portion of the corresponding read segment in the control sample. r N is the normalized read count of the r-th repeating unit length type in the sample to be tested. r The normalized read count is the number of the r-th repeating unit length type in the control sample.

[0153] In some examples of this application, the aforementioned multiple MSI core sites include at least one of the following: type I core site, type II core site, type III core site, type IV core site, and type V core site.

[0154] The aforementioned type I core sites include:

[0155] Table 1

[0156]

[0157]

[0158] The aforementioned type II core sites include:

[0159] Table 2

[0160] ID CHR START END LEFT RIGHT MSI268 chr2 29449344 29449368 GCCTC GAGGA MSI395 chr2 111884996 111885015 CTTTC CCTCA MSI465 chr3 30691871 30691881 GAAGG GCCTG MSI682 chr4 25680309 25680328 TGTAA ACTGG MSI874 chr6 117642992 117643012 GGCAA GCAGA MSI1097 chr7 55273591 55273604 ATTTG GTATA MSI1153 chr7 116414203 116414214 GAAAT TTCTC MSI2166 chr18 72044569 72044581 TGTGT ATATA MSI2245 chr20 43955023 43955047 ATAGT TTAAT .

[0161] The aforementioned type III core sites include:

[0162] Table 3

[0163]

[0164]

[0165] The aforementioned type IV core sites include:

[0166] Table 4

[0167] ID CHR [[ID= ​ ​ ​ ​ ​ 48030639 48030647 ​ ​ ​ ​ 79970914 79970922 ​ ​ ​ ​ 125505377 125505386 ​ ​ MSI1958 chr17 16029445 16029453 TCTTC CTGTT MSI2243 chr20 43954848 43954857 CTGCT CTATG .

[0168] The aforementioned V-shaped core sites include:

[0169] Table 5

[0170]

[0171]

[0172]

[0173] Module 200 is used to determine the MSI status information of the plurality of MSI core sites respectively using at least one algorithm;

[0174] The algorithm used in this module is the same as that in module 100, and will not be repeated here.

[0175] Module 300 is used to input the MSI state information obtained by the at least one algorithm into multiple trained prediction models respectively, and determine the multiple MSI state prediction results of the at least one algorithm;

[0176] In some examples of this application, the aforementioned prediction model is selected from logistic regression, support vector machine, linear support vector machine, K-nearest neighbor algorithm, decision tree, or ensemble learning methods. Among them, ensemble learning methods include Bagging, Boosting, Stacking, Voting, or Blending.

[0177] The 400 module is used to determine the microsatellite state of the sample to be tested based on multiple MSI state prediction results from the at least one algorithm.

[0178] The aforementioned system can efficiently and accurately detect the microsatellite state of samples with a low tumor ratio, with high detection sensitivity and strong specificity.

[0179] It should be understood that the system embodiments and method embodiments can correspond to each other, and similar descriptions can be found in the method embodiments. To avoid repetition, further details are omitted here. Specifically, Figure 3 The system shown can execute the above-described microsatellite state determination method embodiments, and the operations and / or functions performed by each module in the system correspond to those in the method embodiments. For the sake of brevity, they will not be described in detail here.

[0180] The system of this application embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this application embodiment can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.

[0181] Computer program products, computing devices and computer-readable storage media

[0182] Furthermore, this application proposes a computer program product, a computing device, and a computer-readable storage medium. Based on the aforementioned computer program product, computing device, or computer-readable storage medium, the aforementioned microsatellite state determination method is executed.

[0183] The term "electronic device" is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Computing devices can also refer to various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0184] like Figure 4 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 502 or a computer program loaded from storage unit 508 into RAM (Random Access Memory) 503. The RAM 503 can also store various programs and data required for the operation of the device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0185] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0186] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the microsatellite state determination method. For example, in some embodiments, the microsatellite state determination method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the aforementioned microsatellite state determination method by any other suitable means (e.g., by means of firmware).

[0187] In this application, the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, the computer-readable medium can even be paper or other suitable media on which the aforementioned program can be printed, because the aforementioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or, if necessary, processing in other suitable ways, and then stored in a computer memory. The various computer-readable storage media described in this invention can represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" can include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0188] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0189] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The aforementioned program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0190] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0191] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0192] It should be noted that the features and technical effects described in this article for different aspects can be mutually referenced, and will not be elaborated further here.

[0193] The following examples illustrate this application, but should not be construed as limiting the scope of the subject matter of this application to the following examples. All technologies implemented based on the above content of this application fall within the scope of this application.

[0194] Example 1: External Validation Set Microsatellite Status Assessment (Standard)

[0195] Based on the technical solution of this application, the inventors selected an external validation set of 30 standard samples with different tumor concentrations (5 samples each with tumor cell concentrations of 10%, 20%, 30%, 40%, 50%, and 100%) to perform microsatellite status validation, and the results showed an accuracy rate of 100%.

[0196] Specifically, for positive standards 1, 2, 3, 4, and 5, six gradients of tumor content were prepared (tumor cell content of 10%, 20%, 30%, 40%, 50%, and 100%). The MSI prediction model was validated using the aforementioned MSISensor, MSISensor-pro, Manhattan distance algorithm, a total of 60 core loci for types I, II, III, and IV, and all 15 aforementioned machine learning models.

[0197] In the first algorithm, if the proportion of MSI-H predictions exceeds half among the 15 prediction models, the microsatellite state predicted by the algorithm is considered to be MSI-H; otherwise, the microsatellite state predicted by the algorithm is considered to be MSI-L / MSS.

[0198] In the second algorithm, if the proportion of MSI-H predictions exceeds half among the 15 prediction models, the microsatellite state predicted by this algorithm is considered to be MSI-H; otherwise, the microsatellite state predicted by this algorithm is considered to be MSI-L / MSS.

[0199] In the third algorithm, if the proportion of MSI-H predictions exceeds half of the 15 prediction models, the microsatellite state predicted by this algorithm is considered to be MSI-H; otherwise, the microsatellite state predicted by this algorithm is considered to be MSI-L / MSS.

[0200] The three algorithms predict that more than half of the microsatellites in the test sample are in the MSI-H state, and therefore consider the MSI state of the test sample to be MSI-H.

[0201] Table 8

[0202]

[0203] The final results are shown in Table 8. For the 5 positive (MSI-H) standard samples, even when diluted to 10% tumor content, the model can still accurately determine their status.

[0204] Example 2: External validation set microsatellite status assessment (clinical samples)

[0205] Since most of the standard samples were positive, there was a lack of control over the model's specificity. Therefore, an external validation set 2 (clinical samples, including negative samples) was also used.

[0206] Based on the technical solution of this application, the inventors selected an external validation set of 103 clinical tumor samples (59 positive and 44 negative) that had been verified by a third party for microsatellite status verification. The results showed a sensitivity of 94.92%, a specificity of 97.73%, and an accuracy of 96.12%.

[0207] Specifically, for the 103 clinical samples, the aforementioned MSISensor, MSISensor-pro, Manhattan distance algorithm, the aforementioned 60 core loci of type I, type II, type III, and type IV, and all the aforementioned 15 machine learning models were used to validate the MSI detection model.

[0208] In the first algorithm, if the proportion of MSI-H predictions exceeds half among the 15 prediction models, the microsatellite state predicted by the algorithm is considered to be MSI-H; otherwise, the microsatellite state predicted by the algorithm is considered to be MSI-L / MSS.

[0209] In the second algorithm, if the proportion of MSI-H predictions exceeds half among the 15 prediction models, the microsatellite state predicted by this algorithm is considered to be MSI-H; otherwise, the microsatellite state predicted by this algorithm is considered to be MSI-L / MSS.

[0210] In the third algorithm, if the proportion of MSI-H predictions exceeds half of the 15 prediction models, the microsatellite state predicted by this algorithm is considered to be MSI-H; otherwise, the microsatellite state predicted by this algorithm is considered to be MSI-L / MSS.

[0211] The three algorithms predict that more than half of the microsatellites in the test sample are in the MSI-H state, and therefore consider the MSI state of the test sample to be MSI-H.

[0212] Table 9

[0213]

[0214] The final results are shown in Table 9. For the independent validation set of 103 samples, the following performance was achieved: sensitivity 94.92%, specificity 97.73%, and accuracy 96.12%.

[0215] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0216] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A set of genetic markers characterized in that, comprises: at least one of the sites shown in Table 1: Table 1 。 2. The set of genetic markers according to claim 1, characterized in that, further comprises: at least one of the sites shown in Table 2: Table 2 ; Optionally, further comprises: at least one of the sites shown in Table 3: Table 3 Optionally, further comprises: at least one of the sites shown in Table 4: Table 4 ; Optionally, further comprises: at least one of the sites shown in Table 5: Table 5 3. Use of the genetic marker group of claim 1 or 2 in microsatellite status detection.

4. A microsatellite status determination method, characterized by, comprises: obtaining sequencing data of a sample to be tested, the sequencing data comprising a plurality of MSI core sites; determining MSI status information of the plurality of MSI core sites respectively by at least one algorithm; inputting the MSI status information obtained by the at least one algorithm into a plurality of trained prediction models respectively to determine a plurality of MSI status prediction results of the at least one algorithm; determining the microsatellite status of the sample to be tested based on the plurality of MSI status prediction results of the at least one algorithm.

5. The method of claim 4, wherein, The plurality of MSI core sites are obtained by: obtaining a plurality of candidate MSI sites; determining MSI status of the plurality of candidate MSI sites based on at least one algorithm; inputting the MSI status of the plurality of candidate MSI sites into a trained machine learning model to determine the plurality of MSI core sites; Optionally, the machine learning model is selected from at least one of Lasso and Random Forest.

6. The method according to claim 4 or 5, characterized in that, The algorithm comprises at least one of MSISensor, MSISensor-pro and Manhattan distance algorithm; Optionally, the Manhattan distance is calculated by: where d is the Manhattan distance, R T R is the at least one portion of repeat unit length type for which the corresponding reads exist in the test sample, N T is the at least one portion of repeat unit length type for which the corresponding reads exist in the control sample, r Nris the normalized read count for the test sample that falls into the rth repeat unit length type, r Tris the normalized read count for the control sample that falls into the rth repeat unit length type.

7. The method of claim 4, wherein, The plurality of MSI core sites comprise at least one of type I core sites, type II core sites, type III core sites, type IV core sites and type V core sites.

8. The method of claim 7, wherein, The type I core sites are shown in Table 1; Optionally, the type II core sites are shown in Table 2; Optionally, the type III core sites are shown in Table 3; Optionally, the type IV core sites are shown in Table 4; Optionally, the type V core sites are shown in Table 5.

9. The method of claim 4, wherein, The prediction model is selected from logistic regression, support vector machine, linear support vector machine, K-nearest neighbor algorithm, decision tree or ensemble learning method; Optionally, the ensemble learning method comprises Bagging, Boosting, Stacking, Voting or Blending.

10. The method of claim 5, wherein, The sequencing depth of the MSI core site is not less than 20X.

11. The method according to any one of claims 4-10, characterized in that, The tumor proportion of the sample to be tested is not less than 10%; Optionally, the sample to be tested is selected from a tissue sample or a FFPE sample; Optionally, the tumor comprises hereditary nonpolyposis colorectal cancer, colorectal cancer, gallbladder cancer, ovarian cancer, pancreatic cancer, prostate cancer, liver cancer, kidney cancer, lung cancer, gastric cancer, nasopharyngeal cancer, esophageal cancer or endometrial cancer.

12. A microsatellite status determination system, comprising: comprises: a sequencing data obtaining module for obtaining sequencing data of a sample to be tested, the sequencing data comprising a plurality of MSI core sites; an MSI status information determining module for determining MSI status information of the plurality of MSI core sites respectively by at least one algorithm; The MSI state prediction result determination module is configured to input the MSI state information obtained by the at least one algorithm into a plurality of trained prediction models respectively, and determine a plurality of MSI state prediction results of the at least one algorithm. The microsatellite state determination module is configured to determine the microsatellite state of the sample to be tested based on the plurality of MSI state prediction results of the at least one algorithm.

13. A computer program product, characterised in that, The computer program product comprises computer instructions, and when part or all of the computer instructions run on a computer, the microsatellite state determination method in any one of claims 4-11 is implemented.

14. A computing device, comprising: It comprises: a processor and a memory; the memory is configured to store a computer program; the processor is configured to execute the computer program to implement the microsatellite state determination method in any one of claims 4-11.

15. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions or programs, and when the computer instructions or programs run on a computer, the microsatellite state determination method in any one of claims 4-11 is executed.