Tumor microsatellite instability detection method based on machine learning model
By using a machine learning model-based approach, we screened microsatellite loci with high stability and good discriminative power, extracted multi-dimensional peak distribution features, and constructed a random forest classifier. This approach solved the problem of insufficient sensitivity and specificity in the detection of tumor microsatellite instability in existing technologies, and achieved efficient and reliable detection results.
Patent Information
- Application Number
- CN202511593505.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies lack effective site screening strategies for tumor microsatellite instability detection. Feature extraction methods are limited, classification models are simplistic, and the ability to learn complex feature patterns is lacking, resulting in insufficient detection sensitivity and specificity. Furthermore, there is a lack of systematic quality control and false positive/false negative control mechanisms.
A machine learning-based detection method was adopted. By screening stable and highly discriminative microsatellite loci, multi-dimensional peak distribution features were extracted, a random forest classifier was constructed, and combined with data preprocessing and tumor purity correction, leave-one-out cross-validation was used to evaluate the model performance. A reasonable probability threshold was set to achieve accurate classification.
This improves the sensitivity and specificity of tumor microsatellite instability detection, ensures the stability and reliability of detection results, and solves the problems of low throughput and insufficient feature utilization in existing technologies.
Smart Images

Figure CN121459955A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and tumor diagnostic technology, specifically to a method for detecting tumor microsatellite instability based on a machine learning model. Background Technology
[0002] Microsatellite instability (MSI) is a genetic phenomenon caused by defects in the DNA mismatch repair system in tumor cells. It has important clinical significance in various tumors such as colorectal cancer, gastric cancer, and endometrial cancer. Accurate detection of MSI status is of great value for tumor diagnosis, prognostic assessment, and treatment selection, especially in guiding the use of immunotherapy drugs.
[0003] Currently, MSI detection mainly relies on traditional PCR-electrophoresis methods or bioinformatics analysis methods based on next-generation sequencing (NGS). Traditional methods suffer from problems such as low throughput, high cost, and difficulty in standardization. Existing NGS analysis methods are mainly based on simple read counting or ratio calculation, lacking in-depth analysis of the distribution patterns of microsatellite site sequencing coverage, and most of them use biostatistical models, resulting in insufficient sensitivity and specificity of detection.
[0004] The main problems with existing technologies include: lack of effective marker screening strategies, making it impossible to accurately identify discriminative microsatellite sites; limited feature extraction methods that do not fully utilize peak distribution information in sequencing data; simple classification models, mostly based on statistical models, lacking the ability to learn complex feature patterns; and lack of systematic quality control and false positive / false negative control mechanisms. In view of this, we propose a tumor microsatellite instability detection method based on a machine learning model. Summary of the Invention
[0005] The purpose of this invention is to provide a method for detecting tumor microsatellite instability based on a machine learning model. By analyzing the sequencing coverage distribution pattern of microsatellite loci and combining it with machine learning algorithms, this method achieves accurate classification of MSI-H and MSS samples. To achieve the above objective, this invention adopts the following technical solution:
[0006] A method for detecting tumor microsatellite instability based on machine learning models includes the following steps:
[0007] Step 1: Data Preprocessing
[0008] Read the coverage file output by the MSIsensor, which contains sequencing coverage data of tumor and normal samples for each microsatellite locus at lengths of 1-100 repeat units; standardize the coverage data for each locus and calculate the relative coverage distribution;
[0009] Step 2: Screening for valid loci
[0010] Leukocyte samples from multiple healthy individuals were used as MSS samples, and positive tissue samples from a corresponding number of patients with colorectal cancer, gastric cancer, or endometrial cancer were used as MSI-H samples. Tumor purity correction was performed on MSI-H samples using paired leukocyte data. Loci with high stability in the MSS samples were screened. For each candidate microsatellite locus, the stability of its coverage distribution in the MSS samples was calculated, and the coefficient of variation of the main peak position was calculated. Loci with high differences between the MSS and MSI-H samples were screened. For loci that passed the stability screening in the previous step, a difference analysis between the MSS and MSI-H samples was performed using a permutational multivariate ANOVA method. For the screened marker loci, their MSI-H and MSS characteristic patterns were determined.
[0011] Step 3: Feature Extraction
[0012] The following features are extracted for each marker site: the number of reads in the MSI-H region, the number of reads in the MSS region, the number of peaks, the position of the main peak, the position of the secondary peak, and the proportion of the secondary peak;
[0013] Step 4: Building the Machine Learning Model
[0014] Feature selection was performed using the Boruta algorithm, which determines important features by creating shadow features, training a random forest on an extended dataset, calculating feature importance scores, comparing the importance of the original features with the best shadow features, and iterating. A random forest classifier was constructed for MSI-H / MSS classification. Each decision tree was trained using bootstrap-sampled training data and a randomly selected subset of features. Leave-one-out cross-validation was used to evaluate model performance, and accuracy, AUC, and confusion matrix were calculated. A probability threshold was set based on the leave-one-out cross-validation results. The probability of a sample being MSI-H was calculated, and the microsatellite instability state of the sample was determined by comparing this probability with the threshold.
[0015] Compared with existing technologies, this invention provides a method for detecting tumor microsatellite instability based on a machine learning model, which has the following beneficial effects:
[0016] 1. This tumor microsatellite instability detection method based on machine learning model aims to improve the sensitivity and specificity of microsatellite instability (MSI) detection by screening stable and highly discriminative effective microsatellite loci (first screening for high-stability loci in MSS samples with a coefficient of variation <0.5, and then screening for high-difference loci between MSS and MSI-H samples with an F value >20), extracting multi-dimensional peak distribution features (including the number of reads and peaks in the MSI-H region), and constructing a random forest classification model.
[0017] 2. This tumor microsatellite instability detection method based on machine learning model aims to better address the problems of low throughput and insufficient feature utilization in existing technologies. It achieves high-throughput data processing by performing standardized preprocessing based on the coverage file output by the MSIsensor, while fully mining the peak distribution information in the sequencing coverage distribution as a feature.
[0018] 3. To better ensure the stability and reliability of MSI detection results, this tumor microsatellite instability detection method based on machine learning model is achieved by correcting the tumor purity of MSI-H samples, using leave-one-out cross-validation to evaluate model performance (calculating accuracy, AUC value and constructing confusion matrix), and setting reasonable probability thresholds to determine the MSI status of samples. Attached Figure Description
[0019] Figure 1 This is a flowchart of the overall method of the present invention;
[0020] Figure 2 Heatmap of the distribution of the "chr1 26227608 CAGTC 22[A]GCCTG" site at repeat sites;
[0021] Figure 3 A peak diagram showing the distribution of repeat sites at the locus “chr1 26227608 CAGTC 22[A]GCCTG”;
[0022] Figure 4 PCA diagram of the repeat site for the locus “chr1 26227608 CAGTC 22[A]GCCTG”;
[0023] Figure 5 The image shows the prediction results of LOOCV cross-validation.
[0024] Figure 6 The ROC curve for LOOCV cross-validation. Detailed Implementation
[0025] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0026] Please see Figures 1-6 This invention provides a technical solution: a method for detecting tumor microsatellite instability based on a machine learning model, comprising the following steps:
[0027] Step 1: Data Preprocessing
[0028] The system reads the coverage file output by the MSIsensor, which contains sequencing coverage data for tumor and normal samples at each microsatellite locus with 1-100 repeat units. The coverage data for each locus is standardized, and the relative coverage distribution is calculated by dividing the number of reads per repeat unit by the total number of reads at that locus. Before calculating the relative coverage distribution, the system filters out outliers from the coverage file. The formula for calculating the relative coverage distribution is as follows:
[0029]
[0030] i: Index of the repeating unit length (value range: i = 1, 2, ..., 100, corresponding to 1-100 repeating units); C s,i In sample s, the original sequencing coverage when the repeat unit length is i (s represents tumor sample T or normal sample N, i.e., C T,i For tumor sample coverage, C N,i (For normal sample coverage); C s,total : The total coverage of this microsatellite locus in sample s (the sum of the original coverages across all repeating unit lengths); R s,i : The relative coverage (standardized data) of sample s when the repeating unit length is i;
[0031] Step 2: Screening for valid loci
[0032] Leukocyte samples from at least 20 healthy individuals were used as MSS samples, and positive tissue samples from at least 20 patients with colorectal cancer, gastric cancer, or endometrial cancer with a tumor content of ≥20% (verified using capillary electrophoresis) were used as MSI-H samples. All samples had known tumor content and paired normal tissue (leukocytes). For MSI-H samples, paired leukocyte data was used for tumor purity correction using the following formula:
[0033]
[0034] C corrected Corrected tumor coverage; C tumor : Primary tumor coverage; C normal: Paired normal tissue coverage; p: Tumor purity (0%-100%); Screening for sites with high stability in MSS samples. For each candidate microsatellite site, calculate the stability of its coverage distribution in MSS samples. Let the normal sample set be N = {n1, n2, ..., n}. m For site s, its position in sample n i The main peak is located at the peak. i The formula is:
[0035]
[0036] in Indicates sample n i For the coverage at the j-th position of site s, calculate the coefficient of variation of the main peak position using the following formula:
[0037]
[0038] σ(peaks): Standard deviation of the main peak positions; μ(peaks): Mean of the main peak positions; Screening criteria: CV s <θ cv (usually set θ) cv =0.5); Loci with high differences between MSS and MSI-H samples were screened. For loci that passed the stability screening in the previous step, a difference analysis between MSS and MSI-H samples was performed using the permutation multivariate analysis of variance (PERMANOVA): a sample-locus coverage matrix X was constructed, where X... ij This represents the coverage distribution of sample i at site s; the Bray-Curtis distance matrix is calculated using the following formula:
[0039]
[0040] Perform a PERMANOVA analysis and calculate the F-statistic. The formula is as follows:
[0041]
[0042] SS between Between-group sum of squares; SS within : Within-group sum of squares; g: Number of groups (MSS vs MSI-H, g = 2); n: Total number of samples; Screening criterion: F > θ F (usually set θ) F =20); For the selected marker sites, determine their MSI-H and MSS characteristic patterns; MSI-H pattern determination: Based on the binomial distribution model, determine the unstable region, let the sequencing depth be d, and the minimum depth threshold be d. min If the error rate is p, then the probability threshold for considering a position as unstable is calculated using the following formula:
[0043]
[0044] A suitable p value is obtained through iterative solution. Then, locations with a coverage ratio less than p in all MSS samples are identified as MSI-H regions. MSS pattern determination: The main peak distribution pattern in tumor samples is analyzed. For each tumor sample, the main peak region is identified, and the location with the maximum coverage is found: pos. max =argmax j C j To expand the peak region so that the sum of the peak region coverage reaches 75% of the total coverage, the formula is:
[0045]
[0046] The peak regions of all tumor samples were statistically analyzed, and the locations that appeared more than 25% of the time were defined as MSS pattern regions;
[0047] Step 3: Feature Extraction
[0048] For each marker site, the following features are extracted: MSI-H region read count (msihReads): The normalized number of reads in the MSI-H pattern region of the sample, calculated using the following formula:
[0049]
[0050] Where C normalized (j) represents the normalized coverage of the sample at location j; MSS region reads (mssReads): the normalized reads of the sample in the MSS pattern region, calculated using the following formula:
[0051]
[0052] Peak Number (PeakNumber) is the number of peaks identified by the local maximum detection algorithm, where Peak... j This serves as a marker variable for determining whether position j is a peak using a local maximum detection algorithm.
[0053]
[0054] Main Peak Site: The offset of the highest peak relative to the starting position of the microsatellite site, calculated using the following formula, where pos start This indicates the starting position of the microsatellite locus;
[0055]
[0056] Minor Peak Site: The offset of the second peak relative to the starting position of the microsatellite site. The calculation formula is as follows, where p2 is the position of the second peak and n is the total number of peaks detected.
[0057]
[0058] Minor Peak Ratio: The proportion of the height of the minor peak to the total height of all peaks. The formula is as follows: p i Let i be the position of the i-th peak;
[0059]
[0060] Step 4: Building the Machine Learning Model
[0061] The following features were extracted for each marker site: the number of reads in the MSI-H region, the number of reads in the MSS region, the number of peaks, the position of the main peak, the position of the secondary peak, and the proportion of the secondary peak;
[0062] Step 4: Building the Machine Learning Model
[0063] The Boruta algorithm is used for feature selection. This algorithm determines important features by creating shadow features, training a random forest on an expanded dataset, calculating feature importance scores, comparing the importance of the original features with the best shadow features, and iterating through the process. A random forest classifier is then constructed for MSI-H / MSS classification. Let T be the prediction result of the b-th decision tree. b (x), the number of decision trees B is set to 500, and each decision tree is trained using bootstrap sampled training data and a randomly selected subset of features, as shown in the formula:
[0064]
[0065] Leave-one-out cross-validation was used to evaluate model performance, and accuracy and AUC were calculated to construct a confusion matrix, where accuracy was:
[0066]
[0067] AUC value is calculated based on the ROC curve to evaluate the discriminative ability of the classifier.
[0068] The confusion matrix is:
[0069]
[0070] TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative, respectively.
[0071] The probability threshold is set based on the leave-one-out cross-validation results, typically 50%; using the following formula:
[0072]
[0073] Calculate the probability that the sample is MSI-H, where P b (MSI|x) is the probability that sample x is MSI-H predicted by the b-th decision tree. Based on the comparison between this probability and the threshold, the microsatellite instability state of the sample is determined.
[0074] Example 1
[0075] 1. Experimental Design and Sample Preparation: Sequencing Platform and Data Sources: Targeted sequencing technology was used, covering 1858 potential MSI microsatellite loci, with a sequencing depth ≥1000X; Healthy Control Group: 90 healthy individuals' white blood cell samples (MSS status); Tumor-Positive Group: 25 FFPE slide samples of MSI-H verified by capillary electrophoresis (colorectal cancer, endometrial cancer, gastric cancer), all containing paired tumor tissue and normal tissue (white blood cells); Tumor Purity Requirements: Tumor tissue content ≥20%, which was determined by the percentage of tumor cells as determined by pathologists.
[0076] 2. Data preprocessing: Using the coverage file output by MSIsensor, the relative coverage distribution of each site was calculated, and purity correction was performed on the tumor samples; Quality control: Data from sites with a total coverage depth <100X were filtered out.
[0077] 3. Screening of Effective Loci: Stability Screening (MSS Samples): The coefficient of variation (CV) of the main peak position for each locus in 90 healthy samples was calculated, and loci with CV < 0.5 were retained; Differential Screening (MSI-H vs MSS Samples): PermanoVA analysis was used to calculate the Bray-Curtis distance matrix and F statistic, and loci with F values > 20 were retained; The final list of selected loci is shown in the table below (a total of 162 loci), which includes the following key information: Chromosomal location (e.g., “chr1 110566210 CTGTT 7[AA]GTCTG”, where “chr1 110566210” represents the genomic location, and “CTGTT” represents the chromosomal location). 7[AA]GTCTG” reference genome MSI site has two flanking genes (with a 7 “AA” base repeat in the middle); F value (e.g., 201.475); MSS pattern region (e.g., “7,8” indicates a repeat pattern between 7 and 8 repeat units); MSI-H pattern region (e.g., “5” indicates a repeat pattern between 1 and 5 repeat units), as shown in the table below:
[0078]
[0079]
[0080]
[0081]
[0082]
[0083] Taking the locus “chr126227608CAGTC 22[A]GCCTG” as an example, a heatmap was plotted to show the uniformized coverage of the distribution of 0 to 100 repeating loci. Each row represents an MSI-H or MSS sample. Coverage was calculated for all repeating units, and the coverage reflects the heatmap's intensity. The final determined MSI-H and MSS patterns were marked with blue and red regions, respectively. Figure 2 As shown, MSI-H samples exhibit significant coverage and aggregation within the MSI-H region, while MSS samples are stably distributed within the MSS region, demonstrating clear differentiation. Similarly, a peak plot was created using coverage data, with red representing MSI-H samples and gray representing MSS samples. Figure 3 As shown (red represents MSI-H samples, gray represents MSS samples), there are obvious differences in distribution characteristics. The peaks of the MSI-H samples are more shifted to the left and have a larger number of peaks. PCA dimensionality reduction analysis was performed on all samples, and the PCA plot is shown below. Figure 4 As shown (red represents MSI-H samples, blue represents MSS samples), the MSI-H and MSS samples are clearly clustered, indicating that the features are easy to distinguish and suitable for machine learning models.
[0084] Example 2
[0085] 1. Experimental Design and Sample Preparation: Sequencing Platform and Data Sources: Targeted sequencing technology was used, covering 162 effective MSI microsatellite loci screened in Example 1, with a sequencing depth ≥500X; Negative control group: plasma samples from 113 healthy individuals and leukocyte samples from 60 individuals; Positive group: MSI-H positive colorectal cancer patient samples verified by capillary electrophoresis, including 25 FFPE slide samples and 2 plasma samples; Sample quality control standards: Sequencing quality: Q30 ≥80%, target region 500X coverage ≥95%.
[0086] 2. Data preprocessing and feature extraction: Feature extraction method (same as in Example 1, extracting the following features for each sample for 162 valid sites): number of reads in the MSI-H region (msihReads), number of reads in the MSS region (mssReads), number of peaks (PeakNumber), main peak position (MainPeakSite), minor peak position (MinorPeakSite), and minor peak ratio (MinorPeakRatio).
[0087] 3. Model Training and Validation: Feature matrix dimension: 200 samples × 972 features (162 sites × 6 features / site); Feature selection: Boruta algorithm was used (parameter settings: ntree = 500, maxRuns = 100), and 86 significant features were retained in the end; Random forest model parameters: number of decision trees is 500, minimum number of samples per node is 5, and class weights are balanced.
[0088] 4. Leave-one-out cross-validation (LOOCV) results: Leave-one-out cross-validation was used on all 200 samples to determine their detection performance. The confusion matrix of MSI-H is shown in the table below:
[0089] Predicted positive The prediction was negative. The actual result was positive. 27 0 The actual result was negative. 0 173
[0090] Accuracy: (27 + 173) / 200 = 100.00%
[0091] Specificity: 173 / (173+0) = 100.00%
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098] Using a 50% probability threshold, all samples were correctly predicted for classification, with no misclassifications. The predicted probabilities were sorted from low to high, with samples marked as MSS in blue and MSI-H in orange. A scatter plot was then created. Figure 5 As shown, the predicted probabilities of the two types of samples are clearly distinguishable, such as... Figure 6As shown in the figure, the ROC curves of the prediction results for the two types of samples are as follows. The AUC is 1, indicating that they are completely distinguishable. This suggests that the proposed scheme does not exhibit significant overfitting and has excellent prediction performance. Consequently, it can better improve the sensitivity and specificity of microsatellite instability (MSI) detection, thereby better addressing the problems of low throughput and insufficient feature utilization in existing technologies. This also better ensures the stability and reliability of MSI detection results. The above is only used to illustrate the technical solution of the present invention and is not intended to limit it. Any other modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention, as long as they do not depart from the spirit and scope of the technical solution of the present invention, should be covered within the scope of the claims of the present invention.
Claims
1. A method for detecting tumor microsatellite instability based on a machine learning model, characterized by: The method comprises the following steps: Step one, data preprocessing Read the coverage file output by MSI sensor, which contains the sequencing coverage data of tumor samples and normal samples at 1-100 repeat unit lengths for each microsatellite site; standardize the coverage data for each site and calculate the relative coverage distribution; Step two, screening effective sites Use the white blood cell samples of multiple healthy people as MSS samples, and the positive tissue samples of corresponding cases of colorectal cancer, gastric cancer or endometrial cancer patients as MSI-H samples; use paired white blood cell data to correct the tumor purity for MSI-H samples; screen sites with high stability in MSS samples, calculate the stability of the coverage distribution of each candidate microsatellite site in MSS samples, and calculate the coefficient of variation of the main peak position; screen sites with high difference between MSS samples and MSI-H samples, perform difference analysis between MSS samples and MSI-H samples for the sites that pass the stability screening in the previous step, and use the permutation multivariate variance analysis method; determine the MSI-H and MSS characteristic patterns of the screened marker sites; Step three, feature extraction Extract the following features for each Marker site: MSI-H region reads, MSS region reads, peak number, main peak position, secondary peak position, and secondary peak ratio; Step four, construction of machine learning model Use the Boruta algorithm for feature selection, which determines important features by creating shadow features, training random forests on an expanded dataset, calculating feature importance scores, comparing original features with best shadow feature importance, and iterating execution; build a random forest classifier for MSI-H / MSS classification, train each decision tree using bootstrap sampled training data and a randomly selected feature subset, evaluate model performance using leave-one-out cross-validation, calculate accuracy, AUC value, and construct a confusion matrix; set a probability threshold based on the leave-one-out cross-validation result; calculate the probability that the sample is MSI-H, and determine the microsatellite instability status of the sample based on the comparison result of the probability and the threshold.
2. The tumor microsatellite instability detection method based on a machine learning model according to claim 1, characterized in that: In step one, before calculating the relative coverage distribution, the operation of filtering abnormal data in the coverage file is also included.
3. The tumor microsatellite instability detection method based on a machine learning model according to claim 1, characterized in that: In step two, the tumor content of the positive tissue sample is more than 20%.
4. The tumor microsatellite instability detection method based on a machine learning model according to claim 1, characterized in that: The correction formula in step two is: C corrected : Corrected tumor coverage; C tumor : Original tumor coverage; C normal : Paired normal tissue coverage; p: tumor purity (0%-100%). 5.The method of claim 1, wherein: The positive tissue sample in step two has been verified using capillary electrophoresis, and the tumor content of all samples is known, and all have paired normal tissues. 6.The method of claim 1, wherein: The stability method of the calculation in the step two, which covers the distribution in the MSS sample, is that the normal sample set is N = {n1, n2,..., n m}, and for the site s, the main peak position in the sample n i is peak i , and the formula is: where C ni,s (j) represents sample n i Coverage at the jth position of site s; The formula for calculating the coefficient of variation of the main peak position is: σ(peaks): standard deviation of the main peak position; μ(peaks): mean of the main peak position; filter condition: CV s <θ cv (usually set θ cv = 0.5); The method of substitutional multivariate analysis is: constructing a sample-site coverage matrix X, where X ij represents the coverage distribution of sample i at site s; Calculate the Bray-Curtis distance matrix, the formula is: Perform PERMANOVA analysis and calculate the F statistic, the formula is: SS between : Between groups sum of squares; SS within : Within groups sum of squares; g: number of groups (MSS vs MSI-H, g=2); n: total number of samples; screening condition: F > θ F (θ is usually set to 20). F = 20).
7. The method of claim 6, wherein the machine learning model is trained using a dataset comprising a plurality of tumor samples, wherein each tumor sample is associated with a microsatellite instability (MSI) status. The step two MSI-H pattern determination determines the unstable region based on a binomial distribution model, assuming that the sequencing depth is d, the minimum depth threshold is d min , the error rate is p, and the probability threshold that a certain position is considered unstable is calculated by the following formula: Obtain the appropriate p value by iterative solution, then identify the positions with coverage ratio less than p in all MSS samples as MSI-H regions; MSS pattern determination: analyze the distribution pattern of the main peaks in the tumor samples, for each tumor sample, identify the main peak region, find the position of the maximum coverage: pos max = argmax j C j , extend the peak region so that the sum of the peak region coverage reaches 75% of the total coverage, count all the peak regions of the tumor samples, and define the positions with a frequency of more than 25% as the MSS pattern region. 8.The method of claim 1, wherein: The features extracted in step three are specifically: MSI-H region reads number is the normalized reads number of the sample in MSI-H pattern region; MSS region reads number is the normalized reads number of the sample in MSS pattern region; Peak number is the number of peaks identified by local maximum detection algorithm; Main peak position is the offset of the highest peak relative to the starting position of the microsatellite site; Secondary peak position is the offset of the second highest peak relative to the starting position of the microsatellite site; Secondary peak ratio is the proportion of the secondary peak height to the total height of all peaks.
9. The tumor microsatellite instability detection method based on a machine learning model according to claim 1, characterized in that: The step four in the method is: Creating shadow features, randomly permuting each original feature; training random forest on the expanded dataset (original features + shadow features); calculating importance score for each feature; Using statistical test to compare the importance of original features and best shadow features; iteratively performing until all features are confirmed or rejected.
10. The method of claim 9, wherein the machine learning model is trained using a dataset comprising a plurality of tumor samples, wherein each tumor sample is associated with a microsatellite instability (MSI) status. The accuracy in the step four is: AUC value is: calculating AUC value based on ROC curve, evaluating the discriminant ability of the classifier; Confusion matrix is: Where TP, TN, FP, FN represent true positive, true negative, false positive, false negative respectively.