Method and device for predicting state of microsatellite

By acquiring and analyzing the relevant characteristics of repeating unit length distribution of microsatellite characteristic sites and using machine learning models to predict, the accuracy and robustness of microsatellite instability detection in the prior art are solved, and efficient and accurate microsatellite state prediction is achieved.

CN119943146APending Publication Date: 2025-05-06BGI GENOMICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311447453.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems with accuracy, sensitivity and robustness in detecting microsatellite instability (MSI) states, especially immunohistochemistry, multiple fluorescence PCR-capillary electrophoresis and high-throughput sequencing methods.

Method used

By obtaining sequencing data of the sample to be tested, the repeat unit length distribution-related features of the microsatellite feature sites are determined based on multiple read segments, and these features are input into the trained machine learning model to predict the microsatellite state of the sample.

Benefits of technology

It realizes efficient and accurate prediction of microsatellite status, improves analysis efficiency and reliability, reduces the burden of manual data processing, and enhances the robustness of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943146A_ABST
    Figure CN119943146A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for predicting the state of a microsatellite. The method comprises the following steps: acquiring sequencing data of a sample, wherein the sequencing data comprises a plurality of reads from at least one predetermined microsatellite feature site; determining repetition unit length distribution related features of at least one predetermined microsatellite feature site based on the plurality of reads; and inputting the repetitive unit length distribution related features into a trained machine learning model to predict the microsatellite state of the sample. By utilizing the method, the microsatellite state of the sample can be efficiently and accurately judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gene sequencing, and more specifically to a method and device for predicting microsatellite status, and more specifically to a method for predicting microsatellite status, a machine learning model training method, a device for predicting microsatellite status, a machine learning model training device, a computer program product, an electronic device, and a computer-readable storage medium. Background Art

[0002] Microsatellite instability (MSI) is a molecular feature closely related to tumorigenesis and has important clinical significance in a variety of solid tumors. The formation of MSI is mainly due to abnormalities in the mismatch repair system (MMR) during DNA replication. During DNA replication, DNA polymerase is prone to "slippage" when encountering short tandem repeat sequences, resulting in the insertion or deletion of nucleotides in microsatellite loci. Under normal circumstances, the mismatch repair system can recognize this instability and repair it, but if the MMR gene mutates or is hypermethylated, it will lead to the loss of mismatch repair function, and the spontaneous high-frequency length variation in the microsatellite cannot be repaired in time, thus forming microsatellite instability.

[0003] MSI is closely related to the occurrence of many tumors, especially Lynch syndrome. About 90% of Lynch syndromes show microsatellite instability. MSI detection can be used to diagnose and evaluate the prognosis of various solid tumors, including colorectal cancer and endometrial cancer. Studies have shown that patients with colorectal cancer with high microsatellite instability (MSI-H) have a better survival advantage than patients with microsatellite stability (MSS), especially in stage II / III colorectal cancer. The overall survival and disease-free survival of MSI-H patients are significantly prolonged. In addition, patients with advanced solid tumors with MSI-H usually have significant therapeutic effects on immune checkpoint inhibitor treatment.

[0004] Currently, MSI detection has been widely used in the clinical practice of solid tumors such as colorectal cancer as an important molecular marker for prognosis assessment and treatment planning. However, the means of detecting MSI status still have defects.

[0005] Therefore, the methods for detecting MSI still need to be improved. Summary of the invention

[0006] This application is completed by the inventor based on the following findings:

[0007] The current methods for detecting MSI status include:

[0008] 1) Immunohistochemistry (IHC): The expression of DNA mismatch repair proteins (MLH1, MSH2, MSH6 and PMS2) is used to reflect the microsatellite instability (MSI) status. The accuracy of this detection method is affected by subjective factors of the interpreter, antibody quality and experimental factors, and the repeatability of the results cannot be guaranteed.

[0009] 2) Multiplex fluorescence PCR-capillary electrophoresis: PCR is used to simultaneously amplify the microsatellite loci of tumor tissue and normal tissue from the same individual, and then the amplified products are analyzed for DNA fragments using capillary electrophoresis. The gene maps of tumor tissue and normal tissue are compared. If the peaks of the five microsatellite detection sites of tumor tissue are displaced compared to normal tissue, the sample is judged to be MSI-H. However, the PCR detection method has a low throughput, contains fewer sites, and has a low sensitivity.

[0010] 3) High-throughput sequencing (NGS) detection method: MSI is detected using high-throughput sequencing (such as whole genome sequencing, whole exome sequencing or targeted gene sequencing). However, this method also has some limitations, including a small number of loci, low accuracy, low sensitivity, and lack of evaluation of algorithm robustness.

[0011] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, one purpose of the present application is to propose a means for accurately predicting the state of microsatellites.

[0012] In the first aspect of the present application, the present application proposes a method for predicting microsatellite status. According to an embodiment of the present application, the method comprises: obtaining sequencing data of a sample to be tested, wherein the sequencing data comprises a plurality of reads from at least one predetermined microsatellite characteristic site; determining a repeat unit length distribution-related feature of at least one predetermined microsatellite characteristic site based on the plurality of reads; and inputting the repeat unit length distribution-related feature into a trained machine learning model to predict the microsatellite status of the sample.

[0013] According to the embodiments of the present application, the method can quickly obtain relevant information of microsatellite characteristic sites by obtaining and analyzing the sequencing data of the sample. By determining the relevant characteristics of the distribution of repeat unit lengths based on multiple reads, combined with a trained machine learning model, the microsatellite state of the sample can be efficiently and accurately predicted. In addition, the trained machine learning model is used to realize automatic prediction of the microsatellite state. There is no need for manual cumbersome data processing and analysis, which reduces the workload of researchers and improves analysis efficiency and reliability.

[0014] In the second aspect of the present application, the present application proposes a method for training a machine learning model. According to an embodiment of the present application, the method includes: obtaining training set and test set sequencing data, the training set and test set sequencing data have known microsatellite status results, the training set and test set sequencing data include multiple reads from at least one predetermined microsatellite characteristic site, and the training set includes a sample set to be tested and a control sample set; based on the multiple reads, determining the repeat unit length distribution-related features of at least one of the predetermined microsatellite characteristic sites; inputting the repeat unit length distribution-related features into the machine learning model, using the known microsatellite status results as labels, and performing supervised training on the machine learning model to obtain a trained machine learning model, and the machine learning model is used to predict the microsatellite status.

[0015] According to an embodiment of the present application, the machine learning model uses known microsatellite status results as labels to improve the accuracy and generalization ability of the model. By conducting supervised learning on the training set, the model can better understand and capture the characteristics related to the distribution of repeat unit lengths of microsatellite characteristic sites, thereby improving the prediction performance on unknown data. By analyzing multiple microsatellite characteristic sites, the model can simultaneously consider the impact of multiple factors on the microsatellite status, which helps to improve the complexity and accuracy of the model, so that it can better cope with the diversity and complexity in practical application scenarios.

[0016] In the third aspect of the present application, the present application proposes a device for predicting the state of a microsatellite. According to an embodiment of the present application, the device includes: a sequencing data acquisition unit, which is used to acquire sequencing data of a sample to be tested, wherein the sequencing data includes multiple reads from at least one predetermined microsatellite characteristic site; a feature acquisition unit, which is used to determine the repeat unit length distribution-related features of the predetermined microsatellite characteristic site based on the multiple reads; and a prediction unit, which is used to input the repeat unit length distribution-related features into a trained machine learning model to predict the microsatellite state of the sample.

[0017] According to the embodiments of the present application, the device can automatically, highly sensitively and accurately predict the state of microsatellites, and overcome the problem of poor robustness of existing detection methods by quality control of microsatellite characteristic sites.

[0018] In the fourth aspect of the present application, the present application proposes a machine learning model training device. According to an embodiment of the present application, the device includes: a training data acquisition unit, which is used to acquire training set and test set sequencing data, wherein the training set and test set sequencing data have known microsatellite status results, and the training set and test set sequencing data contain multiple microsatellite characteristic sites; a model feature acquisition unit, which is used to determine the repeat unit length distribution-related features of the predetermined microsatellite characteristic sites based on the training set and test set sequencing data; and a training unit, which is used to input the repeat unit length distribution-related features into the machine learning model, use the known microsatellite status results as labels, and supervise the machine learning model to obtain a trained machine learning model, wherein the machine learning model is used to predict the microsatellite status; wherein the machine learning model is obtained by training the method described in the second aspect of the present application.

[0019] According to the embodiments of the present application, the device acquires a large amount of data and adopts a supervised training method to obtain a model that can effectively improve the prediction accuracy of the microsatellite status.

[0020] In a fifth aspect of the present application, the present application proposes a computer program product. According to an embodiment of the present application, the computer program product includes computer instructions, and when part or all of the computer instructions are executed on a computer, the methods described in the first and second aspects of the present application are implemented.

[0021] In the sixth aspect of the present application, the present application proposes an electronic device. According to an embodiment of the present application, the electronic device includes: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the first and second aspects of the present application.

[0022] In a seventh aspect of the present application, the present application proposes a computer-readable storage medium containing a computer program. According to an embodiment of the present application, when the computer program is executed by one or more processors, the method described in the first and second aspects of the present application is implemented.

[0023] It should be noted that in the present application, the features and advantages described for the methods, devices and systems in any of the above aspects, implementation modes or examples are also applicable to other aspects and will not be repeated here.

[0024] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0026] Figure 1 is a flow chart of a method for predicting microsatellite status according to an embodiment of the present application;

[0027] Figure 2 is a schematic diagram of a random forest model training process according to an embodiment of the present application;

[0028] Figure 3 is a schematic diagram of the distribution of MSI-H support numbers of 100 samples according to an embodiment of the present application;

[0029] Figure 4 is a schematic diagram of the accuracy prediction of results of different numbers of sites according to an embodiment of the present application;

[0030] Figure 5 is a schematic diagram of a device for predicting microsatellite status according to an embodiment of the present application; and

[0031] Figure 6 is a schematic diagram of a machine learning model training device according to an embodiment of the present application;

[0032] Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present application.

[0033] The above drawings are only schematic and have no limiting effect. In the drawings, for illustrative purposes, the sizes of some components may be exaggerated and not drawn to scale. The sizes and relative sizes do not necessarily correspond to the actual restoration when the present invention is implemented. DETAILED DESCRIPTION

[0034] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0035] As used herein, the singular forms "a", "an", etc. include plural referents (one or more). "A set" or "a plurality" refers to two or more.

[0036] In this document, "include" or "comprising" is an open expression, including the content or circumstances specified or exemplified thereafter, and also including content or circumstances that are applicable or consistent with the stated circumstances but are not specifically listed.

[0037] In this article, the term "reads" refers to the sequence measured by sequencing. "Sequencing" refers to the determination of the order of bases in the primary structure of nucleic acid molecules, which can be achieved by first-generation, second-generation, and third-generation sequencing. Sequencing can be performed by a sequencing platform. In some examples, the sequencing platform that can be used includes but is not limited to the Chromium platform of 10x Genomics, the DNBelab C Series platform of MGI, and the Rhapsody platform of BD Biosciences. TM Platform, PacBio's REVIO platform, OxfordNanopore Technologies' PromethION platform; the sequencing method can be selected as single-end sequencing, double-end sequencing, or a sequencing method supported by the selected automated sequencing platform, etc.

[0038] In this article, the term "microsatellite (MS)" refers to a short tandem repeat sequence in the genome, the number of which is highly variable between individuals.

[0039] Herein, the term “microsatellite instability (MSI)” refers to the phenomenon that microsatellites in tumor cells undergo abnormal changes in length due to insertion or deletion of repetitive sequences compared with normal cells.

[0040] In this article, the term "MSI status" refers to the presence of MSI or unstable microsatellites (sites), i.e., the number of repetitive DNA nucleotide units in cell colonies (clonal) or somatic cells (somatic) in the microsatellites changes. In this application, the prediction of MSI status includes MSS (microsatellite stability) or MSI-H (microsatellite high instability). MSS refers to the situation where there is no functional defect in DNA mismatch repair, and the number of repeat fragments in the microsatellite site is not significantly different between tumors and normal cells, and MSI-H refers to the situation where the number of repeat fragments present in the microsatellite site is significantly different from the number of repeat fragments in the normal cell DNA.

[0041] In this article, term " crest (peak) " refers to the distribution pattern of microsatellite in microsatellite loci.By analyzing sequencing data, the counting of allele repeat length in each microsatellite loci can be obtained, which is called crest width (peak width), and the reading number of the most common allele observed in this site, which is called crest height (peak height).And compared with reference genome, the position difference of the crest height of a certain microsatellite locus in tumor tissue is called crest position (peak position).In certain embodiments of the present invention, crest width, crest height or crest position can be used as an index of the microsatellite instability feature (MSI feature) estimating the unstable state of microsatellite.

[0042] In this article, the term "microsatellite characteristic site" refers to a specific site in the genome with a certain repetitive sequence. The "predetermined microsatellite characteristic site" refers to a predetermined microsatellite characteristic site. In this application, the predetermined microsatellite characteristic site is obtained by a microsatellite (MS) characteristic site acquisition method.

[0043] In this paper, the term “repeat unit length distribution related features” is characterized by the Manhattan distance.

[0044] In this article, the term "reference sequence (reference, ref)" is a determined sequence, which can be a DNA and / or RNA sequence pre-determined and assembled by oneself, or a DNA and / or RNA sequence determined and disclosed by others, and can be any reference template in the biological category to which the sample source individual / target individual belongs, for example, all or at least a part of the disclosed genome assembly sequence of the same biological category. If the sample source individual or the target individual is human, its genome reference sequence (also referred to as a reference genome or reference chromosome group) can select the human reference genome provided by the UCSC, NCBI or ENSEMBL database, such as hg19, hg38, GRCh36, GRCh37, GRCh38, etc., and those skilled in the art can understand the corresponding relationship of the above-mentioned reference genome versions through the description of the database and select the version to be used. Further, a resource library containing more reference sequences can also be pre-configured, for example, before comparison, a sequence that is closer or more characteristic in a certain aspect is selected or determined and assembled according to factors such as the gender, race, and region of the target individual as a reference sequence, which helps to obtain more accurate sequence analysis results later.

[0045] The so-called reference sequence can be constructed when testing the target sample, or it can be pre-constructed and stored and called when the prepared sample is tested. In some embodiments, the sample to be tested is from the human body, and the reference sequence is a human reference genome or a human autosome group.

[0046] In this article, the term "machine learning model" refers to a computing model or algorithm that can automatically perform tasks such as prediction, classification, recognition or decision-making by learning and analyzing input data. The learning process of the model is based on statistical principles and data pattern recognition, and uses training data sets to adjust parameters and optimize models to improve its prediction or reasoning ability. Machine learning models can use various algorithms and techniques, such as neural networks, support vector machines, decision trees, random forests, deep learning, etc. These models can be trained and optimized by supervised learning, unsupervised learning or reinforcement learning. In practical applications, machine learning models can be used in various fields, such as natural language processing, image recognition, pattern recognition, data mining, recommendation systems, predictive analysis, etc. It has important application potential in processing large-scale data, automated decision-making and intelligent systems. However, it should be noted that in the specific application of machine learning models, it is necessary to conduct in-depth research on the features and models used for prediction in order to obtain a more satisfactory prediction effect, otherwise various problems will arise, such as overfitting problems, underfitting problems, and poor generalization problems. After years of research and combined with research experience, the inventors of the present application unexpectedly discovered a set of repeat unit length distribution-related characteristics of microsatellite characteristic sites, which can be combined with machine learning models to predict microsatellite status.

[0047] This paper selects random forest as the machine learning model for supervised training. The random forest model is trained by using the training set data, and the trained random forest model is tested by using the test set. Among them, "Training Set" refers to the data set used to train the machine learning model. During the training phase, the machine learning model uses the samples in the training set to learn and adjust the parameters of the model to minimize the error between the predicted results and the true labels. By constantly trying different parameters and algorithms, the model gradually learns the rules and characteristics between samples, thereby obtaining more accurate prediction capabilities. "Test Set" is a data set used to evaluate the performance of the machine learning model. Moreover, the samples in the test set have not been used in the training phase of the model, which ensures that the data in the test set is unknown to the model. After the training is completed, the test set is used to evaluate the performance of the model, that is, the model predicts the samples in the test set and compares them with the true labels. By comparing the predicted results with the true labels, the performance of the model on unknown data is evaluated to determine its generalization ability and prediction accuracy.

[0048] In this article, MSI Score (microsatellite instability score) is an indicator used to quantify the degree of microsatellite instability (MSI) in tumor cells. In molecular biology and cancer research, MSI refers to the abnormal amplification or reduction of the DNA microsatellite region, usually due to defects in the DNA repair mechanism. In tumor research, MSI is often used to assess the genetic instability of tumors, which may be related to the occurrence and development of tumors. In the technical solution of the present application, MSI Score is a numerical value obtained by analyzing the difference in the length of the repetitive sequence at a specific microsatellite site between tumor cells and normal control cells. This score represents the number or frequency of mismatch events observed at a given microsatellite site. A higher MSI Score indicates that the microsatellite instability of tumor cells is higher, that is, there are more microsatellite sites that have undergone abnormal amplification or reduction.

[0049] The existing MSI detection technology methods have the following problems:

[0050] The immunohistochemistry (IHC) method only detects four MMR proteins, has low sensitivity and is easily affected by problems with manual interpretation. IHC has strict requirements for sample processing, and the results need to be combined with machine (automated) methods and manual interpretation, and rely on good quality control and internal references. Manual interpretation is easily affected by subjective factors and may lead to misjudgment. In addition, in addition to the four main mismatch repair proteins, other missing mismatch repair proteins can also affect the stability of microsatellites. Due to problems such as accuracy, tumor heterogeneity and subjectivity of interpretation, the accuracy of IHC positive results is low. One study showed that the accuracy of positive (dMMR) patients was only 43.68%.

[0051] The MSI-PCR method detects fewer sites for MSI, has low sensitivity and limited throughput. This method manually interprets microsatellite sites in normal control samples and tumor samples to determine the stability of breakpoints to determine the microsatellite status, which may result in subjective judgment errors.

[0052] The NGS methods currently on the market have a small number of microsatellite loci when detecting MSI, and cannot effectively detect the microsatellite loci included in the panel, and their sensitivity and success rate are low. For example, the VisualMSI software may fail during operation, resulting in an invalid MSI result status, with a failure rate of 0.47%. The robustness of the MSI detection method is poor. If there are 100 microsatellite loci in the method and several of them fail quality control, then the MSI result of the sample cannot be successfully determined.

[0053] In order to increase the accuracy of the detection results and improve the robustness of the MSI detection method. In a first aspect of the present application, the present application proposes a method for predicting microsatellite status.

[0054] The technical solution of the present application is described in detail below with reference to the accompanying drawings. Figure 1 According to an embodiment of the present application, the method includes:

[0055] S100: Determine the characteristics of repeat unit length distribution based on read segments

[0056] In this step, the sequencing data of the sample to be tested is first obtained. There is no particular restriction on the source and sequencing platform of the sequencing data, and it can be any common first-generation, second-generation or third-generation sequencing platform. These sequencing data are valid sequencing data that are aligned to the reference sequence and remove duplicate reads and indel (insertion and deletion) local re-aligned reads. And the sequencing data includes multiple reads from at least one predetermined microsatellite characteristic site, and the sequencing depth at the predetermined microsatellite characteristic site is not less than 20X, and 30X, 50X, 80X or 100X can be selected.

[0057] In some examples, the sample can be from a cell line, a biopsy, a primary tissue, a frozen tissue, a formalin fixed paraffin embedded (FFPE) tissue, a liquid biopsy, blood, serum, plasma, buffy coat, body fluid, visceral fluid, ascites, paracentesis, cerebrospinal fluid, saliva, urine, tears, semen, vaginal secretions, aspirate, lavage, buccal swab, circulating tumor cells (CTC), cell free DNA (cfDNA), circulating tumor DNA (ctDNA), DNA, RNA, nucleic acid, purified nucleic acid, purified DNA, or purified RNA.

[0058] Determine the repeat unit length distribution related characteristics of at least one predetermined microsatellite characteristic site based on multiple reads. The repeat unit length distribution related characteristics are determined based on the repeat unit length distribution difference comparison between the sample to be tested and the control sample. Specifically, find the repeat unit in the sequencing read, and then perform statistics and distribution analysis on these repeat units according to the length. Compare the repeat unit length distribution of the sample and the control sample, determine the sites with differences, and then determine the repeat unit length distribution related characteristics.

[0059] In some examples, the so-called repeat unit length distribution-related characteristic is determined by the following formula:

[0060]

[0061] Where d is the Manhattan distance, R T is the type of at least a portion of the repeating unit length of the corresponding read segment in the sample to be tested, R N is the length type of at least a portion of the repeating units in the corresponding read segment in the control sample, T r is the normalized read count of the rth repeat unit length type in the sample to be tested, N r is the normalized read count that falls into the rth repeat unit length type in the control sample.

[0062] The so-called “at least a part of the repeating unit length types of the corresponding read segments exist in the sample to be tested” means that in the sample to be tested, there are certain read segments in which the repeating unit lengths belong to a specific type.

[0063] Specifically, it is assumed that in a certain sample to be tested, there are the following two types of repeat unit lengths: 3 bases and 5 bases.

[0064] For a repeating unit type of 3 bases, sequence fragments similar to "AAA" or "GAG" may be observed in the reads. These reads have repeating units of the same length, i.e., 3 bases. By analyzing these reads, it can be determined that a sequence with a repeating unit length type of 3 bases exists in the sample.

[0065] For a repeat unit type of 5 bases, sequence fragments similar to "TTTTT" or "CGCGC" may be observed in the reads. These reads have repeat units of the same length, that is, 5 bases. By analyzing these reads, it can be determined that there are sequences with a repeat unit length type of 5 bases in the sample.

[0066] Among them, the type of at least a portion of the repeat unit length of the corresponding read segment in the control sample is the same as explained above and will not be repeated here.

[0067] The so-called RT ∪R N The method includes the union of all repeating unit types of the test sample and the control sample or the union of some repeating unit types.

[0068] The so-called test sample is from a subject suspected of having a tumor, and the so-called control sample is from a non-tumor tissue. In some examples, the relevant distribution information of the control sample can be obtained in a relevant parallel experiment or can be obtained by pre-sequencing.

[0069] In some examples of the present application, the method for predicting microsatellite status is not limited to a specific type of tumor sample. In some specific examples, the tumors include hereditary non-polyposis colorectal cancer, colorectal cancer, gallbladder cancer, ovarian cancer, pancreatic cancer, prostate cancer, liver cancer, kidney cancer, lung cancer, gastric cancer, nasopharyngeal cancer, esophageal cancer or endometrial cancer, etc.

[0070] In some specific embodiments, the method for obtaining microsatellite (MS) characteristic sites is as follows:

[0071] 1. Obtain all MS sites on the hg19 version of the genome to form the original MS site collection;

[0072] 2. Intersect the original MS site collection with the detection range of Panel (Pan-Cancer Detection Kit IDT_XGEN_CUSTOM_TARGETS_CAPTURE_PROBE) to obtain the MS sites within the intersection interval; perform sequencing depth evaluation on the MS sites within the intersection interval (site depth is not less than 20X), and after evaluation and testing of samples (at least 100 cases), filter out sites that do not meet the depth requirements to obtain the filtered MS site set;

[0073] 3. Perform feature preprocessing on the filtered MS sites, remove feature sites with a variance of 0 (sites with the same Manhattan distance d in all samples), and then classify each feature site after preprocessing by MSI-H and MSS, and select sites with an AUC value greater than 0.65 to retain the MS feature sites.

[0074] In some examples, based on the above method for obtaining microsatellite characteristic sites, a total of 1137 characteristic sites were obtained. In order to reduce the complexity of the model, improve the generalization ability of the model, speed up the training and prediction speed, and avoid the occurrence of overfitting, the inventors sorted the importance of the features, deleted the relevant features based on the importance sorting results, and finally obtained the important microsatellite characteristic sites as shown in Table 1. In some specific embodiments, the prediction of the microsatellite state can also be completed by selecting at least one of the following:

[0075] Table 1

[0076]

[0077]

[0078] In some specific embodiments, the microsatellite status is detected by the following steps:

[0079] 1) Obtain the original Fastq data (raw data) of tumor samples and sample control tissues through high-throughput sequencing (NGS);

[0080] 2) Align the raw data to the reference genome, remove duplicate reads and local realignment of Indels to obtain a BAM file;

[0081] 3) By counting the differences in the length distribution of the repeating units of the tumor samples and the control samples at each MS site, the Manhattan distance between the tumor samples and the control samples was obtained as the MSI score of a single MS site (the calculation formula is the same as above), that is, the MSI scores of 1137 MS characteristic sites were obtained;

[0082] 4) Quality control of site depth and number of sites for 1137 characteristic sites;

[0083] 5) The MSI Score values ​​of the 1137 characteristic loci were used as sample input and input into the MSI determination random forest classification model to determine the MSI status of the sample. The Support_MSI-H result value output by the model result was used to determine the microsatellite instability status of the sample.

[0084] S200: Input to trained machine learning model to predict microsatellite instability

[0085] In this step, the microsatellite status of the sample is predicted by inputting the repeat unit length distribution related features obtained in step S100 into a trained machine learning model.

[0086] The so-called machine learning model is selected from at least one of a decision tree, logistic regression, naive Bayes, random forest, and a neural network. In some examples, a random forest is selected as the machine learning model.

[0087] In the second aspect of the present application, the present application proposes a method for training a machine learning model. The method comprises: obtaining training set and test set sequencing data, wherein the training set and test set sequencing data have known microsatellite status results, wherein the training set and test set sequencing data include multiple reads from at least one predetermined microsatellite characteristic site, wherein the training set includes a sample set to be tested and a control sample set; based on the multiple reads, determining the repeat unit length distribution-related features of at least one predetermined microsatellite characteristic site; inputting the repeat unit length distribution-related features into a machine learning model, using the known microsatellite status results as labels, and conducting supervised training on the machine learning model to obtain a trained machine learning model, wherein the machine learning model is used to predict the microsatellite status.

[0088] In some embodiments, the repeat unit length distribution-related characteristics are determined by the following formula:

[0089]

[0090] Where d is the Manhattan distance, R T is the type of at least a portion of the repeating unit length of the corresponding read segment in the sample to be tested, R N is the length type of at least a portion of the repeating units in the corresponding read segment in the control sample, T r is the normalized read count of the rth repeat unit length type in the sample to be tested, N r is the normalized read count that falls into the rth repeat unit length type in the control sample.

[0091] For ease of understanding, combined Figure 2 ,The model training process is explained in detail below.

[0092] 1) The samples used to build the model used PCR verification results as labels and the MSI-Score of 1137 MS loci as features. The samples were divided into 7:3 ratios, with 70% of the samples used as training sets for model construction and 30% of the samples used as test sets for the model;

[0093] 2) After model parameter adjustment, the positive coincidence rate (sensitivity) of the training set is 100%, the negative coincidence rate (specificity) is 100%, the accuracy is 100%, the F1 score is 100%, the recall rate is 100%, the precision rate is 100%, and the Kappa coefficient is 1.00;

[0094] 3) The positive coincidence rate (sensitivity) of the test set is 100%, the negative coincidence rate (specificity) is 100%, the accuracy is 100%, the F1 score is 100%, the recall rate is 100%, the precision rate is 100%, and the Kappa coefficient is 1.00;

[0095] 4) The model uses the number of decision trees supporting MSI-H as the indicator for the final determination of the MSI status. When the number of decision trees supporting MSI-H is greater than or equal to 51, the MSI status of the sample is determined to be MSI-H, otherwise it is determined to be MSS.

[0096] In some examples, the method further includes performing quality control on the obtained characteristic sites related to the repeat unit length distribution. The characteristic site quality control threshold is obtained by the following steps:

[0097] (1) 100 samples with confirmed MSI status (MSI-H: 53, MSS: 47) were selected, among which the important features (1137) were used for analysis. The distribution of MSI-H support numbers of the samples is as follows: Figure 3 ,The selected samples contain all segments with MSI-H support numbers.

[0098] (2) The number of sites for 100 samples was selected, and the gradient was selected to be 100-1100 with a step length of 50, that is, there were 21 gradients in total (see Table 2):

[0099] Table 2: Selection list of site numbers

[0100] 100 150 200 250 300 350 400 450 500 550 600 650 700 750 800 850 900 950 1000 1050 1100 / / /

[0101] For each sample, the number of sites in the above gradient will be selected, and each gradient will randomly select sites 100 times, so each gradient will obtain 100 random selections of 100 samples, that is, 10000 (100*100) MSI results. Finally, the 10,000 MSI results of each gradient are compared and counted with the original results, and finally the number of gradient sites with an accuracy of 10,000 MSI results greater than 95% is selected as the cutoff value of the number of MSI sites.

[0102] (3) The results are as follows Figure 4 As shown in the figure, when the number of sites is greater than or equal to 1100, the accuracy of all samples is greater than 99%. Therefore, the quality control indicator is: when the sample depth meets the standard, the sequencing depth of the MSI sites is counted. If the sequencing depth of more than or equal to 37 sites is unqualified, the quality control will fail. Conversely, if the sequencing depth of less than 37 sites is unqualified, it will not affect the final determination of the MSI status.

[0103] In the third aspect of the present application, the present application proposes a device for predicting the state of microsatellites. Figure 5 As shown, the device comprises:

[0104] A sequencing data acquisition unit S300 is used to acquire sequencing data of a sample to be tested, wherein the sequencing data includes a plurality of read segments from at least one predetermined microsatellite characteristic site;

[0105] A feature acquisition unit S400, the feature acquisition unit S400 is connected to the sequencing data acquisition unit S300, and is used to determine the repeat unit length distribution-related features of the predetermined microsatellite feature site based on the multiple read segments;

[0106] The prediction unit S500 is connected to the feature acquisition unit S400 and is used to input the repeat unit length distribution related features into a trained machine learning model to predict the microsatellite status of the sample.

[0107] In a fourth aspect of the present application, the present application proposes a machine learning model training device. Figure 6 As shown, the device comprises:

[0108] The training data acquisition unit S600 is used to acquire training set and test set sequencing data, wherein the training set and test set sequencing data have known microsatellite status results, and the training set and test set sequencing data contain multiple microsatellite characteristic sites;

[0109] A model feature acquisition unit S700, the model feature acquisition unit S700 is connected to the training data acquisition unit S600, and is used to determine the repeat unit length distribution-related characteristics of the predetermined microsatellite characteristic site based on the training set and the test set sequencing data;

[0110] A training unit S800, wherein the training unit S800 is connected to the model feature acquisition unit S700, and is used to input the repeat unit length distribution related features into the machine learning model, use the known microsatellite status results as labels, and perform supervised training on the machine learning model to obtain a trained machine learning model, and the machine learning model is used to predict the microsatellite status.

[0111] In another aspect of the present application, the present application further provides a computer program product, an electronic device, and a computer-readable storage medium. This embodiment takes an electronic device as an example for detailed description.

[0112] The so-called electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing devices may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0113] like Figure 7 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 to a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0114] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0115] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as a method for predicting the state of a microsatellite. For example, in some embodiments, the method for predicting the state of a microsatellite may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the aforementioned method for predicting the state of a microsatellite in any other appropriate manner (eg, by means of firmware).

[0116] In the present application, the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program may be printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or processing in other suitable ways as necessary, and then stored in a computer memory. The various computer-readable storage media described herein may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0117] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0118] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0119] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0120] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0121] The microsatellite status prediction method provided in this application comprehensively evaluates the microsatellite status of a sample by using multiple microsatellite characteristic sites and a strong classification model. This method has higher sensitivity and accuracy and can more comprehensively display the MSI status of a sample. In addition, this application also adds a quality control step for the number of MS sites. Even if the sequencing depth of some sites does not meet the quality control standard, the MSI status of the sample can still be accurately determined, greatly improving the analysis success rate and the robustness of the analysis method.

[0122] It should be noted that the features and technical effects described in this article for different aspects can be used as reference for each other and will not be repeated here.

[0123] The present application is described below by way of examples, but this should not be construed as limiting the scope of the present application to the following examples. All technologies implemented based on the above content of the present application belong to the scope of the present application.

[0124] Example 1: Model effect verification

[0125] 1. Perform high-throughput sequencing (NGS) on cancer tissue samples (including cervical cancer, colorectal cancer, ovarian cancer, gastric cancer, and hepatobiliary cancer) that have been verified by PCR as MSI-H or MSS and their control samples, and obtain the raw data Fastq (raw data) of 50 samples, including 25 samples with MSI-H status and 25 samples with MSS status, and each status corresponds to its corresponding control sample;

[0126] 2. Align the raw data to the reference genome, remove duplicate reads and Indel local re-alignment after alignment to obtain BAM files of 50 samples and their control samples;

[0127] 3. The depth of 1137 characteristic loci was counted. All 1137 characteristic loci of 50 samples reached the site depth quality control threshold of 20X, that is, the number of qualified characteristic loci was higher than 1100, and the site number quality control passed;

[0128] 4. By comparing the differences in the length distribution of repeat units at 1137 characteristic sites in 50 tissue samples and their control samples, the MSI-Score of each characteristic site was calculated.

[0129] 5. The MSI-Score of 1137 characteristic loci was used as feature values ​​and input into the random forest classification model to determine the MSI status.

[0130] The results show that the MSI-H decision tree support number (Support_MSI-H result value) of the 50 samples and the final classification results are shown in Table 3. If the predict_MSI-H support number is greater than or equal to 51, it is judged to be MSI-H. The MSI status results of the 50 samples are completely consistent with the MSI status of the sample answer, with an accuracy of 100%.

[0131] Table 3: Classification results

[0132]

[0133]

[0134] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0135] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for predicting microsatellite status, characterized in that: include: Acquiring sequencing data of a sample to be tested, wherein the sequencing data includes a plurality of read segments from at least one predetermined microsatellite characteristic site; Determine, based on the plurality of read segments, a repeat unit length distribution-related feature of at least one of the predetermined microsatellite characteristic sites; The repeat unit length distribution related features are input into a trained machine learning model to predict the microsatellite status of the sample.

2. The method according to claim 1, characterized in that The machine learning model is selected from at least one of a decision tree, logistic regression, naive Bayes, random forest, and a neural network; Preferably, the machine learning model is a random forest model.

3. The method according to claim 1, characterized in that The repeating unit length distribution-related characteristics are determined based on a comparison of the repeating unit length distribution differences between the test sample and the control sample.

4. The method according to claim 3, characterized in that The repeat unit length distribution related characteristics are determined by the following formula: Where d is the Manhattan distance, R T is the type of at least a portion of the repeating unit length of the corresponding read segment in the sample to be tested, R N is the length type of at least a portion of the repeating units in the corresponding read segment in the control sample, T r is the normalized read count of the rth repeat unit length type in the sample to be tested, N r is the normalized read count that falls into the rth repeat unit length type in the control sample.

5. The method according to claim 3 or 4, characterized in that: The sample to be tested is from a subject suspected of having a tumor; the control sample is from a non-tumor tissue; Optionally, the tumor comprises hereditary non-polyposis colorectal cancer, colorectal cancer, gallbladder cancer, ovarian cancer, pancreatic cancer, prostate cancer, liver cancer, kidney cancer, lung cancer, gastric cancer, nasopharyngeal cancer, esophageal cancer or endometrial cancer.

6. The method according to claim 1, characterized in that The important characteristic sites in the microsatellite characteristic sites are shown in the following table:

7. The method according to claim 1, characterized in that The sequencing depth of the predetermined characteristic sites is not less than 20X.

8. A machine learning model training method, characterized in that: include: Acquire training set and test set sequencing data, wherein the training set and test set sequencing data have known microsatellite status results, the training set and test set sequencing data include multiple reads from at least one predetermined microsatellite characteristic site, and the training set includes a sample set to be tested and a control sample set; Based on the multiple read segments, determining a repeat unit length distribution-related feature of at least one of the predetermined microsatellite characteristic sites; The repeat unit length distribution related features are input into the machine learning model, and the known microsatellite status results are used as labels to perform supervised training on the machine learning model to obtain a trained machine learning model, which is used to predict the microsatellite status.

9. The method according to claim 8, characterized in that The repeat unit length distribution related characteristics are determined by the following formula: Where d is the Manhattan distance, R T is the type of at least a portion of the repeating unit length of the corresponding read segment in the sample to be tested, R N is the length type of at least a portion of the repeating units in the corresponding read segment in the control sample, T r is the normalized read count of the rth repeat unit length type in the sample to be tested, N r is the normalized read count that falls into the rth repeat unit length type in the control sample.

10. A device for predicting microsatellite status, characterized in that: include: A sequencing data acquisition unit, used to acquire sequencing data of a sample to be tested, wherein the sequencing data includes a plurality of read segments from at least one predetermined microsatellite characteristic site; A feature acquisition unit, used to determine the repeat unit length distribution-related features of the predetermined microsatellite feature site based on the multiple read segments; and A prediction unit is used to input the repeat unit length distribution related features into a trained machine learning model to predict the microsatellite status of the sample.

11. A machine learning model training device, characterized in that: include: A training data acquisition unit, used to acquire training set and test set sequencing data, wherein the training set and test set sequencing data have known microsatellite status results, and the training set and test set sequencing data contain multiple microsatellite characteristic sites; A model feature acquisition unit, used to determine the repeat unit length distribution-related features of the predetermined microsatellite feature site based on the training set and the test set sequencing data; and A training unit, used for inputting the repeat unit length distribution related features into a machine learning model, using known microsatellite status results as labels, and performing supervised training on the machine learning model to obtain a trained machine learning model, wherein the machine learning model is used to predict the microsatellite status; Wherein, the machine learning model is obtained by training through the method described in claim 8 or 9.

12. A computer program product, characterized in that The computer program product comprises computer instructions, and when part or all of the computer instructions are executed on a computer, the method according to any one of claims 1 to 9 is implemented.

13. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 9.

14. A computer-readable storage medium containing a computer program, characterized in that: When the computer program is executed by one or more processors, the method according to any one of claims 1 to 9 is implemented.