A method for constructing an MSI overall state prediction model
By constructing an MSI overall status prediction model that does not require normal sample pairing, and utilizing the length changes of marker microsatellite loci and machine learning algorithms, the problems of detection complexity and low accuracy in existing technologies are solved, and fast and accurate MSI status prediction is achieved.
Patent Information
- Application Number
- CN202410929914.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-21
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2043-09-21
AI Technical Summary
Existing MSI detection methods require normal sample pairing, which makes the detection process complicated and the accuracy low, making it difficult to predict microsatellite instability status efficiently and accurately.
A method based on liquid phase capture second-generation sequencing is provided. By analyzing the length changes of labeled microsatellite loci, a prediction model for the overall MSI status is constructed without the need for normal sample pairing. Machine learning algorithms such as random forest and logistic regression are used to combine the peak number, kurtosis, skewness, standard deviation and standard shift of microsatellite loci to construct a stability prediction model.
It has achieved rapid and accurate prediction of the overall MSI status in different NGS detection platforms and cancer types without the need for normal sample pairing, improving the stability and repeatability of the test and reducing the detection limitations.
Smart Images

Figure CN118800326B_ABST
Abstract
Description
[0001] Related patents
[0002] This application is a divisional application of the Chinese invention patent application with application number 2023112243588, application date September 21, 2023, and invention name “Microsatellite sites, systems and kits for predicting MSI status”. Technical Field
[0003] The present invention belongs to the technical field of MSI detection, and in particular, relates to a method for constructing an MSI overall status prediction model. Background Art
[0004] Microsatellite instability (MSI) refers to the phenomenon of changes in the length of microsatellite (MS) sequences due to insertion or deletion mutations during DNA replication, often caused by deficient mismatch repair (MMR). MS sequences are short, repetitive DNA sequences, typically consisting of 1 to 6 nucleotides, arranged in tandem repeats. Common repeat types include double-base sequences such as CA / GA / GT or single-base sequences such as A / T. MS sequences can be located in important noncoding regions or coding regions of genes, and polymorphisms are distributed throughout the genome, with significant individual variability. Cells can repair regions with mismatches that occur during genome replication using mismatch repair elements. In certain cases, when mismatch repair elements lose their original function, such as mutations in MLH1, MSH2, and MSH6, cells are unable to repair mismatches, resulting in the development of an MSI phenotype. This phenotype manifests as changes in the number of microsatellite repeat units, which can increase or decrease in length.
[0005] Microsatellite instability is a relatively common phenomenon in tumors. The state of microsatellite instability indicates the cause and development of the tumor, and can also play an important role in auxiliary diagnosis and medication guidance in different types of cancer. Generally, the state of microsatellite instability can be divided into microsatellite high instability (MSI-H), microsatellite low instability (MSI-L) and microsatellite stable (MSS). In colorectal cancer, endometrial cancer, gastric cancer and other cancers, patients with MSI-H status have significant differences in survival, medication preference, and palliative treatment prognosis compared with the other two. MSI testing can help doctors fully understand the type of cancer and propose the correct diagnosis and treatment plan.
[0006] Another important role of MSI detection is to assist in screening Lynch Syndrome, a hereditary syndrome characterized by germline mutations in mismatch repair genes, which is characterized by colorectal cancer, endometrial cancer, gastric cancer, and other cancers in younger patients and their families. MSI status is related to Lynch Syndrome, and most patients with microsatellite stability are not Lynch Syndrome. Traditionally, if a patient has the above symptoms, the doctor will usually consider detailed testing. If using the method of immunohistochemistry, two experienced pathologists need to jointly detect, and the accuracy is low. If MSI detection is used, it is more convenient, and the accuracy is also higher. When the patient is determined to be MSI-H, further genetic diagnosis can be performed to determine whether it is Lynch Syndrome.
[0007] Due to the sequence characteristics of MSI sites, the reliability of single site sequencing is low, and MSI detection based on second-generation sequencing usually detects hundreds of MSI sites. The method of multi-site cloning and genome sequencing can be used to detect the microsatellite state. Multi-site cloning is commonly used in current MSI detection. By judging whether the length of 5-10 microsatellite regions highly related to MSI status changes, it is determined whether the site is unstable. The higher the proportion of unstable sites, the more likely the cell is in the MSI state. This method is cheap and has a short experimental process. However, many detection tools for predicting MSI status by NGS method need to be paired with normal samples, such as mSINGS, MANTIS, etc. SUMMARY
[0008] The main purpose of the present application is to provide a method for predicting MSI overall status without normal sample pairing based on liquid-phase capture of second-generation sequencing microsatellite site data. The method is convenient, stable, fast, accurate, and highly reproducible, reducing the limitations of detection.
[0009] To achieve the above purpose, the technical scheme provided by the present application is as follows:
[0010] The first aspect of the present application provides a marker microsatellite site for determining the MSI overall status of a sample to be tested, the marker microsatellite site comprising:
[0011]
[0012]
[0013] Microsatellite (MS) refers to short tandem repeats (STRs) in the human genome, or repeated base sequences, including single nucleotide repeats, dinucleotide repeats, and even multinucleotide repeats. Microsatellite instability (MSI) refers to the change in the length of microsatellites in tumor tissues due to the insertion or deletion of repeat units relative to normal tissues. In the present invention, the overall MSI status refers to the status of all microsatellites in the sample to be tested; the microsatellite loci refer to specific microsatellite loci, and the marker microsatellite loci refer to the microsatellites used to determine the overall MSI status of the sample to be tested.
[0014] In the present invention, microsatellites containing or including other mononucleotide repeat sequences and / or dinucleotide repeat sequences may also be selected as marker microsatellite loci.
[0015] In the present invention, the number of repeats is the standard value (γ) of the length of the marker microsatellite locus, that is, the number of repeats (standard number of repeats) in the above table.
[0016] In the present invention, since microsatellites can vary in length due to insertions or deletions of repeating units, the length of a microsatellite is the specific number of repeats of the repeating unit, i.e., the length is determined by the number of repeats. For example, even if the repeating unit has 2 bases and is repeated 10 times, resulting in a base length of 20, the length of the microsatellite locus is still considered to be 10.
[0017] For microsatellite loci, each repeat unit can exist in three states: deletion, unchanged, and insertion, that is, the number of each repeat unit can be 0, 1, or 2. Based on this, for each microsatellite locus, the length range is [0, 2γ].
[0018] The second aspect of the present invention provides use of the detection reagent for marking microsatellite loci according to the first aspect of the present invention in the preparation of a kit for determining the overall MSI status of a sample to be tested.
[0019] A third aspect of the present invention provides a system for determining the overall MSI status of a sample to be tested, comprising the following modules:
[0020] A marker microsatellite locus data input module is used to receive stability data of the marker microsatellite loci of the first aspect of the present invention for a sample to be tested, wherein the stability includes stable and unstable;
[0021] an MSI overall status storage module, configured to store stability data of the marker microsatellite loci of the second population sample and MSI overall status data, wherein the MSI overall status comprises MSI-H, MSI-L and MSS;
[0022] an MSI overall status determination module, connected with the data input module and the MSI overall status storage module respectively, configured to construct a prediction model by using the stability data of the marker microsatellite loci of the second population sample, and determine the MSI overall status of the sample to be tested based on the stability data of the marker microsatellite loci of the sample to be tested obtained from the marker microsatellite locus stability prediction module.
[0023] In some embodiments of the present application, the stability data of the marker microsatellite loci refers to whether the length of the marker microsatellite loci changes, if a significant change occurs in a marker microsatellite locus, i.e. a significant change in the length of the marker microsatellite locus due to the insertion or deletion of repeat units, then the marker microsatellite locus is unstable; otherwise, if no change occurs in a marker microsatellite locus, i.e. there is no change in the length of the marker microsatellite locus due to the insertion or deletion of repeat units or even if there is the insertion or deletion of repeat units, the length change is not significant, then the marker microsatellite locus is stable.
[0024] In the present application, the MSI-H, MSI-L and MSS have the following meanings:
[0025] (1) MSS: Microsatellite Stability, microsatellite stability, i.e. all microsatellite loci are in a stable state.
[0026] (2) MSI-L: Low-frequency MSI, low-frequency microsatellite instability, i.e. microsatellite loci less than a preset threshold are in an unstable state.
[0027] (3) MSI-H: High-frequency MSI, high-frequency microsatellite instability, i.e. microsatellite loci not less than a preset threshold are in an unstable state.
[0028] In some embodiments of the present application, the preset threshold can be an absolute value of quantity or a proportion, for example 30%.
[0029] In some embodiments of the present application, in the MSI overall status determination module, the step of constructing a prediction model by using the stability data of the marker microsatellite loci of the population sample comprises the following steps:
[0030] S21, randomly dividing the stability data of the marker microsatellite loci of the second population sample into two groups, one group being a second training set and the other group being a second test set, each group comprising the stability data of the marker microsatellite loci of the MSI-H sample, the MSI-L sample and the MSS sample;
[0031] S22, constructing an MSI overall state prediction model based on a machine learning algorithm using the second training set data;
[0032] S23, verifying the obtained MSI overall state prediction model in the second test set.
[0033] In some embodiments of the present application, the machine learning algorithm is selected from any one of the following algorithms:
[0034] a random forest algorithm, a neural network algorithm, a support vector machine algorithm, a Bayesian classification algorithm, a gradient boosting algorithm, a K-nearest neighbor algorithm and a decision tree algorithm.
[0035] In some preferred embodiments of the present application, the MSI overall state prediction model is obtained by training using a random forest model.
[0036] In some embodiments of the present application, the stability data of the marker microsatellite loci obtained in the marker microsatellite loci data input module is obtained based on a PCR-based method.
[0037] Further, the system further comprises:
[0038] a sequencing data input module for inputting capture sequencing data of a target region comprising the marker microsatellite loci in the sample to be tested, and obtaining the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the marker microsatellite loci;
[0039] a microsatellite loci storage module for storing the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite loci and the stability data of the first population sample;
[0040] a marker microsatellite loci stability prediction module, connected with the sequencing data input module, the marker microsatellite loci storage module and the marker microsatellite loci data input module, respectively, for constructing a marker microsatellite loci stability prediction model using the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite loci and the stability data of the first population sample, and predicting the stability of the marker microsatellite loci based on the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the marker microsatellite loci obtained from the sequencing data input module, and for outputting the stability data of the marker microsatellite loci to the marker microsatellite loci data input module.
[0041] In some embodiments of the present application, the known microsatellite sites include:
[0042]
[0043]
[0044] In some embodiments of the present application, in the marker microsatellite site stability prediction module, the step of constructing a marker microsatellite site stability prediction model using the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite site and stability data includes the following steps:
[0045] S11, the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite site and stability data of the first population sample are randomly divided into two groups, one group is the first training set, and the other group is the first test set;
[0046] S12, using the first training set data, a marker microsatellite site stability prediction model is constructed based on a regression algorithm;
[0047] S13, in the first test set, the obtained marker microsatellite site stability prediction model is verified.
[0048] In some embodiments of the present application, the regression algorithm is selected from any one of the following algorithms: logistic regression algorithm, linear regression algorithm.
[0049] In some embodiments of the present application, for any one of the marker microsatellite sites, first, the peak number of the length of the marker microsatellite site is counted according to the peak searching algorithm, and then the skewness Skew, kurtosis Kurt, standard deviation S and standard deviation P of the length of the marker microsatellite site are calculated respectively:
[0050] Skewness measures the asymmetry of the probability distribution of a random variable, which is a measure of the degree of asymmetry relative to the mean. By measuring the skewness coefficient, the degree of asymmetry and direction of the length distribution of the marker microsatellite site can be determined. The measurement of skewness is relative to the normal distribution, and the skewness of the normal distribution is 0, that is, if the length distribution of the marker microsatellite site is symmetric, the skewness is 0. If the skewness is greater than 0, the distribution is right-skewed, that is, the distribution has a long tail on the right; if the skewness is less than 0, the distribution is left-skewed, that is, the distribution has a long tail on the left; at the same time, the greater the absolute value of skewness, the more serious the degree of skewness.
[0051] The formula for calculating skewness is:
[0052]
[0053] Where, x iThe number of reads of the marker microsatellite site with length i, i = [1, n]; The average of the number of reads of the marker microsatellite site with different lengths; n is the number of length types of the marker microsatellite site, i.e. how many different lengths, n = 2γ; γ is the standard value of the length of the marker microsatellite site.
[0054] Kurtosis is a statistical quantity for studying the steepness or flatness of data distribution. By measuring the kurtosis coefficient, it can be determined whether the length of the marker microsatellite site is steeper or flatter relative to the normal distribution. If kurtosis = 3, the peak state of the length distribution of the marker microsatellite site obeys the normal distribution; if kurtosis > 3, the peak state of the length distribution of the marker microsatellite site is steep (high and sharp); if kurtosis < 3, the peak state of the length distribution of the marker microsatellite site is flat (low and fat). The kurtosis calculation formula is:
[0055]
[0056] Wherein, x i The number of reads of the marker microsatellite site with length i, i = [1, n]; The average of the number of reads of the marker microsatellite site with different lengths; n is the number of length types of the marker microsatellite site, n = 2γ, and γ is the standard value of the length of the marker microsatellite site.
[0057] The standard deviation reflects the dispersion degree of the length of the marker microsatellite site. The greater the value, the more dispersed, i.e. the greater the difference between the different lengths of the marker microsatellite site.
[0058] The standard deviation calculation formula is:
[0059]
[0060] Wherein, x i The number of reads of the marker microsatellite site with length i, i = [1, n]; The average of the number of reads of the marker microsatellite site with different lengths; x′ i The normalized value of the number of reads of the microsatellite site with length i, The average of x′ i ; n is the number of lengths of the marker microsatellite site, n = 2γ, and γ is the standard value of the length of the marker microsatellite site.
[0061] The standard deviation is the level of the length of the marker microsatellite site deviating from the standard value of the length of the marker microsatellite site.
[0062] The standard deviation calculation formula is:
[0063]
[0064] wherein x i is the number of reads of the marker microsatellite site with length i, i = [1, n]; x′ i is the normalized value of the number of reads of the marker microsatellite site with length i, n is the number of length categories of the marker microsatellite site; γ is the standard value of the length of the marker microsatellite site, n = 2γ, γ is the standard value of the length of the marker microsatellite site.
[0065] In some embodiments of the present application, in the sequencing data input module, the peak number, kurtosis, skewness, standard deviation and standard deviation of the marker microsatellite site are calculated only when the depth of the captured sequencing of the target region reaches 400x. Specifically, the number of reads of different lengths of the marker microsatellite site is calculated using the result file of Msisensor2.
[0066] The fourth aspect of the present application provides a kit for determining the MSI overall state of a sample to be tested, comprising the detection reagent of the marker microsatellite site of the first aspect of the present application.
[0067] In the present application, the "first population sample" and the "second population sample" are only formal regions, wherein the data of the first population sample includes the peak number, kurtosis, skewness, standard deviation and standard deviation of the known microsatellite site length of each sample and stability data, which are used to construct a marker microsatellite site stability prediction model based on a regression algorithm; the data of the second population sample includes the stability data of the marker microsatellite site and the MSI overall state data, which are used to construct an MSI overall state prediction model.
[0068] In the present application, the sample to be tested is derived from humans, preferably tumor samples such as fresh tissues, tissue paraffin blocks (FFPE) and the like.
[0069] Advantages of the present application
[0070] Compared with the prior art, the present application has the following advantages:
[0071] By using the present application, the MSI overall state can be predicted without pairing with normal samples in different NGS detection platforms and different cancer types through the analysis of second-generation sequencing data.
[0072] By using the present application to predict the MSI overall state, the process is stable and fast, the result is accurate, the repeatability is high, and the limitation of detection is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0073] Figure 1 The length distribution diagram of one microsatellite site in Example 3 of the present application is shown.
[0074] Figure 2 MSI-PCR capillary electrophoresis detection structure of one microsatellite site in the embodiment 4 of the present application is shown. A: tumor sample; B: blood sample.
[0075] Figure 3 The performance evaluation results of the stability of the microsatellite site predicted by the logistic regression model established in the embodiment 4 of the present application are shown.
[0076] Figure 4 The performance evaluation results of the random forest model established in the embodiment 5 of the present application for predicting the MSI overall state are shown.
[0077] Figure 5 The validation results of the external data set in the embodiment 6 of the present application are shown.
[0078] Figure 6 The schematic diagram of the system for determining the MSI overall state of the sample to be tested constructed in the embodiment 7 of the present application is shown. DETAILED DESCRIPTION
[0079] Unless otherwise stated, implied from context, or consistent with the practice in the art, all parts and percentages contained in this application are based on weight, and the test and characterization methods used are those contemporary with the filing date of this application. To the extent applicable, the contents of any patent, patent application or publication referred to in this application are hereby incorporated by reference in their entirety, and equivalent same family patents are hereby incorporated by reference, particularly as to definitions of terms in the art. If a definition of a term in the prior art is inconsistent with any definition provided in this application, the definition provided in this application shall control.
[0080] Numerical ranges in this application are approximate, thus unless otherwise stated, the numerical values are approximations only, and thus can include numbers that are outside the range. The number range includes all values from and including the lower and the upper values, in increments of one unit. In the case of fractions of one unit, the increments are in fractions of 0.0001, 0.001, 0.01 or 0.1, as appropriate. For ranges including these fractions, the increments are normal, e.g., 1.1, 1.5, etc. These are only examples of what is intended. All such increments and fractions are considered to be explicitly stated in this application. The endpoints of the ranges are provided as a clear limitation. Also, all percentages are calculated by weight unless otherwise indicated.
[0081] The terms "comprising", "including", "having" and their derivatives do not exclude the presence of any other components, steps or processes, and are irrelevant to whether these other components, steps or processes are disclosed in this application. To eliminate any doubt, all compositions using the terms "comprising", "including", or "having" in this application may include any additional additives, excipients or compounds unless expressly stated otherwise. In contrast, the term "essentially consisting of" excludes any other components, steps or processes from the scope of any subsequent description of the term, except those necessary for operational performance. The term "consisting of" does not include any components, steps or processes that are not specifically described or listed. Unless expressly stated otherwise, the term "or" refers to the listed members alone or in any combination.
[0082] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the embodiments.
[0083] Example
[0084] The following examples are provided to illustrate preferred embodiments of the present invention. Those skilled in the art will appreciate that the techniques disclosed in the following examples represent techniques discovered by the inventors that can be used to practice the present invention and, therefore, can be considered preferred embodiments of the present invention. However, those skilled in the art will appreciate from this disclosure that many modifications may be made to the specific embodiments disclosed herein while still achieving the same or similar results without departing from the spirit or scope of the present invention.
[0085] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs, and the disclosure and materials they cite are hereby incorporated by reference.
[0086] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many technical equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the claims.
[0087] The experimental methods in the following examples, unless otherwise specified, are all conventional methods. The instruments and equipment used in the following examples, unless otherwise specified, are all conventional laboratory instruments and equipment; the experimental materials used in the following examples, unless otherwise specified, are all purchased from conventional biochemical reagent stores.
[0088] Example 1 Selection of microsatellite loci
[0089] 800 tissue samples from clinical tumor patients with cancers including colorectal cancer and lung cancer were collected for MSI prediction.
[0090] As for the selection of microsatellite sites, the present embodiment uses a large gene combination panel (as shown in Table 1) containing 169 genes related to solid tumors such as colorectal cancer, lung cancer, endometrial cancer, gastric cancer, etc. developed based on VariantBaitsTM technology. The DNA extracted from the sample is hybridized and captured using the target region probe, and then the corresponding sample library is obtained by library construction. Further, the fastq format of the raw data of the second-generation sequencing is obtained by further sequencing the library.
[0091] Table 1 169 gene list
[0092] AKT1 ALK APC ARAF ARID1A ATM BRAF BRCA1 BRCA2 CTNNB1 DDR2 DNMT3A EGFR ERBB2 ERBB3 FGF19 FGFR1 FGFR2 FGFR3 FGFR4 FH FLCN HFE HOXB13 IDH1 IDH2 JAK1 JAK2 KDM6A KIT KRAS MAP2K1 MDM4 MTOR NF2 NRAS NRG1 NTRK1 NTRK2 NTRK3 PDGFRA PIK3C2G PTEN RB1 RET ROS1 SDHA SDHB SMAD4 STK11 TERT TP53 TSC1 TSC2 UGT1A1 VEGFA VHL WT1 BAP1 BARD1 BRIP1 CDK12 CHEK1 CHEK2 ERCC1 FANCA FANCL MEN1 MLH1 MSH6 MUTYH NBN NF1 PALB2 PMS2 POLD1 POLE PPP2R2A RAD51C RAD51D RAD54L SDHAF2 SDHC SDHD EPCAM HLA-A HLA-B ABCB1 ABCC2 ABCC4 ATIC C8orf34 CBR3 CDA CYP19A1 CYP1B1 CYP2C8 CYP2D6 CYP3A4 DHFR DPYD DYNC2H1 ERCC2 FOLR3 GGH MTHFR MTRR NQO1 RRM1 SEMA3C SLC19A1 SLC22A16 SLC22A2 SLCO1B1 TPMT TYMS UMPS XPC XRCC1 ABL1 CDH1 CDKN2A CSF1R EZH2 FBXW7 FLT3 GNAQ GNAS HNF1A JAK3 KDR NOTCH1 PBRM1 PTPN11 SMARCB1 B2M ESR1 HRAS MET PIK3CA TGFBR2 BMPR1A MSH2 RAD51B HLA-C CYP2B6 GSTP1 SOD2 ERBB4 NPM1
[0093] Sample library sequencing and sequence comparison of Example 2
[0094] The raw data obtained in Example 1 is subjected to quality control using the quality control software picard, filtering sequencing adapters, low-quality bases, and sequencing error fragments, etc. After filtering, high-quality data (clean data) is obtained.
[0095] The clean data is analyzed using the sequence alignment software bwa-MEM to obtain the specific genomic position information of the sequence (reference genome is hg19); samtools and sambamba software are used for sorting and deduplication to obtain the sample raw bam file.
[0096] The sample raw bam file is input into the software Msisensor2 for analysis to obtain the number of reads of different lengths of target region repeat base sequences, i.e. the sample data set.
[0097] Example 3 Microsatellite site stability state prediction
[0098] The target region sequencing depth is at least 400x, and subsequent calculations are performed under the condition of meeting the depth.
[0099] For a specific microsatellite site: chr2_47641559_CAGGT_27[A]_GGGTT, as shown in Table 2.
[0100] Table 2 chr2_47641559_CAGGT_27[A]_GGGTT information
[0101] Chromosome Position Repeat unit Repeat number (gamma) 3' flanking sequence 5' flanking sequence chr2 47641559 A 27 CAGGT GGGTT
[0102] The length distribution is shown in Table 3 and Figure 1
[0103] Table 3 chr2_47641559_CAGGT_27[A]_GGGTT length distribution
[0104]
[0105]
[0106] The embodiment utilizes five characteristics of the peak number N, skew Skew, kurtosis Kurt, standard deviation S, and standard deviation P of the length of the target region repeat base sequence to determine the MSI state.
[0107] For a certain repeat base sequence, the following is calculated respectively.
[0108] Peak number calculation:
[0109] The sequencing depth of the target region is calculated according to the peak searching algorithm to obtain the peak number of the microsatellite site length. Specifically, Scipy is used for implementation, and then Peak Prominence is calculated to calibrate the peak number and remove noise.
[0110] In the embodiment, the peak number of chr2_47641559_CAGGT_27[A]_GGGTT is 1.
[0111] Then, using the known stable length γ of the microsatellite site, the theoretical maximum length of the microsatellite site is 2γ, i.e. 54, so the length type number n of chr2_47641559_CAGGT_27[A]_GGGTT is 54, and i = [1, 54].
[0112] The skew calculation formula is:
[0113]
[0114] The kurtosis calculation formula is:
[0115]
[0116] The standard deviation calculation formula is:
[0117]
[0118] The standard deviation calculation formula is:
[0119]
[0120] Wherein, x i is the number of reads when the length of the microsatellite site is i; is the average number of reads of different lengths of the microsatellite site. x′ i is the normalized value of the number of reads when the length of the microsatellite site is i, is x′ i The mean of .
[0121] According to the above formula, the skewness, kurtosis, standard deviation, and standard shift of the length of chr2_47641559_CAGGT_27[A]_GGGTT are 2.80, 9.87, 0.05, and 19.82, respectively.
[0122] Example 4 Establishing a logistic regression model to predict the stability of microsatellite loci
[0123] This example uses six microsatellite loci with known stability data to establish a model to predict the stability of microsatellite loci with unknown stability data. The six microsatellite loci with known stability data are shown in Table 4.
[0124] Table 4 Six known microsatellite loci
[0125] Site number Chromosome Position Repeat unit Repeat number (gamma) 3' flanking sequence 5' flanking sequence 1 chr11 102193508 A 26 CTGGT GCCAC 2 chr14 23652346 A 21 TTGCT GGCCA 3 chr4 55598211 T 25 TTTGA GAGAA 4 chr2 95849361 T 23 TCCTA GTGAG 5 chr2 47641559 A 27 CAGGT GGGTT 6 chr2 39564893 T 28 TCCAG GAGAC
[0126] To establish a microsatellite locus stability model, validate the MSI prediction results, and evaluate their performance, the inventors performed MSI-PCR capillary electrophoresis on the samples, using the results as the gold standard to compare the consistency of the two results. The MSI-PCR capillary electrophoresis kit used for MSI-PCR was the Tongshu Microsatellite Instability (MSI) Detection Kit (Multiplex Fluorescence PCR-Capillary Electrophoresis). The MSI-PCR capillary electrophoresis assay determined the status of each microsatellite locus and the overall MSI status of each sample.
[0127] Specifically, based on the results of MSI-PCR capillary electrophoresis, each microsatellite locus is divided into two categories: stable microsatellite loci and unstable microsatellite loci. If a microsatellite locus changes, that is, the length of the microsatellite locus changes significantly due to the insertion or deletion of a repeating unit, then the microsatellite locus is unstable; conversely, if a microsatellite locus does not change, that is, there is no change in the length of the microsatellite locus due to the insertion or deletion of a repeating unit, or even if there is an insertion or deletion of a repeating unit, the length change is not significant, then the marker microsatellite locus is stable.
[0128] For example, the results of MSI-PCR capillary electrophoresis at BAT25 locus (chr4 55598211) were as follows: Figure 2 As shown, this point is homozygous. In the comparison between the tumor sample and the paired sample, it is not difficult to find that a new set of displacement peaks has been added in the tumor sample (as shown by the dotted line in the figure). Therefore, this site is in an unstable state of the microsatellite site in the tumor sample, and in a stable state of the microsatellite site in the paired sample (blood).
[0129] Further, according to the stability of the microsatellite site, the MSI overall state of the sample is divided into MSI-H, MSI-L, and MSS, which correspond to microsatellite high instability, microsatellite low instability, and microsatellite stability, respectively.
[0130] MMS: All microsatellite sites are in a stable state.
[0131] MSI-L: <30% of the microsatellite sites are in an unstable state.
[0132] MSI-H: ≥30% of the microsatellite sites are in an unstable state.
[0133] In this embodiment, 800 sample MSI-PCR capillary electrophoresis data were obtained, including the stability data of 6 microsatellite sites and the MSI overall state data of 800 samples.
[0134] The MSI-PCR capillary electrophoresis data of the stability of the 6 microsatellite sites of the above-mentioned 800 samples were processed, and samples with NGS results not meeting the quality control were removed, obtaining a total of 3780 microsatellite site stability data, which were randomly stratified sampled into a training set and a test set according to a 7:3 ratio. For the training set, based on the Msisensor2 result file obtained by Example 2, the length of the 6 microsatellite sites was calculated, and the peak number N, kurtosis Kurt, skewness Skew, standard deviation S, and standard deviation P of the length of each microsatellite site were obtained by the method of Example 3, to generate the training set data, with the microsatellite site unstable sample being 1 and the microsatellite site stable sample being 0, forming the training set 1. The same method was used to obtain the test set 1. The training set 1 data was trained based on the logistic regression model to obtain the logistic regression model of the training set 1.
[0135] Based on the logistic regression model obtained from the training set 1, the microsatellite site stability of the test set 1 was predicted, and the accuracy of the microsatellite site stability of the test set was evaluated according to the MSI-PCR capillary electrophoresis result, and the model was then modified to remove noise and modify the model parameters and solution algorithm. Finally, the model regularization parameter was specified as l2, the maximum number of iterations was 5000, and the classification type was binary. Because the number of microsatellite site unstable samples in the actual training set was much less than that of stable samples (15:1), the penalty mode of the loss function was balance, and the L-BFGS algorithm was used to solve, obtaining the preferred logistic regression model model1, and the performance evaluation is as shown in Figure 3 .
[0136] Based on the logistic regression model model1 obtained from the training set 1, the stability of all the selected 22 microsatellite sites was calculated to obtain the microsatellite site data set, as shown in Table 5.
[0137] Table 5 22 microsatellite sites
[0138]
[0139]
[0140] Example 5: Establishing a random forest model to predict the MSI overall status of samples
[0141] The above-mentioned 22 microsatellite locus data sets were randomly stratified sampled into a training set and a test set in a ratio of 7:3 to obtain training set 2 and test set 2. The training set 2 data was trained based on the random forest model, and the penalty mode of the loss function was specified as balance to obtain the random forest model of the training set 2.
[0142] Based on the random forest model obtained from the training set 2, the MSI overall status of the test set 2 was predicted, and the accuracy of the MSI overall status of the test sample was evaluated according to the MSI-PCR capillary electrophoresis result. The above steps were repeated until the preferred random forest model model2 was obtained, and the performance evaluation result was as shown in Table 5. Figure 4 .
[0143] Example 6: Application of the prediction model
[0144] 143 samples were obtained for NGS sequencing data analysis and calculation to obtain an external data set. The microsatellite locus stability and overall MSI status of the external data set were predicted by using the logistic regression model model1 and the random forest model model2, and the MSI overall status data obtained from the MSI-PCR capillary electrophoresis according to Example 4 was used for verification, and the results are shown in Table 6.
[0145] Table 6: Verification results of 143 external samples
[0146]
[0147] As can be seen, the sensitivity of the prediction model of the present application is 1, and the specificity is 0.92.
[0148] In actual clinical medication, MSI-L and MSS are generally classified into one category, i.e., MSI-L / MSS. The ROC curve of the model obtained therefrom is as shown in Figure 5 , and the AUC is 0.99.
[0149] Example 7: System for determining the MSI overall status of a sample to be tested
[0150] This embodiment establishes a system for determining the MSI overall status of a sample to be tested based on the above-mentioned method, as shown in Figure 6 , which comprises:
[0151] The sequencing data input module is used for inputting the capture sequencing data of the target region containing 22 microsatellite loci in the sample to be tested, and obtaining the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the 22 microsatellite loci.
[0152] The raw data obtained by the capture sequencing data is subjected to quality control by using the quality control software picard, and the sequencing adapter, low-quality base and sequencing error fragments are filtered, and high-quality data (clean data) is obtained after filtering.
[0153] The clean data is subjected to alignment analysis by using the sequence alignment software bwa-MEM, and the specific genomic position information of the sequence (reference genome is hg19) is obtained; the sample raw bam file is obtained by using the samtools and sambamba software for sorting and deduplication.
[0154] The sample raw bam file is input into the software Msisensor2 for analysis, and the length of the repeat base sequence (microsatellite locus) of the target region is obtained.
[0155] For a microsatellite locus, first, the peak number of the length of the microsatellite locus is counted according to the peak searching algorithm, and then the skewness, kurtosis, standard deviation and standard deviation of the length of the microsatellite locus are calculated by using the following formula respectively:
[0156] The skewness calculation formula is:
[0157]
[0158] The kurtosis calculation formula is:
[0159]
[0160] The standard deviation calculation formula is:
[0161]
[0162] The standard deviation calculation formula is:
[0163]
[0164] Wherein, x i is the number of reads when the length of the microsatellite locus is i, i=[1, n]; is the average of the number of reads of different lengths of the microsatellite locus; x′ i is the normalized value of the number of reads when the length of the microsatellite locus is i, is x′ ithe mean value of the peak number of the microsatellite site; n is the number of length categories of the microsatellite site, n = 2γ; γ is the standard value of the length of the microsatellite site.
[0165] ②a microsatellite site storage module, configured to store the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the six microsatellite sites of the first population sample, and stability data;
[0166] ③a marker microsatellite site stability prediction module, connected with the sequencing data input module, the microsatellite site storage module, and the marker microsatellite site data input module, configured to construct 22 microsatellite site stability prediction models by using the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the six microsatellite sites of the first population sample, and stability data, and predict the stability of the 22 microsatellite sites based on the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the 22 microsatellite sites obtained from the sequencing data input module, and output the stability data of the 22 microsatellite sites to the marker microsatellite site data input module.
[0167] Constructing the marker microsatellite site stability prediction model by using the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the known microsatellite site, and stability data comprises the following steps:
[0168] S11, randomly dividing the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the known microsatellite site of the first population sample, and stability data into two groups, one group being a first training set and the other group being a first test set;
[0169] S12, constructing a marker microsatellite site stability prediction model based on a logistic regression algorithm by using the first training set data;
[0170] S13, verifying the obtained marker microsatellite site stability prediction model in the first test set.
[0171] ④a marker microsatellite site data input module, configured to receive the stability data of the 22 microsatellite sites of a sample to be tested, the stability including stable and unstable;
[0172] ⑤an MSI overall state storage module, configured to store the stability data of the marker microsatellite site of a second population sample and MSI overall state data, the MSI overall state including MSI-H, MSI-L, and MSS;
[0173] ⑥an MSI overall state determination module, connected with the data input module and the MSI overall state storage module, configured to construct a prediction model by using the stability data of the marker microsatellite site of the second population sample, and determine the MSI overall state of the sample to be tested based on the stability data of the marker microsatellite site of the sample to be tested obtained from the marker microsatellite site stability prediction module.
[0174] The constructing a prediction model by using the stability data of the marker microsatellite loci of the population sample comprises the following steps:
[0175] S21, the stability data of the marker microsatellite loci of the second population sample is randomly divided into two groups, one group is a second training set, and the other group is a second test set, and the stability data of the marker microsatellite loci of MSI-H samples, MSI-L samples and MSS samples are included in each group;
[0176] S22, by using the second training set data, an MSI overall state prediction model is constructed based on a random forest algorithm;
[0177] S23, in the second test set, the obtained MSI overall state prediction model is verified.
[0178] All the documents mentioned in the present application are cited as references in the present application, as if each document is cited as a reference individually. In addition, it should be understood that, after reading the above teaching of the present application, those skilled in the art can make various modifications or amendments to the present application, and these equivalent forms also fall within the scope defined by the claims attached to the present application.
Claims
1. A method for constructing an MSI overall status prediction model, characterized in that: The stability data of the marker microsatellite loci and the MSI overall status data of the second population samples were used to construct the MSI overall status, wherein the MSI overall status includes MSI-H, MSI-L and MSS, and the marker microsatellite loci are composed of the following loci: Wherein, the position refers to the position on the reference genome hg19, The method comprises the following steps: S21, randomly dividing the stability data of the marker microsatellite loci of the second population samples into two groups, one group being a second training set and the other group being a second test set, each group including the stability data of the marker microsatellite loci of MSI-H samples, MSI-L samples, and MSS samples; S22, using the second training set data, constructing an MSI overall status prediction model based on a machine learning algorithm; S23, in the second test set, the obtained prediction model is verified. The machine learning algorithm is a random forest algorithm.
2. The method according to claim 1, characterized in that The stability data of the marker microsatellite loci are obtained based on the PCR method.
Citation Information
Patent Citations
Method for determining MSI state of colorectal cancer patient based on single sample detection of next-generation sequencing platform and application
CN115595371A
Microsatellite unstable site combination based on NGS technology and application and detection method thereof
CN115786520A
Microsatellite instability detection method and system
CN116438602A
Method and device for detecting microsatellite state of plasma sample
CN116543835A