A method for constructing a prediction model for stability of a marker microsatellite site

By constructing a marker microsatellite locus stability prediction model and using machine learning algorithms to analyze microsatellite locus characteristics, the problem of requiring normal sample pairing in existing MSI detection is solved, and efficient and accurate prediction of the overall MSI status is achieved.

CN118942537BActive Publication Date: 2026-01-23HANGZHOU LIANKANG MEDICAL LAB CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410929917.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-21
Publication Date
2026-01-23
Estimated Expiration
2043-09-21

AI Technical Summary

Technical Problem

Existing MSI detection methods require normal sample pairing, resulting in high detection costs and long procedures. Furthermore, single-site sequencing has low reliability and makes it difficult to accurately predict microsatellite instability.

Method used

This paper presents a liquid-phase capture next-generation sequencing method that predicts the overall MSI status by constructing a marker microsatellite locus stability prediction model without the need for normal sample pairing. The method utilizes machine learning algorithms to analyze the number of peaks, kurtosis, skewness, standard deviation, and standard offset of microsatellite loci.

Benefits of technology

It enables the prediction of overall MSI status across different NGS detection platforms and cancer types without the need for normal sample pairing. The results are accurate, stable, and fast, reducing detection limitations and improving detection reproducibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118942537B_ABST
    Figure CN118942537B_ABST
Patent Text Reader

Abstract

The application discloses a method for constructing a marker microsatellite site stability prediction model, and belongs to the technical field of MSI detection. The application establishes a regression model through known microsatellite sites to predict the stability of 22 marker microsatellite sites. The stability data of the marker microsatellite sites is obtained by using the method, and a machine learning model can be further established to predict the MSI overall state. By using the application, the MSI overall state can be predicted without pairing normal samples in different NGS detection platforms and different cancer types through analysis of second-generation sequencing data, and the method is stable, fast, accurate, and highly repeatable, and the detection limitation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patents

[0002] This application is a divisional application of Chinese invention patent application No. 2023112243588, filed on September 21, 2023, entitled "Microsatellite loci, system and reagent kit for predicting MSI status". Technical Field

[0003] This invention belongs to the field of MSI detection technology, specifically, it relates to a method for constructing a marker microsatellite site stability prediction model. Background Technology

[0004] Microsatellite instability (MSI) refers to the alteration of microsatellite (MS) sequence length caused by insertion or deletion mutations during DNA replication, often due to defects in mismatch repair (MMR). MS sequences are short, repetitive DNA sequences, typically consisting of 1–6 nucleotides arranged in tandem repeats, commonly of dibasic (CA / GA / GT) or monobasic (A / T) types. MS sequences can be located in important non-coding regions or coding regions of genes, exhibiting polymorphism throughout the genome and significant individual variability. Cells can repair mismatches occurring during genome replication using mismatch repair elements. In some cases, mismatch repair elements lose their original function; for example, mutations in MLH1, MSH2, and MSH6 can prevent cell repair of mismatches, resulting in an MSI phenotype. This phenotype manifests as changes in the number of repeating units in microsatellites, which may become longer or shorter.

[0005] Microsatellite instability is a relatively common phenomenon in tumors. Its state indicates the cause and development of tumors and plays an important role in auxiliary diagnosis and medication guidance in different cancer types. Generally, microsatellite instability can be classified into highly unstable microsatellite disease (MSI-H), poorly unstable microsatellite disease (MSI-L), and stable microsatellite disease (MSS). In colorectal cancer, endometrial cancer, and gastric cancer, patients with MSI-H status show significant differences in survival, medication preference, and palliative treatment prognosis compared to the other two types. MSI testing helps physicians fully understand the cancer type and propose appropriate treatment plans.

[0006] Another important role of MSI detection is to assist in screening Lynch Syndrome, a hereditary syndrome characterized by germline mutations in mismatch repair genes, which is characterized by colorectal cancer, endometrial cancer, gastric cancer, and other cancers in younger patients and their families. MSI status is related to Lynch Syndrome, and most patients with microsatellite stability are not Lynch Syndrome. Traditionally, if a patient has the above symptoms, the doctor will usually consider detailed testing. If using the method of immunohistochemistry, two experienced pathologists need to jointly detect, and the accuracy is low. If MSI detection is used, it is more convenient, and the accuracy is also higher. When the patient is determined to be MSI-H, further genetic diagnosis can be performed to determine whether it is Lynch Syndrome.

[0007] Due to the sequence characteristics of MSI sites, the reliability of single site sequencing is low, and MSI detection based on second-generation sequencing usually detects hundreds of MSI sites. The method of multi-site cloning and genome sequencing can be used to detect the microsatellite state. Multi-site cloning is commonly used in current MSI detection. By judging whether the length of 5-10 microsatellite regions highly related to MSI status changes, it is determined whether the site is unstable. The higher the proportion of unstable sites, the more likely the cell is in the MSI state. This method is cheap and has a short experimental process. However, many detection tools for predicting MSI status by NGS method need to be paired with normal samples, such as mSINGS, MANTIS, etc. SUMMARY

[0008] The main purpose of the present application is to provide a method for predicting MSI overall status without normal sample pairing based on liquid-phase capture of second-generation sequencing microsatellite site data. The method is convenient, stable, fast, accurate, and highly reproducible, reducing the limitations of detection.

[0009] To achieve the above purpose, the technical scheme provided by the present application is as follows:

[0010] The first aspect of the present application provides a marker microsatellite site for determining the MSI overall status of a test sample, the marker microsatellite site comprising:

[0011]

[0012]

[0013] Microsatellite (MS) refers to short tandem repeat sequence (Short tandem repeat, STR) in human genome, or repeat base sequence, including single nucleotide repeat, double nucleotide repeat, and even more nucleotide repeat. Microsatellite instability (Microsatellite instability, MSI) refers to the change of arbitrary length of microsatellite in tumor tissue due to insertion or deletion of repeat unit compared with normal tissue. In the present application, the MSI overall state refers to the state of all microsatellites in the sample to be tested; the microsatellite site refers to a specific microsatellite site, and the marker microsatellite site refers to a microsatellite used to determine the MSI overall state of the sample to be tested.

[0014] In the present application, other microsatellites of single nucleotide repeat sequence and / or double nucleotide repeat sequence can also be selected or included as marker microsatellite sites.

[0015] In the present application, the number of repeats is the standard value (γ) of the length of the marker microsatellite site, that is, the number of repeats (standard number of repeats) in the above table.

[0016] In the present application, since the length of the microsatellite changes arbitrarily due to the insertion or deletion of the repeat unit, the length of the microsatellite is the specific number of repeats of the repeat unit, that is, the length is how many times the repeat is. For example, even if the number of bases of the repeat unit is 2 and the repeat is 10 times, the length of the base is 20, but it is still considered that the length of the microsatellite site is 10.

[0017] For each repeat unit of the microsatellite site, there are three states of deletion, no change and insertion, that is, the number of each repeat unit can be 0, 1 or 2. Based on this, for each microsatellite site, the length value range is [0, 2γ].

[0018] The second aspect of the present application provides the use of the detection reagent of the marker microsatellite site of the first aspect of the present application in the preparation of a kit for determining the MSI overall state of the sample to be tested.

[0019] The third aspect of the present application provides a system for determining the MSI overall state of the sample to be tested, comprising the following modules:

[0020] The marker microsatellite site data input module is used to receive the stability data of the marker microsatellite site of the first aspect of the present application of the sample to be tested, and the stability includes stable and unstable;

[0021] an MSI overall status storage module, configured to store stability data of the marker microsatellite loci of the second population sample and MSI overall status data, wherein the MSI overall status comprises MSI-H, MSI-L and MSS;

[0022] an MSI overall status determination module, connected with the data input module and the MSI overall status storage module respectively, configured to construct a prediction model by using the stability data of the marker microsatellite loci of the second population sample, and determine the MSI overall status of the sample to be tested based on the stability data of the marker microsatellite loci of the sample to be tested obtained from the marker microsatellite locus stability prediction module.

[0023] In some embodiments of the present application, the stability data of the marker microsatellite loci refers to whether the length of the marker microsatellite loci changes, if a significant change occurs in a marker microsatellite locus, i.e. a significant change in the length of the marker microsatellite locus due to the insertion or deletion of repeat units, then the marker microsatellite locus is unstable; otherwise, if no change occurs in a marker microsatellite locus, i.e. there is no change in the length of the marker microsatellite locus due to the insertion or deletion of repeat units or even if there is the insertion or deletion of repeat units, the length change is not significant, then the marker microsatellite locus is stable.

[0024] In the present application, the MSI-H, MSI-L and MSS have the following meanings:

[0025] (1) MSS: Microsatellite Stability, microsatellite stability, i.e. all microsatellite loci are in a stable state.

[0026] (2) MSI-L: Low-frequency MSI, low-frequency microsatellite instability, i.e. microsatellite loci less than a preset threshold are in an unstable state.

[0027] (3) MSI-H: High-frequency MSI, high-frequency microsatellite instability, i.e. microsatellite loci not less than a preset threshold are in an unstable state.

[0028] In some embodiments of the present application, the preset threshold can be an absolute value of quantity or a proportion, for example 30%.

[0029] In some embodiments of the present application, in the MSI overall status determination module, the step of constructing a prediction model by using the stability data of the marker microsatellite loci of the population sample comprises the following steps:

[0030] S21, randomly dividing the stability data of the marker microsatellite loci of the second population sample into two groups, one group being a second training set and the other group being a second test set, each group comprising the stability data of the marker microsatellite loci of the MSI-H sample, the MSI-L sample and the MSS sample;

[0031] S22, constructing an MSI overall state prediction model based on a machine learning algorithm using the second training set data;

[0032] S23, verifying the obtained MSI overall state prediction model in the second test set.

[0033] In some embodiments of the present application, the machine learning algorithm is selected from any one of the following algorithms:

[0034] a random forest algorithm, a neural network algorithm, a support vector machine algorithm, a Bayesian classification algorithm, a gradient boosting algorithm, a K-nearest neighbor algorithm and a decision tree algorithm.

[0035] In some preferred embodiments of the present application, the MSI overall state prediction model is obtained by training using a random forest model.

[0036] In some embodiments of the present application, the stability data of the marker microsatellite loci obtained in the marker microsatellite loci data input module is obtained based on a PCR-based method.

[0037] Further, the system further comprises:

[0038] a sequencing data input module for inputting capture sequencing data of a target region comprising the marker microsatellite loci in the sample to be tested, and obtaining the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the marker microsatellite loci;

[0039] a microsatellite loci storage module for storing the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite loci and the stability data of the first population sample;

[0040] a marker microsatellite loci stability prediction module connected with the sequencing data input module, the marker microsatellite loci storage module and the marker microsatellite loci data input module, respectively, for constructing a marker microsatellite loci stability prediction model using the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite loci and the stability data of the first population sample, predicting the stability of the marker microsatellite loci based on the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the marker microsatellite loci obtained from the sequencing data input module, and outputting the stability data of the marker microsatellite loci to the marker microsatellite loci data input module.

[0041] In some embodiments of the present application, the known microsatellite sites include:

[0042]

[0043]

[0044] In some embodiments of the present application, in the marker microsatellite site stability prediction module, the step of constructing a marker microsatellite site stability prediction model using the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite site and stability data includes the following steps:

[0045] S11, the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the known microsatellite site and stability data of the first population sample are randomly divided into two groups, one group is the first training set, and the other group is the first test set;

[0046] S12, using the first training set data, a marker microsatellite site stability prediction model is constructed based on a regression algorithm;

[0047] S13, in the first test set, the obtained marker microsatellite site stability prediction model is verified.

[0048] In some embodiments of the present application, the regression algorithm is selected from any one of the following algorithms: logistic regression algorithm, linear regression algorithm.

[0049] In some embodiments of the present application, for any one of the marker microsatellite sites, first, the peak number of the length of the marker microsatellite site is counted according to the peak searching algorithm, and then the skewness Skew, kurtosis Kurt, standard deviation S and standard deviation P of the length of the marker microsatellite site are calculated respectively:

[0050] Skewness measures the asymmetry of the probability distribution of a random variable, which is a measure of the degree of asymmetry relative to the mean. By measuring the skewness coefficient, the degree of asymmetry and direction of the length distribution of the marker microsatellite site can be determined. The measurement of skewness is relative to the normal distribution, and the skewness of the normal distribution is 0, that is, if the length distribution of the marker microsatellite site is symmetric, the skewness is 0. If the skewness is greater than 0, the distribution is right-skewed, that is, the distribution has a long tail on the right; if the skewness is less than 0, the distribution is left-skewed, that is, the distribution has a long tail on the left; at the same time, the greater the absolute value of skewness, the more serious the degree of skewness.

[0051] The formula for calculating skewness is:

[0052]

[0053] Where, x iThe number of reads of the marker microsatellite site with length i, i = [1, n]; The average of the number of reads of the marker microsatellite site with different lengths; n is the number of length types of the marker microsatellite site, i.e., how many different lengths, n = 2γ; γ is the standard value of the length of the marker microsatellite site.

[0054] Kurtosis is a statistical quantity for studying the steepness or flatness of data distribution. By measuring the kurtosis coefficient, it can be determined whether the length of the marker microsatellite site is steeper or flatter relative to the normal distribution. If kurtosis = 3, the peak state of the length distribution of the marker microsatellite site obeys the normal distribution; if kurtosis > 3, the peak state of the length distribution of the marker microsatellite site is steep (high and sharp); if kurtosis < 3, the peak state of the length distribution of the marker microsatellite site is flat (low and fat). The kurtosis calculation formula is:

[0055]

[0056] wherein x i The number of reads of the marker microsatellite site with length i, i = [1, n]; The average of the number of reads of the marker microsatellite site with different lengths; n is the number of length types of the marker microsatellite site, n = 2γ, and γ is the standard value of the length of the marker microsatellite site.

[0057] The standard deviation reflects the dispersion degree of the length of the marker microsatellite site. The greater the value, the more dispersed, i.e., the greater the difference between the different lengths of the marker microsatellite site.

[0058] The standard deviation calculation formula is:

[0059]

[0060] wherein x i The number of reads of the marker microsatellite site with length i, i = [1, n]; The average of the number of reads of the marker microsatellite site with different lengths; x i ′ The normalized value of the number of reads of the microsatellite site with length i, The average of x i ′ ; n is the number of length types of the marker microsatellite site, n = 2γ, and γ is the standard value of the length of the marker microsatellite site.

[0061] The standard deviation is the level of the length of the marker microsatellite site deviating from the standard value of the length of the marker microsatellite site.

[0062] The standard deviation calculation formula is:

[0063]

[0064] wherein, x i is the number of reads of the marker microsatellite site with length i, i = [1, n]; x i is the normalized value of the number of reads of the marker microsatellite site with length i, n is the number of length categories of the marker microsatellite site; γ is the standard value of the length of the marker microsatellite site, n = 2γ, γ is the standard value of the length of the marker microsatellite site.

[0065] In some embodiments of the present application, in the sequencing data input module, the peak number, kurtosis, skewness, standard deviation and standard deviation of the marker microsatellite site are calculated only when the depth of the captured sequencing of the target region reaches 400x. Specifically, the number of reads of different lengths of the marker microsatellite site is calculated using the result file of Msisensor2.

[0066] The fourth aspect of the present application provides a kit for determining the MSI global state of a test sample, comprising the detection reagent of the marker microsatellite site of the first aspect of the present application.

[0067] In the present application, the "first population sample" and the "second population sample" are only formal regions, wherein the data of the first population sample includes the peak number, kurtosis, skewness, standard deviation and standard deviation of the known microsatellite site length of each sample and stability data, which are used to construct a marker microsatellite site stability prediction model based on a regression algorithm; the data of the second population sample includes the stability data of the marker microsatellite site and the MSI global state data, which are used to construct an MSI global state prediction model.

[0068] In the present application, the test sample is derived from human, preferably a tumor sample such as fresh tissue, tissue paraffin block (FFPE) and the like.

[0069] Advantages of the present application

[0070] Compared with the prior art, the present application has the following advantages:

[0071] By using the present application, the MSI global state can be predicted without pairing with normal samples in different NGS detection platforms and different cancer types through analysis of second-generation sequencing data.

[0072] By using the present application to predict the MSI global state, the method is stable, fast, accurate and highly reproducible, and reduces the limitations of detection. BRIEF DESCRIPTION OF DRAWINGS

[0073] Figure 1The length distribution of one microsatellite site in Example 3 of the present application is shown.

[0074] Figure 2 The MSI-PCR capillary electrophoresis detection structure of one microsatellite site in Example 4 of the present application is shown. A: tumor sample; B: blood sample.

[0075] Figure 3 The performance evaluation results of the stability of the microsatellite site predicted by the logistic regression model established in Example 4 of the present application are shown.

[0076] Figure 4 The performance evaluation results of the random forest model established in Example 5 of the present application for predicting the MSI overall state are shown.

[0077] Figure 5 The validation results of the external data set in Example 6 of the present application are shown.

[0078] Figure 6 The schematic diagram of the system for determining the MSI overall state of the sample to be tested constructed in Example 7 of the present application is shown. DETAILED DESCRIPTION

[0079] Unless otherwise indicated, all parts and percentages in the present application are based on weight, and the test and characterization methods used are those in existence at the filing date of the present application. To the extent applicable, the contents of any patent, patent application or publication referred to in the present application are incorporated herein by reference in their entirety, and equivalent same family patents are introduced by reference, especially the definitions of the relevant terms disclosed in these documents. If the definition of a specific term disclosed in the prior art is inconsistent with any definition provided in the present application, the definition provided in the present application shall prevail.

[0080] The numerical ranges in the present application are approximate, and thus can include values outside the stated ranges, unless otherwise indicated. A range includes all values from the lower value to the upper value, increasing by 1 unit, provided that there is at least 2 units of separation between any lower value and any higher value. For ranges containing values less than 1 or containing fractional numbers (e.g. 1.1, 1.5, etc.), the unit of 1 is suitably treated as 0.0001, 0.001, 0.01, or 0.1. For ranges containing decimal numbers less than 10 (e.g. 1 to 5), the unit of 1 is generally treated as 0.1. These are merely specific examples of what is meant by the expression, and all possible combinations of numerical values between the lowest value and the highest value are considered to be clearly described in the present application.

[0081] The terms "comprising," "including," "containing," and "having," and their derivatives, are not intended to exclude any component, step or procedure not specified, and are used to allow for the presence of additional components, steps or procedures, even though such are not necessarily present. For the avoidance of doubt, unless clearly specified otherwise, the use of the term "comprising" in the present application is taken to mean the inclusion of all elements and / or steps set forth in the claim, or the inclusion of any additional components, steps or procedures, to the extent that such are not already expressly set forth in the claim. In contrast, the term "consisting essentially of is used to exclude any additional components, steps or procedures not necessary to the operation of the application, as set forth in the following claims. The term "consisting of is used to exclude any component, step or procedure not specifically described or listed in the claim. The term "or" as used in the claims is to be interpreted as inclusive or meaning any one or any combination of the listed members.

[0082] In order to make the technical problems solved by the present application, the technical solutions and the beneficial effects clearer, the present application will be further explained in details in combination with the embodiments.

[0083] Embodiment

[0084] The following examples are given to illustrate preferred embodiments of the application. Those skilled in the art will readily understand that the techniques given in the following examples represent techniques discovered by the inventors to function well in the practice of the application, and can be varied in various ways. Therefore, the following examples are to be considered in all respects as being given by way of illustration and not of limitation.

[0085] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The materials, methods, and examples provided herein are illustrative only and are not intended to be limiting.

[0086] Those of skill will recognize or be able to ascertain using no more than routine experimentation many equivalents of the specific embodiments of the application described herein. Such equivalents are intended to be encompassed by the following claims.

[0087] The experimental methods in the following examples are routine methods unless specifically indicated otherwise. The instruments and equipment used in the following examples are routine laboratory instruments and equipment unless specifically indicated otherwise; the test materials used in the following examples are commercially available from routine biochemical reagent stores unless specifically indicated otherwise.

[0088] Example 1 Selection of microsatellite sites

[0089] 800 clinical tumor tissue samples from patients with colorectal cancer, lung cancer and other cancer types were collected for MSI prediction.

[0090] As for the selection of microsatellite sites, the present embodiment uses a large gene combination panel (as shown in Table 1) containing 169 genes related to solid tumors such as colorectal cancer, lung cancer, endometrial cancer, gastric cancer, etc. developed based on VariantBaitsTM technology. The DNA extracted from the sample is hybridized and captured using the target region probe, and then the corresponding sample library is obtained by library construction. Further, the fastq format of the raw data of the second-generation sequencing is obtained by further sequencing the library.

[0091] Table 1 169 gene list

[0092] AKT1 ALK APC ARAF ARID1A ATM BRAF BRCA1 BRCA2 CTNNB1 DDR2 DNMT3A EGFR ERBB2 ERBB3 FGF19 FGFR1 FGFR2 FGFR3 FGFR4 FH FLCN HFE HOXB13 IDH1 IDH2 JAK1 JAK2 KDM6A KIT KRAS MAP2K1 MDM4 MTOR NF2 NRAS NRG1 NTRK1 NTRK2 NTRK3 PDGFRA PIK3C2G PTEN RB1 RET ROS1 SDHA SDHB SMAD4 STK11 TERT TP53 TSC1 TSC2 UGT1A1 VEGFA VHL WT1 BAP1 BARD1 BRIP1 CDK12 CHEK1 CHEK2 ERCC1 FANCA FANCL MEN1 MLH1 MSH6 MUTYH NBN NF1 PALB2 PMS2 POLD1 POLE PPP2R2A RAD51C RAD51D RAD54L SDHAF2 SDHC SDHD EPCAM HLA-A HLA-B ABCB1 ABCC2 ABCC4 ATIC C8orf34 CBR3 CDA CYP19A1 CYP1B1 CYP2C8 CYP2D6 CYP3A4 DHFR DPYD DYNC2H1 ERCC2 FOLR3 GGH MTHFR MTRR NQO1 RRM1 SEMA3C SLC19A1 SLC22A16 SLC22A2 SLCO1B1 TPMT TYMS UMPS XPC XRCC1 ABL1 CDH1 CDKN2A CSF1R EZH2 FBXW7 FLT3 GNAQ GNAS HNF1A JAK3 KDR NOTCH1 PBRM1 PTPN11 SMARCB1 B2M ESR1 HRAS MET PIK3CA TGFBR2 BMPR1A MSH2 RAD51B HLA-C CYP2B6 GSTP1 SOD2 ERBB4 NPM1

[0093] Sample library sequencing and sequence comparison of Example 2

[0094] The raw data obtained in Example 1 is subjected to quality control using the quality control software picard, filtering sequencing adapters, low-quality bases, and sequencing error fragments, etc. After filtering, high-quality data (clean data) is obtained.

[0095] The clean data is analyzed using the sequence alignment software bwa-MEM to obtain the specific genomic position information of the sequence (reference genome is hg19); samtools and sambamba software are used for sorting and deduplication to obtain the sample raw bam file.

[0096] The sample raw bam file is input into the software Msisensor2 for analysis to obtain the number of reads of different lengths of target region repeat base sequences, i.e. the sample data set.

[0097] Example 3 Microsatellite site stability state prediction

[0098] The target region sequencing depth is at least 400x, and subsequent calculations are performed under the condition of meeting the depth.

[0099] For a specific microsatellite site: chr2_47641559_CAGGT_27[A]_GGGTT, as shown in Table 2.

[0100] Table 2 chr2_47641559_CAGGT_27[A]_GGGTT information

[0101] Chromosome Position Repeat unit Repeat number (gamma) 3' flanking sequence 5' flanking sequence chr2 47641559 A 27 CAGGT GGGTT

[0102] The length distribution is shown in Table 3 and Figure 1 .

[0103] Table 3 chr2_47641559_CAGGT_27[A]_GGGTT length distribution

[0104]

[0105]

[0106] The embodiment utilizes five characteristics of the peak number N, skew Skew, kurtosis Kurt, standard deviation S, and standard deviation P of the length of the target region repeat base sequence to determine the MSI state.

[0107] For a certain repeat base sequence, the following is calculated respectively.

[0108] Peak number calculation:

[0109] The sequencing depth of the target region is calculated, and the peak number of the length of the microsatellite site obtained is calculated according to the peak searching algorithm. Specifically, Scipy is used to calculate the peak prominence to calibrate the peak number and remove noise.

[0110] In the embodiment, the peak number of chr2_47641559_CAGGT_27[A]_GGGTT is 1.

[0111] Then, using the known stable length γ of the microsatellite site, the theoretical maximum length of the microsatellite site is 2γ, that is, 54, so the length type number n of chr2_47641559_CAGGT_27[A]_GGGTT is 54, and i = [1, 54].

[0112] The skew calculation formula is:

[0113]

[0114] The kurtosis calculation formula is:

[0115]

[0116] The standard deviation calculation formula is:

[0117]

[0118] The standard deviation calculation formula is:

[0119]

[0120] Wherein, x i is the number of reads when the length of the microsatellite site is i; is the average number of reads of different lengths of the microsatellite site. x i ′the normalized value of the number of reads with the length of i for the microsatellite site, x = 0.5 i ′ the mean value of x.

[0121] According to the above formula, the skewness, kurtosis, standard deviation and standard deviation of the length of chr2_47641559_CAGGT_27[A]_GGGTT are 2.80, 9.87, 0.05 and 19.82, respectively.

[0122] Example 4: Establishing a logistic regression model to predict the stability of a microsatellite site

[0123] In this example, a model is established using 6 microsatellite sites with known stability data to predict the stability of microsatellite sites with unknown stability data. The 6 known microsatellite sites are shown in Table 4.

[0124] Table 4: 6 known microsatellite sites

[0125] Site number Chromosome Position Repeat unit Repeat number (gamma) 3' flanking sequence 5' flanking sequence 1 chr11 102193508 A 26 CTGGT GCCAC 2 chr14 23652346 A 21 TTGCT GGCCA 3 chr4 55598211 T 25 TTTGA GAGAA 4 chr2 95849361 T 23 TCCTA GTGAG 5 chr2 47641559 A 27 CAGGT GGGTT 6 chr2 39564893 T 28 TCCAG GAGAC

[0126] In order to establish a model for the stability of microsatellite sites, verify the MSI prediction results and evaluate their performance, the inventors performed MSI-PCR capillary electrophoresis detection on the samples, and the results were used as the gold standard to compare the consistency between the two results. The kit used for MSI-PCR capillary electrophoresis detection was the Tung Tree Microsatellite Instability (MSI) Detection Kit (Multiplex Fluorescence PCR-Capillary Electrophoresis Method). According to the MSI-PCR capillary electrophoresis detection, the state of each microsatellite site in each sample and the overall MSI state of the sample were obtained.

[0127] Specifically, according to the results of MSI-PCR capillary electrophoresis, each microsatellite site is divided into two categories: microsatellite site stable and microsatellite site unstable. If a microsatellite site changes, i.e., there is a significant change in the length of the microsatellite site due to the insertion or deletion of repeat units, then the microsatellite site is unstable. Conversely, if a microsatellite site does not change, i.e., there is no change in the length of the microsatellite site due to the insertion or deletion of repeat units or even if there is an insertion or deletion of repeat units, the length change is not significant, then the marker microsatellite site is stable.

[0128] For example, as shown in Figure 1, the MSI-PCR capillary electrophoresis detection result of the BAT25 site (chr4 55598211) is shown. This is a homozygous site, and as can be seen in the comparison between the tumor sample and the paired sample, a new set of displacement peaks has been added in the tumor sample (indicated by the dashed line in the figure). Therefore, this site is in a microsatellite site unstable state in the tumor sample, and in a microsatellite site stable state in the paired sample (blood). Figure 2 ​

[0129] Further, according to the stability of the microsatellite site, the MSI overall state of the sample is divided into MSI-H, MSI-L, and MSS, which correspond to microsatellite high instability, microsatellite low instability, and microsatellite stability, respectively.

[0130] MMS: All microsatellite sites are in a stable state.

[0131] MSI-L: < 30% of the microsatellite sites are in an unstable state.

[0132] MSI-H: ≥ 30% of the microsatellite sites are in an unstable state.

[0133] In this embodiment, 800 sample MSI-PCR capillary electrophoresis data were obtained, including the stability data of 6 microsatellite sites and the MSI overall state data of 800 samples.

[0134] The MSI-PCR capillary electrophoresis data of the stability of the 6 microsatellite sites of the above-mentioned 800 samples were processed, and samples with NGS results not meeting the quality control were removed, obtaining a total of 3780 microsatellite site stability data, which were randomly stratified sampled into a training set and a test set according to a 7:3 ratio. For the training set, based on the Msisensor2 result file obtained by Example 2, the length of the 6 microsatellite sites was calculated, and the number of peaks N, kurtosis Kurt, skewness Skew, standard deviation S, and standard deviation P of the length of each microsatellite site were obtained by the method of Example 3, to generate training set data, with microsatellite site unstable samples as 1 and microsatellite site stable samples as 0, forming training set 1. The same method was used to obtain test set 1. The training set 1 data was trained based on the logistic regression model to obtain the logistic regression model of the training set 1.

[0135] Based on the logistic regression model obtained from the training set 1, the microsatellite site stability of the test set 1 was predicted, and the accuracy of the microsatellite site stability of the test set was evaluated according to the MSI-PCR capillary electrophoresis result, and the model was modified to remove noise and modify the model parameters and solution algorithm. Finally, the model regularization parameter is specified as l2, the maximum number of iterations is 5000, and the classification type is binary. Because the number of microsatellite site unstable samples in the actual training set is much less than that of stable samples (15:1), the penalty mode of the loss function is balance, and the L-BFGS algorithm is used to solve, to obtain the preferred logistic regression model model1, and the performance evaluation is as shown in Figure 3

[0136] Based on the logistic regression model model1 obtained from the training set 1, the stability of all the selected 22 microsatellite sites was calculated to obtain the microsatellite site data set, as shown in Table 5. ​

[0137] Table 5 22 microsatellite loci

[0138]

[0139]

[0140] Example 5 Establishing a random forest model to predict the MSI overall status of samples

[0141] The above 22 microsatellite loci data set is randomly stratified sampled into a training set and a test set in a ratio of 7:3 to obtain training set 2 and test set 2. The training set 2 data is trained based on the random forest model, and the penalty mode of the loss function is specified as balance to obtain the random forest model of the training set 2.

[0142] Based on the random forest model obtained from the training set 2, the MSI overall status of the test set 2 is predicted, and the accuracy of the MSI overall status of the test sample is evaluated according to the MSI-PCR capillary electrophoresis result. The above steps are repeated until the preferred random forest model model2 is obtained, and the performance evaluation result is as shown in Figure 4 .

[0143] Example 6 Application of the prediction model

[0144] Obtain 143 samples for NGS sequencing data analysis and calculation to obtain an external data set. The microsatellite locus stability and overall MSI status of the external data set are predicted by using the logistic regression model model1 and the random forest model model2, and the MSI overall status data obtained from the MSI-PCR capillary electrophoresis according to Example 4 is verified, and the result is as shown in Table 6.

[0145] Table 6 Verification results of 143 external samples

[0146]

[0147] It can be seen that the sensitivity of the prediction model of the application is 1, and the specificity is 0.92.

[0148] In actual clinical medication, MSI-L and MSS are generally classified into one category, i.e. MSI-L / MSS. The ROC curve of the model obtained therefrom is as shown in Figure 5 , and the AUC is 0.99.

[0149] Example 7 System for determining the MSI overall status of a sample to be tested

[0150] This embodiment establishes a system for determining the MSI overall status of a sample to be tested based on the above method, as shown in Figure 6 , which comprises:

[0151] The ① sequencing data input module is used for inputting the capture sequencing data of the target region containing 22 microsatellite loci in the sample to be tested, and obtaining the peak number, kurtosis, skewness, standard deviation and standard deviation of the length of the 22 microsatellite loci.

[0152] The raw data obtained by the capture sequencing data is subjected to quality control by using the quality control software picard, and the sequencing adapter, low-quality base and sequencing error fragments are filtered, and high-quality data (clean data) is obtained after filtering.

[0153] The clean data is subjected to alignment analysis by using the sequence alignment software bwa-MEM, and the specific genomic position information of the sequence (reference genome is hg19) is obtained; samtools and sambamba software are used for sorting and deduplication, and the sample raw bam file is obtained.

[0154] The sample raw bam file is input into the software Msisensor2 for analysis, and the length of the repeat base sequence (microsatellite locus) of the target region is obtained.

[0155] For a certain microsatellite locus, first, the peak number of the length of the microsatellite locus is counted according to the peak searching algorithm, and then the skewness, kurtosis, standard deviation and standard deviation of the length of the microsatellite locus are calculated by using the following formula respectively:

[0156] The skewness calculation formula is:

[0157]

[0158] The kurtosis calculation formula is:

[0159]

[0160] The standard deviation calculation formula is:

[0161]

[0162] The standard deviation calculation formula is:

[0163]

[0164] Wherein, x i is the number of reads when the length of the microsatellite locus is i, i=[1, n]; is the average of the number of reads of different lengths of the microsatellite locus; x i ′ is the normalized value of the number of reads when the length of the microsatellite locus is i, is x i ′the mean value of the peak number of the microsatellite site; n is the number of length categories of the microsatellite site, n = 2γ; γ is the standard value of the length of the microsatellite site.

[0165] ②a microsatellite site storage module, configured to store the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the six microsatellite sites of the first population sample, and stability data;

[0166] ③a marker microsatellite site stability prediction module, connected with the sequencing data input module, the microsatellite site storage module, and the marker microsatellite site data input module, configured to construct 22 microsatellite site stability prediction models by using the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the six microsatellite sites of the first population sample, and stability data, and predict the stability of the 22 microsatellite sites based on the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the 22 microsatellite sites obtained from the sequencing data input module, and output the stability data of the 22 microsatellite sites to the marker microsatellite site data input module.

[0167] Constructing the marker microsatellite site stability prediction model by using the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the known microsatellite site, and stability data comprises the following steps:

[0168] S11, randomly dividing the peak number, kurtosis, skewness, standard deviation, and standard deviation of the length of the known microsatellite site of the first population sample, and stability data into two groups, one group as a first training set, and the other group as a first test set;

[0169] S12, constructing a marker microsatellite site stability prediction model based on a logistic regression algorithm by using the first training set data;

[0170] S13, verifying the obtained marker microsatellite site stability prediction model in the first test set.

[0171] ④a marker microsatellite site data input module, configured to receive the stability data of the 22 microsatellite sites of a sample to be tested, the stability including stable and unstable;

[0172] ⑤an MSI overall state storage module, configured to store the stability data of the marker microsatellite site of a second population sample and MSI overall state data, the MSI overall state including MSI-H, MSI-L, and MSS;

[0173] ⑥an MSI overall state determination module, connected with the data input module and the MSI overall state storage module, configured to construct a prediction model by using the stability data of the marker microsatellite site of the second population sample, and determine the MSI overall state of the sample to be tested based on the stability data of the marker microsatellite site of the sample to be tested obtained from the marker microsatellite site stability prediction module.

[0174] The constructing a prediction model by using the stability data of the marker microsatellite loci of the population sample comprises the following steps:

[0175] S21, the stability data of the marker microsatellite loci of the second population sample is randomly divided into two groups, one group is a second training set, and the other group is a second test set, and the stability data of the marker microsatellite loci of MSI-H samples, MSI-L samples and MSS samples are included in each group;

[0176] S22, by using the second training set data, an MSI overall state prediction model is constructed based on a random forest algorithm;

[0177] S23, in the second test set, the obtained MSI overall state prediction model is verified.

[0178] All the documents mentioned in the present application are cited as references in the present application, as if each document is cited as a reference individually. In addition, it should be understood that, after reading the above teaching of the present application, those skilled in the art can make various modifications or amendments to the present application, and these equivalent forms also fall within the scope defined by the claims attached to the present application.

Claims

1. A method for constructing a marker microsatellite site stability prediction model, characterized in that, A stability prediction model for the labeled microsatellite sites is constructed using the peak count, kurtosis, skewness, standard deviation, standard shift, and stability data of the known microsatellite site lengths from the first population sample. The known microsatellite loci consist of the following loci: The labeled microsatellite loci consist of the following loci: The location referred to here is the position on the reference genome hg19. The method includes the following steps: S11, the number of peaks, kurtosis, skewness, standard deviation and standard offset, and stability data of the known microsatellite loci lengths of the first group of samples are randomly divided into two groups, one as the first training set and the other as the first test set; S12, using the first training set data, construct a marker microsatellite site stability prediction model based on a regression algorithm; S13, in the first test set, the obtained marker microsatellite site stability prediction model is validated.

2. The method according to claim 1, characterized in that, For specific marker microsatellite loci: The peak count of the marker microsatellite locus length is counted using the peak-finding algorithm, and then the skewness, kurtosis, standard deviation, and standard offset of the marker microsatellite locus length are calculated using the following formulas: The formula for calculating skewness is: The formula for calculating kurtosis is: The formula for calculating standard deviation is: The standard offset calculation formula is: in, The length of the marked microsatellite locus is The number of reads at that time =[1, ]; This represents the average number of reads of different lengths for the marked microsatellite locus. The length of the microsatellite locus is The normalized value of the number of reads at that time. ; for The mean; The length of the marked microsatellite locus is the number of different types. =2 ; This is the standard value for the length of the marked microsatellite locus.

Citation Information

Patent Citations

  • Detection of microsatellite instability

    CN112639983A

  • Method for determining MSI state of colorectal cancer patient based on single sample detection of next-generation sequencing platform and application

    CN115595371A

  • Microsatellite instability detection method and system

    CN116438602A