Baseline modeling method, MSI state detection method and storage medium
By employing baseline modeling and unique molecular tag grouping, the problems of sequencing platform compatibility and sample type applicability were solved, enabling highly accurate MSI detection in different sequencing platforms and tissue and blood samples, especially in blood samples with low tumor content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AMOY DIAGNOSTICS CO LTD
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are not compatible with different sequencing platforms, cannot be applied to both tissue and blood sample types simultaneously, and their detection accuracy is limited to a single cancer type, making them ineffective for detecting blood samples with low tumor content.
By employing baseline modeling methods, baseline samples are selected for next-generation sequencing. Optimal sequences are screened using unique molecular tags for grouping and base correction, and a detection statistic threshold is constructed to form a baseline model suitable for hybridization capture and amplicon sequencing, covering tissue and blood sample types.
It achieves applicability across different sequencing platforms and sample types, improves detection accuracy, enables stable detection at low tumor concentrations, and is suitable for various cancer types, especially blood samples.
Smart Images

Figure CN122024840A_ABST
Abstract
Description
Technical Field
[0001] This application relates primarily to the field of biological detection technology, and in particular to a baseline modeling method, an MSI state detection method, and a storage medium. Background Technology
[0002] In cases of defective mismatch repair (dMMR) gene dysfunction, microsatellite DNA replication can lead to base mismatches, insertions, or deletions in the double-stranded molecules due to factors such as "strand slippage," resulting in structural changes in the microsatellite. This structurally altered microsatellite is known as microsatellite instability (MSI).
[0003] MSI is an important biomarker for PD-1 / PD-L1 inhibitor therapy and can independently predict the efficacy of immunotherapy. MSI-H / dMMR is present in various solid tumors, with higher positive rates in endometrial cancer, colorectal cancer, and gastric cancer patients (16%–33%, 6%–22%, and 9%–22%, respectively). Other tumor types, such as liver cancer, ampullary cancer, ovarian cancer, cervical cancer, esophageal adenocarcinoma, soft tissue tumors, head and neck cancer, renal cancer, and Ewing's sarcoma, have MSI-H / dMMR positive rates greater than 2%, while prostate cancer, lung cancer, and breast cancer have positive rates less than 2%.
[0004] Current MSI detection based on next-generation sequencing technology has the following problems: (1) A Pon (panel of normal) is needed to build a baseline, but since different sequencing platforms have different sequencing error backgrounds, different baselines are usually needed for different sequencing platforms.
[0005] (2) It has certain requirements on the tumor content of the sample and is usually only applicable to the detection of tissue samples. It cannot be truly applicable to the detection of blood samples with a tumor content as low as 0.5%.
[0006] (3) The verification of accuracy usually focuses only on a single cancer type.
[0007] (4) Due to the different data characteristics of hybridization capture and amplicon sequencing, there are very few second-generation sequencing MSI detection algorithms that can be used for both hybridization capture and amplicon sequencing data. Summary of the Invention
[0008] One objective of this application is to provide a baseline modeling method, an MSI state detection method, and a storage medium to address the problems in existing technologies that cannot be adapted to different sequencing platforms and cannot be applied simultaneously to the two main sample types of tissue and blood.
[0009] According to one aspect of this application, a baseline modeling method is provided, the method comprising: A baseline sample and a target MSI site were selected, and the baseline sample was subjected to next-generation sequencing to obtain the sequencing sequence. The sequencing sequences are grouped according to a unique molecular tag, and the sequencing sequences within each group are base-corrected to select the optimal sequence for the target MSI site in each group. Extract multiple deletion lengths from the optimal sequence of the target MSI sites in each group, and determine the deletion ratio of each deletion length for each target MSI site; A probability quality function is determined based on the proportion of each missing length for each target MSI site, and a detection statistic threshold is constructed based on the probability quality function to form a baseline model for MSI state detection.
[0010] Optionally, the step of grouping the sequencing sequences according to a unique molecular tag and performing base correction on the sequencing sequences within each group includes: The sequencing sequences are grouped according to their unique molecular tags and sequencing start coordinates, and valid groups whose number of sequencing sequences meets a preset threshold are selected. Base correction is performed on the sequencing sequences within each valid group.
[0011] Optionally, the process of selecting the optimal sequence for each group of target MSI sites includes: For each valid group, a multi-level priority selection operation is performed within the group to filter out the optimal sequence of the target MSI site in each group. The multi-level priority selection includes a progressive comparison within each group according to the priority order of multiple preset indicators.
[0012] Optionally, the plurality of preset indicators include sequence number, alignment quality value, average base quality value, and alignment length, and the multi-level priority selection includes: Select the sequencing sequence that supports the largest number of sequences with the same repeat base length; When the number of sequences is the same, the sequencing sequence with the higher alignment quality value is selected; When the alignment quality values are the same, the sequencing sequence with the higher average base quality value is selected. When the average base quality values are the same, the sequencing sequence with the longer alignment length is selected.
[0013] Optionally, the step of determining the probability quality function based on the proportion of each deletion length for each target MSI locus, and constructing a detection statistic threshold based on the probability quality function, includes: Determine the mean of the deletion proportion for each deletion length at each target MSI site, and use the mean as a binomial distribution probability parameter to determine the probability quality function; The probability mass function is statistically processed to obtain the statistics for each target MSI site; Determine the mean and standard deviation of the statistic for each target MSI locus in the baseline sample, and calculate the detection statistic for each baseline sample; The mean and standard deviation of the detection statistic for each baseline sample are calculated to determine the threshold for the detection statistic.
[0014] Optionally, statistical processing is performed on the probability quality function to obtain statistics for each target MSI locus, including: The cumulative distribution function is determined based on the probability mass function and logarithmic processing is performed to obtain the processing result; By summing the processing results for different deletion lengths, the statistics for each target MSI locus are determined.
[0015] Optionally, the step of extracting multiple deletion lengths from the optimal sequence of the target MSI sites in each group and determining the deletion ratio of each deletion length for each target MSI site includes: Determine the optimal repeat sequence length for the target MSI site in each group; The length of the repeated sequence is compared with the benchmark reference sequence, and each optimal sequence is classified to obtain multiple missing lengths; The proportion of each deletion length in the target MSI site is calculated to represent the proportion of the missing sequence at each deletion length among all the optimal sequences at that target MSI site.
[0016] Optionally, the step of comparing the length of the repeating sequence with a benchmark reference sequence and classifying each optimal sequence yields multiple missing lengths, including: The length of the repeated sequence is compared with the benchmark reference sequence, and each optimal sequence is classified to obtain normal sequences and missing sequences; The deletion length is determined based on the difference between the length of the repeat sequence corresponding to the deletion sequence and the length of the repeat sequence of the reference sequence, wherein the deletion length ranges from 1 to 9 bases.
[0017] Optionally, the method further includes: Effective sequences that completely cover the target MSI site are selected from the sequencing sequences obtained from second-generation sequencing. The effective sequences are grouped according to the unique molecular tag, and the effective sequences in each group are base-corrected to select the optimal sequence for the target MSI site in each group.
[0018] According to another aspect of this application, a method for MSI state detection is also provided, the method comprising: Extract multiple deletion lengths for each MSI site in the sample to be tested, and determine the deletion ratio of each deletion length for each MSI site; A probability quality function is determined based on the proportion of each missing length at each MSI site, and an MSI detection statistic for the sample to be tested is constructed based on the probability quality function. The detection statistic threshold is compared with the MSI detection statistic of the sample to be tested to determine the MSI status result of the sample to be tested.
[0019] Optionally, the method includes: The test samples were subjected to next-generation sequencing to screen out effective sequences that completely cover the target MSI sites; The effective sequences are grouped according to the unique molecular tag, the effective sequences in each group are base corrected, and the optimal sequence for each target MSI site is selected.
[0020] According to another aspect of this application, an electronic device is also provided, the device comprising: One or more processors; and A memory storing computer-readable instructions that, when executed, cause the processor to perform operations as described above in the baseline modeling method, or operations as described above in the MSI state detection method.
[0021] According to another aspect of this application, a computer-readable storage medium is also provided, having stored thereon computer-readable instructions that can be executed by a processor to implement the baseline modeling method as described above, or the MSI state detection method as described above.
[0022] Compared with existing technologies, this application selects a baseline sample and a target MSI locus, performs next-generation sequencing on the baseline sample to obtain sequencing sequences; groups the sequencing sequences according to a unique molecular tag, performs base correction on the sequencing sequences within each group, and screens out the optimal sequences for the target MSI locus in each group; extracts multiple deletion lengths from the optimal sequences of the target MSI locus in each group, determines the deletion ratio of each deletion length for each target MSI locus; determines a probability quality function based on the deletion ratio of each deletion length for each target MSI locus, and constructs a detection statistic threshold based on the probability quality function to form a baseline model for MSI status detection. The MSI detection statistic of the sample to be tested is calculated; the detection statistic threshold is compared with the MSI detection statistic of the sample to be tested to determine the MSI status result of the sample to be tested. Therefore, using this baseline model, MSI detection can be performed on any sample to be tested without providing a corresponding control sample for each sample. Furthermore, by using a single method, it can be simultaneously applied to both capture sequencing and amplicon sequencing, the two main next-generation sequencing library preparation methods, and to both tissue and blood, the two main clinical sample types. Attached Figure Description
[0023] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 A schematic flowchart of a baseline modeling method provided according to one aspect of this application is shown. Figure 2 This diagram illustrates a single-end base correction in one embodiment of the present application. Figure 3 This diagram illustrates a double-ended base correction in one embodiment of the present application. Figure 4 This diagram illustrates a method flow chart for MSI status detection according to another aspect of this application. Figure 5 This diagram illustrates a structural schematic of an electronic device according to another aspect of this application.
[0024] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0025] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0026] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein, and therefore this application is not limited to the specific embodiments disclosed below.
[0027] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0028] Figure 1 The diagram illustrates a baseline modeling method according to one aspect of this application, comprising steps S11 to S14. This allows for the establishment of a baseline using only a small sample size of PON from a single sequencing platform, applicable to both hybridization capture and amplicon sequencing, and simultaneously applicable to MSI detection in tissues and blood. Specifically: Step S11: Select a baseline sample and the target MSI site, and perform second-generation sequencing on the baseline sample to obtain the sequencing sequence.
[0029] Select a set of MSI sites with repeat sequence lengths within the target range as target MSI sites. If the repeat sequence length is too short, the true MSI positive signal cannot be fully displayed. The longer the length, the greater the noise background, such as sequencing errors, PCR errors in the library construction process, etc., which can easily drown out the true positive signal. The target range can be between 10-20 bp.
[0030] Capture probe design and / or amplicon primer design are based on any combination of MSI sites, where the number of MSI sites is greater than 50, such as using any combination of 100 MSI sites.
[0031] S out of 100 MSI sites are selected, for example, S = 50. 50 samples that are judged to be in the MSS state are selected as baseline samples, for example, by using PCR method. Each baseline sample is subjected to next-generation sequencing by a sequencing platform to obtain the sequencing sequence. The next-generation sequencing platform includes, but is not limited to, Novaseq6000, Nextseq500, MGI2000, MGI200, Miseq, and Miniseq.
[0032] It should be noted that the PCR method uses fluorescently labeled primers and capillary electrophoresis to determine the fragment length polymorphism (FLP) at five microsatellite loci: NR-21, NR-24, BAT-25, BAT-26, and MONO-27. By comparing the polymorphism at these five microsatellite loci with those in normal control samples, it is possible to confirm whether the loci are unstable. If two or more of the five microsatellite loci show microsatellite instability, it is classified as highly unstable (MSI-H); otherwise, it is classified as low-instable (MSI-L) or stable (MSS).
[0033] Step S12: Group the sequencing sequences according to the unique molecular tag, perform base correction on the sequencing sequences in each group, and screen out the optimal sequence for the target MSI site in each group.
[0034] Sequencing sequences that fully cover each MSI locus are grouped using unique molecular tags (UMIs). Base correction is then performed on the sequences within each group, and duplicate sequences are removed to select the optimal sequences from each group. UMIs correct sequencing errors, ultimately correcting errors across different sequencing platforms. The more sequences contained in each UMI, the better the correction effect. This allows for the use of a single sequencing platform baseline sample applicable to any sequencing platform.
[0035] In one embodiment of this application, effective sequences that completely cover the target MSI site can be screened from the sequencing sequences obtained by second-generation sequencing; the effective sequences are grouped according to the unique molecular tag, and the effective sequences in each group are base-corrected to screen out the optimal sequence for the target MSI site in each group.
[0036] When performing base correction, the sequencing sequences can be screened first to obtain valid sequences, and then base correction can be performed on the valid sequences; the steps to obtain valid sequences are as follows: The sequenced base sequence is compared with the human genome (hg19) reference sequence using alignment software. The best sequences are those that completely cover the MSI site and those that completely cover the region of the MSI and at least three bases on each side. These are considered as effective sequences. Using effective sequences can reduce the impact of miscalculating the length of repeating bases when DNA molecules break in the MSI repeating sequence.
[0037] Step S13: Extract multiple deletion lengths from the optimal sequence of the target MSI sites in each group, and determine the deletion ratio of each deletion length for each target MSI site.
[0038] For the optimal sequence obtained after processing, the repeat sequence length of each target MSI site is counted, thereby classifying each optimal sequence and counting various deletion lengths, and calculating the proportion of each deletion length in all deletion sequences.
[0039] Step S14: Determine the probability quality function based on the missing proportion of each missing length for each target MSI site, construct a detection statistic threshold based on the probability quality function, and form a baseline model for MSI state detection.
[0040] The probability quality function can be determined by using the proportion of each deletion length at a single site. This probability quality function is then used to construct the MSI detection statistic MSI_Score for each site. The MSI_Score for each baseline sample is calculated, and a detection statistic threshold (MSI_Score) is then established based on the baseline samples. cutoff The baseline model for MSI state detection is formed by using the detection statistic threshold. When using the baseline model to detect the MSI state of the test sample, it is only necessary to compare the detection statistic threshold of the baseline model with the detection statistic of the test sample to determine the MSI state.
[0041] A baseline for MSI detection was constructed using a small sample size of MSS samples. Once the baseline was constructed, it could be used to perform MSI detection on any sample without the need for a corresponding control sample for each sample. Furthermore, this method is applicable to both capture sequencing and amplicon sequencing, the two main next-generation sequencing library preparation methods, as well as to both tissue and blood, the two main clinical sample types.
[0042] In one embodiment of this application, in step S12, the sequencing sequences are grouped according to the unique molecular tag and sequencing start coordinates, and valid groups whose number of sequencing sequences meets the preset number threshold are selected; base correction is performed on the sequencing sequences in each valid group.
[0043] UMI sequences are extracted from the sequencing reads. Only reads containing at least three sequences of each UMI sequence are retained. The original molecular origin is determined based on the UMI and start coordinates of each sequence. For repetitive reads originating from the same original molecular fragment, they are grouped, and the number of sequences in each group is determined. If a preset threshold (e.g., three reads) is met, the group is considered valid. Base correction is performed on the sequences within the valid group to obtain the optimal sequence, which can be the unique corrected sequence. UMI correction can be performed on either single-end or paired-end UMI sequences. Paired-end UMI correction requires that the corrected sequence be identical to the single-end sequence; otherwise, the returned sequence after paired-end UMI correction is a wild-type sequence.
[0044] like Figure 2 As shown, the colored segments at the single ends of the sequence represent UMIs. All sequencing sequences are grouped according to their start coordinate information and single-end UMIs. Figure 2 Each box represents a group, and the colored lines inside the box represent multiple sequencing reads within that group, corresponding to sequences of different microsatellite lengths, such as wild-type 15T, mutant 14T, and mutant 16T. All sequencing reads obtained from PCR amplification of the same original DNA fragment are clustered together using single-end UMI and start coordinates. The number of sequencing reads in each group is counted. If the number of sequencing reads in a group is at least 3, the group is considered a valid group and base correction can be performed. If the number of valid reads is less than 3, no correction is performed and the group can be directly filtered. Figure 2 The sequence is divided into four effective groups, each containing three sequences. Sequence correction is performed according to the rules of base correction, which can be progressively compared based on multiple indicators such as quantity, quality, and length. Each group retains a unique sequence, resulting in four optimal sequences. If the sequence type with the most occurrences is retained, it is used as the corrected base sequence. The corrected base sequence in each group is the optimal sequence for the target MSI site in that group.
[0045] like Figure 3 As shown, when using paired-end correction (i.e., when both ends of the sequence have UMIs), the start coordinates and paired-end UMIs are compared to see if they are consistent. Grouping is done according to UMI1 and UMI2 combined with the start coordinates of paired-end sequencing. Each box represents a group, and the multiple lines within the box represent all sequencing reads obtained from the same original DNA fragment after PCR amplification. The number of sequencing reads in each group is counted; if the number is at least 3, it is considered a valid group and base correction is performed. Figure 3 Two valid groups were obtained, and the base correction was performed according to the rules. Each group retained a unique sequence, resulting in two optimal sequences.
[0046] Following the above embodiments, when performing base correction on each group of sequences, a multi-level priority selection operation can be performed within each effective group to screen out the optimal sequence of the target MSI site in each group. The multi-level priority selection includes a progressive comparison within each group according to the priority order of multiple preset indicators.
[0047] For each valid group containing multiple sequences, a progressive comparison is performed according to the priority order of multiple preset indicators to gradually select the optimal sequence that meets the conditions. The remaining sequences in the group are all repetitive sequences generated during the PCR process and are eventually removed.
[0048] Specifically, the multiple preset indicators include sequence number, alignment quality value, average base quality value, and alignment length. The multi-level priority selection includes: selecting the sequencing sequence with the largest number of sequences supporting the same repeat base length; when the number of sequences is the same, selecting the sequencing sequence with a higher alignment quality value; when the alignment quality values are the same, selecting the sequencing sequence with a higher average base quality value; and when the average base quality values are the same, selecting the sequencing sequence with a longer alignment length.
[0049] The rules for base correction within each valid group are as follows: When the repeat base lengths of MSI repeat sequences within the same valid group are different, the sequence with the most sequences supporting a certain repeat base length is selected; when the number of sequences supporting a certain repeat base length is the same, the alignment quality values of the sequences are compared, and the sequence with the higher alignment quality value is selected; when the alignment quality values of the sequences are also the same, the average base quality values of the sequences are compared, and the sequence with the higher average base quality value is selected; if the average base quality values are still the same, the alignment lengths of the sequences are compared, and the sequence with the longer alignment length is selected.
[0050] Different library preparation reagents can produce varying degrees of PCR errors, and different sequencers have different sequencing error rates. Existing technologies typically use extremely large sample sizes for constructing Pons, encompassing different sequencers and reagent batches. This application, however, uses UMI correction to remove background from different batches (different sequencers, different reagents), correcting sequencing errors from different sequencing platforms. This allows for the construction of Pons using only a small number of samples from a single batch, making it applicable to any sequencing platform. This simplifies Pon construction, eliminating the need to create corresponding Pons for each sequencing platform, saving costs while improving applicability.
[0051] In one embodiment of this application, in step S13, the repeat sequence length of the optimal sequence of the target MSI site in each group is determined; the repeat sequence length is compared with the benchmark reference sequence, and each optimal sequence is classified to obtain multiple deletion lengths; the proportion of each deletion length of the deletion sequence in each target MSI site to all optimal sequences of that target MSI site is calculated to obtain the deletion ratio of each deletion length.
[0052] For the optimal sequences obtained after base correction and repetitive sequence removal, the repetitive sequence length of each target MSI locus is calculated. Based on the calculated repetitive sequence length, it is compared with a benchmark reference sequence (such as the human genome reference sequence). Each optimal sequence is then classified into normal sequences and deleted sequences, thereby statistically determining different deletion lengths for deleted sequences. The proportion of deleted sequences with different deletion lengths at each target MSI locus is calculated out of all deleted sequences at that MSI locus, denoted as del{i}_ratio{s,n}, where s represents the MSI locus, n represents the sample, and i represents the deletion length.
[0053] Specifically, when calculating various deletion lengths, the length of the repeat sequence can be compared with the benchmark reference sequence, and each optimal sequence can be classified to obtain normal sequences and deletion sequences; the deletion length is determined according to the difference between the length of the repeat sequence corresponding to the deletion sequence and the length of the repeat sequence of the benchmark reference sequence, wherein the deletion length ranges from 1 to 9 bases.
[0054] When classifying each optimal sequence, if the length of the repeating sequence is the same as the length of the reference sequence, it is classified as a normal sequence; if the length of the repeating sequence is shorter than the length of the reference sequence, it is classified as a missing sequence. The missing sequence can be further subdivided, and the missing length can be determined according to the difference between the comparisons. For example, -1, -2, -3, -4, -5, -6, -7, -8 and -9 represent the specific missing lengths.
[0055] Assuming a given MSI locus has a total of 1000 sequences (sequences with 15 T+, sequences with less than 15 T+, and sequences with more than 15 T+), and the number of sequences with -1, -2, -3, -4, and -5 T+ lengths are 10, 20, 30, 40, and 50 respectively, then the proportions of each type of deletion are as follows: sequences with a 1 T+ deletion account for 1% (10 / 1000), sequences with a 2 T+ deletion account for 2% (20 / 1000), sequences with a 3 T+ deletion account for 3% (30 / 1000), sequences with a 4 T+ deletion account for 4% (40 / 1000), and sequences with a 5 T+ deletion account for 5% (50 / 1000). This application further subdivides different deletion lengths into different categories, performs five binomial distribution tests on each category, and accumulates the results to obtain the statistic for each locus. Then, it calculates the mean and standard deviation, sums the results for all MSI loci, and calculates the MSI detection statistic for each baseline sample.
[0056] Instead of simply classifying them into two categories, the distribution of each deletion length at each MSI site in the baseline sample is analyzed for each deletion length. Only then does the distribution of reads for each deletion length at each MSI site truly conform to a binomial distribution. Therefore, the MSI detection statistic constructed based on this method has extremely high detection performance and can stably detect even tumors with a content as low as 0.4%.
[0057] In one embodiment of this application, in step S14, the mean of the deletion proportion of each deletion length for each target MSI locus is determined, and the mean is used as a binomial distribution probability parameter to determine the probability quality function; the probability quality function is statistically processed to obtain the statistic for each target MSI locus; the mean and standard deviation of the statistic for each target MSI locus in the baseline sample are determined, and the detection statistic for each baseline sample is calculated; the mean and standard deviation of the detection statistic for each baseline sample are statistically analyzed to determine the detection statistic threshold.
[0058] Based on the deletion proportion del{i}_ratio{s,n} of each target MSI site and each deletion length in the baseline samples, the average del{i}_ratio{s,n} of all baseline samples is calculated and denoted as mean_del{i}_ratio{s}, which is simplified to p. i .
[0059] For each read of a certain deletion length at each target MSI site, the probability of its distribution being consistent with p is... i The binomial distribution has a probability mass function of , Where K represents the total number of reads covering the MSI site, and k represents the number of reads with a missing length of i; the probability mass function is processed to obtain the statistic H for each site. s Calculate H at MSI sites s s The mean and standard deviation of the baseline samples are used to sum all MSI loci, and the MSI detection statistic for each baseline sample is calculated. Based on the MSI detection statistic MSI_Score, a detection statistic threshold (MSI_Score) is then established based on the baseline samples. cutoff value).
[0060] Specifically, when calculating the statistic for each target MSI locus, the cumulative distribution function is determined based on the probability quality function and logarithmic processing is performed to obtain the processing result; the processing results of different missing lengths are accumulated to determine the statistic for each target MSI locus.
[0061] Calculate the cumulative distribution function based on the probability mass function. Taking its logarithm, we get H is calculated by summing different missing lengths. i The statistics for each MSI locus were obtained. Calculate the MSI at site s. The mean and standard deviation in the baseline sample are denoted as . and ; Sum all MSI loci and calculate the MSI_Score statistic for each baseline sample, where, .
[0062] MSI_Score is established based on baseline samples. cutoff The value is denoted as MSI_Score cutoff = mean pon (MSI_Score) + 3 × sd pon (MSI_Score). This forms a baseline model for MSI state detection, where MSI_Score cutoff The value is used as a statistical threshold to compare with the MSI detection statistic of the sample to be detected.
[0063] For each MSI locus, a statistic was constructed. Instead of determining its own positivity or positivity, the statistical information of all MSI loci was integrated. The integrated statistic was then used to determine positivity or positivity, thus avoiding the problem of encountering a positivity or positivity near the judgment threshold when judging a single locus. Furthermore, different MSI loci had different weights in determining the positivity or positivity of the sample.
[0064] Figure 4 The diagram shows a flowchart of a method for MSI status detection according to another aspect of this application, which includes steps S21 to S23.
[0065] Step S21: Extract multiple deletion lengths for each MSI site in the sample to be tested, and determine the deletion ratio of each deletion length for each MSI site.
[0066] To detect the MSI status of the sample to be tested, the logic for calculating the deletion ratio in step S13 can be followed. First, calculate the proportion of each deletion length at each MSI site in the sample to be tested relative to all deletion sequences at that MSI site, i.e., the deletion ratio, denoted as del. {i} _ratio {s,m} s represents the MSI site, m represents the sample to be tested, and i represents the length of the missing information.
[0067] Step S22: Determine the probability quality function based on the missing proportion of each missing length for each MSI site, and construct the MSI detection statistic for the sample to be tested based on the probability quality function.
[0068] Calculate the statistic H for each MSI locus in the sample to be tested. s Thus, the MSI detection statistic MSI_Score of the sample to be tested is obtained. 待测样本 The calculation of the MSI detection statistic for the sample to be tested in this step can be performed in accordance with the step S14 in the baseline sample to construct the detection statistic threshold.
[0069] Step S23: Compare the detection statistic threshold with the MSI detection statistic of the sample to be tested to determine the MSI status result of the sample to be tested.
[0070] Based on the detection statistic threshold MSI_Score in the baseline model cutoff The value is used to determine the MSI status of the sample to be tested. If MSI_Score 待测样本 ≥ MSI_Score cutoff If the sample is in an unstable MSI state, it is denoted as MSI; otherwise, the sample is in a stable MSI state, it is denoted as MSS.
[0071] The MSI status detection method described in this application can be applied to both tissue and blood clinical samples. For tissue samples, it covers not only the three major MSI-detected cancer types—colorectal cancer, gastric cancer, and endometrial cancer—but also a variety of other cancer types, all with extremely high detection accuracy, making it truly applicable to pan-cancer MSI detection; it also exhibits extremely high detection accuracy across any cancer type. For blood samples, it can accurately detect samples with tumor content as low as 0.4%. For most clinical blood samples, the tumor content is below 1%, truly meeting the clinical needs of blood testing.
[0072] In one embodiment of this application, the sample to be tested is preprocessed, and the sample to be tested is subjected to next-generation sequencing to screen out effective sequences that completely cover the target MSI site; the effective sequences are grouped according to a unique molecular tag, the effective sequences in each group are base-corrected, and the optimal sequence for each target MSI site is screened based on the base-correction results.
[0073] Next-generation sequencing is performed on the test samples, and a sequence alignment process completely consistent with the baseline samples is executed, along with effective sequence screening, UMI base correction, and repetitive sequence removal to obtain the optimal sequence set for each target MSI site in the test samples. This avoids bias caused by differences in preprocessing.
[0074] This application establishes a baseline using Pon samples from healthy individuals, solving the problem of needing paired samples for MSI detection. By correcting sequences covering MSI sites using UMI, sequencing errors from different sequencing platforms can be corrected, achieving the goal of adapting to any sequencing platform using only a small sample size of Pon samples from a single sequencing platform. Through the constructed MSI detection statistics, a method is developed that is applicable to both capture sequencing and amplicon sequencing (two major next-generation sequencing library preparation methods), as well as both tissue and blood (two major clinical sample types). For tissue sample detection, a large number of validation samples cover both capture sequencing and amplicon sequencing platforms, and multiple cancer types, showing high consistency with the gold standard detection results. For blood sample detection, it can stably detect tumor content as low as 0.4%, making it truly applicable to real-world blood sample testing scenarios.
[0075] The performance of the MSI detection method described in this application is verified through the following embodiments.
[0076] Example 1: Hybridization capture sequencing data of tissue samples.
[0077] (1) Sample preparation and sample information: Baseline samples: 45 white blood cell (WBC) samples from healthy individuals, all of which were sequenced using Novaseq6000. Cell line simulated tissue validation samples: LS174T and SW48 cell lines and one healthy individual's white blood cell (WBC) sample were pooled. The MSI (Microsatellite Instability) status of WBCs was MSS (microsatellite stable), while the MSI status of LS174T and SW48 was MSI-H (high microsatellite instability). LS174T and SW48 were incorporated into the WBC samples at different proportions, setting 5 pooling ratios, resulting in 5 pools with LS174T and SW48 mass percentages of 40%, 30%, 20%, 10%, and 5%, respectively. Each of the above 5 pools was used as MSI-H tissue simulated samples with tumor purity of 40%, 30%, 20%, 10%, and 5%, respectively.
[0078] For the MSI-H pooled samples and the unpooled WBC samples, six repeated library construction experiments were performed using corresponding fragmentation conditions to simulate tissue samples, resulting in a total of 66 libraries.
[0079] Clinical validation tissue samples: a total of 156 cases, including 34 MSS samples and 122 MSI samples. The MSI results were all validated by PCR. The samples included 40 cases of colorectal cancer, 38 cases of endometrial cancer, 31 cases of gastric cancer, and the remaining 47 cases were distributed in small amounts in various cancer types.
[0080] (2) Acquiring and preprocessing capture sequencing data: Capture probes were designed for 100 MSI sites. Library construction experiments were performed on 45 baseline clinical samples, 66 cell line simulated tissue validation samples, and 152 clinical validation samples, and hybridization capture was performed using capture probes. Paired-end sequencing was performed on the captured libraries using the Novaseq 6000 sequencing platform to obtain raw bcl files. The bcl files were split using bcl2fastq to obtain paired-end sequencing fastq data for each sample. The deduplication depth for each sample was approximately 1000x.
[0081] For each sample's FASTQ data, adapter sequences and low-sequencing-quality sequences in each read are removed using software (such as open-source software). Then, bwa is used to align the clean reads to the human reference genome hg19 to obtain the original BAM file.
[0082] (3) Constructing tissue MSI baselines using 45 baseline samples: For each MSI site s, according to the method described in step S12, extract reads from the bam of each sample that completely cover the region where the site is located and at least 3 bases long on each side, and perform sequence correction and deduplication; for each MSI site s, according to the method described in step S13, calculate the proportion of deletion reads of different deletion lengths to all reads (deletion ratio), denoted as del {i} _ratio {s,n} s represents the MSI locus, n represents the sample, and i represents the length of the deletion; the mean deletion ratio for each MSI locus of 45 baseline samples with different deletion lengths is taken as the average value, denoted as mean_del. {i} _ratio {s} Simplified notation For all baseline samples, the mean baseline data for each site is obtained according to the method described in step S14. pon (H s ) and sd pon (H s Calculate the MSI detection statistic for each baseline sample, denoted as MSI_Score; construct the MSI detection threshold based on the baseline samples according to the method described in step S14, denoted as MSI_Score. cutoff Get MSI_Score cutoff =200.
[0083] (4) Use the MSI baseline obtained in the previous step to detect the MSI status of the sample to be tested: For each MSI site s in each test sample, the following steps are performed with reference to the baseline sample: For each MSI site s, reads with a length of at least 3 bases on each side of the site's bam are extracted, and sequence correction and deduplication are performed; For each MSI site s, the proportion of deletion reads with different deletion lengths to all reads in the test sample is calculated (deletion ratio), denoted as del. {i} _ratio {s,n} s represents the MSI locus, n represents the sample, and i represents the length of the missing data; based on the baseline data mean pon (H s ) and sd pon (H s ), calculate the MSI detection statistic for each test sample, denoted as MSI_Score; and determine the MSI status of the test sample.
[0084] (a) The results of MSI detection in the cell line simulated tissue validation samples are shown in Table 1:
[0085] Table 1
[0086] In this embodiment, cell lines were used to simulate tissue samples and five different tumor content gradients were set. According to the detection results, it was demonstrated that this embodiment can accurately detect MSI states with tumor content as low as 5% in tissue samples. Since the tumor content of tissue samples is usually required to be greater than 10% in clinical applications, this embodiment will not explore the detection performance of lower tumor content in tissue samples.
[0087] (b) Results of MSI testing of clinical samples from statistical tissues: In this embodiment, a total of 156 clinical tissue samples were tested, including 34 MSS samples and 122 MSI samples. The MSI results were all verified by PCR. The samples included 40 cases of colorectal cancer, 38 cases of endometrial cancer, 31 cases of gastric cancer, and the remaining 47 cases were distributed in small quantities in various cancer types. The final result was 100% accurate, which proves that this embodiment has extremely high detection performance in clinical samples of different cancer types.
[0088] Example 2: Hybridization capture sequencing data of blood samples.
[0089] (1) Sample preparation and sample information: Baseline samples: cfDNA samples from healthy individuals, totaling 47 samples, all of which were sequenced using Novaseq6000. Cell line simulated blood validation samples: HCT116 cell line sample and one healthy human leukocyte (WBC) sample were mixed. Among them, the MSI (Microsatellite Instability) status of WBC was MSS (microsatellite stable), and the MSI status of HCT116 was MSI-H (high microsatellite instability). HCT116 was incorporated into the WBC sample at different proportions, setting 5 mixing ratios, and finally obtaining 5 mixtures with HCT116 mass percentages of 8%, 0.8%, 0.4%, 0.2%, and 0.1%, respectively. The above 5 mass percentage mixtures were used as MSI-H blood simulated samples with tumor purity of 8%, 0.8%, 0.4%, 0.2%, and 0.1%, respectively.
[0090] For the MSI-H pooled samples and the unpooled WBC samples, blood samples were simulated using corresponding fragmentation conditions, and six repeated library construction experiments were performed to obtain a total of 36 libraries.
[0091] Clinical validation blood samples: a total of 16 cases, all of which had ctDNA from colorectal cancer patients. The MSI status of their tissue samples was confirmed by PCR, including 10 MSS samples and 6 MSI samples.
[0092] (2) Acquire capture sequencing data and perform data preprocessing: Capture probes were designed for 100 MSI sites. Library construction experiments were performed on 47 baseline samples, 36 cell line simulated blood validation samples, and 12 blood clinical validation samples. Hybridization capture was performed using capture probes. The original BAM file was obtained according to the method in Example 1. The depth of each sample after deduplication was approximately 4000X.
[0093] (3) Constructing a blood MSI baseline using 47 baseline samples: Following the method in Example 1, the mean baseline data for each locus was obtained. pon (H s ) and sd pon (H s ), calculate the MSI statistic for each baseline sample, denoted as MSI_Score; construct the threshold for blood MSI detection based on 47 baseline samples, denoted as MSI_Score. cutoff Get MSI_Score cutoff =100.
[0094] (4) Use the MSI baseline obtained in the previous step to detect the MSI status of the sample to be tested: Based on baseline data meanpon (H s ) and sd pon (H s ), calculate the MSI detection statistic for each test sample, denoted as MSI_Score, and determine the MSI status of the test sample.
[0095] (a) The results of MSI detection in the cell line simulated blood validation samples are shown in Table 2:
[0096] Table 2
[0097] In this embodiment, a cell line was used to simulate a blood sample and five different tumor content gradients were set. According to the detection results, it was demonstrated that this embodiment can accurately detect MSI status with tumor content as low as 0.4% in blood samples. Since more than 50% of real clinical blood samples have a tumor content of less than 1% and more than 30% of samples have a tumor content of less than 0.5%, this embodiment demonstrates its application prospects in real clinical blood samples.
[0098] (b) Statistical analysis of MSI results in clinical blood samples: This embodiment tested 16 ctDNA samples from colorectal cancer patients. The MSI status of the tissue samples was confirmed by PCR, including 10 MSS samples and 6 MSI samples. Ultimately, the positive concordance rate was 83.3% (5 / 6), the negative concordance rate was 100% (10 / 10), and the overall concordance rate was 93.75% (15 / 16). One sample with a negative concordance rate was confirmed by tumor content assessment to have a tumor content below 0.2%, which is below the detection limit, demonstrating that this embodiment also has extremely high detection performance for clinical blood samples.
[0099] Example 3: Amplicon sequencing data of tissue samples.
[0100] (1) Preparation and sample information: Baseline samples: WBC samples from healthy individuals, a total of 50 samples, all of which were sequenced by Novaseq6000.
[0101] Clinical validation tissue samples: 1347 clinical samples, including 3 types of cancer, including 880 cases of colorectal cancer, 340 cases of gastric cancer, and 127 cases of endometrial cancer; from 6 sequencing platforms, including 426 cases from Novaseq6000, 371 cases from Nextseq500, 241 cases from MGI2000, 118 cases from MGI200, 118 cases from Miseq, and 73 cases from Miniseq.
[0102] (2) Acquired and preprocessed the captured sequencing data: Amplification primers were designed for 55 MSI sites. Library construction experiments were performed on 50 baseline samples and 1347 tissue clinical validation samples. The original BAM files were obtained according to the method in Example 1. The depth of each sample after deduplication was about 1000X.
[0103] (3) Constructing an amplicon sequencing tissue MSI baseline using 50 baseline samples: Following the method in Example 1, the mean baseline data for each locus was obtained. pon (H s ) and sd pon (H s ), calculate the MSI statistic for each baseline sample, denoted as MSI_Score; construct the threshold for MSI detection in amplicon sequencing tissue based on 50 baseline samples, denoted as MSI_Score. cutoff Get MSI_Score cutoff =200.
[0104] (4) Use the MSI baseline obtained in the previous step to detect the MSI status of the sample to be tested: Based on baseline data mean pon (H s ) and sd pon (H s ), calculate the MSI detection statistic for each test sample, denoted as MSI_Score, and determine the MSI status of the test sample.
[0105] Statistical analysis of MSI results in amplicon sequencing of clinical tissue samples: This embodiment analyzed 1347 clinical tissue samples obtained through amplicon sequencing, including 228 MSI samples and 1119 MSS samples, covering three cancer types: 880 cases of colorectal cancer, 340 cases of gastric cancer, and 127 cases of endometrial cancer. All samples were sequenced using six sequencing platforms: Novaseq 6000 (426 samples), Nextseq 500 (371 samples), MGI2000 (241 samples), MGI200 (118 samples), Miseq (118 samples), and Miniseq (118 samples). Of the 73 cases, the positive concordance rate was 99.56% (227 / 228), the negative concordance rate was 99.82% (1117 / 1119), and the overall concordance rate was 99.80% (1344 / 1347). The three cases of inconsistency were all near the threshold, with one case each of colorectal cancer, gastric cancer, and endometrial cancer, originating from Nextseq500, MGI2000, and Miniseq, respectively. This demonstrates that this embodiment has extremely high detection performance for clinical samples from different cancer types and different sequencing platforms.
[0106] Figure 5The diagram illustrates an electronic device structure according to another aspect of this application, the device including at least a processor 501 and a memory 502.
[0107] Processor 501 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 501 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 501 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0108] Memory 502 may include one or more computer-readable storage media, which may be non-transitory. Memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 502 is used to store at least one instruction, which is executed by processor 501 to implement a baseline modeling method or an MSI state detection method provided in the method embodiments of this application.
[0109] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 501, memory 502, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to: radio frequency circuits, touch displays, audio circuits, and power supplies.
[0110] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.
[0111] This application also provides a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to implement a baseline modeling method as described above, or an MSI state detection method.
[0112] The baseline modeling method, or the MSI state detection method, when implemented as a computer program, can also be stored as an article of art in a computer-readable storage medium. For example, computer-readable storage media can include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic stripes), optical disks (e.g., compact discs (CDs), digital multifunction discs (DVDs)), smart cards, and flash memory devices (e.g., electrically erasable programmable read-only memory (EPROM), cards, sticks, key drives). Furthermore, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" can include, but is not limited to, wireless channels and various other media (and / or storage media) capable of storing, containing, and / or carrying code and / or instructions and / or data.
[0113] It should be understood that the embodiments described above are merely illustrative. The embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or any combination thereof. For hardware implementation, the processor may be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and / or other electronic units designed to perform the functions described herein, or combinations thereof.
[0114] Some aspects of this application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The aforementioned hardware or software may be referred to as a "data block," "module," "engine," "unit," "component," or "system." The processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or combinations thereof. Furthermore, aspects of this application may manifest as computer products residing in one or more computer-readable media, including computer-readable program code. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), optical discs (e.g., compressed CDs, digital multifunction DVDs, etc.), smart cards, and flash memory devices (e.g., cards, sticks, key drives, etc.).
[0115] A computer-readable medium may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and so on, or suitable combinations thereof. A computer-readable medium can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer-readable medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, radio frequency signals, or similar media, or any combination of the above media.
[0116] The basic concepts have been described above. Obviously, for those skilled in the art, the above disclosure is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the exemplary embodiments of this application.
[0117] Furthermore, this application uses specific terms to describe embodiments of the application. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of the application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined.
[0118] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of scope in some embodiments of this application are approximate values, in specific embodiments, such values are set as precisely as feasible.
Claims
1. A baseline modeling method, characterized in that, The method includes: A baseline sample and a target MSI site were selected, and the baseline sample was subjected to next-generation sequencing to obtain the sequencing sequence. The sequencing sequences are grouped according to a unique molecular tag, and the sequencing sequences within each group are base-corrected to select the optimal sequence for the target MSI site in each group. Extract multiple deletion lengths from the optimal sequence of the target MSI sites in each group, and determine the deletion ratio of each deletion length for each target MSI site; A probability quality function is determined based on the proportion of each missing length for each target MSI site, and a detection statistic threshold is constructed based on the probability quality function to form a baseline model for MSI state detection.
2. The method according to claim 1, characterized in that, The step of grouping the sequencing sequences according to a unique molecular tag and performing base correction on the sequencing sequences within each group includes: The sequencing sequences are grouped according to their unique molecular tags and sequencing start coordinates, and valid groups whose number of sequencing sequences meets a preset threshold are selected. Base correction was performed on the sequencing sequences within each valid group.
3. The method according to claim 2, characterized in that, The optimal sequences for selecting target MSI sites in each group include: For each valid group, a multi-level priority selection operation is performed within the group to filter out the optimal sequence of the target MSI site in each group. The multi-level priority selection includes a progressive comparison within each group according to the priority order of multiple preset indicators.
4. The method according to claim 3, characterized in that, The multiple preset indicators include sequence number, alignment quality value, average base quality value, and alignment length. The multi-level priority selection includes: Select the sequencing sequence that supports the largest number of sequences with the same repeat base length; When the number of sequences is the same, the sequencing sequence with the higher alignment quality value is selected; When the alignment quality values are the same, the sequencing sequence with the higher average base quality value is selected. When the average base quality values are the same, the sequencing sequence with the longer alignment length is selected.
5. The method according to claim 1, characterized in that, The step of determining a probability quality function based on the proportion of each deletion length for each target MSI locus, and constructing a detection statistic threshold based on the probability quality function, includes: Determine the mean of the deletion proportion for each deletion length at each target MSI site, and use the mean as a binomial distribution probability parameter to determine the probability quality function; The probability mass function is statistically processed to obtain the statistics for each target MSI site; Determine the mean and standard deviation of the statistic for each target MSI locus in the baseline sample, and calculate the detection statistic for each baseline sample; The mean and standard deviation of the detection statistic for each baseline sample are calculated to determine the threshold for the detection statistic.
6. The method according to claim 5, characterized in that, Statistical processing is performed on the probability quality function to obtain statistics for each target MSI locus, including: The cumulative distribution function is determined based on the probability mass function and logarithmic processing is performed to obtain the processing result; By summing the processing results for different deletion lengths, the statistics for each target MSI locus are determined.
7. The method according to claim 1, characterized in that, The step of extracting multiple deletion lengths from the optimal sequence of the target MSI sites in each group and determining the deletion ratio of each deletion length for each target MSI site includes: Determine the optimal repeat sequence length for the target MSI site in each group; The length of the repeated sequence is compared with the benchmark reference sequence, and each optimal sequence is classified to obtain multiple missing lengths; The proportion of each deletion length in the target MSI site is calculated to represent the proportion of the missing sequence at each deletion length among all the optimal sequences at that target MSI site.
8. The method according to claim 7, characterized in that, The process involves comparing the length of the repeating sequence with a benchmark reference sequence, classifying each optimal sequence, and obtaining various missing lengths, including: The length of the repeated sequence is compared with the benchmark reference sequence, and each optimal sequence is classified to obtain normal sequences and missing sequences; The deletion length is determined based on the difference between the length of the repeat sequence corresponding to the deletion sequence and the length of the repeat sequence of the reference sequence, wherein the deletion length ranges from 1 to 9 bases.
9. The method according to claim 1, characterized in that, The method further includes: Effective sequences that completely cover the target MSI site are selected from the sequencing sequences obtained from second-generation sequencing. The effective sequences are grouped according to the unique molecular tag, and the effective sequences in each group are base-corrected to select the optimal sequence for the target MSI site in each group.
10. A method for MSI state detection, characterized in that, The method includes: Extract multiple deletion lengths for each MSI site in the sample to be tested, and determine the deletion ratio of each deletion length for each MSI site; A probability quality function is determined based on the proportion of each missing length at each MSI site, and an MSI detection statistic for the sample to be tested is constructed based on the probability quality function. The detection statistic threshold is compared with the MSI detection statistic of the sample to be tested to determine the MSI status result of the sample to be tested.
11. The method according to claim 10, characterized in that, The method includes: The test samples were subjected to next-generation sequencing to screen out effective sequences that completely cover the target MSI sites; The effective sequences are grouped according to the unique molecular tag, the effective sequences in each group are base corrected, and the optimal sequence for each target MSI site is selected.
12. An electronic device, characterized in that, The device includes: One or more processors; and A memory storing computer-readable instructions that, when executed, cause the processor to perform the operation of the baseline modeling method as described in any one of claims 1 to 9, or the operation of the MSI state detection method as described in claim 10 or 11.
13. A computer-readable storage medium having stored thereon computer-readable instructions that can be executed by a processor to implement the baseline modeling method as described in any one of claims 1 to 9, or the MSI state detection method as described in claim 10 or 11.