Method, device, and program product for predicting met14 skipping mutation
Patent Information
- Application Number
- PCT/CN2025/112642
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-19
- Filing Date
- 2025-08-05
- Publication Date
- 2026-01-08
AI Technical Summary
Existing technologies have a high false negative rate when detecting MET 14 skipping mutations, especially in the inaccurate identification of hidden splicing mutations. Furthermore, traditional methods rely on RNA quality, resulting in high detection costs and low efficiency.
An AI-trained predictive model was used to identify MET 14 skipping mutations, particularly classical and hidden splicing, by acquiring chromosome number, physical coordinates, number of base changes in the reference genome, and number of mutated base changes from the patient's MET gene data and using machine learning algorithms such as random forests, thus avoiding dependence on RNA quality.
It improves the detection rate of MET 14 skipping mutations, reduces testing costs, improves doctors' work efficiency, and can accurately predict hidden splicing mutations at the DNA level even when RNA data is unavailable.
Smart Images

Figure CN2025112642_08012026_PF_FP_ABST
Abstract
Description
Method, device, and program product for predicting met 14 skipping mutations TECHNICAL FIELD
[0001] The present application relates to the field of intelligent medical treatment, in particular to a method, device, program product and computer readable storage medium for predicting MET 14 skipping mutations. BACKGROUND
[0002] Cancer has become the leading cause of death in China with increasing morbidity and mortality, and is a major public health problem. Lung cancer is the most common cancer. Non-small cell lung cancer (NSCLC) is the most common pathological type of lung cancer, and most patients are in advanced stage at the time of diagnosis. NSCLC is a genetically driven disease, and the development of lung cancer is closely related to driver genes. The mutation status of driver genes is an important predictor of the efficacy of targeted therapy. After EGFR gene mutation and ALK gene fusion, MET gene is another important driver gene in NSCLC, and has become a hot spot for targeted therapy. There are mainly three forms of abnormal activation of MET gene in NSCLC: MET exon 14 (MET 14) skipping mutation, MET gene amplification and protein overexpression. MET inhibitors have achieved good antitumor effect in NSCLC patients with MET 14 skipping mutation. In recent years, with the rapid development of drugs, FDA has approved Tepotinib and Capmatinib, and NMPA has approved Savolitinib for non-small cell lung cancer patients with MET 14 skipping mutation. Currently, the methods for detecting MET 14 skipping mutation include DNA next generation sequencing (NGS), Sanger sequencing of exon 14 and its flanking introns, reverse transcription-PCR (RT-PCR) and RNA-based NGS detection. The approved companion diagnostic methods for detecting MET 14 skipping mutation are FoundationOne NGS analysis in the United States and Archer MET analysis in Japan. The NCCN guidelines (v1.2024) have deleted the detection recommendation of IHC, and therefore MET 14 skipping mutation is best detected by FISH and NGS. Even so, a study found that DNA-level detection has a risk of missing MET 14 skipping mutation, with a detection rate of 1.3%. RNA-based analysis detects a higher proportion of MET 14 skipping cases, with an RNA detection rate of 4.2%. However, the problem is that RNA-based analysis is highly dependent on RNA quality, which may not be ideal in some clinical samples. In addition, after FDA approved Tepotinib and Capmatinib for metastatic NSCLC patients carrying MET 14 skipping mutation, NMPA approved Savolitinib (Vosaroxin) on June 22, 2021, and included it in the medical insurance drug directory on March 1, 2023. However, in clinical practice, due to incomplete understanding of the splicing codon, it is difficult to accurately identify hidden mutations other than the classic GT and AG splicing, and cryptic splice variants are often overlooked. SUMMARY
[0003] To solve the above problems, the application provides a method for predicting MET 14 skip mutation, specifically comprising: obtaining patient MET gene data, including chromosome number, physical coordinates, reference genome base number of changes, and mutation base number of changes;
[0004] The chromosome number, physical coordinates, reference genome base number of changes, and mutation base number of changes are input into a prediction model to obtain a prediction result of whether it is a MET 14 skip mutation.
[0005] Optionally, the MET gene data further includes reference genome bases and mutation bases, and the chromosome number, physical coordinates, reference genome bases, mutation bases, reference genome base number of changes, and mutation base number of changes are input into a prediction model to obtain a prediction result of whether it is a MET 14 skip mutation.
[0006] Optionally, the MET 14 skip mutation includes classic splicing and / or hidden splicing, and the chromosome number, physical coordinates, reference genome base number of changes, and mutation base number of changes are input into a prediction model to obtain a prediction result of whether it is a classic splicing and / or hidden splicing MET 14 skip mutation.
[0007] Optionally, the classic splicing includes GT-AG.
[0008] Optionally, the hidden splicing includes one or more of the following: indel of more than 50 bp, polypyrimidine region of intron region, Branch AA, and splicing auxiliary factor ESE of exon region.
[0009] Optionally, the hidden splicing includes: indel of more than 50 bp, polypyrimidine region of intron region, Branch AA, and splicing auxiliary factor ESE of exon region.
[0010] Optionally, the training process of the prediction model is: obtaining a patient MET gene data set and a label, including chromosome number, physical coordinates, reference genome base number, and mutation base number of changes, inputting the chromosome number, physical coordinates, reference genome base number, and base number of changes into a to-be-trained model for training to obtain a prediction model, wherein the label is an RNA positive sample or other positive sample.
[0011] Optionally, the prediction model uses one or more of the following: random forest, decision tree, support vector machine, logistic regression model, convolutional neural network, XGBoost, and AdaBoost.
[0012] Optionally, the patient MET gene data acquisition process includes:
[0013] Obtaining patient MET gene basic data, including: chromosome number, physical coordinate, reference genome base, mutant base;
[0014] Calculating the number of base changes of the reference genome base and the mutant base to obtain the number of reference genome base changes and the number of mutant base changes; obtaining patient MET gene data.
[0015] Optionally, the obtaining of the MET gene basic data comprises:
[0016] Obtaining NGS sequencing data of the patient;
[0017] Aligning the NGS sequencing data with the MET reference genome to obtain an alignment result;
[0018] Performing variation detection on the alignment result to obtain a detection result;
[0019] Extracting the MET gene basic data based on the detection result.
[0020] The NGS sequencing data is DNA sequencing data and / or RNA sequencing data.
[0021] The purpose of the present application is to provide a computer program product, which comprises a computer program or instructions thereon, and the computer program or instructions are executed by a processor to realize the above-mentioned method for predicting MET 14 skipping mutation.
[0022] The purpose of the present application is to provide a computer device, which comprises a memory, a processor and a computer program or instructions stored on the memory, and the computer program or instructions are executed by the processor to realize the above-mentioned method for predicting MET 14 skipping mutation.
[0023] The purpose of the present application is to provide a computer readable storage medium, which stores a computer program or instructions thereon, and the computer program or instructions are executed by a processor to realize the above-mentioned method for predicting MET 14 skipping mutation.
[0024] The advantages of the present application are:
[0025] 1. Prior art manual experience review, RNA supplement test method. But these methods have the following shortcomings: ① The subjective factor is unstable, and the experience and cognitive difference has a great influence on the big mistake. ② For the technical party and the patient, the time cost and the detection cost increase. ③ The probability of personalized patient cases is not large enough, especially rare hidden mutations. But the variability of hidden type splicing mutations is also very high, and the rules are not easy to summarize and not convenient to promote. ④ Collecting cases needs long time, multi-region, large data accumulation and special mining. Therefore, the present application proposes a method for predicting MET 14 skip mutation, which uses artificial intelligence to train a prediction model to form an objective, efficient and accurate MET 14 skip mutation detection method, which helps to improve the detection rate of patient gene variation and improve the work efficiency of doctors.
[0026] 2. For the missed detection of hidden type splicing mutations, the present application focuses on considering the damage to the stability of snRNPs when the number of base changes of MET gene exons and introns or the structural changes are large, so as to find the hidden type splicing mutation rule of MET 14 in addition to the classic splicing. In order to achieve one purpose: even if the tissue is not accessible or the RNA data cannot be used, the hidden type splicing of MET 14 exons occurring at the DNA level can also be accurately predicted. Specifically, the number of base changes is used as the main feature for training when training the prediction model, so as to realize the detection of classic detection mutations and hidden type splicing mutations of MET 14, avoid missing, and improve the disease detection rate.
[0027] 3. There are some problems in the existing artificial intelligence for MET 14 skip splicing prediction, which cannot identify same base substitution delins, too large indel type mutations (>50bp), seems to have a poor tendency for predicting variable subtypes of ESE, and the sensitivity of Branch AA is not reflected. In view of this problem, the present application is based on the fact that RNA splicing is catalyzed by the assembly of snRNPs and other proteins, which together constitute the principle of spliceosome, which is different from the previous base processing process, but uses the number of base changes to identify same base substitution delins, >50bp indel type mutations, variable subtypes of ESE, and Branch AA. BRIEF DESCRIPTION OF DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0029] FIG. 1 is a flowchart of a method for predicting MET 14 skipping mutation according to an embodiment of the present application;
[0030] FIG. 2 is a schematic diagram of a system for predicting MET 14 skipping mutation according to an embodiment of the present application;
[0031] FIG. 3 is a schematic diagram of a device for predicting MET 14 skipping mutation according to an embodiment of the present application;
[0032] FIG. 4 is a mutation type detection rate of NSCLC research and internal data according to an embodiment of the present application;
[0033] FIG. 5 is a comparison result of mutation type of NSCLC research and internal data according to an embodiment of the present application;
[0034] FIG. 6 is a comparison result of DNA and RNA double detection mutation of NSCLC research and internal data according to an embodiment of the present application;
[0035] FIG. 7 is a flowchart of a method for predicting MET 14 skipping mutation based on MET gene basic data according to an embodiment of the present application;
[0036] FIG. 8 is a schematic diagram of a system for predicting MET 14 skipping mutation based on MET gene basic data according to an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the person skilled in the art better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application.
[0038] In some of the descriptions in the specification and claims of the present application and the above-mentioned drawings, a plurality of operations are included which occur in a specific order, but it should be clearly understood that these operations can be executed or in parallel without the order in which they appear in the text. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this text are used to distinguish different messages, devices, modules, etc., and do not represent the order of sequence, nor do "first" and "second" represent different types.
[0039] FIG. 1 is a flowchart of a method for predicting MET 14 skipping mutation according to an embodiment of the present application, which specifically includes:
[0040] S1: Obtain MET gene data of a patient, including chromosome number, physical coordinates, reference genome base change number, and mutation base change number;
[0041] In one embodiment, MET 14 skipping mutation prediction is performed by acquiring MET gene data, wherein the base data of the MET gene data, the data processed from the base data constitute the MET gene data, and the specific steps are:
[0042] Acquiring patient MET gene data, including chromosome number, physical coordinates, reference genome base number of changes, and mutation base number of changes;
[0043] Based on the chromosome number, physical coordinates, reference genome base number of changes, and mutation base number of changes, whether it is a MET 14 skipping mutation prediction is performed to obtain a prediction result.
[0044] Optionally, the base data of the MET gene data includes one or more of the following: chromosome number, physical coordinates; the data processed from the original data includes: reference genome base number of changes, and mutation base number of changes.
[0045] In one specific embodiment, regarding the method for predicting splicing, although scientists have explored modeling, statistical, and other prediction methods for many years, it is well known that the prediction accuracy has not yet reached the level of evidence-based medicine. However, a recent influential and highly credible study from a 2019 cell report describes SpliceAI, a convolutional neural network that can accurately predict gene mutations that lead to cryptic splicing. The study found that synonymous mutations and intronic mutations that affect splicing have a high validation rate on RNA-seq and are highly damaging in the human population.
[0046] In recent years, high-quality clinical trials have been carried out in China and are at the forefront of the international community, but the understanding of MET exon 14 skipping is not complete. The current approach is to establish an industry consensus through well-known experts and organizations in the industry. For example, the purpose of establishing a consensus process is to discuss controversial issues related to MET mutation detection, including MET exon 14 skipping mutations in non-small cell lung cancer, MET gene amplification, and MET protein expression. This consensus was formed after two rounds of in-depth discussions at a virtual meeting, involving 20 pathologists and 19 clinical experts.
[0047] In the cBioPortal big data platform, 15 recent studies on NSCLC were selected, a total of 8352 samples, and 11086 NSCLC patient tissue samples from January 2021 to April 2023 were selected as internal data (In House) representatives. Although there is no big difference in the detection rate of MET 14 mutations between the two types of databases (as shown in FIG. 4), further comparison is made to each patient sample and specific to mutation type (as shown in FIG. 5), the number of mutations detected in the Exon range in the public database is not much different, but interestingly, there is no detection of mutations in the Intron range.
[0048] According to the internal data research, under the condition of DNA-RNA double detection, the sample with RNA positive result is selected to observe its DNA mutation (as shown in FIG. 6), if the mutation occurs in the intron region, there is a significant difference between DNA and RNA positive results, which indicates that DNA has a loophole in the detection performance of the intron region. However, in the real world, the sample condition is limited or the tissue sample cannot be obtained, so the detection of RNA cannot be carried out, which means that the patient's treatment process is at a bottleneck, and such a situation is common in clinical practice. And if the NSCLC patient has additional detection experiments, time and cost, it is not only inconvenient but also very uneconomical, so it is particularly important to predict the occurrence of splicing through non-invasive DNA or DNA detection. Internal data testing also found that SpliceAI can predict most hidden splicing mutations, especially solving deletion type mutations. It is slightly regrettable that through real-world data testing, it is found that it cannot identify same-base substitution delins, and indel type mutations (> 50bp) that are too large. It seems to have a poor tendency to predict variable subtypes of ESE, and the sensitivity to Branch AA has not been shown.
[0049] In one embodiment, the acquisition of the MET gene data (acquisition of the MET gene basic data) comprises: acquiring NGS sequencing data of a patient;
[0050] Based on the NGS sequencing data, a comparison with a MET reference genome is performed to obtain a comparison result;
[0051] A variation detection is performed on the comparison result to obtain a detection result;
[0052] Based on the detection result, MET gene data is extracted.
[0053] The NGS sequencing data is DNA sequencing data and / or RNA sequencing data.
[0054] In one embodiment, the method for acquiring gene data comprises:
[0055] When the rsID is known, directly query dbSNP or Ensembl; dbSNP (NCBI): input rsID (such as rs123456) query, get chromosome position, reference / mutation base and reference genome version. Ensembl VEP: input chromosome coordinates or rsID, get detailed annotation. Extract the chromosome number, coordinates, reference sequence and mutation sequence directly from the data returned by the result page or API.
[0056] When there is a VCF file, extract gene data through tools, including Python, R language, etc. When obtained from raw sequencing data, data preprocessing: use tools such as FastQC to assess the quality of raw NGS sequencing data (usually in FASTQ format), view quality distribution, base content, GC content, sequence repetition rate, etc. to understand the overall quality of the data; use one or more of the following tools: fastp Trimmomatic, Cutadapt, remove adapter sequences, low-quality bases and short sequence fragments in sequencing data. Generally, quality threshold (such as Phred quality value less than 20) and length threshold (such as sequence length less than 30bp) will be set, and sequences or bases that do not meet the requirements will be removed to improve the accuracy of subsequent analysis.
[0057] Sequence alignment: download a suitable reference genome containing the MET gene, ensure the integrity and accuracy of the reference genome, so as to accurately align the sequencing sequence to the genome later; use alignment tools such as BWA (Burrows-Wheeler Aligner), Bowtie2, etc. to align the preprocessed sequencing data with the reference genome. These tools can quickly and accurately find the best matching position of the sequencing reads on the reference genome and generate a SAM (Sequence Alignment / Map) format alignment file; use tools such as SAMtools to convert the SAM file to a BAM (Binary Alignment / Map) file, which is a binary compressed form of the SAM file, occupying small space and being convenient for subsequent processing. At the same time, sort the BAM file by chromosome position for subsequent variant detection and analysis. In one specific embodiment, bwa mem is used to align the quality-controlled split reads to the reference genome hg19 (GRCh37), samtools view is used to filter out multiple alignments and unaligned reads, and samtools sort is used to sort the alignment results to generate sort.bam.
[0058] Using gencore to perform read deduplication on tumor samples, base correction on low-quality and error bases, and then using samtools sort to sort. The deduplicated and sorted bam is used for somatic snv indel detection; using sambamba to perform read deduplication on tumor samples and white blood cell control samples, and using the white blood cell bam for somatic snv indel detection.
[0059] Variant detection: Select a variant detection tool, common variant detection tools include GATK (Genome Analysis Toolkit), FreeBayes. These tools can identify single nucleotide variations (SNVs), insertions and deletions (InDels), and other variation information in sequencing data based on alignment results. Through these steps, the variation sites in the sample can be accurately detected, and a VCF file containing all variation information can be generated. In a specific embodiment, the variant detection tool samtools mpileup is used to establish a tumor and white blood cell sample pileup format file, MutLoc (SNPIndel) and Varcidt (longindel) are used to detect original mutations, and a VCF file containing all variation information is generated.
[0060] Screening MET gene related variants: Using tools such as VCFtools, Bcftools, etc., based on the location information of the MET gene on the reference genome (such as chromosome location, gene region), screen out the variation sites related to the MET gene from the whole genome VCF file, and generate a VCF file containing only MET gene variation information.
[0061] In a specific embodiment, the present application is based on the NGS capture panel design, library construction, and sequencing method, which can be considered to obtain approximately or the same data at the experimental level. Through background pool filtering or paired analysis methods, etc., tumor somatic mutation information is obtained. As shown in Table 1, the model input data is obtained from the VCF file, including chromosome, physical coordinates, reference genome base, and mutation base information. However, unlike previous operations, ATCG bases are not OneHot encoded. Through data preprocessing, base changes are converted to digital statistics as input elements. This is because OneHot encoding has a length limitation on base sequences, and is not suitable for scenarios where DNA sequences are damaged, such as deletions or base substitutions of more than 50 bp.
[0062] Table 1 VCF file structure (excerpts)
[0063] In one specific embodiment, according to internal non-small cell lung cancer patient NGS data in the last 5 years, sample types include tissue slides and wax rolls, plasma, pleural effusion, cerebrospinal fluid, etc. DNA data, and RNA positive results as label labels to build supervised machine learning models: random forest models. The extracted features need to be converted into a format suitable for machine learning model input. This includes converting text data to numerical data, handling missing values, encoding categorical features, etc. Generally, when building a model to learn, the DNA sequence ATCG base is One-Hot encoded, which is converted into information that can be processed by a neural network. One-Hot encoding is very useful when processing DNA sequence data, as it can convert sequence data into numerical data, allowing machine learning models to better understand and process this data. For example, A: [1, 0, 0, 0]; T: [0, 1, 0, 0]; C: [0, 0, 1, 0]; G: [0, 0, 0, 1] However, for longer sequence strings, One-Hot encoding greatly increases the dimensionality of the data, and according to SpliceAI testing and literature recommendations, the length of the base sequence should not exceed 50bp. Considering that some patterns and rules in DNA sequences may be lost during One-Hot encoding, such as periodic patterns of certain nucleotides or interaction information between nucleotides. As previously understood in the literature review of RNA splicing mechanisms, RNA splicing is catalyzed by the assembly of snRNPs plus other proteins, which collectively make up the spliceosome. However, base loss, substitution greater than 50bp is very important to the damage of snRNPs, and must not be missed in analysis and annotation. Based on the above, count the ATCG bases of REF and ALT respectively as shown in Table 2 and Table 3, and use the number of base changes as a feature element as the training and validation dataset content.
[0064] Table 2 Training and validation dataset (example)
[0065] Table 3 Training and validation dataset after data preprocessing (example)
[0066] S2: Input the chromosome number, physical coordinates, reference genome base change number, and mutant base change number into the prediction model to get the prediction result of whether it is a MET 14 skip mutation.
[0067] In one embodiment, the MET gene basic data further comprises: reference genome base, mutant base, the chromosome number, physical coordinate, reference genome base, mutant base, reference genome base change number, mutant base change number are input to the prediction model to obtain the prediction result of whether it is MET 14 skip mutation. Optionally, the training process of the prediction model is: obtaining patient MET gene dataset and label, including chromosome number, physical coordinate, reference genome base number and mutant base change number, inputting the chromosome number, physical coordinate, reference genome base number, mutant base change number to the model to be trained to train to obtain the prediction model, wherein the label is RNA positive sample or other positive sample.
[0068] Optionally, the training process of the prediction model is: obtaining patient MET gene dataset and label, including chromosome number, physical coordinate, reference genome base number and mutant base change number, reference genome base, mutant base, inputting the chromosome number, physical coordinate, reference genome base, mutant base, reference genome base number, mutant base change number to the model to be trained to train to obtain the prediction model, wherein the label is RNA positive sample or other positive sample.
[0069] In one embodiment, the MET 14 skip mutation includes classical splicing and / or hidden splicing; the chromosome number, physical coordinate, reference genome base change number, mutant base change number are input to the prediction model to obtain the prediction result of whether it is classical splicing and / or hidden splicing MET 14 skip mutation.
[0070] Optionally, the classical splicing includes GT-AG.
[0071] In one embodiment, the chromosome number, physical coordinate, reference genome base change number, mutant base change number are input to the prediction model to obtain the prediction result of whether it is classical splicing MET 14 skip mutation.
[0072] In one embodiment, the chromosome number, physical coordinate, reference genome base change number, mutant base change number are input to the prediction model to obtain the prediction result of whether it is hidden splicing MET 14 skip mutation.
[0073] In one embodiment, the chromosome number, physical coordinate, reference genome base change number, mutant base change number are input to the prediction model to obtain the prediction result of whether it is classical splicing and hidden splicing MET 14 skip mutation.
[0074] In one embodiment, the chromosome number, physical coordinate, reference genome base number, mutant base number are input into the prediction model to obtain the prediction result of whether it is a MET 14 skip mutation, and the prediction probability of classic splicing MET 14 skip mutation and hidden splicing MET 14 skip mutation is included.
[0075] In one embodiment, the hidden splicing includes one or more of the following: indel of more than 50 bp, poly-pyrimidine region of intron region, Branch AA, splicing auxiliary factor ESE of exon region.
[0076] Optionally, the hidden splicing includes: indel of more than 50 bp, poly-pyrimidine region of intron region, Branch AA, splicing auxiliary factor ESE of exon region.
[0077] In one embodiment, the prediction model uses one or more of the following: random forest, decision tree, support vector machine, logistic regression model, convolutional neural network, XGBoost, AdaBoost.
[0078] In a specific embodiment, the prediction of random forest can be represented as: y^i = mode{y^i1, y^i2, …, y^iB}
[0079] Where y^i is the prediction of the i-th sample by the random forest, and y^ij is the prediction of the i-th sample by the j-th tree. A class that implements the random forest algorithm is established to solve the classification problem. X = data[['CHROM', 'POS', 'lenR', 'lenA']] y = data['label']
[0080] Where X: feature data. y: label data, containing the label of each sample. RNA positive is 1 and negative is 0. The proportion of the test set is set to 0.3, and 30% of the data will be used as the test set. random_state is set to 42 to ensure that the division is consistent every time the code is run, which helps the reproducibility of the results. Create a random forest model, use GridSearchCV to iterate through different parameter combinations, and evaluate the performance of each parameter combination through 5-fold cross-validation. Finally, the best parameter combination is used to train the model and evaluate it on the test set. Finally, the trained model is saved for new data prediction.
[0081] In one embodiment, the method for obtaining the specific chromosome number, physical coordinate, reference genome sequence, and mutant genome sequence of the mutation does not include PCR and amplicon methods, or causes the MET 14 hidden splicing mutation to be missed.
[0082] In one specific embodiment, the best parameter combination of the model Rfmodel is optimized as follows: max_depth: None, the maximum depth of the decision tree has no limit, and the tree will grow until all leaf nodes are pure. Best cross-validation score: 0.98, the best score obtained in cross-validation is 0.98, which is a very high score, indicating that the model has good performance on the training data. As shown in Table 4, Accuracy: 0.99, the accuracy is 99%, which means that the model can correctly predict 99% of the samples. Sensitivity: 1.00, also known as recall or true positive rate, indicates that the model's ability to correctly identify positive samples is 100%. Specificity: 0.98, indicating that the model's ability to correctly identify negative samples is 98%, as shown in Table 4. The prediction performance of the random forest model of the present application is better than that of the DNN neural network (Accuracy: 0.67), the support vector machine (Accuracy: 0.81), and the Logist regression model (Accuracy: 0.70).
[0083] Table 4 Rfmodel model evaluation indicators
[0084] As shown in Table 5, the model Rfmodel can predict: classic splicing GT-AG, and hidden splicing including: indel of more than 50bp, poly-pyrimidine region (pypy) in intron region and Branch AA, and splicing auxiliary factor ESE in exon region. In terms of MET 14 jump mutation detection, this aspect can well make up for or replace SpliceAI.
[0085] Table 5 Prediction and comparison of other analysis methods on new negative and positive data collected in the past year
[0086] In one specific embodiment, the process of training and applying the prediction model to predict MET 14 jump mutations is as follows:
[0087] Step A: Obtain the lung cancer DNA data of non-small cell lung cancer patients in previous years, analyze the MET gene VCF analysis file content: chromosome number chr, physical coordinate pos, reference genome base ref, and mutation base alt information as input layer pre-information.
[0088] Step B: data preprocessing, the information of step A is digitally converted into chr, pos, lenR, lenA. Another MET 14 jump mutation with RNA positive result or other ways to confirm positive sample is labeled as label = 1, and the other is labeled as label = 0, which are together as the input layer data of the random forest model.
[0089] Step C: load data into the model, use the train_test_split function from sklearn.model_selection module to divide the data set into training set and test set, and specify that the test set accounts for 30% of the total data set, random_state: set the seed of the random number generator to ensure that the result of each division is the same, increase the reproducibility of the code. Build the model and realize GridSearchCV parameter tuning to get the best parameters and model. Start training the random forest model, and save the trained model as Rfmodel.
[0090] Step D: obtain new NGS data, and analyze the content of VCF file including chromosome number chr, physical coordinate pos, reference genome base ref, and mutation base alt information as input pre-information.
[0091] Step E: data preprocessing, the input layer and information are digitally converted into chr, pos, lenR, lenA, and loaded into the random forest model Rfmodel for prediction. Get Table 1 Rfmodel binary classification prediction results. "0" represents not MET 14 splicing mutation, and "1" represents MET 14 splicing mutation.
[0092] Figure 7 is an embodiment of another method for predicting MET 14 jump mutation provided by the application, which predicts MET 14 jump mutation by obtaining MET gene data basic data and processing the obtained data, specifically:
[0093] Obtain patient MET gene data (obtain patient MET gene basic data), including chromosome number, physical coordinate, reference genome base, and mutation base;
[0094] Calculate the number of base changes of the reference genome base and the mutation base to obtain the number of reference genome base changes and the number of mutation base changes;
[0095] Input the chromosome number, physical coordinate, reference genome base change number, and mutation base change number into the (first) prediction model to obtain the prediction result of whether it is a MET 14 jump mutation.
[0096] Optionally, the input of the (first) prediction model further comprises a reference genome base and a mutant base, and the chromosome number, the physical coordinate, the reference genome base, the mutant base, the number of changes of the reference genome base, and the number of changes of the mutant base are input into the prediction model to obtain a prediction result of whether it is a MET 14 skipping mutation.
[0097] Optionally, the training process of the (first) prediction model is: obtaining a patient MET gene dataset and a label, including a chromosome number, a physical coordinate, a reference genome base, and a mutant base; calculating the number of changes of the reference genome base and the mutant base, and then inputting the chromosome number, the physical coordinate, the number of changes of the reference genome base, and the number of changes of the mutant base into the model to be trained to obtain the prediction model, wherein the label is an RNA positive sample or other positive sample.
[0098] In another embodiment, the training process of the (first) prediction model is: obtaining a patient MET gene dataset and a label, including a chromosome number, a physical coordinate, a reference genome base, and a mutant base; calculating the number of changes of the reference genome base and the mutant base, and then inputting the chromosome number, the physical coordinate, the number of changes of the reference genome base, the number of changes of the mutant base, the reference genome base, and the mutant base into the model to be trained to obtain the prediction model, wherein the label is an RNA positive sample or other positive sample.
[0099] In one embodiment, the chromosome number, the physical coordinate, the reference genome base, the mutant base, the number of changes of the reference genome base, and the number of changes of the mutant base are input into the prediction model to obtain a prediction result of whether it is a classic cleavage and hidden cleavage MET 14 skipping mutation.
[0100] In one embodiment, the chromosome number, the physical coordinate, the reference genome base, the mutant base, the number of changes of the reference genome base, and the number of changes of the mutant base are input into the prediction model to obtain a prediction result of whether it is a classic cleavage MET 14 skipping mutation.
[0101] In another embodiment, the chromosome number, the physical coordinate, the reference genome base, the mutant base, the number of changes of the reference genome base, and the number of changes of the mutant base are input into the prediction model to obtain a prediction result of whether it is a hidden cleavage MET 14 skipping mutation.
[0102] In one embodiment, the MET 14 skipping mutation includes classic cleavage and / or hidden cleavage, and the chromosome number, the physical coordinate, the number of changes of the reference genome base, and the number of changes of the mutant base are input into the prediction model to obtain a prediction result of whether it is a classic cleavage and / or hidden cleavage MET 14 skipping mutation.
[0103] Optionally, the classic cleavage includes GT-AG.
[0104] In one embodiment, the chromosome number, physical coordinate, number of reference genome base changes, and number of mutation base changes are input into a prediction model to obtain a prediction result of whether it is a classic splice MET 14 skipping mutation.
[0105] In one embodiment, the chromosome number, physical coordinate, number of reference genome base changes, and number of mutation base changes are input into a prediction model to obtain a prediction result of whether it is a hidden splice MET 14 skipping mutation.
[0106] In one embodiment, the chromosome number, physical coordinate, number of reference genome base changes, and number of mutation base changes are input into a prediction model to obtain a prediction result of whether it is a classic splice MET 14 skipping mutation.
[0107] In one embodiment, the chromosome number, physical coordinate, number of reference genome base changes, and number of mutation base changes are input into a prediction model to obtain a prediction result of whether it is a MET 14 skipping mutation, and the prediction probability of classic splice MET 14 skipping mutation and hidden splice MET 14 skipping mutation is included.
[0108] In one embodiment, the hidden splice includes one or more of the following: an indel of more than 50 bp, a polypyrimidine region in an intron region, Branch AA, a splice auxiliary factor ESE in an exon region.
[0109] Optionally, the hidden splice includes: an indel of more than 50 bp, a polypyrimidine region in an intron region, Branch AA, a splice auxiliary factor ESE in an exon region.
[0110] In one embodiment, the embodiment of FIG. 1 and FIG. 7 are different in that the data obtained is different. The embodiment of FIG. 1 is to obtain MET gene basic data and processed gene data, and then directly input into a prediction model for prediction. The embodiment of FIG. 7 is to obtain MET gene basic data, and then input the MET gene basic data and calculated data into a first prediction model for prediction after calculating and processing the MET gene basic data. The prediction model and the first prediction model are different in that the input data is different, and then different prediction models are obtained by training.
[0111] The present application discloses an embodiment of a computer program product or system, comprising a computer program, which, when executed by a processor, implements the method steps of predicting the MET 14 skipping mutation described above.
[0112] Figure 2 is a schematic diagram of a system for predicting MET 14 skipping mutation according to an embodiment of the present application, which specifically comprises: an acquisition unit configured to acquire patient MET gene data, including chromosome number, physical coordinate, reference genome base number of changes, and mutation base number of changes;
[0113] a prediction unit configured to input the chromosome number, the physical coordinate, the reference genome base number of changes, and the mutation base number of changes into a prediction model to obtain a prediction result of whether the MET 14 skipping mutation is present.
[0114] Figure 3 is a schematic diagram of a device for predicting MET 14 skipping mutation according to an embodiment of the present application, which specifically comprises: a memory and a processor; the memory is configured to store program instructions; the processor is configured to invoke the program instructions, and when the program instructions are executed, any one of the above-mentioned methods for predicting MET 14 skipping mutation is performed.
[0115] The present application discloses an embodiment of a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, any one of the above-mentioned methods for predicting MET 14 skipping mutation is performed.
[0116] Figure 8 is a schematic diagram of a system for predicting MET 14 skipping mutation according to an embodiment of the present application, which specifically comprises: an acquisition unit configured to acquire patient MET gene data, including chromosome number, physical coordinate, reference genome base, and mutation base;
[0117] a calculation unit configured to calculate the reference genome base number of changes and the mutation base number of changes by calculating the base number of changes of the reference genome base and the mutation base;
[0118] a prediction unit configured to input the chromosome number, the physical coordinate, the reference genome base number of changes, and the mutation base number of changes into a prediction model to obtain a prediction result of whether the MET 14 skipping mutation is present.
[0119] The verification result of the verification embodiment shows that assigning inherent weights to the indications can improve the performance of the method compared with the default setting. It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described here. In the several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or in other forms. The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e. can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme. In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be in the form of hardware or software function unit. Those skilled in the art can understand that all or part of the steps of the various methods in the above embodiments can be completed by a program instructing related hardware, and the program can be stored in a computer readable storage medium, which can include read only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0120] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiment methods can be completed by a program instructing related hardware, and the program can be stored in a computer readable storage medium, and the above-mentioned medium storage can be read only memory, magnetic disk or optical disk, etc.
[0121] The computer device provided by the present application has been described in detail above. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in specific implementation manners and application ranges. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method of predicting a MET 14 skipping mutation, characterized in that, The method comprises the following steps: obtaining patient MET gene data, including chromosome number, physical coordinates, reference genome base number, and mutation base number; inputting the chromosome number, physical coordinates, reference genome base number, and mutation base number into a prediction model to obtain a prediction result of whether it is a MET 14 skipping mutation.
2. The method of predicting a MET 14 skipping mutation of claim 1, wherein, The MET gene data further comprises reference genome base and mutation base, and inputting the chromosome number, physical coordinates, reference genome base, mutation base, reference genome base number, and mutation base number into the prediction model to obtain the prediction result of whether it is a MET 14 skipping mutation.
3. The method of predicting a MET 14 skipping mutation of claim 1, wherein, The MET 14 skipping mutation comprises classical splicing and / or cryptic splicing, and inputting the chromosome number, physical coordinates, reference genome base number, and mutation base number into the prediction model to obtain a prediction result of whether it is a classical splicing and / or cryptic splicing MET 14 skipping mutation.
4. The method of predicting a MET 14 jump mutation of claim 3, wherein, The classical splicing comprises GT-AG.
5. The method of predicting a MET 14 skipping mutation of claim 3, wherein, The cryptic splicing comprises one or more of the following: indel of more than 50 bp, polypyrimidine tract in intron region, Branch AA, and splicing auxiliary factor ESE in exon region.
6. The method of predicting a MET 14 jump mutation of claim 3, wherein, The cryptic splicing comprises: indel of more than 50 bp, polypyrimidine tract in intron region, Branch AA, and splicing auxiliary factor ESE in exon region.
7. The method of predicting a MET 14 skipping mutation of claim 1, wherein, The training process of the prediction model comprises: obtaining a patient MET gene data set and a label, including chromosome number, physical coordinates, reference genome base number, and mutation base number, inputting the chromosome number, physical coordinates, reference genome base number, and base number into a to-be-trained model for training to obtain a prediction model, wherein the label is an RNA positive sample or other positive sample.
8. The method of predicting a MET 14 skipping mutation of claim 1, wherein, The prediction model adopts one or more of the following: random forest, decision tree, support vector machine, logistic regression model, convolutional neural network, XGBoost, and AdaBoost.
9. The method of predicting a MET 14 skipping mutation of claim 1, wherein, The patient MET gene data obtaining process comprises: obtaining patient MET gene basic data, including chromosome number, physical coordinates, reference genome base, and mutation base; calculating the base number of the reference genome base and the mutation base to obtain the reference genome base number and the mutation base number, and obtaining patient MET gene data.
10. The method of predicting a MET 14 skipping mutation of claim 9, wherein, The MET gene basic data obtaining process comprises: obtaining patient NGS sequencing data; performing alignment based on the NGS sequencing data and the MET reference genome to obtain an alignment result; performing variation detection on the alignment result to obtain a detection result; extracting the MET gene basic data based on the detection result.
11. The method of predicting a MET 14 jump mutation of claim 10, wherein, The NGS sequencing data is DNA sequencing data and / or RNA sequencing data.
12. A computer program product comprising a computer program or instructions embodied therein, characterized in that, The computer program or instruction is executed by the processor to implement the method for predicting the MET 14 skipping mutation according to any one of claims 1-11.
13. A computer device comprising a memory, a processor, and a computer program or instructions stored on the memory, wherein, The computer program or instruction is executed by the processor to implement the method for predicting the MET 14 skipping mutation according to any one of claims 1-11.
14. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions are executed by a processor to implement the method of predicting a MET 14 jump mutation according to any one of claims 1-11.
Citation Information
Patent Citations
Method for detecting mutation, electronic equipment and computer storage medium
CN111292802A
Method and device for detecting mutation
CN114596918A
Prediction method, system and platform based on influence of base mutation on mRNA splicing
CN114613431A
Method and device for detecting skipping mutation of 14 # exon of MET gene, storage medium and equipment
CN117305448A
Method, equipment and program product for predicting MET 14 jump mutation
CN120048347A