Method, equipment and program product for predicting MET 14 jump mutation

Through the artificial intelligence prediction model, the number of base changes in MET gene data is used to accurately predict MET 14 jump mutations, solving the problems of missed detection risks and RNA quality dependence in the existing technology, and improving detection efficiency and accuracy.

CN120048347APending Publication Date: 2025-05-27LIAONING KANGHUI BIOTECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510183538.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has a risk of missed detection when detecting MET 14 jump mutations, especially the identification of hidden shear mutations is not accurate enough, and it depends on RNA quality, which may not be ideal in clinical samples.

Method used

Using a prediction model trained by artificial intelligence, the number of base changes is calculated by obtaining patient MET gene data, and relevant information is input to the prediction model to predict whether it is a MET 14 jump mutation, including classical shear and hidden shear.

Benefits of technology

It improves the detection accuracy and efficiency of MET 14 jump mutations, and can accurately predict hidden shear mutations at the DNA level, reduce the risk of missed detection, and reduce detection costs and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048347A_ABST
    Figure CN120048347A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent medical treatment, in particular to a prediction method for predicting MET 14 jump mutation. Comprising the following steps: S1, acquiring MET gene data of a patient, including a chromosome number, a physical coordinate, a reference genome basic group and a mutant basic group; s2, calculating the basic group change number of the reference genome basic groups and the mutation basic groups to obtain the basic group change number of the reference genome and the change number of the mutation basic groups; s3, the chromosome number, the physical coordinates, the reference genome base change number and the mutation base change number are input into a prediction model, and a prediction result whether MET 14 jump mutation exists or not is obtained. The method can be used for detecting classical shearing and hidden shearing of MET 14 jump mutation, the mutation detection rate is increased, and the method has good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent medicine, and particularly relates to a method, device, program product and computer-readable storage medium for predicting MET 14 exon skipping mutations. Background Art

[0002] With the continuous increase in morbidity and mortality, cancer has become the leading cause of death in China and a major public health problem. Lung cancer is the most common cancer. Non-small cell lung cancer (NSCLC) is the most common pathological type of lung cancer, and most patients are already in the advanced stage at the time of diagnosis. NSCLC belongs to genotype diseases, and the development process of lung cancer is closely related to driver genes. The mutation status of driver genes is an important predictor of the efficacy of targeted therapy. After EGFR gene mutation and ALK gene fusion, the MET gene is another important driver gene in NSCLC and has now become a hot topic of targeted therapy. There are mainly three forms of abnormal activation of the MET gene in NSCLC: exon 14 (MET 14) skipping mutation of the MET gene, MET gene amplification, and protein overexpression. MET inhibitors have achieved good anti-tumor effects in NSCLC patients with MET 14 skipping mutations. In recent years, with the rapid development of drugs, the FDA has approved Tepotinib and Capmatinib, and the NMPA has also approved Savolitinib for non-small cell lung cancer patients with MET 14 skipping mutations. Currently, the methods used to detect MET 14 skipping mutations include DNA next-generation sequencing (NGS), Sanger sequencing of exon 14 and its flanking introns, reverse transcription-PCR (RT-PCR), and RNA-based NGS detection, etc. The companion diagnostic methods that have been approved to detect MET 14 skipping mutations are the FoundationOne NGS analysis in the United States and the Archer MET analysis in Japan. The NCCN guidelines (v1.2024) removed the IHC detection recommendation, and therefore MET 14 skipping mutations are best detected by FISH and NGS. Even so, some studies have found that there is a risk of missed detection in DNA-level detection for MET 14 skipping mutations, with a detection rate of 1.3%. RNA-based analysis detected a higher proportion of MET 14 skipping cases, with an RNA detection rate of 4.2%. However, the problem is that RNA-based analysis is highly dependent on RNA quality, which may not be ideal in some clinical samples. In addition, following the FDA's approval of Tepotinib and Capmatinib for metastatic NSCLC patients with MET 14 skipping mutations, the NMPA also approved Savolitinib (Volrasa) on June 22, 2021, and included it in the national medical insurance drug list on March 1, 2023. However, in clinical practice, due to the incomplete understanding of splicing codons, it is difficult to accurately identify hidden mutations other than the classical GT and AG splices, and cryptic splice variants are often overlooked. Summary of the Invention

[0003] In view of the above problems, the present invention provides a method for predicting MET 14 skipping mutations, specifically including: S1. Obtain the MET gene data of the patient, including chromosome number, physical coordinates, reference genome bases, and mutant bases; S2. Calculate the number of base changes of the reference genome bases and mutant bases to obtain the number of reference genome base changes and the number of mutant base changes; S3. Input the chromosome number, physical coordinates, number of reference genome base changes, and number of mutant base changes into the prediction model to obtain the prediction result of whether it is a MET 14 skipping mutation.

[0004] The MET 14 skipping mutation includes classical splicing and / or cryptic splicing.

[0005] Optionally, the classical splicing includes GT-AG.

[0006] The cryptic splicing includes: indels of more than 50 bp, polypyrimidine tracts in the intron region, Branch AA, and exon splicing enhancer ESE.

[0007] S3 is replaced by: Input the chromosome number, physical coordinates, reference genome bases, mutant bases, number of reference genome base changes, and number of mutant base changes into the prediction model to obtain the prediction result of whether it is a MET 14 skipping mutation.

[0008] The training process of the prediction model is: Obtain the MET gene data set and labels of the patient, including chromosome number, physical coordinates, reference genome bases, and mutant bases; Calculate the number of changes in the reference genome bases and mutant bases, and then input the chromosome number, physical coordinates, and number of base changes into the model to be trained for training to obtain the prediction model, where the label is an RNA positive sample or other positive sample.

[0009] The prediction model adopts one or more of the following: random forest, decision tree, support vector machine, logistic regression model, convolutional neural network, XGBoost, AdaBoost.

[0010] The acquisition of the MET gene data includes: Obtain the NGS sequencing data of the patient; Based on the NGS sequencing data, perform alignment with the MET reference genome to obtain the alignment result; Perform variant detection on the alignment result to obtain the detection result; Extract the MET gene data based on the detection result.

[0011] Optionally, the NGS sequencing data is DNA sequencing data and / or RNA sequencing data.

[0012] The object of the present invention is to provide a computer program product, which includes a computer program or instruction, and the computer program or instruction is executed by a processor to implement the method for predicting MET 14 skipping mutations described above.

[0013] The object of the present invention is to provide a computer device, which includes a memory, a processor, and a computer program or instruction stored on the memory, and the computer program or instruction is executed by the processor to implement the method for predicting MET 14 skipping mutations described above.

[0014] The object of the present invention is to provide a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is executed by a processor to implement the method for predicting MET 14 skipping mutations described above.

[0015] Advantages of the present invention: 1. In the prior art, there are manual experience review, RNA retesting method, etc. However, these methods have the following disadvantages: ① The subjective factors of people are unstable, and the influence of experience and cognitive differences is large, and the mistakes are also large. ② For both technology providers and patients, they both face increased time costs and detection costs. ③ The probability of occurrence of personalized patient cases is not large enough, especially for rare hidden mutations. However, organisms have diversity, and the variability of hidden splicing mutation forms is also very high, and the rules are not easy to summarize and are not convenient to promote. ④ Collecting cases requires long-term, multi-regional, large-data accumulation and special excavation. Therefore, the present invention proposes a method for predicting MET 14 skipping mutations, using artificial intelligence to train a prediction model to form an objective, efficient, and accurate method for detecting MET 14 skipping mutations, which helps to improve the detection rate of gene mutations in patients and improve the work efficiency of doctors.

[0016] 2. Aiming at the missed detection of hidden splicing mutations, the present invention focuses on considering the damage to the binding stability of snRNPs when the number of base changes in the exons and introns of the MET gene or large structural changes occur, so as to discover the rules of hidden splicing mutations other than classical splicing on MET 14. In this way, an objective is achieved: even when the tissue is inaccessible or RNA data cannot be used, the hidden splicing of the MET 14 exon occurring at the DNA level can be accurately predicted. Specifically, when training the prediction model, the number of base changes is used as the main feature for training to achieve the detection of classical detection mutations and hidden splicing mutations of MET 14, avoid omission, and improve the disease detection rate.

[0017] 3. Existing artificial intelligence has some problems in predicting MET 14 skipping mutations. It cannot identify the same base substitution delins and indel type mutations that are too large (>50bp). It seems to have a tendency to poorly predict the ESE of variable subtypes, and the sensitivity to Branch AA is not reflected. To address this problem, based on the principle that RNA splicing is catalyzed by the assembly of snRNPs plus other proteins, which together constitute the spliceosome, different from the previous base processing process, this invention uses the number of base changes to be able to complete the identification of the same base substitution delins, indel type mutations >50bp, variable subtype ESE, and Branch AA. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 Schematic flowchart of the method for predicting MET 14 skipping mutations provided by the embodiments of the present invention; Figure 2 Schematic diagram of the system for predicting MET 14 skipping mutations provided by the embodiments of the present invention; Figure 3 Schematic diagram of the device for predicting MET 14 skipping mutations provided by the embodiments of the present invention; Figure 4 Mutation detection rate of the research and internal data of NSCLC provided by the embodiments of the present invention; Figure 5 Comparison result of the mutation types of the research and internal data of NSCLC provided by the embodiments of the present invention; Figure 6 Comparison result of the mutations detected by double detection of DNA and RNA in the research and internal data of NSCLC provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention.

[0021] In some of the processes described in the specification, claims, and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0022] Figure 1 The schematic diagram of the method for predicting MET exon 14 skipping mutations provided by the embodiments of the present invention specifically includes: S1: Obtain the MET gene data of the patient, including chromosome number, physical coordinates, reference genome bases, and mutant bases; In a specific embodiment, regarding the method of splicing prediction, although cross-border scientists have explored modeling, statistics and other prediction methods for many years, it is well known that the prediction accuracy has not reached the level of evidence-based medicine. However, a recently influential and convincing study came from a 2019 Cell report describing SpliceAI as a convolutional neural network that can accurately predict gene mutations leading to cryptic splicing. The study found that synonymous mutations and intronic mutations affecting splicing have a high verification rate on RNA-seq and are highly harmful in the human population.

[0023] In recent years, high-quality clinical trials have been carried out in China and are at the international forefront. However, the understanding of MET exon 14 skipping mutations is not complete. The current method is to establish an industry consensus through well-known experts and organizations in the industry. For example, the purpose of the consensus-building process is to discuss controversial issues related to MET mutation detection, including MET exon 14 skipping mutations, MET gene amplification, and MET protein expression in non-small cell lung cancer. This consensus was formed through two rounds of in-depth discussions at a virtual conference, involving 20 pathologists and 19 clinical experts.

[0024] Select 15 recent studies on NSCLC (15 Studies) with a total of 8352 samples on the cBioPortal big data platform, and excerpt 11086 NSCLC patient tissue samples from January 2021 to April 2023 as internal data (In House) representatives. Although there is not much difference in the detection rate of MET 14 mutations between the two types of databases ( Figure 4 as shown), but when delving into each patient sample and specifically comparing the mutation types again ( Figure 5As shown, the difference in the number of mutations detected in the Exon range in the public database is not significant. However, interestingly, no mutations were detected in the Intron range.

[0025] According to the internal data study, under the condition of double detection of DNA-RNA, samples with positive RNA results were selected to observe their DNA mutations ( Figure 6 As shown), if the mutation occurs in the intron region, there is a significant difference between the DNA and RNA positive results, indicating that there are loopholes in the detection performance of DNA in the intron region. However, in the real world, when the sample conditions are limited or tissue samples cannot be obtained, the detection of RNA cannot be carried out, which means a bottleneck for the patient's treatment process. Such situations are relatively common in clinical practice. And if NSCLC patients have to undergo additional detection experiments, time, and costs, it is not only inconvenient but also very uneconomical. Therefore, it is particularly important to predict the occurrence of splicing through non-invasive DNA or DNA detection.

[0026] The internal data test also found that SpliceAI can predict most hidden splicing mutations, especially solving deletion-type mutations. It is slightly regrettable that through real-world data testing, it was found that it really cannot identify the same base substitution delins and overly large indel-type mutations (>50bp). It seems to have a tendency to predict poorly for variable subtypes of ESE, and the sensitivity to Branch AA is not reflected.

[0027] In one embodiment, the acquisition of the MET gene data includes: Obtaining the NGS sequencing data of the patient; Based on the comparison of the NGS sequencing data with the MET reference genome to obtain a comparison result; Performing variant detection on the comparison result to obtain a detection result; Based on the detection result, extracting the MET gene data.

[0028] The NGS sequencing data is DNA sequencing data and / or RNA sequencing data.

[0029] In one embodiment, the method for obtaining gene data includes: When the rsID is known, directly query dbSNP or Ensembl; dbSNP (NCBI): Enter the rsID (such as rs123456) for query, and obtain the chromosome position, reference / mutated base, and reference genome version. Ensembl VEP: Enter the chromosome coordinates or rsID to obtain detailed annotations. Directly extract the chromosome number, coordinates, reference sequence, and mutated sequence from the data returned on the result page or API.

[0030] When there is an existing VCF file, gene data is extracted through tools, including Python, R language, etc.

[0031] When obtaining from raw sequencing data, data preprocessing: Use tools such as FastQC to perform quality assessment on the raw NGS sequencing data (usually in FASTQ format), view indicators such as the quality distribution, base content, GC content, and sequence duplication rate of the sequencing data to understand the overall quality of the data; Use one or several of the following tools: fastp, Trimmomatic, Cutadapt to remove adapter sequences, low-quality bases, and short sequence fragments from the sequencing data. Generally, a quality threshold (such as a Phred quality value below 20) and a length threshold (such as a sequence length less than 30bp) will be set to remove sequences or bases that do not meet the requirements to improve the accuracy of subsequent analysis.

[0032] Sequence alignment: Download a suitable reference genome containing the MET gene to ensure the integrity and accuracy of the reference genome for subsequent accurate alignment of the sequencing sequences to the genome; Use alignment tools such as BWA (Burrows-Wheeler Aligner), Bowtie2 to align the preprocessed sequencing data with the reference genome. These tools can quickly and accurately find the best matching positions of the sequencing reads on the reference genome and generate an alignment file in SAM (Sequence Alignment / Map) format; Use tools such as SAMtools to convert the SAM file to a BAM (Binary Alignment / Map) file. The BAM file is a binary compressed form of the SAM file, which occupies less space and is convenient for subsequent processing. At the same time, sort the BAM file by chromosome position for subsequent variant detection and analysis. In a specific embodiment, use bwa mem to align the reads after quality control and splitting to the reference genome hg19 (GRCh37), use samtools view to filter out multiply aligned and unaligned reads, and use samtools sort to sort the alignment results to generate sort.bam.

[0033] Use gencore to remove duplicates from the reads of the tumor sample, correct the low-quality and incorrect bases, and then use samtools sort for sorting. The deduplicated and sorted bam is used for somatic snv indel detection; Use sambamba to remove duplicates from the reads of the tumor sample and the white blood cell control sample, and the white blood cell bam is used for somatic snv indel detection.

[0034] Variant Detection: Select variant detection tools. Commonly used variant detection tools include GATK (Genome Analysis Toolkit) and FreeBayes. These tools can identify variant information such as single nucleotide variants (SNVs) and insertions / deletions (InDels) in the sequencing data based on the alignment results. Through these steps, the variant sites in the sample can be accurately detected, and a VCF file containing all variant information can be generated. In a specific embodiment, the variant detection tool samtools mpileup is used to establish pileup format files for tumor and leukocyte samples, and MutLoc (SNPIndel) and Varcidt (longindel) are used to detect the original mutations. And a VCF file containing all variant information is generated.

[0035] Screening MET Gene-Related Variants: Use tools such as VCFtools and Bcftools to screen out the variant sites related to the MET gene from the genome-wide VCF file according to the position information of the MET gene on the reference genome (such as chromosome position, gene region), and generate a VCF file containing only the variant information of the MET gene.

[0036] In a specific embodiment, based on the NGS capture panel design, library construction, and sequencing methods of the present invention, it can be considered that approximate or identical off-machine data can be obtained at the experimental level. Tumor somatic mutation information is obtained through background pool filtering or paired analysis methods. As shown in Table 1, the model input data obtains chromosome, physical coordinates, reference genome bases, and mutant base information from the VCF file, but unlike previous operations, the ATCG bases are not OneHot encoded. The base changes are converted into digital statistics through data preprocessing as input elements. This is because the OneHot encoding has a restrictive effect on the length of the base sequence and is not applicable to scenarios such as detecting DNA sequence disruptions such as deletions or base substitutions of more than 50 bp.

[0037] Table 1 VCF File Structure (Excerpt)

[0038]

[0039] S2: Calculate the number of base changes between the reference genome base and the mutant base to obtain the number of reference genome base changes and the number of mutant base changes; In a specific embodiment, based on the NGS data of non-small cell lung cancer patients in the past 5 years, the sample types include DNA data such as tissue slides and paraffin rolls, plasma, pleural effusion, cerebrospinal fluid, etc., and RNA positive results as label tags to construct a supervised machine learning model: a random forest model. The extracted features need to be converted into a format suitable for input into the machine learning model. This includes converting text data into numerical data, handling missing values, encoding categorical features, etc. Usually, when building a model for learning, One-Hot encoding is performed on the ATCG bases of the DNA sequence to convert it into information that can be processed by a neural network. One-Hot encoding is very useful when processing DNA sequence data. It can convert sequence data into numerical data, enabling the machine learning model to better understand and process this data. For example, A: [1, 0, 0, 0]; T: [0, 1, 0, 0]; C: [0, 0, 1, 0]; G: [0, 0, 0, 1]. However, for longer sequence strings, One-Hot encoding will greatly increase the dimensionality of the data. According to the SpliceAI test and literature suggestions, the base sequence length should not exceed 50bp. Considering that some patterns and regularities in the DNA sequence may be lost during the One-Hot encoding process, such as the periodic patterns of certain nucleotides or the interaction information between nucleotides. As previously understood in the literature review of the RNA splicing mechanism, RNA splicing is catalyzed by the assembly of snRNPs plus other proteins, which together form the spliceosome. However, deletions and substitutions of more than 50bp bases that damage the snRNPs are extremely important features and must not be missed for analysis and annotation. To sum up, the counts of ATCG bases for REF and ALT are shown in Tables 2 and 3 respectively, with the change in the number of bases as the characteristic element as the content of the training and validation datasets.

[0040] Table 2 Training and Validation Datasets (Example)

[0041]

[0042] Table 3 Training and Validation Datasets after Data Preprocessing (Example)

[0043]

[0044] S3: Input the chromosome number, physical coordinates, number of base changes in the reference genome, and number of mutant base changes into the prediction model to obtain the prediction result of whether it is a MET 14 skipping mutation.

[0045] In one embodiment, the MET 14 skipping mutation includes classical splicing and / or cryptic splicing; Optionally, the classical splicing includes GT-AG.

[0046] In one embodiment, the cryptic splicing includes: indels of more than 50 bp, polypyrimidine tracts in the intron region, Branch AA, and exon splicing enhancer (ESE).

[0047] In one embodiment, the training process of the prediction model is as follows: Obtain a patient MET gene dataset and labels, including chromosome number, physical coordinates, reference genome bases, and mutant bases; calculate the number of changes in the reference genome bases and mutant bases, and then input the chromosome number, physical coordinates, number of changes in the reference genome bases, and number of changes in the mutant bases into the model to be trained for training to obtain the prediction model, where the label is an RNA-positive sample or other positive sample.

[0048] In one embodiment, S3 is replaced with: Input the chromosome number, physical coordinates, reference genome bases, mutant bases, number of changes in the reference genome bases, and number of changes in the mutant bases into the prediction model to obtain a prediction result of whether it is a MET exon 14 skipping mutation.

[0049] The training process of the prediction model is as follows: Obtain a patient MET gene dataset and labels, including chromosome number, physical coordinates, reference genome bases, and mutant bases; calculate the number of changes in the reference genome bases and mutant bases, and then input the chromosome number, physical coordinates, number of changes in the reference genome bases, number of changes in the mutant bases, reference genome bases, and mutant bases into the model to be trained for training to obtain the prediction model, where the label is an RNA-positive sample or other positive sample.

[0050] In one embodiment, the prediction model uses one or more of the following: random forest, decision tree, support vector machine, logistic regression model, convolutional neural network, XGBoost, AdaBoost.

[0051] In a specific embodiment, the prediction of the random forest can be expressed as: y^i = mode{y^i1, y^i2, …, y^iB} Among them, y^i is the prediction of the random forest for the i-th sample, and y^ij is the prediction of the j-th tree for the i-th sample. A class implementing the random forest algorithm is established to solve the classification problem.

[0052] X = data[['CHROM', 'POS', 'lenR', 'lenA']] y = data['label'] Among them, X: characteristic data. y: label data, which contains the labels of each sample. RNA positive is 1 and negative is 0. The proportion of the test set is set to 0.3, and 30% of the data will be used as the test set. random_state is set to 42 to ensure that the split is consistent each time the code is run, which helps with the reproducibility of the results. A random forest model is created, and GridSearchCV is used to traverse different parameter combinations, and 5-fold cross-validation is used to evaluate the performance of each parameter combination. Finally, the best parameter combination is used to train the model and evaluate it on the test set. Finally, the trained model is saved for prediction on new data.

[0053] In one embodiment, the method for obtaining the specific chromosome number, physical coordinates, reference genome sequence, and mutant genome sequence of a mutation does not include the PCR method and the amplicon method, or result in the missed detection of MET 14 cryptic splicing mutations.

[0054] In a specific embodiment, the optimal parameter combination of the Rfmodel of this model is optimized as follows: max_depth: None, the maximum depth of the decision tree is not limited, and the tree will grow until all leaf nodes are pure. Best cross-validation score: 0.98, the best score obtained in cross-validation is 0.98, which is a very high score, indicating that the model has good performance on the training data. As shown in Table 4, Accuracy: 0.99, the accuracy rate is 99%, which means that the model can correctly predict 99% of the samples. Sensitivity: 1.00, also known as the recall rate or true positive rate, indicating that the model's ability to correctly identify positive samples is 100%. Specificity: 0.98, indicating that the model's ability to correctly identify negative samples is 98%, as shown in Table 4. The prediction performance of the random forest model of the present invention is tested year-on-year and is superior to the DNN neural network (Accuracy: 0.67), support vector machine (Accuracy: 0.81), and Logist regression model (Accuracy: 0.70).

[0055] Table 4 Evaluation Metrics of the Rfmodel

[0056]

[0057] Testing and comparison with new data outside the dataset found that, as shown in Table 5, the Rfmodel of this model can predict: classical splicing GT-AG, and hidden splicing including: indels over 50bp, polypyrimidine tracts (pypy) in the intron region, Branch AA, and exon splicing enhancer ESE. In terms of detecting MET 14 skipping mutations, this can well complement or replace SpliceAI.

[0058] Table 5 Comparison of prediction and other analysis methods with newly collected negative and positive data in the past year

[0059]

[0060] In a specific embodiment, the process of training a prediction model and applying the prediction model to predict MET 14 skipping mutations: Step A: Obtain the content of the MET gene VCF analysis file by analyzing the lung cancer DNA data of non-small cell lung cancer patients over the years: chromosome number chr, physical coordinate pos, reference genome base ref, and mutant base alt information as the pre-information of the input layer.

[0061] Step B: Data preprocessing, convert the information in Step A into chr, pos, lenR, and lenA numerically. Additionally, label the samples with RNA positive results or confirmed as positive by other means of MET14 skipping mutations as label = 1, and others as label = 0, and together use them as the input layer data of this random forest model.

[0062] Step C: Load the data into the model, use the train_test_split function in Python from the sklearn.model_selection module to divide the dataset into a training set and a test set, and specify that the proportion of the test set in the total dataset is 30%. random_state: Set the seed of the random number generator to ensure that the results of each split are the same, increasing the reproducibility of the code. Establish a model and implement GridSearchCV for parameter tuning to obtain the best parameters and model. Start training the random forest model and save the trained model as Rfmodel.

[0063] Step D: Obtain the new NGS off-machine data, and the content of the bioinformatics analysis VCF file includes: chromosome number chr, physical coordinate pos, reference genome base ref, and mutant base alt information as the input pre-information.

[0064] Step E: Data preprocessing. Convert the input layer and information numbers into chr, pos, lenR, and lenA, and load them into this random forest model Rfmodel for prediction. Obtain the binary classification prediction results of Table 1 Rfmodel. "0" indicates not a MET 14 splicing mutation, and "1" indicates a MET 14 splicing mutation.

[0065] The disclosed embodiments of the present invention also provide a computer program product or system, including a computer program that, when executed by a processor, implements the method steps for predicting MET 14 skipping mutations described above.

[0066] Figure 2 Schematic diagram of the system for predicting MET 14 skipping mutations provided by the embodiments of the present invention, specifically including: Acquisition unit: Acquire patient MET gene data, including chromosome number, physical coordinates, reference genome bases, and mutant bases; Calculation unit: Calculate the number of base changes in the reference genome bases and mutant bases to obtain the number of reference genome base changes and the number of mutant base changes; Prediction unit: Input the chromosome number, physical coordinates, number of reference genome base changes, and number of mutant base changes into the prediction model to obtain the prediction result of whether it is a MET 14 skipping mutation.

[0067] Figure 3 Schematic diagram of the device for predicting MET 14 skipping mutations provided by the embodiments of the present invention, specifically including: Memory and processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it performs any of the above methods for predicting MET 14 skipping mutations.

[0068] The disclosed embodiments of the present invention also provide a computer-readable storage medium that stores a computer program, and when the computer program is executed by a processor, it performs any of the above methods for predicting MET 14 skipping mutations.

[0069] The verification results of this verification embodiment show that allocating fixed weights for indications can improve the performance of this method compared to the default settings. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, or optical disc, etc.

[0070] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned medium storage can be read-only memory, magnetic disk, or optical disc, etc.

[0071] The above has introduced in detail a computer device provided by the present invention. For those of ordinary skill in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for predicting MET 14 jumping mutations, characterized in that: include: S1. Obtain the patient's MET gene data, including chromosome number, physical coordinates, reference genome bases, and mutant bases; S2. Calculate the number of base changes of the reference genome bases and mutant bases to obtain the number of base changes of the reference genome and the number of base changes of the mutant bases; S3. Input the chromosome number, physical coordinates, number of base changes in the reference genome, and number of mutant base changes into the prediction model to obtain a prediction result of whether it is a MET 14 jump mutation.

2. The method for predicting MET 14 jumping mutation according to claim 1, characterized in that: The MET 14 jumping mutation includes classical splicing and / or hidden splicing; Optionally, the classical shear comprises GT-AG.

3. The method for predicting MET 14 jumping mutation according to claim 2, characterized in that: The hidden shearing includes: indels of more than 50 bp, polypyrimidine regions in the intron region, Branch AA, and shearing auxiliary factor ESE in the exon region.

4. The method for predicting MET 14 jumping mutation according to claim 1, characterized in that: The S3 is replaced by: inputting the chromosome number, physical coordinates, reference genome base, mutant base, number of reference genome base changes, and number of mutant base changes into the prediction model to obtain a prediction result of whether it is a MET 14 jump mutation.

5. The method for predicting MET 14 jumping mutation according to claim 1, characterized in that: The training process of the prediction model is: obtaining the patient's MET gene data set and label, including chromosome number, physical coordinates, reference genome bases, and mutant bases; calculating the number of changes in reference genome bases and mutant bases, and then inputting the chromosome number, physical coordinates, and number of base changes into the model to be trained for training to obtain a prediction model, wherein the label is an RNA positive sample or other positive sample.

6. The method for predicting MET 14 jumping mutation according to claim 1, characterized in that: The prediction model adopts one or more of the following: random forest, decision tree, support vector machine, logistic regression model, convolutional neural network, XGBoost, AdaBoost.

7. The method for predicting MET 14 jumping mutation according to claim 1, characterized in that: The acquisition of the MET gene data includes: Obtain NGS sequencing data of patients; Comparing the NGS sequencing data with the MET reference genome to obtain a comparison result; Performing variation detection on the comparison result to obtain a detection result; Extracting and obtaining MET gene data based on the test results; Optionally, the NGS sequencing data is DNA sequencing data and / or RNA sequencing data.

8. A computer program product comprising a computer program or instructions, characterized in that: The computer program or instructions are executed by a processor to implement the method for predicting MET 14 jumping mutations according to any one of claims 1 to 7.

9. A computer device comprising a memory, a processor and a computer program or instruction stored in the memory, characterized in that: The computer program or instructions are executed by a processor to implement the method for predicting MET 14 jumping mutations according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: The computer program or instructions are executed by a processor to implement the method for predicting MET 14 jumping mutations according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method, device, and program product for predicting met14 skipping mutation

    WO2025237449A3