Harmful initiation codon loss mutation prediction method based on cross-modal feature fusion
By integrating DNA, RNA, protein, and epigenetic modification features through cross-modal feature fusion and a deep learning framework, this approach overcomes the limitations of existing tools in predicting start codon loss mutations and achieves efficient and accurate identification of mutation harmfulness.
Patent Information
- Application Number
- CN202511703754.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-02-13
AI Technical Summary
Existing start codon loss mutation prediction tools have limitations in data coverage, feature utilization depth, and model generalization ability, making it difficult to effectively identify the harmfulness of mutations, especially for newly emerging and difficult-to-quantify mutation features.
A prediction method based on cross-modal feature fusion is constructed, which integrates multi-level features such as DNA sequence, RNA structure, protein function and epigenetic modification. Feature extraction and interactive modeling are performed through a deep learning framework. Bidirectional attention fusion and multi-layer fully connected layers are used for prediction. The Adam optimization algorithm is used to optimize the model parameters.
It significantly improves the prediction accuracy of harmful start codon loss mutations, enhances the objectivity and comprehensiveness of feature extraction, and can systematically reveal the differences and intrinsic connections between different biomolecular levels, providing a new analytical method for the functional impact mechanism of genetic variation.
Smart Images

Figure CN121528323A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biological information calculation, and relates to pre-training language models, feature fusion, gene mutation prediction classification and the like technologies, in particular to a harmfulness start codon loss mutation prediction method based on cross-modal feature fusion. BACKGROUND
[0002] Start-loss mutation is a type of gene mutation, which refers to the change of nucleotides in DNA sequence, thus changing the basic composition of codon, interfering with the recognition of ribosome to the codon, causing abnormal translation initiation, leading to the production of abnormal proteins, and then interfering with the normal physiological function of the human body. With the development of sequencing technology, a large number of start-loss mutations have been identified, however, the analysis of these mutations is still insufficient. Abad-Navarro et al. conducted a statistical analysis on Ensembl database (data as of December 2017), and found 11261 genetic variations in the start codon AUG of 7205 genes, but most of these mutations (99.5%) have not been clearly related to harmful traits. Given the long time and high cost of biological experiments, researchers have developed a variety of computational tools to assess the harmfulness of mutations, such as CADD (Schubach M, Maass T, Nazaretyan L, et al. CADD v1. 7: using protein language models, regulatory CNNs and other nucleotide-level scores to improve genome-wide variant predictions[J]. Nucleic acids research, 2024, 52(D1): D1143-D1154.), CAPICE (Li S, van der Velde K J, De Ridder D, et al. CAPICE: a computational method for Consequence-Agnostic Pathogenicity Interpretation of Clinical Exome variations[J]. Genome Medicine, 2020, 12: 1-11.) and GPN-MSA (Benegas G, Albors C, Aw A J, et al. GPN-MSA: an alignment-based DNA language model for genome-wide variant effect prediction[J]. bioRxiv, 2023.). However, the start-loss mutation data used in the training process of these tools is relatively small, which limits their effectiveness in predicting such mutations.Furthermore, tools focusing on the harmfulness prediction of start codon loss mutations are extremely rare. The only two available tools are PoStaL (Takata A, Hamanaka K, Matsumoto N. Refinement of the clinical variant interpretation framework by statistical evidence and machine learning[J]. Med, 2021, 2(5): 611-632. e9.) and initiationMutationPredictor (Castell-Díaz J, Abad-Navarro F, de la Morena-Barrio ME, et al. Using machine learning for predicting the effect of mutations in the initiation codon[J]. IEEE journal of biomedical and health informatics, 2022, 26(11): 5750-5756.). These tools require artificially defined biological characteristics as input and are difficult to effectively predict emerging and difficult-to-quantify mutational features. The deep learning supervised model StartPred (Liu J, Wang L, Su Y, et al. Prediction of Human Pathogenic Start Loss Variants Based on Multi-channel Features[J]. Journal of Chemical Information and Modeling, 2025, 65(17): 9330-9341.) has limitations in feature construction. Its feature system only integrates conserved features and epigenetic modification features, and fails to fully explore the interaction relationships of multi-level features of biomolecules. Summary of the Invention
[0003] This invention addresses the shortcomings of existing technologies by proposing a method for predicting harmful start codon loss mutations based on cross-modal feature fusion. This method integrates cross-modal biological information by fusing features from DNA sequence, RNA structure, protein function, and epigenetic modifications to assess the harmfulness of start codon loss mutations. This enables efficient and accurate identification of the harmfulness of start codon loss mutations and overcomes the limitations of existing start codon loss mutation harmfulness prediction tools in terms of data coverage, feature utilization depth, and model generalization ability.
[0004] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The method for predicting harmful start codon loss mutations based on cross-modal feature fusion, as described in this invention, is characterized by the following steps: Step 1: Obtain the dataset of mutation samples with missing start codons and perform preprocessing to obtain the preprocessed mutation sample dataset. },in, Represents the first after preprocessing A start codon loss mutation sample, The number of mutation samples representing the loss of the start codon is given by: The true category label is ,and ;when When, it means This is a benign mutation sample; when When, it means This is a sample with a harmful mutation. Step 2: Construct a multi-channel feature set containing DNA sequence features, RNA structural features, protein functional features, and epigenetic modification features; Step 3: Use M different sizes of convolution kernels to respectively... Multi-channel feature input set The first in Channel characteristics Feature extraction is performed to obtain the first... The M local features of the channel feature are concatenated to form the first... intermediate feature vectors ;right After processing with the Mish activation function, and then through batch normalization and max pooling layers, the th... The deep derived feature vector of each channel ;in, The output dimension represents the feature of a single channel; Step 4: Utilize bidirectional attention fusion and concatenation methods to process the deep-derived feature vectors. By merging, we obtain the first... fusion feature vectors ; Step 5: Utilize three fully connected layers to... Processing is performed to obtain Predicted probability ; Step 6: Construct the loss function of the classification method using equation (6). : (6) Step 7: Use the Adam optimization algorithm to optimize the start codon harmfulness loss mutation prediction and classification network, and calculate the optimal solution. The network parameters are updated until convergence is achieved, thus obtaining a prediction classification model with optimal parameters, which is used to predict harmful loss mutations of the start codon.
[0005] The method for predicting harmful start codon loss mutations based on cross-modal feature fusion described in this invention is also characterized in that step 2 is performed as follows: Step 2.1: Extract using a large model. DNA short-range sequence characteristics DNA long-distance sequence characteristics RNA structural features Related to protein functional characteristics and will Any feature of is denoted as the first. Channel characteristics ,in, This represents the first segment of the large model. The length of each channel sequence The first output of the large model represents the... Each channel is embedded in a dimension; ; Step 2.2, Construction Epigenetic modification characteristics ,in, Dimensions representing epigenetic modification characteristics; Step 2.3, Construction Multi-channel feature input set .
[0006] Furthermore, step 4 is performed as follows: Step 4.1: Use equations (1)-(3) to derive feature vectors from the RNA structure depth. Deeply derived feature vectors of protein function Perform bidirectional attention fusion to obtain the joint features after fusion of the i-th BAN. : (1) (2) (3) In equations (1), (2), and (3), and This represents the two bilinear weight matrices to be learned. express right Attention weights express right Reverse attention weights; Indicates transpose; Step 4.2, Deeply derived feature vectors from short-range DNA sequences DNA long-range sequence deep-derived feature vector , After splicing, we get the first... Multimodal feature vectors ; Step 4.3, using equation (4) to... Processing yields the first... The fusion feature vector of start codon loss mutation samples : (4) In equation (4), This indicates a batch normalization operation. This represents the activation function. and These represent the weights and biases of the fully connected layer, respectively.
[0007] Furthermore, step 5 involves using equation (5) to... Process the data to obtain the output of the i-th hidden layer. And after processing by the Sigmoid mapping function, we obtain Predicted probability : + (5) In equation (5), , , This represents the weight matrix of each layer in a three-layer fully connected layer. , , This represents the bias terms for each of the three fully connected layers.
[0008] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program supporting the processor in executing the harmful start codon loss mutation prediction method, and the processor is configured to execute the program stored in the memory.
[0009] The present invention provides a computer-readable storage medium on which a computer program is stored, characterized in that the computer program, when executed by a processor, performs the steps of the harmful start codon loss mutation prediction method.
[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Existing studies on the harmfulness prediction of start codon loss mutations mostly rely on single-level features or simple combinations of features, making it difficult to comprehensively capture the complex effects of mutations at multiple levels, including DNA, RNA, proteins, and epigenetics. This invention, for the first time, constructs a four-dimensional cross-modal feature fusion system encompassing DNA sequence features, epigenetic modification features, RNA structural features, and protein functional features. By deeply integrating embedded features extracted from multiple pre-trained models, it systematically reveals the differences and intrinsic connections between different biomolecular levels, significantly improving the prediction accuracy of harmful start codon loss mutations.
[0011] 2. The MSCPred model proposed in this invention adopts an end-to-end deep learning framework, which can automatically learn and extract distributed representations containing multi-scale biological laws from the original sequence data. This avoids the tedious and subjective manual feature engineering in traditional methods, effectively reduces the model development cost, and enhances the objectivity and comprehensiveness of feature extraction.
[0012] 3. This invention, through a cross-modal feature interaction modeling strategy, can not only capture the contrast features between mutant sequences and reference sequences, but also analyze the conformational changes within mutant sequences and the synergistic relationships of features at different biomolecular levels. This systematic modeling of cross-level biological influence networks provides a new methodological paradigm for understanding the functional impact mechanisms of genetic variation, and has significant theoretical innovation value. Attached Figure Description
[0013] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0014] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0015] In this embodiment, a method for predicting harmful start codon loss mutations based on cross-modal feature fusion is used to accurately identify harmful start codon loss mutations and their potential consequences, such as... Figure 1 As shown, the method includes the following steps: Step 1: Obtain start codon loss mutation sample data from multiple gene mutation databases. Preprocess the harmful mutation samples (positive samples) and benign mutation samples (negative samples) by removing duplicates from different sources and excluding mutation samples with contradictory labels. Then, based on the Close-by principle, divide the preprocessed start codon loss mutation dataset into a training set and an independent test set to obtain the dataset. },in This represents the mutation sample where the i-th start codon is lost. The number of mutation samples representing the loss of the start codon is given by: The category label is ,and ;when When, it means This is a benign mutation sample; when When, it means This is a sample with harmful mutations.
[0016] Step 2: Construct a multi-channel input set containing DNA sequence features, RNA structural features, protein functional features, and epigenetic modification features; Step 2.1: Extract using a large model. DNA short-range sequence characteristics DNA long-distance sequence characteristics RNA structural features Related to protein functional characteristics and will Any feature of is denoted as the first. Channel characteristics ,in, This represents the first segment of the large model. The length of each channel sequence The first output of the large model represents the... Each channel is embedded in a dimension; .
[0017] Step 2.2, Construction Epigenetic modification characteristics ,in, Dimensions representing epigenetic modification characteristics; Step 2.3, Construction Multi-channel feature input set .
[0018] Step 3: Use M different sizes of convolution kernels to respectively... Feature extraction is performed to obtain the first... The M local features of each channel are concatenated into an intermediate feature vector. ;right After processing with the Mish activation function, and then through batch normalization and max pooling layers, the th... The deep derived feature vector of each channel ;in, This represents the output dimension of a single channel feature.
[0019] Step 4: Utilize bidirectional attention fusion and concatenation methods to process the deep-derived feature vectors. Perform fusion to obtain the i-th fused feature vector. ; Step 4.1: Use equations (1)-(3) to derive feature vectors from the RNA structure depth. Deeply derived feature vectors of protein function Perform bidirectional attention fusion to obtain the joint features after fusion of the i-th BAN. : (1) (2) (3) In equations (1), (2), and (3), and This represents the two bilinear weight matrices to be learned. express right Attention weights express right Reverse attention weights; This indicates transpose.
[0020] Step 4.2, Deeply derived feature vectors from short-range DNA sequences DNA long-range sequence deep-derived feature vector , After concatenation, a multimodal feature vector is obtained. .
[0021] Step 4.3, using equation (4) to... The process is performed to obtain the fusion feature vector of the i-th start codon loss mutation sample. : (4) In equation (4), This indicates a batch normalization operation. This represents the activation function. and These represent the weights and biases of the fully connected layer, respectively.
[0022] Step 5: Use equation (5) to... Process the data to obtain the output of the i-th hidden layer. After processing by the Sigmoid mapping function, we obtain Predicted probability : + (5) In equation (5), , , This represents the weight matrix of a three-layer fully connected layer. , , This represents the corresponding bias term. The Sigmoid layer then outputs the predicted probability, representing the probability that a sample belongs to the positive class. A threshold (e.g., 0.5) is used to determine the class: a value greater than or equal to 0.5 indicates a pathogenic start codon loss mutation, while a value less than 0.5 indicates a benign start codon loss mutation. (6) In equation (6), for Mapping function.
[0023] Step 6: Construct the loss function of the classification method using equation (7). And use equation (8) to construct the update formula for the network parameters: (7) (8) In equation (8), This represents the current learning rate. This represents the set of parameters of the network. This indicates assignment. This represents the gradient.
[0024] Step 7: Use the Adam optimization algorithm to optimize the start codon harmfulness loss mutation prediction and classification network, and calculate the optimal solution. The network parameters are updated until convergence, thus obtaining the optimal prediction and classification model, which is used to predict the harmful loss mutation of the start codon. This solves the problems of insufficient feature utilization, weak cross-level interactive modeling ability, and insufficient prediction accuracy in the existing technology.
[0025] The dataset for this experiment consists of two parts: a training set and a test set. The training set contains 2826 samples, and the test set contains 1557 samples. The dataset used in this experiment is shown in Table 1.
[0026] Table 1. Dataset Composition
[0027] The validation criteria used include recall, specificity (SPE), precision (PR), accuracy (ACC), Matthews correlation coefficient (MCC), and F1 score (F1), which are calculated as shown in equations (9) to (14): (9) (10) (11) (12) (13) (14) In equations (9)-(14), TP (True positive) represents the number of true positives, i.e., the number of true harmful start codon loss mutations correctly predicted as harmful start codon loss mutations; TN (True negative) represents the number of true negatives, i.e., the number of true benign start codon loss mutations correctly predicted as benign start codon loss mutations; FP (False positive) is the number of false positives, i.e., the number of mutations that were originally harmful start codon loss mutations but were predicted as benign start codon loss mutations; and FN (False negative) is the number of false negatives, i.e., the number of mutations that were originally benign start codon loss mutations but were predicted as harmful start codon loss mutations. In addition, AUC and AUPR were used in this experiment to measure the overall performance of the model. Generally, the six indicators given in the above formulas are affected by thresholds; samples greater than or equal to the threshold are predicted as positive samples, while samples less than the threshold are considered negative samples. The default value for the threshold is 0.5, but it can be manually adjusted. AUC and AUPR are not affected by thresholds and range from 0 to 1. The closer they are to 1, the better the overall performance of the model. Therefore, they are often considered to be more important evaluation metrics.
[0028] To verify the superiority of the model, several excellent tools were selected for comparison, including: StartPred, Evo2 (Brixi G, Durrant MG, Ku J, et al. Genome modeling and design across all domains of life with Evo 2[J]. BioRxiv, 2025: 2025.02. 18.638918.), PoStaL, CADD, DANN (Quang D, Chen YF, Xie X H. DANN: a deep learning approach for annotating the pathogenicity of genetic variants[J]. Bioinformatics, 2015, 31(5): 761-763.), MutationTaster2 (Schwarz JM, Cooper DN, Schuelke M, et al. MutationTaster2: mutation prediction for the deep-sequencing age[J]. NatureMethods, 2014, 11(4): 361-362.), PROVEAN (Choi Y, Chan A P. PROVEAN web server: a tool to predict the functional effect of amino acid substitutions andindels[J]. Bioinformatics, 2015, 31(16): 2745-2747.), CAPICE, GPN-MSA, PhD-SNPg (Capriotti E, Fariselli P. PhD-SNPg: updating a webserver and lightweighttool for scoring nucleotide variants[J]. Nucleic Acids Research, 2023, 51(W1): W451-W458.), FATHMM-XF (Rogers MF, Shihab HA, Mort M, et al.FATHMM-XF: accurate prediction of pathogenic point mutations via extended features[J].Bioinformatics, 2018, 34(3): 511-513.), MetaSVM (Dong CL, Wei P, Jian XQ, et al. Comparison and integration of deleteriousness prediction methods for nonsynonymous SNVs in whole exome sequencing studies[J]. Human Molecular Genetics, 2015, 24(8): 2125-2137.), and MetaLR, among which StartPred and PoStaL are specific tools, while the remaining tools are broad-spectrum tools. Table 2 shows a complete comparison of the performance of MSCPred with other tools under eight evaluation metrics (RECALL, SPE, PRE, F1, MCC, ACC, AUC, and AUPR). Specifically, the following results are presented to facilitate the comparison and analysis of differences between the various tools. Because some methods have missing values on the test set, the Null column indicates the number of missing mutations for which each comparison method does not return a prediction result. The results are shown in Table 2. Table 2 Comparison results of MSCPred and 13 mutation harmfulness prediction methods based on independent test sets.
[0029] As shown in Table 2, except for the MSCPred, GPN-MSA, CADD, Evo2, and StartPred models, the other methods all suffer from varying degrees of data gaps, failing to cover all test samples. This lack of predictive coverage significantly limits the applicability of these methods in human genome start codon loss mutation analysis. Methods like MSCPred, by directly using genomic sequences as input and automatically extracting features using a deep learning framework, effectively circumvent the aforementioned limitations of insufficient predictive coverage. Furthermore, MSCPred demonstrates superior performance across multiple evaluation metrics, with AUC and AUPR of 0.922 and 0.958, respectively, outperforming the other 13 mutation pathogenicity prediction methods. This indicates that MSCPred significantly improves its predictive ability for harmful start codon loss mutations by effectively fusing cross-modal features, making it a powerful and reliable prediction tool in this field.
Claims
1. A method for predicting harmful start codon loss mutations based on cross-modal feature fusion, characterized in that, The procedure is as follows: Step 1: Obtain the dataset of mutation samples with missing start codons and perform preprocessing to obtain the preprocessed mutation sample dataset. },in, Represents the first after preprocessing A start codon loss mutation sample, The number of mutation samples representing the loss of the start codon is given by: The true category label is ,and ;when When, it means This is a benign mutation sample; when When, it means This is a sample with a harmful mutation. Step 2: Construct a multi-channel feature set containing DNA sequence features, RNA structural features, protein functional features, and epigenetic modification features; Step 3: Use M different sizes of convolution kernels to respectively... Multi-channel feature input set The first in Channel characteristics Feature extraction is performed to obtain the first... The M local features of the channel feature are concatenated to form the first... intermediate feature vectors ;right After processing with the Mish activation function, and then through batch normalization and max pooling layers, the th... The deep derived feature vector of each channel ;in, The output dimension represents the feature of a single channel; Step 4: Utilize bidirectional attention fusion and concatenation methods to process the deep-derived feature vectors. By merging, we obtain the first... fusion feature vectors ; Step 5: Utilize three fully connected layers to... Processing is performed to obtain Predicted probability ; Step 6: Construct the loss function of the classification method using equation (6). : (6) Step 7: Use the Adam optimization algorithm to optimize the start codon harmfulness loss mutation prediction and classification network, and calculate the optimal solution. The network parameters are updated until convergence is achieved, thus obtaining a prediction classification model with optimal parameters, which is used to predict harmful loss mutations of the start codon.
2. The method for predicting harmful start codon loss mutations based on cross-modal feature fusion according to claim 1, characterized in that, Step 2 is performed as follows: Step 2.1: Extract using a large model. DNA short-range sequence characteristics DNA long-distance sequence characteristics RNA structural features Related to protein functional characteristics and will Any feature of is denoted as the first. Channel characteristics ,in, This represents the first segment of the large model. The length of each channel sequence The first output of the large model represents the... Each channel is embedded in a dimension; ; Step 2.2, Construction Epigenetic modification characteristics ,in, Dimensions representing epigenetic modification characteristics; Step 2.3, Construction Multi-channel feature input set .
3. The method for predicting harmful start codon loss mutations based on cross-modal feature fusion according to claim 2, characterized in that, Step 4 is performed as follows: Step 4.1: Use equations (1)-(3) to derive feature vectors from the RNA structure depth. Deeply derived feature vectors of protein function Perform bidirectional attention fusion to obtain the joint features after fusion of the i-th BAN. : (1) (2) (3) In equations (1), (2), and (3), and This represents the two bilinear weight matrices to be learned. express right Attention weights express right Reverse attention weights; Indicates transpose; Step 4.2, Deeply derived feature vectors from short-range DNA sequences DNA long-range sequence deep-derived feature vector , After splicing, we get the first... Multimodal feature vectors ; Step 4.3, using equation (4) to... Processing is performed to obtain the first... The fusion feature vector of start codon loss mutation samples : (4) In equation (4), This indicates a batch normalization operation. This represents the activation function. and These represent the weights and biases of the fully connected layer, respectively.
4. The method for predicting harmful start codon loss mutations based on cross-modal feature fusion according to claim 3, characterized in that, Step 5 involves using equation (5) to... Process the data to obtain the output of the i-th hidden layer. And after processing by the Sigmoid mapping function, we obtain Predicted probability : + (5) In equation (5), , , This represents the weight matrix of each layer in a three-layer fully connected layer. , , This represents the bias terms for each of the three fully connected layers.
5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports a processor in executing the harmful start codon loss mutation prediction method according to any one of claims 1-4, the processor being configured to execute the program stored in the memory.
6. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is run by the processor, it performs the steps of the method for predicting harmful start codon loss mutations as described in any one of claims 1-4.