An analysis method for predicting pathogenicity of missense mutations

CN117393042BActive Publication Date: 2026-09-08NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311339424.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-16
Publication Date
2026-09-08
Estimated Expiration
2043-10-16

AI Technical Summary

Technical Problem

目前对这两种方法的优劣和差异性的研究还相对有限

Benefits of technology

[0064](1)通过多尺度残差神经网络捕获不同尺度范围内的信息,从而增强对错义突变的表示能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117393042B_ABST
    Figure CN117393042B_ABST
Patent Text Reader

Abstract

The application discloses an analysis method for predicting pathogenicity of missense mutations, comprising obtaining missense mutations of a gene to be tested, a protein sequence and a missense mutation protein sequence mapped with the missense mutations; extracting an amino acid feature vector of a mutation site of the missense mutation protein sequence, and capturing multi-scale features of the amino acid feature vector by using a multi-scale residual neural network; obtaining an amino acid feature matrix by using ESM-1b and ProtT5-XL-U50 respectively, and preprocessing each amino acid feature matrix; capturing mapping features of the preprocessed amino acid feature matrix by using a multi-head attention network multi-path, and processing the mapping features captured by the multi-head attention network by using multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix; combining the multi-scale features and amino acid mapping features of each amino acid feature matrix to obtain amino acid features of the mutation site, and determining whether the missense mutations are pathogenic mutations or benign mutations by using the amino acid features. The missense mutations can be predicted to be pathogenic mutations or benign mutations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an analytical method for predicting the pathogenicity of missense mutations, belonging to the field of genomic variation prediction technology. Background Technology

[0002] Understanding the pathogenicity of missense mutations is crucial for revealing genetic diseases, gene function, and individual differences.

[0003] First, establishing a comprehensive and non-redundant large-scale dataset is crucial for understanding the pathogenicity of missense mutations. Such a dataset should include information related to known mutational effects and requires in-depth discussion of data quality, integration methods, and potential biases. The current dataset construction work needs further refinement to ensure the accuracy and reliability of the data.

[0004] Secondly, developing computational models and analyzing feature importance, interactions, and their contributions to the model are crucial for understanding the pathogenicity of missense mutations. Currently, our understanding of the importance and interactions of features in prediction models needs further improvement. We need to further explore and optimize computational models of mutation effects to improve their accuracy and interpretability.

[0005] Finally, comparing the differences in classification ability between individual outputs and embedding representations based on large-scale protein language models is crucial for understanding the pathogenicity of missense mutations. Currently, research on the advantages, disadvantages, and differences between these two approaches is relatively limited. We need to conduct more in-depth comparative studies to assess their differences in classification ability and determine which method is more suitable for predicting mutation effects.

[0006] In summary, existing methods for predicting the pathogenicity of missense mutations need further improvement in dataset construction, optimization of computational models, and in-depth research into the differences between individual outputs and embedding representations based on large-scale protein language models.

[0007] Therefore, this application proposes an analytical method for predicting the pathogenicity of missense mutations. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide an analytical method for predicting the pathogenicity of missense mutations, which can predict whether a missense mutation is a pathogenic mutation or a benign mutation.

[0009] To achieve the above objectives, the present invention is implemented using the following technical solution:

[0010] This invention introduces an analytical method for predicting the pathogenicity of missense mutations, comprising the following steps:

[0011] Obtain the missense mutations and protein sequences of the gene to be tested;

[0012] Missense mutations are mapped onto the protein sequence of the gene to be tested, resulting in a missense mutant protein sequence mapped with the missense mutation.

[0013] The Ensembl mutation annotator was used to extract the amino acid feature vectors of the mutation sites in missense mutant protein sequences, and a multi-scale residual neural network was used to capture the multi-scale features of the amino acid feature vectors.

[0014] ESM-1b and ProtT5-XL-U50 were used to insert adjacent residues into each mutation site of the missense mutant protein sequence to obtain amino acid feature matrix 1b and amino acid feature matrix U50, respectively, and each amino acid feature matrix was preprocessed.

[0015] After capturing the mapping features of the preprocessed amino acid feature matrix using a multi-head attention network, the mapping features captured by the multi-head attention network are processed by multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix.

[0016] The amino acid features of the mutation site are obtained by concatenating and merging the multi-scale features with the amino acid mapping features of each amino acid feature matrix, and the missense mutation is determined to be a pathogenic mutation or a benign mutation by using the amino acid features.

[0017] Furthermore, the multi-scale residual neural network includes multiple cascaded multi-scale residual neural network modules;

[0018] Each multi-scale residual neural network module includes an input layer, a convolutional unit, and a residual connection layer arranged sequentially.

[0019] The convolutional unit comprises multiple parallel convolutional layers, each with a different convolutional kernel.

[0020] Furthermore, the missense mutant protein sequence includes wild-type protein sequences and mutant protein sequences;

[0021] The amino acid feature matrix 1b includes wild-type amino acid 1b embedding features and mutant amino acid 1b embedding features;

[0022] The amino acid feature matrix U50 includes wild-type amino acid U50 embedding features and mutant amino acid U50 embedding features;

[0023] The step of embedding adjacent residues at each mutation site of the missense mutant protein sequence using ESM-1b and ProtT5-XL-U50 respectively to obtain amino acid feature matrix 1b and amino acid feature matrix U50 includes:

[0024] Based on ESM-1b, wild-type protein sequences and mutant protein sequences were used to obtain wild-type amino acid 1b embedding features and mutant amino acid 1b embedding features, respectively.

[0025] Based on ProtT5-XL-U50, wild-type and mutant protein sequences were used to obtain U50 embedding features of amino acids, respectively.

[0026] Furthermore, the preprocessed amino acid feature matrix includes:

[0027] Let the absolute value of the difference between the wild-type amino acid 1b embedding feature and the mutant amino acid 1b embedding feature be the difference between the wild-type and mutant 1b embedding features.

[0028] Furthermore, the preprocessed amino acid feature matrix includes:

[0029] Let the absolute value of the difference between the wild-type amino acid U50 embedding characteristics and the mutant amino acid U50 embedding characteristics be the difference between the wild-type and mutant U50 embedding characteristics.

[0030] Furthermore, after capturing the mapping features of the preprocessed amino acid feature matrix using a multi-head attention network, the mapping features captured by the multi-head attention network are processed through multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix, including:

[0031] The preprocessed amino acid feature matrix 1b and the preprocessed amino acid feature matrix U50 are respectively input into the input layer of different multi-head attention networks.

[0032] Furthermore, the multi-head attention network includes multiple attention units connected in parallel;

[0033] After capturing the mapping features of the preprocessed amino acid feature matrix using a multi-head attention network, the mapping features captured by the multi-head attention network are processed through multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix, including:

[0034] By using each attention unit to weight the preprocessed amino acid feature matrix, multiple mapping features of the amino acid feature matrix are obtained.

[0035] By fusing the features of each mapping, multi-head fused features are obtained;

[0036] The multi-head fusion features are residually connected to the input amino acid feature matrix and then subjected to layer normalization to obtain the amino acid mapping features of the input amino acid feature matrix.

[0037] Furthermore, the step of using each attention unit to perform weighted processing on the preprocessed amino acid feature matrix to obtain multiple mapping features of the amino acid feature matrix includes:

[0038] When the preprocessed amino acid feature matrix 1b is input into a multi-head attention network:

[0039]

[0040]

[0041] In the formula, ΔF Eb The dimension representing the difference in b1 embedding features between wild-type and mutant types. Let be the self-attention score of the i-th attention unit in a multi-head attention network. Let be the weight matrix used for querying in the i-th attention unit of a multi-head attention network. Let be the weight matrix of the keys of the i-th attention unit in a multi-head attention network. Let be the dimension of the key in the i-th attention unit of a multi-head attention network, and let softmax() be the softmax activation function. Let be the mapping feature of the i-th attention unit in a multi-head attention network. Let T be the weight matrix of the i-th attention unit, where the superscript T is the transpose.

[0042] When the preprocessed amino acid feature matrix U50 is input into another multi-head attention network:

[0043]

[0044]

[0045] In the formula, ΔF T5 For the dimensions of the differences in embedding features between wild-type and mutant U50, Let be the self-attention score of the i-th attention unit in another multi-head attention network. Let be the weight matrix used for querying in the i-th attention unit of another multi-head attention network. Let be the weight matrix of the keys of the i-th attention unit in another multi-head attention network. Let be the dimension of the key in the i-th attention unit of another multi-head attention network. Let be the mapping feature of the i-th attention unit in another multi-head attention network.

[0046] Furthermore, the step of performing residual connection and layer normalization processing on the multi-head fusion features and the input amino acid feature matrix to obtain the amino acid mapping features of the input amino acid feature matrix includes:

[0047] The amino acid mapping features of amino acid feature matrix b1 include the following formula:

[0048]

[0049] In the formula: Y Eb For the amino acid feature matrix b1, ΔF represents the amino acid mapping feature. Eb Differences in 1b embedding features between wild-type and mutant types, The multi-head fusion feature of the amino acid feature matrix b1 is defined by LayerNorm(), which is the layer normalization function.

[0050] The amino acid mapping features of the amino acid feature matrix U50 include the following formula:

[0051]

[0052] In the formula: Y T5 For the amino acid feature matrix U50, ΔF represents the amino acid mapping features. T5 Differences in U50 embedding characteristics between wild-type and mutant types This represents the multi-head fusion feature of the amino acid feature matrix U50.

[0053] Furthermore, the method of determining whether a missense mutation is pathogenic or benign based on amino acid characteristics includes:

[0054] By using a tandem feedforward network and layer normalization layers to process amino acid features, the pathogenicity of missense mutations can be predicted.

[0055] When the pathogenicity rate of a missense mutation is greater than a preset threshold, the prediction result is a pathogenic mutation.

[0056] When the pathogenicity rate of a missense mutation is less than or equal to a preset threshold, the prediction result is a benign mutation.

[0057] The feedforward network consists of two cascaded fully connected layers.

[0058] Furthermore, the feedforward network includes a cascaded first fully connected layer and a second fully connected layer;

[0059] The feedforward network includes the following formula:

[0060] X fc1 =ReLU(X) concat ·W fc1 +b fc1 )

[0061] X fc2 =ReLU(X) fc1 ·W fc2 +b fc2 )

[0062] In the formula: X fc1 For the prediction of the first fully connected layer, X concat Characterized by amino acids, X fc2 For the prediction of the second fully connected layer, ReLU() is the ReLU activation function, W fc1 Here is the weight matrix of the first fully connected layer, and W. fc2 Let b be the weight matrix of the second fully connected layer. fc1 For the bias of the first fully connected layer, b fc2 This is the bias for the second fully connected layer.

[0063] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0064] (1) By capturing information within different scale ranges through multi-scale residual neural networks, the ability to represent missense mutations is enhanced.

[0065] (2) By using two different models, ESM-1b and ProtT5-XL-U50, amino acid feature matrix 1b and amino acid feature matrix U50 are obtained to capture the structural and functional information of missense mutations and the complex relationships in missense mutation data, which can potentially enhance the analytical ability to predict the pathogenicity of missense mutations.

[0066] (3) By calculating the differences between wild-type amino acid embedding features of wild-type proteins and mutant amino acid embedding features of mutant proteins, more potential differential properties of missense mutations can be obtained.

[0067] (4) Each attention unit in a multi-head attention network has a weight matrix with different parameters. The different parameters of the weight matrix enable the multi-head attention network to capture different feature patterns and associations more flexibly. Furthermore, the parallel processing of different mapping features by each attention unit can improve the performance of the multi-head attention network in capturing missense mutation features. The multi-head attention network obtains amino acid mapping features by fusing and residual connecting the mapping features captured by each attention unit. This makes the amino acid mapping features contain the correlation and importance of mutation features at different positions, thereby improving the expressive power of the multi-head attention network.

[0068] (5) By fusing feature information from different network paths, the complex relationships between different features can be learned using feedforward networks and layer normalization layers, and the pathogenicity of mutations can be predicted, which can improve the accuracy of the prediction results. Attached Figure Description

[0069] Figure 1 The diagram shown is an embodiment of an analytical method for predicting the pathogenicity of missense mutations according to the present invention.

[0070] Figure 2The diagram shown is an embodiment of an analytical method for predicting the pathogenicity of missense mutations according to the present invention.

[0071] Figure 3 The diagram shown is an embodiment of an analytical method for predicting the pathogenicity of missense mutations according to the present invention.

[0072] Figure 4 The diagram shown is an embodiment of an analytical method for predicting the pathogenicity of missense mutations according to the present invention.

[0073] Figure 5 The diagram shown is an embodiment of an analytical method for predicting the pathogenicity of missense mutations according to the present invention. Detailed Implementation

[0074] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0075] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0076] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0077] Example 1

[0078] This embodiment introduces an analytical method for predicting the pathogenicity of missense mutations.

[0079] The analytical method for predicting the pathogenicity of missense mutations in this embodiment includes the following steps, referring to... Figure 5 :

[0080] S1 obtains the missense mutation and protein sequence of the gene to be tested;

[0081] S2 maps missense mutations onto the protein sequence of the gene to be tested, obtaining the missense mutant protein sequence mapped with the missense mutation;

[0082] S3 uses the Ensembl mutation annotator to extract the amino acid feature vectors of the mutation sites in missense mutant protein sequences, and uses a multi-scale residual neural network to capture the multi-scale features of the amino acid feature vectors.

[0083] S4 uses ESM-1b and ProtT5-XL-U50 to insert adjacent residues at each mutation site of the missense mutant protein sequence to obtain amino acid feature matrix 1b and amino acid feature matrix U50, and preprocesses each amino acid feature matrix.

[0084] S5 uses a multi-head attention network to capture the mapping features of the preprocessed amino acid feature matrix through multiple paths, and then processes the mapping features captured by the multi-head attention network through multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix.

[0085] S6 combines multi-scale features with amino acid mapping features of each amino acid feature matrix to obtain amino acid features of mutation sites, and uses amino acid features to determine whether missense mutations are pathogenic or benign mutations.

[0086] This invention captures information across different scale ranges using a multi-scale residual neural network, thereby enhancing the ability to represent missense mutations.

[0087] This invention uses two different models, ESM-1b and ProtT5-XL-U50, to obtain amino acid feature matrix 1b and amino acid feature matrix U50, in order to capture the structural and functional information of missense mutations and the complex relationships in missense mutation data, which can potentially enhance the analytical ability to predict the pathogenicity of missense mutations.

[0088] The multi-head attention network of this invention obtains amino acid mapping features by fusing and residual connecting various mapping features, so that the amino acid mapping features contain the correlation and importance of mutation features at different positions, thereby improving the expressive power of the multi-head attention network.

[0089] This invention improves the accuracy of prediction results by fusing feature information from different network paths, using feedforward networks and layer normalization layers to learn the complex relationships between different features, and then predicting the pathogenicity of mutations.

[0090] Example 2

[0091] Based on Example 1, this example details an analytical method for predicting the pathogenicity of missense mutations.

[0092] The analytical method for predicting the pathogenicity of missense mutations in this embodiment includes the following steps:

[0093] S1 obtains the missense mutation and protein sequence of the gene to be tested.

[0094] S2 maps missense mutations onto the protein sequence of the gene to be tested, obtaining the missense mutant protein sequence mapped with the missense mutation.

[0095] S3 uses the Ensembl mutation annotator to extract amino acid feature vectors of mutation sites in missense mutant protein sequences and uses a multi-scale residual neural network to capture the multi-scale features of the amino acid feature vectors.

[0096] In application, the multi-scale residual neural network in step S3 includes multiple cascaded multi-scale residual neural network modules. Each multi-scale residual neural network module includes an input layer, a convolutional unit, and a residual connection layer arranged sequentially. Among them, the convolutional unit includes multiple parallel convolutional layers, and each convolutional layer has a different convolutional kernel, using convolutional kernels of different sizes to capture features at different scales.

[0097] The multi-scale residual neural network in this embodiment includes three cascaded multi-scale residual neural network modules. Each convolutional unit includes three parallel convolutional layers with convolutional kernels of 3, 5, and 7 respectively, which are used to capture information within different scale ranges, thereby enhancing the multi-scale residual neural network's ability to represent missense mutations.

[0098] In practical applications, the amino acid feature vector is used as the input of the multi-scale residual neural network. The residual connection layer uses the connection function to merge the outputs of each convolutional layer as the output of the current multi-scale residual neural network module and as the input of the next multi-scale residual neural network module. The output of the last multi-scale residual neural network module is used as the output of the multi-scale residual neural network.

[0099] S4 uses ESM-1b and ProtT5-XL-U50 to insert adjacent residues at each mutation site of the missense mutant protein sequence to obtain amino acid feature matrix 1b and amino acid feature matrix U50, and preprocesses each amino acid feature matrix.

[0100] Among them, the missense mutant protein sequence includes wild-type protein sequence and mutant protein sequence; the amino acid feature matrix 1b includes wild-type amino acid 1b embedding features of wild-type protein and mutant amino acid 1b embedding features of mutant protein; the protein feature matrix U50 includes wild-type amino acid U50 embedding features of wild-type protein and mutant amino acid U50 embedding features of mutant protein.

[0101] By using two different models, ESM-1b and ProtT5-XL-U50, to obtain amino acid feature matrices 1b and U50, the structural and functional information of missense mutations and the complex relationships in missense mutation data can be captured, potentially enhancing the analytical ability to predict the pathogenicity of missense mutations.

[0102] In application, ESM-1b and ProtT5-XL-U50 are used to insert adjacent residues at each mutation site of the missense mutant protein sequence to obtain amino acid feature matrix 1b and amino acid feature matrix U50, respectively, including:

[0103] S41, based on ESM-1b, uses wild-type protein sequences and mutant protein sequences respectively to obtain the wild-type amino acid 1b embedding features of wild-type proteins and the mutant amino acid 1b embedding features of mutant proteins.

[0104] S42, based on ProtT5-XL-U50, uses wild-type protein sequences and mutant protein sequences respectively to obtain the U50 embedding features of wild-type amino acids in wild-type proteins and the U50 embedding features of mutant amino acids in mutant proteins.

[0105] This embodiment of ESM-1b references Rives, A.; Meier, J.; Sercu, T.; Goyal, S.; Lin, Z.; Liu, J.; Guo, D.; Ott, M.; Zitnick, CL; Ma, J., Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences. Proc. Natl. Acad. Sci. USA. 2021, 118, e2016239118.

[0106] ProtT5-XL-U500 in this embodiment refers to Elnaggar, A.; Heinzinger, M.; Dallago, C.; Rehawi, G.; Wang, Y.; Jones, L.; Gibbs, T.; Feher, T.; Angerer, C.; Steinegger, M., Prottrans: Toward Understanding the Language of Life through Self-supervisedLearning.ITPAM.2021,44,7112-7127.

[0107] When applying this method, the preprocessing of each protein feature matrix includes:

[0108] S43 Let the absolute value of the difference between the wild-type amino acid 1b embedding feature of the wild-type protein and the mutant amino acid 1b embedding feature of the mutant protein be the difference between the wild-type and mutant 1b embedding features.

[0109] S44 sets the absolute value of the difference between the wild-type amino acid U50 embedding characteristics of the wild-type protein and the mutant amino acid U50 embedding characteristics of the mutant protein as the difference between the wild-type and mutant U50 embedding characteristics.

[0110] By calculating the differences between wild-type and mutant amino acid embedding features, more potential differential properties about missense mutations can be obtained.

[0111] S5 uses a multi-head attention network to capture the mapping features of the preprocessed amino acid feature matrix via multiple paths. Then, it processes the mapping features captured by the multi-head attention network through multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix.

[0112] In application, the preprocessed amino acid feature matrix 1b and the preprocessed amino acid feature matrix U50 are respectively input into the input layer of different multi-head attention networks.

[0113] In this embodiment, the differences in embedding features between wild-type and mutant 1b and the differences in embedding features between wild-type and mutant U50 are input into two different multi-head attention networks to better capture missense mutation information.

[0114] The multi-head attention network in this embodiment includes multiple attention units connected in parallel.

[0115] In practical applications, step S5 includes:

[0116] S51 uses each attention unit to weight the preprocessed amino acid feature matrix to obtain multiple mapping features of the amino acid feature matrix.

[0117] (1) When the preprocessed amino acid feature matrix 1b is input into a multi-head attention network:

[0118]

[0119]

[0120] In the formula, ΔF Eb The dimension representing the difference in b1 embedding features between wild-type and mutant types. Let be the self-attention score of the i-th attention unit in a multi-head attention network. Let be the weight matrix used for querying in the i-th attention unit of a multi-head attention network. Let be the weight matrix of the keys of the i-th attention unit in a multi-head attention network. Let be the dimension of the key in the i-th attention unit of a multi-head attention network, and let softmax() be the softmax activation function. Let be the mapping feature of the i-th attention unit in a multi-head attention network. Let T be the weight matrix of the i-th attention unit, where the superscript T is the transpose.

[0121] (2) When the preprocessed amino acid feature matrix U50 is input into another multi-head attention network:

[0122]

[0123]

[0124] In the formula, ΔF T5 For the dimensions of the differences in embedding features between wild-type and mutant U50, Let be the self-attention score of the i-th attention unit in another multi-head attention network. Let be the weight matrix used for querying in the i-th attention unit of another multi-head attention network. Let be the weight matrix of the keys of the i-th attention unit in another multi-head attention network. Let be the dimension of the key in the i-th attention unit of another multi-head attention network. Let be the mapping feature of the i-th attention unit in another multi-head attention network.

[0125] S52 fuses various mapping features to obtain multi-head fused features.

[0126] (1) The multi-head fusion features of amino acid feature matrix b1 include the following formula:

[0127]

[0128] In the formula: b1 represents the multi-head fusion feature of the amino acid feature matrix, concat() is the connection function, and H is the number of attention units.

[0129] (2) The multi-head fusion features of the amino acid feature matrix U50 include the following formula:

[0130]

[0131] In the formula: This represents the multi-head fusion feature of the amino acid feature matrix U50.

[0132] S53 performs residual connection and layer normalization on the multi-head fusion features and the input amino acid feature matrix to obtain the amino acid mapping features of the input amino acid feature matrix.

[0133] (1) The amino acid mapping features of amino acid feature matrix b1 include the following formula:

[0134]

[0135] In the formula: Y Eb For the amino acid feature matrix b1, ΔF represents the amino acid mapping feature. Eb LayerNorm() is the layer normalization function, representing the difference in 1b embedding features between wild-type and mutant types.

[0136] (2) The amino acid mapping features of the amino acid feature matrix U50 include the following formula:

[0137]

[0138] In the formula: Y T5 For the amino acid feature matrix U50, ΔF represents the amino acid mapping features. T5 Differences in U50 embedding characteristics between wild-type and mutant types.

[0139] In this embodiment, each attention unit of the multi-head attention network has a weight matrix with different parameters. The different parameters of the weight matrix enable the multi-head attention network to capture different feature patterns and associations more flexibly. Furthermore, each attention unit processes different mapping features in parallel, which can improve the performance of the multi-head attention network in capturing missense mutation features.

[0140] The multi-head attention network in this embodiment obtains amino acid mapping features by fusing and residually connecting the mapping features captured by each attention unit. This makes the amino acid mapping features contain the correlation and importance of mutation features at different positions, thereby improving the expressive power of the multi-head attention network.

[0141] S6 combines multi-scale features with amino acid mapping features of each amino acid feature matrix to obtain amino acid features, and uses amino acid features to determine whether missense mutations are pathogenic or benign mutations.

[0142] Step S6 includes:

[0143] S61 concatenates and merges the multi-scale features with the amino acid mapping features of each amino acid feature matrix, including the following formula:

[0144]

[0145]

[0146]

[0147]

[0148] In the formula, Y Eb Y represents the amino acid mapping features of amino acid feature matrix b1. T5 X represents the amino acid mapping features of the amino acid feature matrix U50. ResNet For multi-scale features, d TE1 Let d be the dimension of the amino acid mapping features of the amino acid feature matrix b1. TE2 Let d be the dimension of the amino acid mapping features of the amino acid feature matrix U50. ResNet X represents the dimension of the multi-scale feature, where N is the sequence length of the missense mutant protein sequence, and X is the multi-scale feature dimension. concat It is characterized by amino acids.

[0149] S62 uses a tandem feedforward network and layer normalization layers to process amino acid features and predict the pathogenicity of missense mutations.

[0150] In application, after fusing feature information from different network paths, the pathogenicity of mutations is predicted through a feedforward network and a layer normalization layer.

[0151] S621 uses the fully connected layers of the feedforward network to learn the complex relationships between different features, and then uses the sigmoid activation function to map the output of the feedforward network to the probability value of binary classification, that is, the pathogenicity rate of missense mutations.

[0152] S622 uses a preset threshold to mark prediction results where the pathogenicity rate of a missense mutation is greater than the preset threshold as 1, meaning the prediction result is a pathogenic mutation; and marks prediction results where the pathogenicity rate of a missense mutation is less than or equal to the preset threshold as 0, meaning the prediction result is a benign mutation.

[0153] The feedforward network in this embodiment includes two cascaded fully connected layers.

[0154] In application, the feedforward network consists of a cascaded first fully connected layer and a second fully connected layer.

[0155] The feedforward network in this embodiment includes the following formula:

[0156] X fc1 =ReLU(X) concat ·W fc1 +b fc1 ),

[0157] X fc2 =ReLU(X) fc1 ·W fc2 +b fc2 ),

[0158] In the formula: X fc1 For the prediction of the first fully connected layer, X fc2 For the prediction of the second fully connected layer, ReLU() is the ReLU activation function, W fc1 Here is the weight matrix of the first fully connected layer, and W. fc2 Let b be the weight matrix of the second fully connected layer. fc1 For the bias of the first fully connected layer, b fc2 This is the bias for the second fully connected layer.

[0159] In practical applications, the output of the last layer uses the sigmoid function for binary classification prediction.

[0160] This embodiment improves the accuracy of prediction results by fusing feature information from different network paths, using feedforward networks and layer normalization layers to learn the complex relationships between different features, and predicting the pathogenicity of mutations.

[0161] Example 3

[0162] Based on Example 1 or 2, this example details a method for predicting the pathogenicity of missense mutations.

[0163] The missense mutation pathogenicity prediction method in this embodiment includes: a method for constructing a large-scale non-redundant missense mutation benchmark dataset, a method for feature representation of mutation sites, a data partitioning method, a deep network parallel optimization and prediction method, and references. Figure 1 .

[0164] First, a large-scale non-redundant missense mutation benchmark dataset was established based on the entire Ensembl database. For each mutation in this dataset, features at the variant level, amino acid level, individual output, and genome level were extracted using the Ensembl Variant Effect Predictor (VEP) v104 and the Database for Nonsynonymous SNPs' Functional Predictions (dbNSFP) v4.1a. Wild-type protein sequences were obtained using the Ensembl API.

[0165] Then, the ESM-1b and ProtTrans-T5 embedding feature matrices of the mutation sites and adjacent amino acids were extracted.

[0166] Finally, the pathogenicity of missense mutations is predicted using the above features.

[0167] Combination Figure 2 and Figure 1 The method for constructing a large-scale non-redundant missense mutation benchmark dataset includes the following steps:

[0168] Step 1: Download the entire Ensembl database (6.18GB) from https: / / web.expasy.org / swissvar.html. This database contains mutations from multiple databases, such as ClinVar, dbSNP, gnomAD, 1000Genomes, UniProt, ExAC, COSMIC, etc.

[0169] Step 2: To screen for relevant missense mutations, two filters were applied: for mutation type, records with values ​​of "Missense Mutation" or "Missense" were selected; for clinical significance, records with values ​​of "Benign / Likely Benign" or "Pathogenic / Likely Pathogenic" were selected. After this screening process, 622,270 mutations were obtained.

[0170] Step 3: Then apply five additional filters to further remove unwanted mutations, specifically including:

[0171] (1) Mutations with contradictory interpretations, such as those simultaneously labeled as benign and pathogenic mutations;

[0172] (2) Mutations with uncertain significance, i.e., clinical significance is marked as unknown significance;

[0173] (3) Repeated and inconsistent mutations from different sources, such as NC_000001.11:g.1022225G>A and NC_000001.11:g.1022225T>C, which are inconsistent mutations;

[0174] (4) Mutations that are classified as benign and pathogenic / likely pathogenic, i.e. mutations that are labeled as benign and also as pathogenic / likely pathogenic.

[0175] (5) Mutations that are simultaneously labeled as both likely benign and pathogenic / likely pathogenic, i.e. mutations that are labeled as benign (Likely Benign) and simultaneously labeled as pathogenic (Pathogenic) / likely pathogenic (Likely Pathogenic);

[0176] Through the above five screening processes, a total of 91,072 mutations were obtained.

[0177] Step 4: Using the Ensembl mutation effect predictor, the mutation VCF format file information is mapped to GRch38 genome coordinates. Inconsistent chromosomes / locations (such as the chromosome location of the mutation in the VCF being 1-949422, resulting in a change in the mapped value) and mismatches (such as 1-949422-GA and 1-949422-CT) are excluded, resulting in 77,700 mutations.

[0178] according to Figure 1 The data processing procedure shown indicates that the benchmark dataset contains 37,317 benign mutations in 2,595 proteins and 40,383 pathogenic mutations in 3,294 proteins. However, during processing with Ensembl VEP v104, some mutations lacked annotation. Therefore, these mutations and their corresponding proteins were excluded. After the above procedure, a large-scale, non-redundant missense mutation benchmark dataset was finally obtained.

[0179] Combination Figure 3 and Figure 1 The method for characterizing mutation sites specifically includes the following steps:

[0180] Step 1: Obtain the protein encoding sequence:

[0181] Use Ensembl API functions and ENSP-id information to retrieve FASTA format protein sequences from the Ensembl protein database.

[0182] Then, check if the length of the obtained protein sequence matches the length of the protein provided by Ensembl VEP, and also check if the amino acids at the mutation sites of the obtained protein match the wild-type amino acids at the corresponding sites in the protein provided by Ensembl VEP. Delete any inconsistent mutations.

[0183] Step 2: Feature extraction is performed using the Ensembl mutation annotator. Ensembl VEP v104 is used to extract feature vectors, including variant level features such as the percentage of protein length affected by the mutation (Protein_position), mutation influence (IMPACT), and mutated gene phenotype (GENE_PHENO); amino acid level features, mainly referring to the physicochemical properties of amino acids; individual outputs, i.e., the outputs and annotation information of individual prediction methods such as SIFT and PolyPhen2; and genome level features, such as the frequency of alternative alleles 1000Gp3_AFR_AF in African-descended samples from the 1000Gp3 database.

[0184] For reference, SIFT is found in Ng, PC; Henikoff, S., SIFT: Predicting Amino Acid Changes that Affect Protein Function. Nucleic Acids Res. 2003, 31, 3812-3814.

[0185] PolyPhen2 Reference Adzhubei, IA; Schmidt, S.; Peshkin, L.; Ramensky, VE; Gerasimova, A.; Bork, P.; Kondrashov, AS; Sunyaev, SR, A Method and Server for Predicting Damaging Missense Mutations. Nat. Methods. 2010, 7, 248-249.

[0186] Step 3: For each mutation site, ESM-1b and ProtT5-XL-U50 are embedded in the adjacent residues of the joint mutation site to generate the mutation feature matrix.

[0187] Combination Figure 1 The data partitioning method specifically includes the following steps:

[0188] Step 1: Use the proteins in the benchmark dataset as the sole criterion for splitting the training and testing data, and store them in a set;

[0189] Step 2: Use a data partitioning function to divide the baseline data into two parts, namely training data (80%) and test data (20%), while ensuring that there is no overlap of mutated proteins and mutated samples between the two parts;

[0190] Step 3: Manually check the completed training and test data to ensure that there is no overlap between protein and mutant samples.

[0191] Combination Figure 4 and Figure 1 The aforementioned deep network parallel optimization and prediction method specifically includes the following steps:

[0192] Step 1: Baseline data splitting. The baseline data is split into 80% training set and 20% test set.

[0193] Step 2: Use ESM-1b and ProtT5-XL-U50 to insert adjacent residues into each mutation site of the missense mutant protein sequence to obtain amino acid feature matrix 1b and amino acid feature matrix U50.

[0194] (1) By using ESM-1b, the wild-type protein sequence was used to obtain the wild-type amino acid 1b embedding feature of the wild-type protein.

[0195] The feature dimensions of wild-type amino acid 1b embedding features of wild-type proteins are:

[0196]

[0197] in: The feature dimension of wild-type amino acid 1b embedding features for wild-type proteins. It is 1280. The length of the 1b neighboring amino acid centered at the mutant amino acid site in the wild-type protein sequence.

[0198] (2) By using ESM-1b, the mutant protein sequence is used to obtain the mutant amino acid 1b embedding feature of the mutant protein.

[0199] The feature dimension of the mutant amino acid 1b embedding feature of mutant proteins is:

[0200]

[0201] in: The feature dimension of the 1b embedding feature of mutant amino acids in mutant proteins. It is 1280. The length of the 1b neighboring amino acid centered at the mutant amino acid site in the mutant protein sequence.

[0202] (3) By using ProtT5-XL-U50, the wild-type protein sequence was used to obtain the wild-type amino acid U50 embedding feature of the wild-type protein.

[0203] The feature dimensions of the wild-type amino acid U50 embedding feature of wild-type proteins are:

[0204]

[0205] in, The feature dimension for embedding wild-type amino acid U50 of wild-type protein. It is 1024. The length of the U50 adjacent amino acid centered at the mutant amino acid site in the wild-type protein sequence.

[0206] (4) By using ProtT5-XL-U50, the U50 embedding feature of mutant amino acids of mutant protein is obtained using mutant protein sequence.

[0207] The feature dimensions of the mutant amino acid U50 embedding feature of mutant proteins are:

[0208]

[0209] in, The feature dimension for the U50 embedding feature of mutant amino acids in mutant proteins. It is 1024. This represents the length of the U50 neighboring amino acid centered at the mutant amino acid site in the mutant protein sequence.

[0210] By using two different models, ESM-1b and ProtT5-XL-U50, to obtain amino acid feature matrices 1b and U50, the structural and functional information of missense mutations and the complex relationships in missense mutation data can be captured, potentially enhancing the analytical ability to predict the pathogenicity of missense mutations.

[0211] Step 3: Preprocess the feature matrix of each amino acid

[0212] By calculating the differences between the wild-type amino acid embedding characteristics of wild-type proteins and the mutant amino acid embedding characteristics of mutant proteins, more potential differential properties about missense mutations can be obtained.

[0213] (1) Let the absolute value of the difference between the wild-type amino acid 1b embedding feature of the wild-type protein and the mutant amino acid 1b embedding feature of the mutant protein be the difference between the wild-type and mutant 1b embedding features.

[0214] The dimensions of the differences in 1b embedding features between wild-type and mutant types are:

[0215]

[0216] In the formula: ΔF Eb The dimension of the difference in 1b embedding features between wild-type and mutant types.

[0217] (2) Let the absolute value of the difference between the wild-type amino acid U50 embedding feature of the wild-type protein and the mutant amino acid U50 embedding feature of the mutant protein be the difference between the wild-type and mutant U50 embedding features.

[0218] The dimensions of the differences in U50 embedding features between wild-type and mutant types are:

[0219]

[0220] In the formula: ΔF T5 The dimension of the difference in embedding features between wild-type and mutant U50.

[0221] Step 4: After capturing the mapping features of the preprocessed amino acid feature matrix using a multi-head attention network, the mapping features captured by the multi-head attention network are processed by multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix.

[0222] The differences in embedding features between wild-type and mutant 1b and U50 embedding features between wild-type and mutant are respectively input into two Transformer Encoder multi-head attention networks to better capture missense mutation information.

[0223] Each attention unit in a multi-head attention network (MNB) has a weight matrix with different parameters. These different parameters allow the MNB to more flexibly capture different feature patterns and associations. Furthermore, the parallel processing of different mapping features by each attention unit improves the performance of the MNB in ​​capturing missense mutation features.

[0224] Multi-head attention networks obtain amino acid mapping features by fusing and residually connecting the mapping features captured by each attention unit. This allows the amino acid mapping features to contain the correlation and importance of mutation features at different positions, thereby improving the expressive power of multi-head attention networks.

[0225] Step 5: Capture the multi-scale features of amino acid feature vectors using a multi-scale residual neural network.

[0226] In multi-scale residual neural networks, convolutional kernels of different sizes can be used to capture features at different scales. The following are feature processing formulas for three sets of multi-scale one-dimensional residual neural networks with different neurons, using 3, 5, and 7 convolutional kernels respectively to capture information within different scale ranges, thereby enhancing the ability of multi-scale residual neural networks to represent missense mutations.

[0227] The input amino acid feature vector is Where N is the sequence length of the missense mutant protein sequence, and d is the feature dimension of the amino acid feature vector.

[0228] (1) In the first group of multi-scale one-dimensional residual neural networks, the number of neurons is [128, 128, 512]. Convolution operations are performed using kernels of 3, 5, and 7, and residual connections and layer normalization operations are performed. Then, the outputs are merged using the connection function concat(Output3, Output5, Output7) and used as the input of the next group of multi-scale one-dimensional residual neural networks.

[0229] (2) In the second group of one-dimensional multi-scale residual neural networks, the output of the previous group is used as its input, and after convolution operation using kernels of 3, 5, and 7, as well as residual connection and normalization operation, the output is merged using the connection function and used as the input of the next group of multi-scale one-dimensional residual neural networks, and so on. The neurons of the second and third groups of multi-scale residual neural networks are [256,256,1024] and [512,512,2048], respectively.

[0230] Step 6: Concatenate and merge the multi-scale features with the amino acid mapping features of each amino acid feature matrix to obtain amino acid features, and use the amino acid features to determine whether the missense mutation is a pathogenic mutation or a benign mutation.

[0231] Step 7: SHAP Feature Analysis

[0232] The Shapley value based on game theory can be used to analyze the contribution of each feature to the output of the prediction model in this embodiment, which is based on the consensus evolutionary information from steps 1-6 and the parallel connection of deep networks.

[0233] The SHAP value represents the influence of each feature on the prediction result of a particular mutation. In a prediction model, it is necessary to interpret its predicted output for the input sample x. And the influence of each feature in the input sample x.

[0234] Calculate the SHAP value for each feature:

[0235]

[0236] Where, φ i SHAP value; p is the total number of features, i.e., mutant-level features, amino acid-level features, genome-level features, individual outputs, and the total number of features based on ESM1b and ProtT5 feature embeddings; S is the feature set; x S This represents a sample that contains only features from the feature set S; x S∪{i}Let φ represent a sample containing feature i. The smaller the influence of feature i on the sample, the higher its SHAP value φ. i The smaller.

[0237] Determining the minimum SHAP value at which a feature can be considered to have almost no impact on the model's output is a relative question; there is no fixed threshold that applies to all situations. The magnitude of the SHAP value is influenced by the specific data, the model, and the problem.

[0238] Therefore, SHAP values ​​can quantify the contribution of each feature to the model, helping to understand the model's decision-making process. Furthermore, calculated SHAP values ​​can be used to visualize, interpret, and analyze the predictions of the proposed deep learning model, as well as to understand the model's dependencies on different features.

[0239] Step 8: Train the above prediction model using the training data;

[0240] Step 9: Use the trained model to make predictions on the test data.

[0241] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0242] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0243] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0244] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0245] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. An analytical method for predicting the pathogenicity of missense mutations, characterized in that, Includes the following steps: Obtain the missense mutations and protein sequences of the gene to be tested; Missense mutations are mapped onto the protein sequence of the gene to be tested, resulting in a missense mutant protein sequence mapped with the missense mutation. The Ensembl mutation annotator was used to extract the amino acid feature vectors of the mutation sites in missense mutant protein sequences, and a multi-scale residual neural network was used to capture the multi-scale features of the amino acid feature vectors. ESM-1b and ProtT5-XL-U50 were used to insert adjacent residues into each mutation site of the missense mutant protein sequence to obtain amino acid feature matrix 1b and amino acid feature matrix U50, respectively, and each amino acid feature matrix was preprocessed. After capturing the mapping features of the preprocessed amino acid feature matrix using a multi-head attention network, the mapping features captured by the multi-head attention network are processed through multi-head fusion and residual connections to determine the mapping features of each amino acid feature matrix, including: The multi-head attention network comprises multiple attention units connected in parallel; By using each attention unit to weight the preprocessed amino acid feature matrix, multiple mapping features of the amino acid feature matrix are obtained. By fusing the features of each mapping, multi-head fused features are obtained; The multi-head fusion features are residually connected and layer normalized with the input amino acid feature matrix to obtain the amino acid mapping features of the input amino acid feature matrix. The amino acid features of the mutation site are obtained by concatenating and merging the multi-scale features with the amino acid mapping features of each amino acid feature matrix, and the missense mutation is determined to be a pathogenic mutation or a benign mutation by using the amino acid features. The missense mutant protein sequences include wild-type protein sequences and mutant protein sequences; The amino acid feature matrix 1b includes wild-type amino acid 1b embedding features and mutant amino acid 1b embedding features; The amino acid feature matrix U50 includes wild-type amino acid U50 embedding features and mutant amino acid U50 embedding features; The step of embedding adjacent residues at each mutation site of the missense mutant protein sequence using ESM-1b and ProtT5-XL-U50 respectively to obtain amino acid feature matrix 1b and amino acid feature matrix U50 includes: Based on ESM-1b, wild-type protein sequences and mutant protein sequences were used to obtain wild-type amino acid 1b embedding features and mutant amino acid 1b embedding features, respectively. Based on ProtT5-XL-U50, wild-type and mutant protein sequences were used to obtain U50 embedding features of amino acids, respectively.

2. The analytical method for predicting the pathogenicity of missense mutations according to claim 1, characterized in that, The multi-scale residual neural network includes multiple cascaded multi-scale residual neural network modules; Each multi-scale residual neural network module includes an input layer, a convolutional unit, and a residual connection layer arranged sequentially. The convolutional unit comprises multiple parallel convolutional layers, each with a different convolutional kernel.

3. The analytical method for predicting the pathogenicity of missense mutations according to claim 1, characterized in that, The preprocessed amino acid feature matrix includes: Let the absolute value of the difference between the wild-type amino acid 1b embedding feature and the mutant amino acid 1b embedding feature be the difference between the wild-type and mutant 1b embedding features; Let the absolute value of the difference between the wild-type amino acid U50 embedding characteristics and the mutant amino acid U50 embedding characteristics be the difference between the wild-type and mutant U50 embedding characteristics.

4. The analytical method for predicting the pathogenicity of missense mutations according to claim 1, characterized in that, After capturing the mapping features of the preprocessed amino acid feature matrix using a multi-head attention network, the mapping features captured by the multi-head attention network are processed through multi-head fusion and residual connection to determine the mapping features of each amino acid feature matrix, including: The preprocessed amino acid feature matrix 1b and the preprocessed amino acid feature matrix U50 are respectively input into the input layer of different multi-head attention networks.

5. The analytical method for predicting the pathogenicity of missense mutations according to claim 1, characterized in that, The step of using each attention unit to weight the preprocessed amino acid feature matrix to obtain multiple mapping features of the amino acid feature matrix includes: When the preprocessed amino acid feature matrix 1b is input into a multi-head attention network: ; ; In the formula, The dimension representing the difference in b1 embedding features between wild-type and mutant types. For a multi-head attention network i Self-attention score of each attention unit. For the first in a multi-head attention network i The weight matrix for querying each attention unit. For the first in a multi-head attention network i The weight matrix of the keys in each attention unit. For the first in a multi-head attention network i The dimension of the keys in each attention unit. for Activation function For the first in a multi-head attention network i Mapping features of each attention unit For the first i The weight matrix of each attention unit, with superscript... T For transpose; When the preprocessed amino acid feature matrix U50 is input into another multi-head attention network: ; ; In the formula, For the dimensions of the differences in embedding features between wild-type and mutant U50, For another multi-head attention network, the first i Self-attention score of each attention unit. For another multi-head attention network, the first i The weight matrix for querying each attention unit. For another multi-head attention network, the first i The weight matrix of the keys in each attention unit. For another multi-head attention network, the first i The dimension of the keys in each attention unit. For another multi-head attention network i Mapping features of attention units.

6. The analytical method for predicting the pathogenicity of missense mutations according to claim 1, characterized in that, The step of performing residual connection and layer normalization processing on the multi-head fusion features and the input amino acid feature matrix to obtain the amino acid mapping features of the input amino acid feature matrix includes: The amino acid mapping features of amino acid feature matrix b1 include the following formula: ; In the formula: The amino acid mapping features of amino acid feature matrix b1, Differences in 1b embedding features between wild-type and mutant types, This represents the multi-head fusion feature of the amino acid feature matrix b1. For layer normalization function; The amino acid mapping features of the amino acid feature matrix U50 include the following formula: ; In the formula: The amino acid mapping features of the amino acid feature matrix U50. Differences in U50 embedding characteristics between wild-type and mutant types This represents the multi-head fusion feature of the amino acid feature matrix U50.

7. The analytical method for predicting the pathogenicity of missense mutations according to claim 1, characterized in that, The method of determining whether a missense mutation is pathogenic or benign based on amino acid characteristics includes: By using a tandem feedforward network and layer normalization layers to process amino acid features, the pathogenicity of missense mutations can be predicted. When the pathogenicity rate of a missense mutation is greater than a preset threshold, the prediction result is a pathogenic mutation. When the pathogenicity rate of a missense mutation is less than or equal to a preset threshold, the prediction result is a benign mutation. The feedforward network consists of two cascaded fully connected layers.

8. The analytical method for predicting the pathogenicity of missense mutations according to claim 7, characterized in that, The feedforward network includes a cascaded first fully connected layer and a second fully connected layer. The feedforward network includes the following formula: ; ; In the formula: For the prediction of the first fully connected layer, Characterized by amino acids, For the prediction of the second fully connected layer, For ReLU activation function, This is the weight matrix of the first fully connected layer. This is the weight matrix of the second fully connected layer. For the bias of the first fully connected layer, This is the bias for the second fully connected layer.