A System and Method for Predicting the Association between Post-Translational Modification of Proteins and Diseases

By constructing a data set of protein post-translation modified sequences and disease labels, using Word2Vec and Transformer structures for feature extraction and prediction, the inadequate understanding of protein function changes and the difficulty of simple machine learning models to deal with complexity is solved, and disease prediction with high accuracy and generalization ability is achieved.

CN119560009BActive Publication Date: 2025-06-24ZHEJIANG UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510096736.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-24
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art ignores the understanding of protein function changes when predicting the association of protein post-translational modifications and relies on simple machine learning models to make it difficult to deal with the high complexity and diversity of protein post-translational modification sites.

Method used

A method that integrates the fields of bioinformatics and medical diagnosis is proposed to predict the occurrence and development of diseases by analyzing bioinformatics data modified by protein translation. The specific steps include constructing a sequence dataset containing protein post-translation modified sequences and their disease tags, using Word2Vec structure for feature extraction, and using the multi-head attention mechanism of the Transformer structure to predict disease associations, combining cross-entropy loss function and L2 regularization optimization model.

Benefits of technology

The mining and prediction of disease information that may be carried by protein translation modification sites is achieved, which improves the accuracy of disease prediction and generalizes the model, and provides a new perspective and method for medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119560009B_ABST
    Figure CN119560009B_ABST
Patent Text Reader

Abstract

The present invention discloses a system and method for predicting the association between protein post-translational modifications and diseases, which relates to the fields of bioinformatics and medical diagnosis. The system includes: a data cleaning module, which is used to obtain sequence data related to protein post-translational modifications and corresponding disease label data, and clean the sequence data; a feature extraction module, which is used to learn the feature embeddings of the sequence data and extract feature vectors containing feature information; an association prediction module, which is used to transform the feature vectors through a multi-head attention mechanism and perform disease association prediction through a Transformer structure; a function definition module, which is used to define a loss function according to the feature complexity and purpose of biological information; a model evaluation module, which is used to input the sequence data into a trained network model and output an evaluation result. According to the technical solution of the present application, the association between protein post-translational modifications and disease development can be predicted, which has high application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and medical diagnosis, and in particular to a system and method for predicting the association between protein post-translational modification and disease. Background Art

[0002] Nowadays, human research on life sciences has entered the "protein era". Protein is the material basis of all life. Its functional expression not only depends on its amino acid sequence, but also is affected by post-translational modification. In the biomedical field, protein post-translational modification is a key mechanism for regulating protein function. It plays a vital role in maintaining normal physiological functions of the human body and preventing diseases. With the occurrence and development of certain cancers, the phosphorylation of some specific proteins will show abnormal states. These abnormalities may cause uncontrolled cell growth and eventually form tumors.

[0003] By detecting the modification status of specific proteins, it is possible to provide a strong basis for the early diagnosis of tumors and other diseases. The construction of relevant databases, the maturity of bioinformatics, and the diversification of natural language models have made this type of method a reality. Traditional research on the prediction of disease functions by post-translational modification of proteins has only focused on the sequence information of post-translational modification sites of proteins, ignoring the understanding of changes in protein function. The physiological or pathological state of proteins largely reflects the pathological or abnormal conditions of the body. Most existing disease prediction methods rely on simple machine learning models that can handle the high complexity and diversity of post-translational modification sites of proteins in maintaining protein function.

[0004] Therefore, based on this problem, the present application proposes a system and method for predicting the association between protein post-translational modification and disease. Summary of the invention

[0005] Technical Purpose

[0006] In order to solve the above problems, the purpose of the present invention is to provide a protein post-translational modification and disease association prediction system and method, which can not only mine the disease information that may be carried by protein post-translational modification sites, but also establish a prediction bridge between protein post-translational modification and disease through bioinformatics tools and machine learning models, while ensuring rationality and timeliness. The system and method have a wide range of application scenarios, provide new perspectives and methods for medical diagnosis, and promote the research on disease diagnosis and pathogenesis.

[0007] Technical Solution

[0008] To achieve the above object, the present invention provides a system and method for predicting the association between protein post-translational modifications and diseases, and proposes an innovative method that combines bioinformatics and the field of medical diagnosis. The aim is to predict the occurrence and development of diseases by analyzing the bioinformatics data of protein post-translational modifications, pre-construct a sequence data set containing protein post-translational modification sequences and their disease labels, extract features from the sequence data through the Word2Vec structure, and process the extracted feature vectors according to the multi-head attention mechanism of the Transformer structure. Combine the cross-entropy loss function and L2 regularization to optimize the prediction model, and use the Adam algorithm to optimize the model parameters, so as to realize the mining and prediction of the possible disease information carried by protein post-translational modification sites, and provide a reference basis for bioinformatics analysis and disease diagnosis.

[0009] In the first aspect, the present invention provides a system for predicting the association between protein post-translational modifications and diseases, including:

[0010] A data cleaning module, which is used to obtain sequence data related to protein post-translational modifications and corresponding disease label data, and clean the sequence data;

[0011] A feature extraction module, which is used to learn the feature embedding of the sequence data and extract feature vectors containing feature information;

[0012] An association prediction module, which is used to transform the feature vectors through the multi-head attention mechanism and perform disease association prediction through the Transformer structure;

[0013] A function definition module, which is used to define a loss function according to the feature complexity and purpose of biological information, and optimize the loss function through the Adam algorithm;

[0014] A model evaluation module, which is used to input the sequence data into the trained network model and output the evaluation results.

[0015] Further, the sequence data is represented as X = , and its data distribution is , the disease label data is represented as Y = , and its data distribution is ;

[0016] In the formula, is the number of sequences of protein post-translational modifications; is the number of relevant disease labels; X is the sequence of protein post-translational modifications; is the relevant disease label.

[0017] Further, for each post-translational modification event of a protein, the data cleaning module clips and expands the sequence data centered on the modified amino acid.

[0018] Further, the basis for the data cleaning module to clip and expand the sequence data is a length of 15 amino acids, with 7 amino acids before and after the modified amino acid as the center. During the expansion process, the blank positions are filled with the placeholder "_".

[0019] Further, the methods for the data cleaning module to perform data cleaning also include operations such as standardizing the types of post-translational modifications of proteins and deleting redundant data.

[0020] Further, the feature extraction module learns the feature embedding of the sequence data by predicting the context through a feature extractor. In the corpus of the feature extractor, in addition to words with a single amino acid as a group, there are also words with three consecutive amino acids as a group. The features are extracted by converting the words in the corpus into feature vectors containing feature information. Among them, the extraction process of the feature vectors is shown in the following formula:

[0021]

[0022] In the formula, h is the feature vector, representing the vector representation of a specific word in the feature space; T(x) represents the original input feature set; is an N-dimensional vector, representing the i-th row in the hidden layer W.

[0023] Further, the feature extractor is a Word2Vec structure.

[0024] Further, the specified vector size in the feature extractor is 100, and the sliding window size is 5.

[0025] Converting the amino acid sequence into a feature vector rich in semantic information through the Word2Vec structure can better represent the context and biological characteristics of the sequence.

[0026] Further, the specific way for the association prediction module to transform the feature vector is to transform the large-dimensional feature vector into multiple batches of small-dimensional feature vectors. The feature vector passes through multiple layers of encoders in a loop and is output as an association prediction result after passing through a classifier composed of linear layers. The calculation process is shown in the following formula:

[0027]

[0028] In the formula, is the probability value predicted by the model; classifier is the classifier; F represents the feature transformation part, which transforms the original input T(x) into a form that the model can process; T(x) represents the original input feature set.

[0029] Further, the predictor in the association prediction module is specifically a Transformer structure.

[0030] Further, the hidden layer dimension of the position encoding part in the predictor is 256, and the number of heads is 10.

[0031] The Transformer structure is used to capture the complex patterns and relationships in the post-translational modification sequence of proteins, improving the accuracy of system disease prediction; the multi-head attention mechanism is used to process high-dimensional feature vectors, reducing noise and improving the generalization ability of the model.

[0032] Further, the main body of the loss function is the cross-entropy function, as shown in the following formula:

[0033]

[0034] In the formula, is the probability value predicted by the model; is the value of the true label;

[0035] Quantifying the difference between the model's predicted probability distribution and the true disease label through the cross-entropy loss function improves the accuracy of disease prediction.

[0036] L2 regularization is added during the calculation of the loss function, and the calculation formula of the L2 regularization is as follows:

[0037]

[0038] In the formula, is the weight decay coefficient; is the total number of features; is the weight vector of the model.

[0039] Further, the decay coefficient of the L2 regularization is 0.25, and the cycle consistency loss is realized through the L1 norm.

[0040] Penalizing the overly large weight values through L2 regularization reduces the overfitting of the model to the training data and improves the generalization ability of the model.

[0041] Further, the function definition module optimizes the loss function based on precision, and updates and optimizes the weights of the parameter matrices of each layer during the backpropagation training process of the model.

[0042] Further, the weight decay loss of the learner is updated through the Adam algorithm, and the learning rate is decreased step by step according to the learning rate adjustment strategy, where the gamma value is 0.25 and the step size is 4.

[0043] Furthermore, during the process of the model evaluation module inputting the sequence data into the network model, the Encoders of the Transformer structure perform multi-head learning and gradually evolve to determine the optimal values of the network model parameters, ultimately achieving dynamic balance, completing the training of the network model, and outputting the evaluation results.

[0044] Furthermore, in each round of loop, feature extraction and association prediction are performed, and the loss value is calculated. After each round of backpropagation ends, the gradient clipping strategy is executed to update the parameters of the predictor.

[0045] Furthermore, the settings of the hyperparameters include: the batch size is 32; training for 50 epochs; the learning rate is 0.00025; the number of input channels is 15; the number of threads is 4.

[0046] Furthermore, the already authenticated sequence data is input into the trained restoration network model to obtain the final evaluation result.

[0047] Furthermore, the system further includes a time series prediction module for learning the feature representation of protein sequences by capturing the time dependence and context information in the sequence data;

[0048] Among them, the update formula of the hidden state is as follows:

[0049] In the formula, is the hidden state at time step ; is the weight matrix between hidden states for transmitting the information of the previous time step; is the hidden state at time step ; is the weight matrix between the input and the hidden state; is the input feature vector at time step ; is the weight matrix of the unique features of protein post-translational modifications for capturing the specific biological characteristics of protein post-translational modification sites; is the unique feature vector of protein post-translational modifications at time step ; is the bias term; is the activation function for introducing non-linearity.

[0050] By capturing the time dependence and context information to identify the key functional regions and domains in the protein sequence, it helps the model understand the function of the protein and its role in diseases; by analyzing the context information in the sequence, the model can more accurately predict the protein post-translational modification sites, thereby improving the accuracy of disease association prediction.

[0051] In a second aspect, the present invention further provides a method for predicting the association between protein post-translational modification and disease. The method is based on the system described in the first aspect above and includes:

[0052] Step 1: Obtain a sequence data set related to protein post-translational modification and a corresponding disease label data set;

[0053] Step 2: Clean the sequence data set, perform sequence shearing and expansion on the sequence data set. If there are blank positions, use the placeholder "_";

[0054] Step 3: Learn the feature embedding of the sequence data and extract feature vectors containing feature information;

[0055] Step 4: Transform the feature vectors through a multi-head attention mechanism and perform disease association prediction through a Transformer structure;

[0056] Step 5: Define a loss function according to the feature complexity and purpose of biological information, and optimize the loss function through the Adam algorithm;

[0057] Step 6: Input the sequence test set into the network model for training, gradually evolve, determine the optimal value of the network model parameters, and finally reach a dynamic balance to complete the training of the network model;

[0058] Step 7: Input the sequence data to be predicted into the network model and output the evaluation result.

[0059] In a third aspect, the present invention further provides a computer device, including a processor and a memory. The processor is connected to the memory. The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the computer device executes at least one step in the method for predicting the association between protein post-translational modification and disease described above.

[0060] In a fourth aspect, the present invention further provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, it realizes at least one step in the method for predicting the association between protein post-translational modification and disease described above.

[0061] The present invention pre - constructs a sequence data set containing post - translational modification sequences of proteins and their disease labels, extracts features from the sequence data through the Word2Vec structure, and processes the extracted feature vectors according to the multi - head attention mechanism of the Transformer structure; combines the cross - entropy loss function and L2 regularization to optimize the prediction model, and uses the Adam algorithm to optimize the model parameters; and learns the feature representation of protein sequences by capturing the temporal dependence and context information in the sequence data. The application scenarios of this system and method are extensive, providing new perspectives and methods for medical diagnosis and promoting the research on disease diagnosis and pathogenesis.

[0062] Beneficial effects

[0063] By implementing the protein post - translational modification and disease association prediction system and method provided by the present invention, the following technical effects are achieved:

[0064] (1) The present invention pre - constructs a sequence data set containing post - translational modification sequences of proteins and their disease labels, extracts features from the sequence data through the Word2Vec structure, and processes the extracted feature vectors according to the multi - head attention mechanism of the Transformer structure. Converting the amino acid sequence into a feature vector rich in semantic information through the Word2Vec structure can better represent the context and biological characteristics of the sequence; capturing the complex patterns and relationships in the post - translational modification sequence of proteins through the Transformer structure improves the accuracy of disease prediction of the system; processing high - dimensional feature vectors through the multi - head attention mechanism reduces noise and improves the generalization ability of the model.

[0065] (2) Combine the cross - entropy loss function and L2 regularization to optimize the prediction model, and use the Adam algorithm to optimize the model parameters. Quantifying the difference between the model prediction probability distribution and the true disease label through the cross - entropy loss function improves the accuracy of disease prediction; punishing overly large weight values through L2 regularization reduces the over - fitting of the model to the training data and improves the generalization ability of the model.

[0066] (3) Learn the feature representation of protein sequences by capturing the temporal dependence and context information in the sequence data. Identifying key functional regions and domains in protein sequences by capturing temporal dependence and context information helps the model understand the function of proteins and their roles in diseases; by analyzing the context information in the sequence, the model can more accurately predict post - translational modification sites of proteins, thereby improving the accuracy of disease association prediction. Brief description of the drawings

[0067] To make the above-mentioned protein post-translational modification and disease association prediction system and method of the present invention more clearly understandable, the following will briefly introduce the drawings required for the specific implementation of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0068] Figure 1 It represents a schematic flowchart of the method for predicting the association between protein post-translational modification and disease;

[0069] Figure 2 It represents a schematic diagram of the data acquisition and cleaning process;

[0070] Figure 3 It represents a schematic diagram of the Word2Vec structure;

[0071] Figure 4 It represents a schematic diagram of the association prediction principle. Specific implementation mode

[0072] Example 1:

[0073] A protein post-translational modification and disease association prediction system and method are provided. The protein post-translational modification and disease association prediction system includes: a data cleaning module, a feature extraction module, an association prediction module, a function definition module, and a model evaluation module. The flowchart of the protein post-translational modification and disease association prediction method is as Figure 1 shown, and the protein post-translational modification and disease association prediction system and method of the following steps are specifically described.

[0074] The data cleaning module is used to obtain sequence data related to protein post-translational modification and corresponding disease label data, and clean the sequence data. The data acquisition and cleaning process is as Figure 2 shown.

[0075] The sequence data is expressed as X = , and its data distribution is . The disease label data is expressed as Y = , and its data distribution is ;

[0076] In the formula, is the number of sequences of protein post-translational modification; is the number of relevant disease labels; X is the sequence of protein post-translational modification; is the relevant disease label.

[0077] For each protein post-translational modification event, the data cleaning module cuts and expands the sequence data with the modified amino acid as the center.

[0078] The basis for the data cleaning module to clip and expand the sequence data is a length of 15 amino acids, with 7 amino acids before and after the modified amino acid as the center. During the expansion process, the blank positions are filled with the placeholder "_".

[0079] The data cleaning methods of the data cleaning module also include operations such as post-translational modification type specification of proteins and redundant data deletion.

[0080] The feature extraction module is used to learn the feature embedding of the sequence data and extract the feature vectors containing feature information.

[0081] The feature extraction module learns the feature embedding of the sequence data by predicting the context through a feature extractor. In the corpus of the feature extractor, in addition to the words with a single amino acid as a group, there are also words with three consecutive amino acids as a group. The feature extraction is realized by converting the words in the corpus into feature vectors containing feature information. Among them, the extraction process of the feature vector is shown in the following formula:

[0082]

[0083] In the formula, h is the feature vector, representing the vector representation of a specific word in the feature space; T(x) represents the original input feature set; is an N-dimensional vector, representing the i-th row in the hidden layer W.

[0084] The feature extractor is a Word2Vec structure, and the Word2Vec structure is as Figure 3 shown.

[0085] The specified vector size in the feature extractor is 100, and the sliding window size is 5.

[0086] The association prediction module is used to transform the feature vectors through the multi-head attention mechanism and perform disease association prediction through the Transformer structure. The association prediction principle is as Figure 4 shown.

[0087] The specific way for the association prediction module to transform the feature vectors is to transform the feature vectors with large dimensions into multi-batch feature vectors with small dimensions. The feature vectors pass through multiple layers of encoder loops and then pass through a classifier composed of linear layers and output as the association prediction result. The calculation process is shown in the following formula:

[0088]

[0089] In the formula, is the probability value predicted by the model; classifier is the classifier; F represents the feature transformation part, which converts the original input T(x) into a form that the model can process; T(x) represents the original input feature set.

[0090] The predictor in the association prediction module is specifically a Transformer structure.

[0091] The hidden layer dimension of the position encoding part in the predictor is 256, and the number of multi-heads is 10.

[0092] The function definition module is used to define the loss function according to the feature complexity and purpose of the biological information, and optimize the loss function through the Adam algorithm.

[0093] The main body of the loss function is the cross-entropy function, as shown in the following formula:

[0094]

[0095] In the formula, is the probability value predicted by the model; is the value of the true label;

[0096] L2 regularization is added during the calculation of the loss function, and the calculation formula of the L2 regularization is as follows:

[0097]

[0098] In the formula, is the weight decay coefficient; is the total number of features; is the weight vector of the model.

[0099] The decay coefficient of the L2 regularization is 0.25, and the cycle consistency loss is realized through the L1 norm.

[0100] The function definition module optimizes the loss function based on the accuracy, and updates and optimizes the weights of the parameter matrices of each layer during the backpropagation training process of the model.

[0101] The weight decay loss of the learner is updated through the Adam algorithm, and the learning rate is decreased step by step according to the learning rate adjustment strategy, where the gamma value is 0.25 and the step size is 4.

[0102] The model evaluation module is used to input the sequence data into the trained network model and output the evaluation result.

[0103] When the model evaluation module inputs the sequence data into the network model, the Encoders of the Transformer structure perform multi-head learning and gradually evolve to determine the optimal values of the network model parameters, ultimately achieving dynamic balance, completing the training of the network model, and outputting the evaluation results.

[0104] In each round of loop, feature extraction and association prediction are performed, and the loss value is calculated. After each round of backpropagation ends, the gradient clipping strategy is executed to update the parameters of the predictor.

[0105] The settings of the hyperparameters include: the batch size is 32; training for 50 epochs; the learning rate is 0.00025; the number of input channels is 15; the number of threads is 4.

[0106] The authenticated sequence data is input into the trained reduction network model to obtain the final evaluation result.

[0107] Example 2:

[0108] On the basis of the foregoing embodiment, the system adds a time series prediction module for learning the feature representation of protein sequences by capturing the time dependence and context information in the sequence data.

[0109] Collect protein sequences with post-translational modification sites of proteins, remove sequences with similarity exceeding a certain threshold through a protein sequence clustering tool, and perform data cleaning on the de-redundant protein sequences to delete low-confidence site annotation information.

[0110] Normalize the length of the protein sequences by padding with zeros or truncating, and encode the protein sequences.

[0111] The feature extraction module extracts features from the protein sequences.

[0112] A recurrent neural network model is designed with a dynamic "cascade" structure to improve the average sensitivity of the model. The update formula for the hidden state is as follows:

[0113] In the formula, is the hidden state at time step ; is the weight matrix between hidden states, used to transmit information from the previous time step; is the hidden state at time step ; is the weight matrix between the input and the hidden state; is the input feature vector at time step ; A weight matrix that is a unique feature of post - translational modification of proteins, used to capture the specific biological characteristics of post - translational modification sites of proteins; is the time step unique feature vector of post - translational modification of proteins; is the bias term; is the activation function, used to introduce non - linearity.

[0114] The gradient of the recurrent neural network is updated by the Adam optimizer in the function definition module, and the loss function still uses the cross - entropy function.

[0115] Prevent overfitting of the model by early stopping.

[0116] Input the protein sequence to be tested into the recurrent neural network model to predict the association between post - translational modification of proteins and diseases.

[0117] For example, assume the following data:

[0118] The protein sequence is A, represented by one - hot encoding; the post - translational modification feature is none, represented by a zero vector; the disease label is 1, indicating the patient has the disease.

[0119] Assume the following parameters:

[0120] ; ; ; ; ; .

[0121] Initialize the hidden state , at time step 1:

[0122] Input feature : of encoding, assume it is (a 20 - dimensional vector, only the first position is 1);

[0123] Post - translational modification feature of protein : ;

[0124] For the calculation of the hidden state : ;

[0125] Calculate the output according to the hidden state : .

[0126] According to the output result, the probability of the patient having the disease is about 70%.

[0127] The effect of the time series prediction module is shown in Table 1.

[0128] Table 1. Summary of the time series prediction module

[0129] Index Before the application of the deep learning module After the application of the deep learning module Test mean squared error High Low Prediction accuracy Low High Ability to process long sequences Limited Strong Short-term memory ability None Yes Training complexity Low High Degree of demand for large-scale data Low High

[0130] According to the experimental table, the deep learning model requires a relatively large-scale dataset and a relatively complex training process to learn features. However, its mean square error during testing is lower, which means that the application of this model can make the prediction results closer to the true values, improving the prediction accuracy. After the model is applied, the system has stronger capabilities in processing long sequence data and has short-term memory capabilities, which helps to process and analyze sequence data.

[0131] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media containing computer-usable program code.

[0132] The present invention can provide computer program instructions to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the system.

[0133] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions of the system.

[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions of the system.

Claims

1. A protein post-translational modification and disease association prediction system, characterized in that: include: A data cleaning module is used to obtain sequence data related to protein post-translational modification and corresponding disease label data, and clean the sequence data; Feature extraction module, used to learn feature embedding of sequence data and extract feature vectors containing feature information; The association prediction module is used to transform feature vectors through a multi-head attention mechanism and perform disease association prediction through a Transformer structure; Function definition module, which is used to define the loss function according to the feature complexity and purpose of biological information, and optimize the loss function through the Adam algorithm; The model evaluation module is used to input sequence data into the trained network model and output the evaluation results.

2. The system according to claim 1, characterized in that: For each protein post-translational modification event, the data cleaning module cuts and expands the sequence data centered on the modified amino acid.

3. The system according to claim 1, characterized in that: The feature extraction module learns feature embedding of sequence data by predicting context through a feature extractor. In addition to words that are a group of single amino acids, the corpus of the feature extractor also contains words that are a group of three consecutive amino acids. The feature extraction is achieved by converting the words in the corpus into feature vectors containing feature information. The feature vector extraction process is shown in the following formula: In the formula, h is the feature vector, which represents the vector representation of a specific word in the feature space; T(x) represents the original input feature set; is an N-dimensional vector representing the i-th row in the hidden layer W.

4. The system according to claim 1, characterized in that: The association prediction module converts the feature vector in a specific way by converting the feature vector of large dimension into multiple batches of feature vectors of small dimension. The feature vector is output as the association prediction result after passing through a multi-layer encoder cycle and a classifier composed of linear layers. The calculation process is shown in the following formula: In the formula, is the probability value predicted by the model; classifier is the classifier; F represents the feature transformation part, which converts the original input T(x) into a form that the model can process; T(x) represents the original input feature set.

5. The system according to claim 1, characterized in that: The main body of the loss function is the cross entropy function, as shown in the following formula: In the formula, is the probability value predicted by the model; is the value of the true label; L2 regularization is added during the calculation of the loss function. The calculation formula of L2 regularization is as follows: In the formula, is the weight attenuation coefficient; is the total number of features; is the weight vector of the model.

6. The system according to claim 1, characterized in that: The system also includes a temporal prediction module for learning feature representations of protein sequences by capturing temporal dependencies and contextual information in sequence data.

7. A method for predicting the association between protein post-translational modification and disease, characterized in that: The method is implemented based on the system according to any one of claims 1 to 6: The method comprises: Step 1: Obtain sequence datasets related to protein post-translational modifications and corresponding disease label datasets; Step 2: Clean the sequence data set, cut and expand the sequence data set, and use the placeholder "_" if there is a blank position; Step 3: Learn the feature embedding of the sequence data and extract the feature vector containing the feature information; Step 4: Transform the feature vector through the multi-head attention mechanism and perform disease association prediction through the Transformer structure; Step 5: Define the loss function according to the feature complexity and purpose of biological information, and optimize the loss function through the Adam algorithm; Step 6: Input the sequence test set into the network model for training, gradually evolve, determine the optimal values ​​of the network model parameters, and finally reach a dynamic balance to complete the training of the network model; Step 7: Input the sequence data to be predicted into the network model and output the evaluation results.

8. A computer device comprising a processor and a memory, wherein the processor is connected to the memory, and the memory is used to store a computer program, wherein: The processor is configured to execute the computer program stored in the memory, so that the computer device executes the method of claim 7.

9. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, characterized in that: The computer program implements the method of claim 7 when executed.

Citation Information

Patent Citations

  • Protein phosphorylation modification site-disease relationship recognition method, system and device and storage medium

    CN111696621A

  • Deep learning method for predicting protein post-translational modification sites

    CN114724630A