A protein methylation site prediction method, device and equipment
By converting the three-dimensional structure of a protein into a graph structure and extracting sub-graphs, inputting it into a graph convolutional neural network for prediction, the problem of low prediction stability and accuracy of protein methylation sites based on one-dimensional structure in the prior art is solved, and a more accurate and stable prediction effect is achieved.
Patent Information
- Application Number
- CN202210468759.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-04-29
AI Technical Summary
In the prior art, when predicting protein methylation sites, methods based on one-dimensional structures have poor stability and low accuracy.
By converting the three-dimensional structure of the protein to be tested into a graph structure with amino acids as nodes, and extracting the sub-graphs based on preset parameters, inputting them into the trained graph convolution neural network for methylation site prediction.
Improve the accuracy and robustness of protein methylation site prediction, and use more comprehensive amino acid information in the three-dimensional structure to predict.
Smart Images

Figure CN114882952B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of bioinformatics, and in particular to a method, device and equipment for predicting protein methylation sites. Background Art
[0002] Methylation is an important modification of proteins and nucleic acids. It regulates gene expression and shutdown. It is closely related to many diseases such as cancer, aging, and Alzheimer's disease. It is one of the important research contents of epigenetics. The most common methylation modifications are DNA methylation and histone methylation. Methylation refers to the process of catalytic transfer of methyl groups from active methyl compounds to other compounds, which can form various methyl compounds, or chemically modify certain proteins or nucleic acids to form methylated products.
[0003] Protein methylation mainly occurs at lysine and arginine. Currently, most of the research on predicting protein methylation sites uses traditional machine learning and deep learning methods, such as the PseAAC method that quickly generates pseudo amino acid groups based on protein sequence data and integrates sequence coupling information to extract features of protein sequences, and then uses the Support Vector Machine (SVM) method to predict methylation sites. For example, the LSTM network is used to predict lysine post-translational modification sites. However, these methods are based on the one-dimensional structure of proteins, that is, protein sequences, to predict methylation sites, which are less stable and have lower accuracy. Summary of the invention
[0004] The embodiments of the present application provide a method and device for predicting protein methylation sites, which can more accurately predict protein methylation sites with higher robustness.
[0005] In a first aspect, the present application provides a method for predicting protein methylation sites, comprising:
[0006] The three-dimensional structure data of the protein to be tested is obtained, and a graph structure with amino acids as nodes is constructed based on the three-dimensional structure data, where each node in the graph structure corresponds to an amino acid; a subgraph is extracted from the graph structure based on preset parameters, where the preset parameters include the number of subgraph hops and the distance threshold; the subgraph is input into a trained graph convolutional neural network for methylation site prediction.
[0007] In an embodiment of the present application, the three-dimensional structure of the protein to be tested is converted into a graph structure with amino acids as nodes, and then a subgraph is extracted from the graph structure and input into a trained graph convolutional neural network model for methylation site prediction. Since the subgraph extracted from the graph structure constructed based on the three-dimensional structure data of the protein contains more comprehensive amino acid information, the graph convolutional neural network can more accurately predict protein methylation sites with higher robustness.
[0008] In an optional manner provided in the first aspect, extracting a subgraph from the graph structure according to preset parameters includes:
[0009] According to the protein sequence, any amino acid in the above protein sequence is used as the starting node of the above subgraph;
[0010] Taking the above starting node as the center, according to the above preset parameters, a subgraph of the amino acid corresponding to the above starting node is extracted from the above graph structure.
[0011] The above subgraph is composed of an adjacency matrix A and a feature matrix X. The adjacency matrix A represents the matrix of adjacent relationships between vertices, A∈(0,1) L×L , L is the number of amino acid nodes in the above subgraph; the above feature matrix X is composed of feature matrix X o and the feature matrix X w Composition, the above feature matrix X o The feature matrix is obtained based on the sequence position features of the amino acid nodes obtained by one-hot encoding. The feature matrix X w It is a feature matrix obtained based on the sequence position features of the amino acid nodes obtained by word vectors.
[0012] In another optional manner provided in the first aspect, extracting a subgraph from the graph structure according to preset parameters includes:
[0013] The above feature matrix X is transformed by the corrected linear unit ReLU mapping function o And the above feature matrix X w Perform a learnable nonlinear mapping to obtain the above feature matrix X. The above ReLU mapping function is:
[0014] X = Relu(X w W w +X o W o +b X )
[0015] Where X∈R L×C , L is the number of amino acid nodes in the subgraph, C is the feature dimension of the subgraph nodes; W w Represents the above feature matrix Xw The weight matrix, W o Represents the above feature matrix X o The weight matrix, b represents the bias, b X Represented as the hyperparameter of the ReLU mapping function, where b X ∈R L×C .
[0016] In another optional manner provided in the first aspect, the above-mentioned construction of a graph structure with amino acids as nodes based on the above-mentioned three-dimensional structure data includes:
[0017] Get the distance between the central carbon atoms of any two amino acids;
[0018] When the distance value is less than the distance threshold, it is determined that the arbitrary two amino acids are connected.
[0019] In another optional manner provided in the first aspect, before inputting the subgraph into a trained graph convolutional neural network for methylation site prediction, the method further includes:
[0020] The trained graph convolutional neural network is subjected to Bayesian optimization to optimize the subgraph hop count and distance threshold used in subgraph extraction.
[0021] In another optional manner provided in the first aspect, the inputting the subgraph into a trained graph convolutional neural network for methylation site prediction includes:
[0022] Perform a global pooling operation on the feature matrix output by each convolutional layer in the above graph convolutional neural network;
[0023] The output of the pooling layer after the global pooling operation is calculated using a fully connected layer and a ReLU activation function.
[0024] In a second aspect, the present application provides a protein methylation site prediction device, comprising:
[0025] A three-dimensional structure data acquisition unit, used to acquire the three-dimensional structure data of the protein to be tested, wherein the protein is composed of a plurality of amino acids;
[0026] A graph structure construction unit, used to construct a graph structure with amino acids as nodes according to the three-dimensional structure data, wherein each node in the graph structure corresponds to an amino acid;
[0027] A subgraph extraction unit, configured to extract a subgraph from the graph structure according to preset parameters, wherein the preset parameters include a subgraph hop count and a distance threshold;
[0028] The methylation site prediction unit is used to input the above subgraph into the trained graph convolutional neural network for methylation site prediction.
[0029] In an optional manner provided in the second aspect, the sub-graph extraction unit is specifically used to:
[0030] According to the protein sequence, any amino acid in the above protein sequence is used as the starting node of the above subgraph;
[0031] Taking the above starting node as the center, according to the above preset parameters, a subgraph of the amino acid corresponding to the above starting node is extracted from the above graph structure.
[0032] The above subgraph is composed of an adjacency matrix A and a feature matrix X. The adjacency matrix A represents the matrix of adjacent relationships between vertices, A∈(0,1) L×L , L is the number of amino acid nodes in the above subgraph; the above feature matrix X is composed of feature matrix X o and the feature matrix X w Composition, the above feature matrix X o The feature matrix is obtained based on the sequence position features of the amino acid nodes obtained by one-hot encoding. The feature matrix X w It is a feature matrix obtained based on the sequence position features of the amino acid nodes obtained by word vectors.
[0033] In another optional manner provided in the second aspect, the sub-graph extraction unit is specifically used to:
[0034] The above feature matrix X is transformed by the corrected linear unit ReLU mapping function o And the above feature matrix X w Perform a learnable nonlinear mapping to obtain the above feature matrix X. The above ReLU mapping function is:
[0035] X = Relu(X w W w +X o W o +b X )
[0036] Where X∈R L×C , L is the number of amino acid nodes in the subgraph, C is the feature dimension of the subgraph nodes; W w Denotes the feature matrix X w The weight matrix, W o Denotes the feature matrix X o The weight matrix, b represents the bias, b X is the hyperparameter of the ReLU mapping function, where b X ∈R L×C .
[0037] In another optional manner provided in the second aspect, the graph structure building unit includes:
[0038] A distance value acquisition subunit is used to obtain the distance value between the central carbon atoms of any two amino acids;
[0039] The connection confirmation subunit is used to determine that the above-mentioned arbitrary two amino acids are connected when the above-mentioned distance value is less than the distance threshold.
[0040] In another optional manner provided in the second aspect, the above-mentioned device also includes:
[0041] The trained graph convolutional neural network is subjected to Bayesian optimization to optimize the subgraph hop count and distance threshold used in subgraph extraction.
[0042] In another optional manner provided in the second aspect, the methylation site prediction unit is specifically used to:
[0043] Perform a global pooling operation on the feature matrix output by each convolutional layer in the above graph convolutional neural network;
[0044] The output of the pooling layer after the global pooling operation is calculated using a fully connected layer and a ReLU activation function.
[0045] In a third aspect, the present application provides a protein methylation site prediction device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method as described in the first aspect or any optional mode of the first aspect when executing the computer program.
[0046] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect or any optional manner of the first aspect is implemented.
[0047] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on a protein methylation site prediction device, the protein methylation site prediction device executes the steps of the protein methylation site prediction method described in the first aspect.
[0048] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0050] Figure 1 It is a schematic diagram of a process of predicting a protein methylation site provided in an embodiment of the present application;
[0051] Figure 2 The embodiment of the present application provides a group of structural schematic diagrams of subgraphs extracted based on different distance thresholds and different subgraph hop numbers;
[0052] Figure 3 is a flow chart of a method for determining whether amino acids are connected provided in an embodiment of the present application;
[0053] Figure 4 It is a schematic diagram of the structure of a protein methylation site prediction device provided in an embodiment of the present application;
[0054] Figure 5 It is a schematic diagram of the structure of a protein methylation site prediction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0056] It should be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0057] It should also be understood that references to "one embodiment" or "some embodiments" etc. described in the specification of the present application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Thus, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in the specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0058] The common point of existing studies on predicting protein methylation is that protein methylation sites are predicted based on the one-dimensional structure of the protein, namely the protein sequence. However, since the information contained in the one-dimensional structure is not comprehensive and complete, the accuracy of the final protein methylation site prediction is not very high.
[0059] The one-dimensional structure is coiled and folded to form a two-dimensional structure, and then further coiled or folded to form a three-dimensional spatial structure with certain rules. After the polypeptide chain is coiled in this way, it can form certain specific areas that play biological functions, such as the active center of the enzyme, so the three-dimensional structure formed by the one-dimensional structure of the protein after coiling can also play some functions of the protein. Therefore, compared with the one-dimensional structure, the information contained in the three-dimensional structure is more comprehensive and complete. Moreover, unlike the traditional prediction of protein post-modification sites through the one-dimensional structure of the protein, the three-dimensional structure is relatively stable, and the biological functions of certain specific areas can also be played. It is also more accurate to predict methylation sites through the three-dimensional structure of the protein.
[0060] Therefore, based on the advantages of the three-dimensional structure of proteins, this application uses the artificial intelligence Alphafold method proposed by Google to predict the one-dimensional structure of proteins into the three-dimensional structure of proteins, and predict methylation sites through the predicted three-dimensional structure of proteins. Methylation site prediction is to predict whether each amino acid on the three-dimensional protein structure is methylated and determine the methylated amino acids.
[0061] See also Figure 1 , Figure 1 This is a schematic diagram of a method for predicting protein methylation sites provided in an embodiment of the present application, which is described in detail as follows:
[0062] Step S101, obtaining the three-dimensional structure data of the protein to be tested.
[0063] In the present embodiment, a protein is composed of multiple amino acids, and a protein is defined as P, an amino acid is defined as a, and a protein sequence composed of n amino acids is defined as P = (a1, a2, ..., a n ), each amino acid has a central carbon atom called C α atom.
[0064] Step S102: constructing a graph structure with amino acids as nodes based on the three-dimensional structure data.
[0065] In the embodiment of the present application, each node in the constructed graph structure corresponds to an amino acid. After the three-dimensional structure of the protein to be tested is converted into a graph structure, the amino acids are nodes of the graph structure.
[0066] Step S103: extracting a subgraph from the above graph structure according to preset parameters.
[0067] In the embodiment of the present application, the preset parameters are hyperparameters set in the trained graph convolutional neural network, including the subgraph hop number H and the distance threshold D. If the distance value between the central carbon atoms of any two amino acids is less than the distance threshold D, the two amino acids can be considered to be connected.
[0068] When converting the three-dimensional structure of protein into a graph structure, amino acids are nodes of the graph structure, and H is the subgraph hop number. If two amino acids are connected, the subgraph hop number H is 1, and the hop number increases with every other amino acid. Therefore, the extraction of the subgraph is determined by the distance threshold D and the subgraph hop number H. Different subgraph hop numbers and different distance thresholds determine different subgraph structures.
[0069] Since each node in the constructed graph structure corresponds to an amino acid, when subgraphs are extracted from amino acids in a protein sequence according to different subgraph hop numbers H and different distance thresholds, the extracted subgraphs correspond to one amino acid.
[0070] like Figure 2 As shown, Figure 2 Schematic diagram of the structure of the subgraph extracted based on different distance thresholds D and different subgraph hops. Specifically, taking the first amino acid of the current protein sequence as lysine K, different subgraph hops H and different distance thresholds D are used to extract the subgraph corresponding to lysine K as an example.
[0071] When the subgraph hop count H is 1 and the distance threshold D is 5, other amino acids with subgraph hop count H of 1 are searched within the range of the distance threshold D, that is, other amino acids connected to the lysine are searched within the range of the distance threshold D of 5. Within the range of the distance threshold D, only serine S and leucine L are connected to lysine K. Then, when the subgraph hop count H is 1 and the distance threshold D is 5, the extracted subgraph of lysine K is as follows: Figure 2As shown in (a) in .
[0072] When the subgraph hop number H is 1 and the distance threshold D is 10, other amino acids with subgraph hop number H of 1 are searched within the range of the distance threshold D, that is, other amino acids connected to the lysine are searched within the range of the distance threshold D of 10. Within the range of the distance threshold D, one each of leucine L, threonine T, arginine R, and glutamine Q and two serine S are connected to lysine K. Then, when the subgraph hop number H is 1 and the distance threshold D is 10, the extracted subgraph of lysine K is as follows: Figure 2 As shown in (b) in .
[0073] When the subgraph hop count H is 2 and the distance threshold D is 5, other amino acids with subgraph hop counts H of 1 and 2 are searched within the range of the distance threshold D, that is, other amino acids connected to the lysine are searched within the range of the distance threshold D of 5. Within the range of the distance threshold D of 5, the subgraph hop count H of 2 is found to have one each of leucine L, threonine T, serine S, and glutamine Q connected to lysine K. Then, when the subgraph hop count H is 2 and the distance threshold D is 10, the extracted subgraph of lysine K is as follows: Figure 2 As shown in (c) in .
[0074] When the subgraph hop count H is 2 and the distance threshold D is 10, the extracted subgraph of lysine K is as follows: Figure 2 As shown in (d) in .
[0075] See also Figure 3 , Figure 3 is a flow chart of a method for determining whether amino acids are connected provided in an embodiment of the present application, which is described in detail as follows
[0076] Step S301, obtaining the distance value between the central carbon atoms of any two amino acids.
[0077] Step S302: when the distance value is less than the distance threshold, it is determined that the arbitrary two amino acids are connected.
[0078] In the present embodiment, the distance value between the central carbon atoms of any two amino acids is obtained, and the distance is compared with a preset distance threshold to determine whether the two amino acids are connected. When the distance value between the central carbon atoms of any two amino acids is less than the distance threshold, the two amino acids can be considered to be connected; when the distance value between the central carbon atoms of any two amino acids is greater than or equal to the distance threshold, the two amino acids can be considered to be not connected.
[0079] Since a protein sequence is composed of multiple amino acids arranged in a certain order, each subgraph corresponds to an amino acid. After determining the preset parameters, the subgraphs corresponding to each amino acid can be extracted from the graph structure corresponding to the protein. According to the preset parameters, extracting subgraphs from the graph structure specifically includes:
[0080] According to the protein sequence, any amino acid in the above protein sequence is used as the starting node of the above subgraph; with the above starting node as the center, according to the preset parameters, the subgraph of the amino acid corresponding to the starting node is extracted from the graph structure, and this step is repeated until the subgraphs corresponding to all the amino acids in the protein sequence are obtained. Here, the number of subgraphs finally extracted is consistent with the number of amino acids in the protein sequence. When the number of amino acids in the protein sequence is n, the number of extracted subgraphs is n.
[0081] In the embodiment of the present application, each subgraph is composed of an adjacency matrix A and a feature matrix X, wherein the adjacency matrix A represents a matrix of adjacent relationships between vertices, A∈(0,1) L×L , L is the number of amino acid nodes in the subgraph, that is, the number of amino acids contained in the extracted subgraph. The feature matrix X is composed of the feature matrix X o and the feature matrix X w Composition, feature matrix X o The feature matrix X is obtained based on the sequence position features of the amino acid nodes obtained by one-hot encoding. w It is a feature matrix obtained based on the sequence position features of the amino acid nodes obtained by word vectors.
[0082] Specifically, the feature matrix X is mapped by the rectified linear unit (ReLU) function o and the feature matrix X w Perform a learnable nonlinear mapping to obtain the feature matrix X. The ReLU mapping function is:
[0083] X = Relu(X w W w +X o W o +b X ) (1)
[0084] Where X∈R L×C , L is the number of amino acid nodes in the subgraph, C is the feature dimension of the subgraph nodes; W w Denotes the feature matrix X w The weight matrix, W o Denotes the feature matrix X o The weight matrix, b represents the bias, b Xis the hyperparameter of the ReLU mapping function, where b X ∈R L×C .
[0085] Step S104: input the above subgraph into the trained graph convolutional neural network to predict methylation sites.
[0086] In the embodiment of the present application, in order to reduce the computational complexity, the Chebyshev polynomial approximation graph convolution kernel is used. The Chebyshev graph convolution layer is a convolution layer that uses the Chebyshev polynomial approximation graph convolution kernel. The adjacency matrix A and the matrix X from the previous layer are combined. l The feature matrix embedding of is taken as input, where C l Embed the dimension of the nodes in the lth layer and output it to the feature matrix embedding X in the next layer l+1 ,in, C l+1 Embed the node dimension of the l+1th layer, and fuse the feature matrix and adjacency matrix of the previous layer into X l As the input layer of the graph convolutional neural network, the output is used as the input of the next network layer l+1, then:
[0087] X l+1 =ChebConv(A,X l ) (2)
[0088] Where l represents the current number of layers of the graph convolutional neural network. Chebyshev polynomial T l(0) T l(1) …T l(k) As shown below:
[0089] T l(0) =X l (3.1)
[0090] T l(0) =X l (3.2)
[0091]
[0092] Here, k represents the number of iterations.
[0093] Substituting into the graph convolution process, the weight matrix is approximated by Chebyshev as follows:
[0094]
[0095]
[0096] where x∈R H , g∈R n×m , H, n, m are different dimensions, θ iis Chebyshev's coefficient, k represents the number of iterations, L laplace is the regularized Laplace matrix, I is the identity matrix, Q is the degree matrix of the adjacency matrix A, λ max for The maximum eigenvalue of , formula (2) is approximated by Chebyshev graph polynomials. The convolution layer can also be mathematically expressed as shown in formula (6), which updates the feature matrix of each subgraph by the weighted sum of the features of the directly connected nodes:
[0097]
[0098] Among them, N K represents the order of Chebyshev expansion, T l(k) represents Chebyshev type, W l(k) and b l(k) is the trainable weight matrix of layer l.
[0099] In some embodiments of the present application, a global pooling operation is performed on the feature matrix output by each convolutional layer in the graph convolutional neural network; then a fully connected layer and a ReLU activation function are used to calculate the output of the pooling layer after the global pooling operation.
[0100] Specifically, due to the different number of amino acid nodes in the subgraphs, the graph convolutional neural network uses convolutional layers with different convolutional kernels. We take the feature matrix X output by each layer of the convolutional neural network as l+1 Concatenate into a feature matrix X output , that is, X output =[X 1 , X 2 , …, X n ], and output X output Perform a global pooling operation, that is
[0101]
[0102] Then use the fully connected layer and ReLU activation function to calculate the output of the pooling layer. Finally, use the softmax activation function to connect the final fully connected layer and the output The softmax function is a normalized exponential function that can map the output result to the interval [0, 1], and finally convert the output Y into a binary classification problem, that is, mapping Y to Z∈{0, 1}. The cross-entropy objective function of the network, that is, the loss function Loss, is defined as:
[0103]
[0104] Among them, N b is the batch size of the training set, Y is the one-hot encoded true label, O is the embedding of the subgraph through one-hot encoding, is mapped to Z∈{0,1}.
[0105] In some embodiments of the present application, after constructing a graph convolutional neural network, the trained graph convolutional neural network is subjected to Bayesian optimization to optimize the number of subgraph hops and distance thresholds used in subgraph extraction. Compared with other optimization methods such as random search and grid search, Bayesian optimization only requires continuous sampling to infer the maximum value of the function, and does not require global sampling, reducing computing time.
[0106] The embodiment of the present application applies Bayesian optimization to the selection of the subgraph hop count H and the distance threshold D, because it is impossible to obtain the optimal subgraph hop count H and the distance threshold D through prior knowledge to determine the optimal subgraph as the input of the graph convolutional neural network, and the subgraphs corresponding to different subgraph hop counts and distance thresholds are quite different, which affects the training speed of the graph convolutional neural network. Therefore, it is very necessary to determine a reasonable subgraph through the Bayesian optimization method.
[0107] The objective function of the graph convolutional neural network is defined as f(·), and the parameter to be optimized is defined as a vector v = [H, D]. The target of optimizing the hyperparameters can be expressed as:
[0108]
[0109] Among them, v represents the parameter to be optimized, v best represents the optimization parameter, and r represents the decision space.
[0110] The objective function is obtained by modeling the Gaussian distribution. The Gaussian distribution N(u(·), k(·, ·)), i.e. f(·)~N(u(·), k(·, ·)), can be determined by a kernel function k(·, ·) and an average function u(·). The radial basis function (RBF) is used as the kernel function, which can be expressed as k RBF =exp(-||v-v'|| 2 / (2δ 2 )), where δ is a hyperparameter and v' is the mean or expectation of v. Assume V = {v1, v2, ..., v n} is an evaluation set, where v n is obtained after v is traversed n times, and V'={v1', v2', ..., v n '} is an unevaluated set, v n ' is v n If we consider the mean or expectation of , then the Gaussian model has the following properties:
[0111]
[0112] Among them, K VV=K T V’V , K VV’ =K T V’V , we can get a prediction distribution f(v') to determine the next point of the hyperparameter to be predicted, which obeys the Gaussian distribution f(v')~N(μ(v'),σ(v')), where Then use an expected improvement function e i (·) Improve network performance, defined as follows:
[0113]
[0114] where φ(·) is a cumulative distribution function, e i (·) is a probability density function that conforms to the standard normal distribution, v best is the goal of optimizing parameters, the next point v to be selected next Calculated by argmax,
[0115]
[0116] Among them, v next represents the next point to be selected, r' belongs to the real number domain, and represents a point searched in the decision space r.
[0117] The next point v to be selected is found next Add to V and o(V), and continue to iterate until the optimal v is found best Thus, the optimal subgraph hop number H and distance threshold D are determined.
[0118] Methylation mostly occurs on lysine and arginine. The protein methylation site prediction method provided in the embodiment of the present application is used to train and test the sites where methylation of two different types of amino acids occurs. The methylation sites occurring on arginine and lysine are from the public data set GPS-MSP. The protein methylation site prediction method provided in the embodiment of the present application is used to train and test the sites of lysine and arginine. The prediction results of the arginine data set and the lysine data set are shown in Table 1 below:
[0119]
[0120] Table 1 Prediction results on two datasets of lysine and arginine
[0121] As shown in Table 1, the graph convolutional neural network was trained and tested on data sets in two different scenarios. According to the results analysis, the overall performance of the graph convolutional neural network has been improved after Bayesian optimization. At the same time, because it is trained on data sets in two different scenarios, the results obtained will also be different. Overall, the model has a certain degree of robustness.
[0122] It should be noted that in different prediction environments or after improving the graph convolutional neural network, the final prediction results will be better than the data provided in Table 1.
[0123] In an embodiment of the present application, the three-dimensional structural data of the protein to be tested is obtained, and a graph structure with amino acids as nodes is constructed based on the three-dimensional structural data. Then, subgraphs are extracted from the constructed graph structure based on the number of subgraph hops and the distance threshold. The extracted subgraphs are then input into a trained graph convolutional neural network for methylation site prediction. Since the subgraphs extracted from the graph structure constructed based on the three-dimensional structural data of the protein contain more comprehensive amino acid information, the graph convolutional neural network can more accurately predict protein methylation sites with higher robustness.
[0124] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0125] Based on the protein methylation site prediction method provided in the above embodiment, the present application further provides an embodiment of a device for implementing the above method embodiment.
[0126] See also Figure 4 , Figure 4 Schematic diagram of a protein methylation site prediction device provided in an embodiment of the present application. The units included are used to execute Figure 1 For details, please refer to the steps in the corresponding embodiment. Figure 1 Related description in the corresponding embodiment: For ease of explanation, only the parts related to this embodiment are shown.
[0127] See also Figure 4 , the protein methylation site prediction device 4 comprises:
[0128] A three-dimensional structure data acquisition unit 41 is used to acquire the three-dimensional structure data of the protein to be tested, wherein the protein is composed of a plurality of amino acids;
[0129] A graph structure construction unit 42, configured to construct a graph structure with amino acids as nodes according to the three-dimensional structure data, wherein each node in the graph structure corresponds to an amino acid;
[0130] A subgraph extraction unit 43, configured to extract a subgraph from the graph structure according to preset parameters, wherein the preset parameters include a subgraph hop count and a distance threshold;
[0131] The methylation site prediction unit 44 is used to input the above subgraph into the trained graph convolutional neural network to predict the methylation sites.
[0132] In some embodiments of the present application, the sub-graph extraction unit 43 is specifically used for:
[0133] According to the protein sequence, any amino acid in the above protein sequence is used as the starting node of the above subgraph;
[0134] Taking the above starting node as the center, according to the above preset parameters, a subgraph of the amino acid corresponding to the above starting node is extracted from the above graph structure.
[0135] The above subgraph is composed of an adjacency matrix A and a feature matrix X. The adjacency matrix A represents the matrix of adjacent relationships between vertices, A∈(0,1) L×L , L is the number of amino acid nodes in the above subgraph; the above feature matrix X is composed of feature matrix X o and the feature matrix X w Composition, the above feature matrix X o The feature matrix is obtained based on the sequence position features of the amino acid nodes obtained by one-hot encoding. The feature matrix X w It is a feature matrix obtained based on the sequence position features of the amino acid nodes obtained by word vectors.
[0136] In some other embodiments of the present application, the sub-graph extraction unit 43 is specifically used for:
[0137] The above feature matrix X is transformed by the corrected linear unit ReLU mapping function o And the above feature matrix X w Perform a learnable nonlinear mapping to obtain the above feature matrix X. The above ReLU mapping function is:
[0138] X = Relu(X w W w +X o W o +b X )
[0139] Where X∈R L×C , L is the number of amino acid nodes in the subgraph, C is the feature dimension of the subgraph nodes; W w Denotes the feature matrix X w The weight matrix, W o Denotes the feature matrix X oThe weight matrix, b represents the bias, b X is the hyperparameter of the ReLU mapping function, where b X ∈R L×C .
[0140] In some other embodiments of the present application, the graph structure building unit 42 includes:
[0141] A distance value acquisition subunit is used to obtain the distance value between the central carbon atoms of any two amino acids;
[0142] The connection confirmation subunit is used to determine that the above-mentioned arbitrary two amino acids are connected when the above-mentioned distance value is less than the distance threshold.
[0143] In some other embodiments of the present application, the above-mentioned device further includes:
[0144] The trained graph convolutional neural network is subjected to Bayesian optimization to optimize the subgraph hop count and distance threshold used in subgraph extraction.
[0145] In other embodiments of the present application, the methylation site prediction unit 44 is specifically used to:
[0146] Perform a global pooling operation on the feature matrix output by each convolutional layer in the above graph convolutional neural network;
[0147] The output of the pooling layer after the global pooling operation is calculated using a fully connected layer and a ReLU activation function.
[0148] In an embodiment of the present application, the three-dimensional structural data of the protein to be tested is obtained, and a graph structure with amino acids as nodes is constructed based on the three-dimensional structural data. Then, subgraphs are extracted from the constructed graph structure based on the number of subgraph hops and the distance threshold. The extracted subgraphs are then input into a trained graph convolutional neural network for methylation site prediction. Since the subgraphs extracted from the graph structure constructed based on the three-dimensional structural data of the protein contain more comprehensive amino acid information, the graph convolutional neural network can more accurately predict protein methylation sites with higher robustness.
[0149] It should be noted that the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0150] Figure 5 Schematic diagram of a protein methylation site prediction device provided in an embodiment of the present application. Figure 5As shown, the protein methylation site prediction device 5 of this embodiment includes: a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50, such as a speech recognition program. When the processor 50 executes the computer program 52, the steps in the above-mentioned protein methylation site prediction method embodiments are implemented, such as Figure 1 Alternatively, when the processor 50 executes the computer program 52, the functions of each module / unit in the above-mentioned device embodiments are implemented, for example Figure 4 The functions of units 41-44 are shown.
[0151] Exemplarily, the computer program 52 may be divided into one or more modules / units, one or more modules / units are stored in the memory 51 and executed by the processor 50 to complete the present application. One or more modules / units may be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 52 in the protein methylation site prediction device 5. For example, the computer program 52 may be divided into a three-dimensional structure data acquisition unit 41, a graph structure construction unit 42, a subgraph extraction unit 43, and a methylation site prediction unit 44. For specific functions of each unit, please refer to Figure 1 The relevant descriptions in the corresponding embodiments are not repeated here.
[0152] The protein methylation site prediction device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will appreciate that Figure 5 It is only an example of the protein methylation site prediction device 5 and does not constitute a limitation of the protein methylation site prediction device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the protein methylation site prediction device may also include input and output devices, network access devices, buses, etc.
[0153] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0154] The memory 51 may be an internal storage unit of the protein methylation site prediction device 5, such as a hard disk or memory of the protein methylation site prediction device 5. The memory 51 may also be an external storage device of the protein methylation site prediction device 5, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the protein methylation site prediction device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the protein methylation site prediction device 5. The memory 51 is used to store computer programs and other programs and data required by the protein methylation site prediction device. The memory 51 may also be used to temporarily store data that has been output or is to be output.
[0155] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned protein methylation site prediction method can be implemented.
[0156] The embodiment of the present application provides a computer program product. When the computer program product is run on a protein methylation site prediction device, the protein methylation site prediction device can implement the above-mentioned protein methylation site prediction method when executed.
[0157] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0158] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0159] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0160] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for predicting protein methylation sites, characterized in that: The method comprises: Acquiring three-dimensional structural data of a protein to be tested, wherein the protein is composed of a plurality of amino acids; According to the three-dimensional structure data, a graph structure is constructed with amino acids as nodes, wherein each node in the graph structure corresponds to an amino acid; Extracting a subgraph from the graph structure according to preset parameters, wherein the preset parameters include a subgraph hop count and a distance threshold; Inputting the subgraph into a trained graph convolutional neural network to predict methylation sites; The extracting a subgraph from the graph structure according to preset parameters includes: According to the protein sequence, any amino acid in the protein sequence is used as the starting node of the subgraph; Taking the starting node as the center, extracting a subgraph of the amino acid corresponding to the starting node from the graph structure according to the preset parameters; The step of constructing a graph structure with amino acids as nodes according to the three-dimensional structure data comprises: Get the distance between the central carbon atoms of any two amino acids; When the distance value is less than the distance threshold, it is determined that the arbitrary two amino acids are connected; Before inputting the subgraph into the trained graph convolutional neural network for methylation site prediction, the method further includes: The trained graph convolutional neural network is subjected to Bayesian optimization to optimize the subgraph hop count and distance threshold used in subgraph extraction.
2. The protein methylation site prediction method according to claim 1, characterized in that: The subgraph is composed of an adjacency matrix A and a feature matrix X. The adjacency matrix A represents a matrix of adjacent relationships between vertices, A∈(0,1) L×L , L is the number of amino acid nodes in the subgraph; The feature matrix X is composed of the feature matrix X o and the feature matrix X w The feature matrix X is composed of o The feature matrix X is obtained based on the sequence position features of the amino acid nodes obtained by one-hot encoding. w It is a feature matrix obtained based on the sequence position features of the amino acid nodes obtained by word vectors.
3. The protein methylation site prediction method according to claim 2, characterized in that: In the extracting the subgraph from the graph structure according to the preset parameters, the step includes: The feature matrix X is mapped by the rectified linear unit ReLU function o and the feature matrix X w A learnable nonlinear mapping is performed to obtain the feature matrix X, and the ReLU mapping function is: X=Relu(X w W w +X o W o +b X ) Where X∈R L×C , L is the number of amino acid nodes in the subgraph, C is the feature dimension of the subgraph nodes; W w Denotes the feature matrix X w The weight matrix, W o Denotes the feature matrix X o The weight matrix, b represents the bias, b X is the hyperparameter of the ReLU mapping function, where b X ∈R L×C .
4. The protein methylation site prediction method according to claim 2, characterized in that: Inputting the subgraph into a trained graph convolutional neural network for methylation site prediction includes: Performing a global pooling operation on the feature matrix output by each convolutional layer in the graph convolutional neural network; The output of the pooling layer after the global pooling operation is calculated using a fully connected layer and a ReLU activation function.
5. A protein methylation site prediction device, characterized in that: The device comprises: A three-dimensional structure data acquisition unit, used to acquire three-dimensional structure data of a protein to be tested, wherein the protein is composed of a plurality of amino acids; A graph structure construction unit, used to construct a graph structure with amino acids as nodes according to the three-dimensional structure data, wherein each node in the graph structure corresponds to an amino acid; A subgraph extraction unit, configured to extract a subgraph from the graph structure according to preset parameters, wherein the preset parameters include a subgraph hop count and a distance threshold; A methylation site prediction unit, used for inputting the subgraph into a trained graph convolutional neural network to perform methylation site prediction; The extracting a subgraph from the graph structure according to preset parameters includes: According to the protein sequence, any amino acid in the protein sequence is used as the starting node of the subgraph; Taking the starting node as the center, extracting a subgraph of the amino acid corresponding to the starting node from the graph structure according to the preset parameters; The step of constructing a graph structure with amino acids as nodes according to the three-dimensional structure data comprises: Get the distance between the central carbon atoms of any two amino acids; When the distance value is less than the distance threshold, it is determined that the arbitrary two amino acids are connected; Before inputting the subgraph into the trained graph convolutional neural network for methylation site prediction, the method further includes: The trained graph convolutional neural network is subjected to Bayesian optimization to optimize the subgraph hop count and distance threshold used in subgraph extraction.
6. A protein methylation site prediction device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the protein methylation site prediction method according to any one of claims 1 to 4 when executing the computer program.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for predicting protein methylation sites according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Prediction method for protein post-translational modification methylation loci
CN105893787A
Protein phosphorylation site prediction method based on inner product self-attention neural network
CN113096722A