Protein DNA binding tendency prediction method based on multi-modal biomolecular model

The multimodal biomolecular model predicts the tendency of protein DNA binding, and the EGNN model is constructed using the AlphaFold3 and DeepPBS training sets to generate the PWM matrix of protein DNA binding tendency, which solves the shortcomings of the existing methods in terms of calculation cost and identification accuracy, and achieves efficient and accurate prediction.

CN120356510APending Publication Date: 2025-07-22GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510318608.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing methods for predicting protein DNA binding tendency are insufficient in terms of calculation cost and identification accuracy, and cannot meet the requirements of practical applications.

Method used

The prediction method based on multimodal biomolecular model is adopted, and the three-dimensional structure of proteins is predicted using the AlphaFold3 tool, combined with DSSP, Shrake-Rupley algorithm and InterProScan tool to obtain the secondary structure and functional annotation information of proteins, extract features through the multimodal protein language model ESM3, construct an isovariant neural network EGNN model, and trained using the DeepPBS training set. Finally, the PWM matrix of protein DNA binding tendency is generated through the convolutional neural network.

Benefits of technology

It reduces calculation costs, improves the accuracy and efficiency of predicting protein DNA binding tendency, enhances user convenience, obtains more constraints and evolutionary information, and improves the accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356510A_ABST
    Figure CN120356510A_ABST
Patent Text Reader

Abstract

The invention discloses a prediction method, and particularly relates to a protein DNA binding tendency prediction method based on a multi-modal biomolecular model, which comprises the following steps: S1, inputting a protein sequence; s2, predicting a three-dimensional structure; s3, secondary structure calculation is carried out; s4, calculating the accessibility of the solvent; s5, performing protein function annotation; s6, generating feature representation; step S7, sample representation; s8, obtaining a training set; s9, generating all protein residue samples; step S10, constructing a protein map; step S11, building a graph neural network; step S12, adjusting through a convolutional layer; step S13, training and storing the model; and step S14, outputting a PWM matrix. The invention relates to a protein DNA binding tendency prediction method based on a multi-modal biomolecular model, which can reduce the calculation cost of the existing protein DNA binding tendency prediction method and improve the recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics, pattern recognition, and computer applications, and particularly relates to a method for predicting protein-DNA binding propensity based on a multimodal biomolecular model. Background Art

[0002] The ligand-binding specificity between proteins and DNA is an important topic in gene regulation research. Protein transcription factors (TFs) regulate gene expression by binding to specific DNA sequences, and this process plays a key role in cell growth, differentiation, and disease occurrence. Predicting protein-DNA binding specificity, especially the binding pattern described by a position weight matrix (PWM), can effectively reveal the mechanism of gene expression regulation and provide new ideas for disease research. The PWM matrix for predicting protein-DNA binding can not only help to reveal the function of specific transcription factors but also provide valuable information for genomics research. By systematically analyzing the binding specificity of transcription factors, new transcription factors and their target genes can be discovered, thus providing a theoretical basis for understanding the molecular mechanism of diseases and having very important biological significance.

[0003] Currently, the methods specifically used for predicting protein-DNA binding propensity are as follows: Mitra R, Li J, Sagendorf JM, et al. DeepPBS: Geometric deep learning for interpretable prediction of protein-DNA binding specificity[J]. bioRxiv, 2023: 2023.12. 15.571942. (DeepPBS is a geometric deep learning method for interpretable prediction of protein-DNA binding specificity). Aizenshtein-Gazit S, Orenstein Y. Deepzf: improved DNA-binding prediction of c2h2-zinc-finger proteins by deep transfer learning[J]. Bioinformatics, 2022, 38(Supplement_2): ii62-ii67. (Deepzf improves the DNA-binding prediction of c2h2-zinc-finger proteins through deep transfer learning). Wetzel JL, Zhang K, Singh M. Learning probabilistic protein–DNA recognition codes from DNA-binding specificities using structural mappings[J]. Genome Research, 2022, 32(9): 1776-1786. (rCLAMPS learns probabilistic protein-DNA recognition codes from DNA-binding specificities using structural mappings).

[0004] DeepPBS captures the physicochemical and geometric context of protein-DNA interactions to predict protein-DNA binding specificity. Deepzf learns the recognition code of ZF-DNA binding through transfer learning, which is used to predict the binding ZF and its DNA-binding preference only given the amino acid sequence of the C2H2-ZF protein. rCLAMPS infers the DNA nucleotide positions contacted by each amino acid when the transcription factor binds to DNA through the three-dimensional structure data of protein-DNA contacts, and also speculates on the preference of these amino acids for specific nucleotides through statistical learning.

[0005] In summary, there is still a large gap between the existing methods for predicting protein-DNA binding propensity and the requirements of practical applications in terms of computational cost and recognition accuracy, and there is an urgent need for improvement. Summary of the Invention

[0006] To overcome the deficiencies of existing protein-DNA binding propensity prediction methods in terms of computational cost and recognition accuracy, the present invention proposes a protein-DNA binding propensity prediction method based on a multimodal biomolecular model with low computational cost and high recognition accuracy.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: A protein-DNA binding propensity prediction method based on a multimodal biomolecular model, comprising the following steps: S1. Input a protein sequence information of a protein to be predicted for protein-DNA binding propensity with a residue number of , denoted as ; S2. For the protein , use the AlphaFold3 tool to predict its three-dimensional structure information to obtain a CIF file, and represent its atomic coordinates as a matrix of size , denoted as , where is the number of atoms of each residue, and is the three-dimensional spatial coordinates of each atom; S3. For the CIF file of the protein obtained in step S2, use the DSSP tool to calculate the eight-state secondary structure of the protein , where represents the secondary structure state of the P th residue of the protein ; S4. Use the Shrake-Rupley algorithm to calculate the solvent-accessible surface area of the CIF file of the protein obtained in step S2. The non-negative real value represents the solvent-accessible surface area of the th residue of the protein ; S5. Search the Pfam, PROSITE, and CATH databases for the protein sequence using the InterProScan tool to obtain a protein function annotation information file , where represents the th functional annotation information in , represents the total number of functional annotation information in Each functional annotation information and the corresponding functional intervals , where ; S6. Input all the features obtained in steps S2 to S5, that is , , , and input them into the multi-modal protein language model ESM3 to obtain an embedding matrix of size as the feature representation of the protein , where is the number of residues in the protein , and is the feature dimension of each residue; S7. The sample corresponding to any residue in the protein sequence is represented as , , and respectively represent the feature vector of the th residue and the protein-DNA binding propensity label information. A, T, G, and C are used to represent the probabilities of binding specific bases at this DNA site; S8. Obtain its training set from DeepPBS. Extract the protein sequences in the DeepPBS protein-DNA complex training set. The total number of sequences is denoted as . The label is the PWM matrix corresponding to the protein sequence in the DeepPBS training set; the new training set is denoted as , , represents the i th protein sequence in the training set D, represents 's label information, is the total number of protein sequences; S9. According to steps S2 to S7, generate the residue sample set of all proteins . The feature matrix and label vector of each protein sample are denoted as and , ; S10. According to the protein structure atom coordinate matrix generated in step S2 and the feature matrix generated in step S9, construct the protein graph , where represents the set of residue nodes. Any node is characterized by a length of feature vector , denoted as the th and th contact residue pairs define edges, and residue pairs with an Euclidean distance less than 14 Å are considered in contact. is the center carbon atom coordinate set; S11. Construct an equivariant graph neural network (EGNN) model, denoted as EGNN, which consists of three layers of equivariant graph convolutional layers, each layer denoted as EGCL. The protein graph obtained in step S10 is input into the EGCL to perform the following transformation:

[0008]

[0009]

[0010]

[0011] where and represent edge and node operations based on a multi-layer perceptron, represents taking the average, converts to a scalar. The set of node features output by each layer of EGCL and the coordinate set are used as the input for the next layer of EGCL. The output of the last layer of EGCL is matrix, and the output linear layer maps to the final output dimension , is , indicating that each node will have output classes; S12. Pass the matrix obtained in step S11 into a convolutional neural network (CNN), which undergoes two layers of convolutional operations and adaptive pooling, and finally obtains matrix through the output layer. is the length of the label , corresponding to classes at each position. The output at each position is converted into a probability distribution through the softmax function to obtain the predicted PWM matrix; S13. Input the protein residue sample set Ω constructed in step S9 into the EGNN model constructed in step S11 for training. Use the cross-entropy function as the loss, AdamW as the parameter optimizer of the model, and use the Dropout algorithm with a dropout rate of 0.1 and L2 regularization with a λ parameter of 0.1 to prevent the model from overfitting; S14. For the protein Generate the protein graph and the feature vectors of the residues through steps S2 - S10, and input them into the EGNN model trained in step S13 to obtain the predicted PWM matrix of the protein DNA binding propensity.

[0012] The technical concept of the present invention is as follows: First, for the protein sequence with the input residue number to be predicted for the protein DNA binding propensity, successively use the AlphaFold2, DSSP, Shrake - Rupley, and InterProScan programs to obtain the protein structure PDB file, the protein secondary structure DSSP file, the protein solvent accessibility SS, and the InterPro functional annotation file; then, input the above feature files into the multi - modal protein language model ESM3 to obtain a feature matrix F and convert it into a feature tensor; Secondly, process the protein sequence into residue samples; then, construct an equivariant graph neural network model based on the protein structure diagram, and use the training set obtained from the known DeepPBS to train the constructed network. Finally, input the protein sequence sample to be predicted for the protein DNA binding propensity into the trained model to obtain the PWM matrix of the protein DNA binding propensity.

[0013] The present invention has the following beneficial effects and advantages: (1) Use the newly released AlphaFold3 tool to predict its three - dimensional structure information to obtain a CIF file, so that the PWM matrix of the protein DNA binding propensity can be obtained only by inputting the protein sequence, increasing the convenience for users.

[0014] (2) Extract multi - modal information of the protein's sequence, structure, and functional annotation, obtain more useful constraints and evolutionary information, and prepare for improving the accuracy of protein DNA binding propensity prediction.

[0015] (3) Use the equivariant graph neural network to maintain the rotational invariance of the protein graph and learn the spatial constraint information of neighboring residues. Further improve the efficiency and accuracy of protein DNA binding propensity prediction. Description of the Drawings

[0016] Figure 1Schematic diagram of a method for predicting protein-DNA binding propensity based on a multimodal biomolecular model; Figure 2 Graph for predicting the protein-DNA binding propensity of Protein 1 using the method for predicting protein-DNA binding propensity based on a multimodal biomolecular model a 0 a_A Detailed implementation manners

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] A method for predicting protein-DNA binding propensity based on a multimodal biomolecular model, referring to Figure 1 and Figure 2 , includes the following steps: S1. Input the protein sequence information of a protein to be predicted for protein-DNA binding propensity with a residue number of , denoted as ; S2. For the protein , use the AlphaFold3 tool to predict its three-dimensional structure information to obtain a CIF file, and represent its atomic coordinates as a matrix of size , denoted as , where is the number of atoms of each residue, and is the three-dimensional spatial coordinates of each atom; S3. For the CIF file of the protein obtained in step S2, use the DSSP tool to calculate the eight-state secondary structure of the protein , where represents the secondary structure state of the th residue of the protein ; S4. Use the Shrake-Rupley algorithm to calculate the solvent-accessible surface area of the CIF file of the protein obtained in step S2. The non-negative real value represents the solvent-accessible surface area of the th residue of the protein ; S5. For the protein ​The sequences were searched against the Pfam, PROSITE, and CATH databases using the InterProScan tool to obtain a protein function annotation information file. , where denotes the th functional annotation information in , and denotes the total number of functional annotation information in . Each piece of functional annotation information contains a protein function annotation text ; S6. Input all the features obtained in steps S2 to S5, that is, , , , and into the multi-modal protein language model ESM3 to obtain an embedding matrix of size as the feature representation of the protein , where is the number of residues in the protein , and is the feature dimension of each residue; S7. The sample representation corresponding to any residue in the protein sequence is denoted as , , and respectively denote the feature vector of the th residue and the protein-DNA binding propensity label information. The probabilities of binding specific bases at this DNA site are represented by A, T, G, and C; S8. Obtain its training set from DeepPBS. Extract the protein sequences in the DeepPBS protein-DNA complex training set. The total number of sequences is denoted as , and the label is the PWM matrix corresponding to the protein sequence in the DeepPBS training set.

[0019] The new training set is denoted as , , denotes the i th protein sequence in the training set D, denotes 's label information, is the total number of protein sequences; S9. Generate all proteins according to steps S2 to S7 Set of residue samples For each protein sample, the feature matrix and label vector are denoted as and , ; S10. Construct a protein graph from the protein structure atom coordinate matrix generated in step S9 and the feature matrix , where represents the set of residue nodes, and the feature of any node is a feature vector of length . is defined as an edge by the contact residue pair of the -th and -th residues. Residue pairs with an Euclidean distance less than 14 Å are considered to be in contact, is the center of each residue carbon atom coordinate set; S11. Construct an equivariant graph neural network model, denoted as EGNN, which consists of three equivariant graph convolutional layers, each layer denoted as EGCL. Input the protein graph obtained in step S10 into EGCL and perform the following transformation:

[0020]

[0021]

[0022]

[0023] where and represent edge and node operations based on a multi-layer perceptron, represents taking the average, converts to a scalar. The node feature set and coordinate set output by each layer of EGCL are used as the input for the next layer of EGCL. The output of the last layer of EGCL is matrix, and the output linear layer maps to the final output dimension . is , indicating that each node will have output classes; S12. Use the result ​The matrix is passed to a convolutional neural network (CNN), undergoes two layers of convolutional operations and adaptive pooling, and finally obtains the matrix is the label length of corresponding to each position categories, and the output of each position is converted into a probability distribution through the softmax function to obtain the predicted PWM matrix; S13. Input the set of protein residue samples Ω constructed in step S9 into the EGNN model constructed in step S11 for training. Use the cross-entropy function as the loss, AdamW as the parameter optimizer of the model, and use the Dropout algorithm with a dropout rate of 0.1 and L2 regularization with a λ parameter of 0.1 to prevent the model from overfitting; S14. The protein generates the protein graph and the feature vectors of the residues through steps S2 - S10, and inputs them into the EGNN model trained in step S13 to obtain the predicted PWM matrix of the protein DNA binding tendency.

[0024] This embodiment takes the prediction of the protein DNA binding tendency of protein as an example. A method for predicting the protein DNA binding tendency based on a multimodal biomolecular model includes the following steps: S1. Input a protein with a residue number of for predicting the protein DNA binding tendency, denoted as ; S2. For the protein , use the AlphaFold3 tool to predict its three-dimensional structure information to obtain a CIF file, and represent its atomic coordinates as a matrix of size , denoted as , where is the number of atoms per residue, is the three-dimensional spatial coordinate of each atom; S3. For the CIF file of the protein obtained in step S2, use the DSSP tool to calculate the eight-state secondary structure of the protein , where represents the secondary structure state of the th residue of the protein ; S4. Use the Shrake-Rupley algorithm to calculate the solvent-accessible surface area of the CIF file of the protein , non - negative real value represents the solvent - accessible surface area of the th residue of the protein; S5. Use the InterProScan tool to search the Pfam, PROSITE, and CATH databases for the sequence of the protein to obtain a protein function annotation information file , where represents the th functional annotation information in , and represents the total number of functional annotation information in . Each functional annotation information contains a protein function annotation text and the corresponding functional interval ; S6. Input all the features obtained in steps S2 to S5, that is , , , and into the multi - modal protein language model ESM3 to obtain an embedding matrix of size as the feature representation of the protein , where is the number of residues in the protein , and is the feature dimension of each residue; S7. The sample representation corresponding to any residue in the protein sequence is denoted as , , , and represent the feature vector of the th residue and the protein - DNA binding propensity label information respectively. The probability of the DNA site binding to a specific base is represented by A, T, G, C; S8. Obtain its training set from DeepPBS. Extract the protein sequences in the DeepPBS protein - DNA complex training set. The total number of sequences is denoted as , and the label is the PWM matrix corresponding to the protein sequence in the DeepPBS training set. The new training set is denoted as , , represents the th protein sequence in the training set D, represents The label information is the total number of protein sequences; S9. Generate all proteins according to steps S2 to S7 to obtain a residue sample set , and the feature matrix and label vector of each protein sample are denoted as and , ; S10. Based on the protein structure atom coordinate matrix generated in step S2 and the feature matrix generated in step S9, construct a protein graph , where represents the set of residue nodes, and the feature of any node is a feature vector with a length of , , is defined as an edge by the contact residue pair of the th and th residues. Residue pairs with an Euclidean distance less than 14 Å are considered to be in contact. is the center of each residue carbon atom coordinate set; S11. Construct an equivariant graph neural network model, denoted as EGNN, which consists of three layers of equivariant graph convolutional layers, each layer denoted as EGCL. Input the protein graph obtained in step S10 into EGCL to perform the following transformation:

[0025]

[0026]

[0027]

[0028] where and represent edge and node operations based on a multi-layer perceptron, represents taking the average, converts into a scalar. The node feature set and coordinate set output by each layer of EGCL are used as the input for the next layer of EGCL. The output of the last layer of EGCL is matrix. The output linear layer maps to the final output dimension , is , indicating that each node will have output categories, where is the th element in S12. Transfer the matrix obtained in step S11 to a convolutional neural network (CNN), undergo two layers of convolutional operations and adaptive pooling, and finally obtain the matrix through the output layer. is the length of the label , corresponding to the categories at each position. Convert the output at each position into a probability distribution through the softmax function to obtain the predicted PWM matrix. S13. Input the set of protein residue samples Ω constructed in step S9 into the EGNN model constructed in step S11 for training. Use the cross-entropy function as the loss, AdamW as the parameter optimizer for the model, and use the Dropout algorithm with a dropout rate of 0.1 and L2 regularization with a λ parameter of 0.1 to prevent the model from overfitting. S14. Pass the protein through steps S2 - S10 to generate the protein graph and the feature vectors of the residues, and input them into the EGNN model trained in step S13 to obtain the predicted PWM matrix of the protein-DNA binding tendency.

[0029] Taking the prediction of the protein-DNA binding tendency of the protein as an example, the DNA binding tendency prediction of the protein is obtained by using the above method as shown in Figure 2 .

[0030] The above description is the classification result obtained by taking the prediction of the protein-DNA binding tendency of the protein as an example in the present invention, and does not limit the scope of implementation of the present invention.

[0031] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting the protein-DNA binding propensity based on a multimodal biomolecular model, characterized in that, Including the following steps: S1. Input the protein sequence information to be predicted for protein-DNA binding tendency, denoted as ; S2. Use the AlphaFold3 tool to predict the three-dimensional structural information of the protein and generate relevant feature data; S3. Extract the protein 's functional annotation information and structural information; S4. Input the multi-modal features of the protein into a multi-modal protein language model to generate an embedding matrix of the protein ; S5. Construct a residue-based protein graph ; S6. Process the predicted feature embedding to generate a predicted PWM matrix; S7. Use the trained model to predict the binding tendency of the protein and obtain the final PWM matrix.

2. The method for predicting the protein-DNA binding propensity based on the multimodal biomolecular model according to claim 1, wherein In the step S2, the protein sequence is input, and the AlphaFold3 tool is used to predict its three-dimensional structure information to obtain a CIF file, and its atomic coordinates are represented as a matrix .

3. The method for predicting the protein-DNA binding propensity based on a multimodal biomolecular model according to claim 1, wherein In the step S3, the DSSP tool is used to calculate the secondary structure information of the protein , and the Shrake-Rupley algorithm is used to calculate the solvent accessibility , and the InterProScan tool is used to extract the functional annotation information .

4. The method for predicting the protein-DNA binding propensity based on the multimodal biomolecular model according to claim 1, wherein In the step S5, according to the protein structural information and the feature matrix, a protein graph of residue nodes is constructed , where represents the residue node set, is the edge set of residue pairs, is the central carbon atom coordinate.

5. The method for predicting the protein-DNA binding propensity based on the multi-modal biomolecular model according to claim 1, wherein, In the step S6, an equivariant graph neural network is used to perform multi-layer feature extraction on the protein graph to obtain the prediction result of the protein binding site.

6. The method for predicting the protein-DNA binding propensity based on the multi-modal biomolecular model according to claim 1, wherein In step S7, the model is trained using a cross-entropy loss function, an AdamW optimizer, and combined with Dropout and L2 regularization methods to prevent model overfitting.

7. The method for predicting the protein-DNA binding propensity based on a multimodal biomolecular model according to claim 1, wherein In step S7, the PWM matrix is converted into a probability distribution through the Softmax function, representing the binding probabilities of the four bases at each DNA binding site.

Citation Information

Cited By

  • Protein-DNA binding site prediction method based on graph neural network and RBF-gating mechanism

    CN122245413A