Drug target binding affinity prediction method based on multi-modal diagram learning driven by large model
Through a multimodal graph learning method driven by a large model, combined with ESM-GearNet and KAN neural network, the shortcomings of existing drug target binding affinity prediction methods in terms of calculation methods and regression accuracy are solved, and more efficient and accurate drug target binding affinity prediction is achieved.
Patent Information
- Application Number
- CN202510212593.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
The existing drug target binding affinity prediction methods have not yet met the practical application requirements in terms of calculation methods and regression accuracy, and improvements are needed to improve prediction efficiency and accuracy.
Using a multimodal graph learning method driven by a large model, the protein sequence is characterized by ESM-GearNet pre-trained large model, and the sequence information and structural information of proteins and molecules are combined, and the KAN neural network is used to predict drug target binding affinity.
It improves the efficiency and accuracy of drug target binding affinity prediction, provides a richer and comprehensive multimodal information expression, and enhances the utilization of features and the prediction of results.
Smart Images

Figure CN120048335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of bioinformatics, artificial intelligence, and drug discovery. More specifically, it particularly relates to a method for predicting the binding affinity of drug targets based on large model-driven multi-modal graph learning. Background Art
[0002] Drug targets are the core objects that act in the body. The characteristics of the targets directly determine the efficacy and selectivity of drugs, and the binding affinity between drug targets is an important indicator for measuring their biological activity. With the rapid accumulation of high-throughput screening technologies and large-scale biological data, methods for predicting the binding affinity of molecules and targets from molecular and target sequence and structure information have rapidly emerged, and the computing power and algorithms of artificial intelligence have been updated rapidly. Therefore, accurately predicting the binding affinity of drug targets is helpful for the design of new drugs, the promotion of personalized medicine, the facilitation of drug repositioning, and the improvement of the success rate of new drug invention, which has very important significance.
[0003] Currently, the methods specifically for predicting the affinity of drug targets are as follows: H, A, Ozkirimli E. DeepDTA: deep drug–target binding affinity prediction[J]. Bioinformatics, 2018, 34(17): i821-i829. That is: H et al. Nguyen T, Le H, Quinn TP, et al. GraphDTA: predicting drug–target binding affinity with graph neural networks[J]. Bioinformatics, 2021, 37(8): 1140-1147. That is, Nguyen T et al. DGraphDTA: Huang X, Yang Y, Wang Y, et al. Jiang M, Li Z, Zhang S, et al. Drug–target affinity prediction using graph neural network and contact maps[J]. RSC advances, 2020, 10(35): 20701-20712. That is, Huang X et al. Nguyen TM, Nguyen T, Le TM, et al. Gefa: early fusion approach in drug-target affinity prediction[J]. IEEE / ACM transactions on computational biology and bioinformatics, 2021, 19(2): 718-728. That is, Nguyen TM et al. Thafar MA, Alshahrani M, Albaradei S, et al. Affinity2Vec: drug-target binding affinity prediction through representation learning, graph mining, and machine learning[J]. Scientific reports, 2022, 12(1): 4751. That is, Thafar MA et al. DeepDTA began to predict drug-target affinity from the traditional machine learning approach to the deep learning method. DGraphDTA further explored the protein structure on the basis of DeepDTA, obtained the protein graph through the contact map, and successfully combined the drug sequence features and protein structure features. Gefa put forward the view of early feature fusion, fused the early features and late features of drugs together, and captured the features before and after the graph structure changed.Affinity2Vec proposes a method based on graph learning and machine learning. By integrating drug-drug similarity, target-target similarity, and drug-target binding affinity data, a weighted heterogeneous graph is constructed for prediction.
[0004] In summary, there is still a large gap between the existing drug-target binding affinity prediction methods and the requirements of practical applications in terms of calculation methods and regression accuracy, and there is an urgent need for improvement. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for predicting drug-target binding affinity based on large model-driven multi-modal graph learning to overcome the defects of the existing technology.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A method for predicting drug-target binding affinity based on large model-driven multi-modal graph learning, comprising the following steps:
[0008] S1. Input a protein sequence information T and a drug SMILES sequence information S with a residue number of L 1 and a length of L 2 for drug-target binding prediction;
[0009] S2. Use the ESM-GearNet pre-trained large model to perform feature representation on the protein sequence information T. ESM-GearNet uses the protein language model ESM combined with the GearNet graph neural network to perform multi-view pre-training using a feature serial fusion strategy to generate a feature matrix F 1 ;
[0010] S3. Download the pdb file corresponding to the protein sequence information T in the protein prediction structure database, and represent the protein structure with a graph G T =(V T , E T ), where represents the set of all residue nodes of the protein sequence, represents the edge set of interacting residues;
[0011] S4. Perform one-hot encoding representation on the protein sequence information T with a residue number of L 1 for drug-target binding prediction. Each residue generates a corresponding 20-dimensional vector representation, the category to which the vector belongs is 1, and the rest are 0. In addition, 3-dimensional physical and chemical properties, 3-dimensional secondary structure, 1-dimensional B factor, 1-dimensional sequence position, and 1-dimensional ASA features are added. Finally, the feature dimension of each residue is 29 dimensions, and the feature matrix F2 Size: L1×29;
[0012] S5. In the protein graph obtained in step S3, residues nodes with close spatial distances are defined as the edges of the protein graph;
[0013] S6. Concatenate the feature matrices F 1 and the feature matrix F 2 to obtain the final residue representation matrix with a size of L 1 ×4381, where each residue node is represented by a vector f i of length 4381;
[0014] S7. Perform truncation and padding processing on the protein sequence information T in step S1 to obtain the processed sequence T T ;
[0015] S8. Input the graph G T =(V T , E T ) constructed in step S3 into the network GinConvNet composed of three layers of GinCon to perform the following transformation on the residue representation matrix F:
[0016]
[0017] In the formula, represents the feature vector of node v in the k-th layer, represents the feature vector of node v in the (k - 1)-th layer, represents the set of neighbor nodes of node v, 1+∈ represents the importance of node v, MLP represents a multi-layer perceptron, and perform global_max_pool processing on the output of the last layer to obtain the 4381-dimensional feature M 1 ;
[0018] S9. Pass the processed sequence T T in step S7 through a neural network stacked with one Embedding layer and three CNN layers for feature extraction. After processing, generate a vector of size 1000×128, and perform max_pool operation on this vector to obtain the 128-dimensional feature M 2 ;
[0019] S10. Use the chemical informatics tool library RDKit in Python to transform the drug SMILES sequence information S in step S1 to obtain the molecular graph structure G D =(V D , E D ), where V DDenote the set of all atomic nodes in the molecular graph as F D Denote the set of all connected edges in the molecular graph;
[0020] S11. Perform one-hot encoding on each atom in the drug SMILES sequence information S, calculate the connectivity of the atom, calculate the total number of hydrogen atoms on the atom, calculate the number of implicit valence electrons of the atom, and determine whether it is part of an aromatic ring, to obtain a feature matrix of size L 2 ×78, and use it as the atomic representation matrix E;
[0021] S12. For the molecular graph structure G D =(V D , E D ), process it in the manner of step S8 to obtain a 78-dimensional feature M 3 ;
[0022] S13. Use the custom function get_selfies to transform the drug SMILES sequence information S to generate a molecular representation SELFIES string S F ;
[0023] S14. Use the tokenizer to segment the SELFIES string S F , and map each word using a dictionary, with the mapping range between 1 and 25, and control the generated digital sequence within 100. For those with a length less than 100, pad with 0, and for those exceeding 100, truncate. After processing, denote it as S T ;
[0024] S15. For the obtained S T , process it in the manner of step S9 to obtain a 128-dimensional feature M 4
[0025] S16. Concatenate the feature M 1 obtained in step S8, the feature M 2 obtained in step S9, the feature M 3 obtained in step S12, and the feature M 4 obtained in step S15 to obtain the binding feature M of the drug molecule and the protein target;
[0026] S17. Use the KAN neural network as the prediction model, and use the binding feature M obtained in step S16 as the input in the KAN neural network for transformation;
[0027] S18. Denote the steps S1 - S17 as EGK - DTA, implement it using the Pytorch and DGL frameworks in python, set the batch_size to 8, use the mean squared error loss as the loss function, train for 100 times, and optimize using the Adam function with a learning rate of 0.0005;
[0028] S19. Input the protein sequence information T and the drug SMILES sequence information S into EGK - DTA to obtain the sequence representations and structural representations of the protein and the drug, fuse them together and input them into the KAN neural network, and then output the binding affinity between the protein and the drug small molecule.
[0029] Furthermore, in step S2, the feature matrix F 1 has a size of L 1 ×4352, and use the library function numpy in python to store the feature matrix F 1 for storage.
[0030] Furthermore, in step S4, encoding is performed according to the residue class.
[0031] Furthermore, step S5 is specifically defined as that the Euclidean distance between the C α atoms of two residues is less than is defined as having contact.
[0032] Furthermore, the transformation formula for step S17 is:
[0033]
[0034] where φ q,p (x p ) belongs to the internal function, and Φ q (x q ) belongs to the external function.
[0035] Compared with the prior art, the advantages of the present invention are as follows: ESM - GearNet in the present invention adopts the protein language model ESM and combines the pre - trained model of graph learning, which makes the representation of proteins more abundant and provides more useful information for subsequent predictions; the present invention combines the sequence information and structural information of proteins and molecules, and the fusion expression of multi - modal information is more rich and comprehensive; the present invention utilizes the latest KAN model, which further enhances the utilization of features and the prediction of results, improving the efficiency and accuracy of drug - target binding affinity prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 It is a flowchart of the method for predicting the drug-target binding affinity based on large model-driven multi-modal graph learning of the present invention.
[0038] Figure 2 It is a graph after the present invention predicts the drug-target binding affinity of the drug CHEMBL373751 and the protein P53350. Detailed implementation manners
[0039] The following elaborates on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0040] Refer to Figure 1 As shown, this embodiment discloses a method for predicting the drug-target binding affinity based on large model-driven multi-modal graph learning, including the following steps:
[0041] Step S1: Input a protein sequence information T with a residue number of L 1 and a drug SMILES (Simplified Molecular Input Line Entry System) sequence information S with a length of L 2 for drug-target binding prediction.
[0042] Step S2: Use the ESM-GearNet pre-trained large model to perform feature representation on the protein sequence information T. ESM-GearNet uses the protein language model ESM combined with the GearNet graph neural network to perform multi-view pre-training using the feature serial fusion strategy, generating a feature matrix F 1 , and use the library function numpy in python to store F 1 .
[0043] Step S3: Download the pdb file corresponding to the protein sequence information T in the protein prediction structure database (AlphaFold Protein Structure Database) composed of the protein structure predictions of about 350,000 proteins by the AlphaFold artificial intelligence system, and represent the protein structure with a graph G T =(V T , ET ) is represented, where represents the set of all residue nodes of the protein sequence, represents the edge set of interacting residues.
[0044] Step S4: Perform one-hot encoding on the protein sequence information T with L 1 residues to be predicted for drug target binding. Each residue generates a corresponding 20-dimensional vector representation, where the category to which the vector belongs is 1 and the rest are 0. Additionally, 3-dimensional physicochemical properties (hydrophobicity, molecular row, and polarity), 3-bit secondary structures (α-helix, β-sheet, and random coil), 1-dimensional B-factor, 1-dimensional sequence position, and 1-dimensional ASA (Accessible Surface Area) features are added. Finally, the feature dimension of each residue is 29-dimensional, and the feature matrix F 2 has a size of L 1 × 29.
[0045] Step S5: In the protein graph obtained in Step S3, residue nodes with close spatial distances are defined as the edges of the protein graph, that is, the Euclidean distance between the C α atoms of two residues is less than is defined as having contact.
[0046] Step S6: Concatenate the feature matrix F 1 obtained in Step S2 and the feature matrix F 2 to obtain the final residue representation matrix with a size of L 1 × 4381. Each residue node is represented by a vector f i of length 4381.
[0047] Step S7: Perform truncation and padding processing on the protein sequence information T in Step S1 to obtain the processed sequence T T , so that the length of the protein sequence meets the condition of being 1000. For sequences with a length less than 1000, they are padded with 0s. Each character is converted into a corresponding number using a dictionary, and finally the sequence is represented by a string of numbers in the range of 0 to 25.
[0048] Step S8: Input the graph G T = (V T , E T ) constructed in Step S3 into the network GinConvNet composed of three layers of GinCon, and perform the following transformation on the residue representation matrix F:
[0049]
[0050] Feature vector represents the set of neighbor nodes of node v, 1+∈ represents the importance of node v, MLP represents a multi-layer perceptron, and the output of the last layer is processed by global_max_pool to obtain features of 4381 dimensions.
[0051] Step S9: The sequence T after being processed in step S7 T is subjected to feature extraction through a neural network stacked with one Embedding layer and three CNNs (Convolutional Neural Networks). After processing, a vector of size 1000×128 is generated, and a max_pool operation is performed on this vector to obtain features M of 128 dimensions. 2 .
[0052] Step S10: Use the cheminformatics toolkit RDKit in Python to transform the drug SMILES sequence information S in step S1 to obtain the molecular graph structure G D =(V D , E D ), where V D represents the set of all atomic nodes of the molecular graph, and E D represents the set of all connected edges of the molecular graph.
[0053] Step S11: Perform one-hot encoding on each atom of the drug SMILES sequence information S, calculate the connectivity of the atoms, calculate the total number of hydrogen atoms on the atoms, calculate the number of implicit valence electrons of the atoms, and determine whether it is part of an aromatic ring, to obtain a feature matrix of size L 2 ×78, which is used as the atomic representation matrix E.
[0054] Step S12: For the molecular graph structure G D =(V D , E D ), process it in the same way as step S8 to obtain features M of 78 dimensions 3 .
[0055] Step S13: Use the custom function get_selfies to transform the drug SMILES sequence information S to generate a more reliable and effective molecular representation SELFIES string S F .
[0056] Step S14: Use the tokenizer tokenizer to process the SELFIES string S FSegment it, map each word using a dictionary, with the mapping range between 1 and 25, and control the generated digital sequence within 100. For those with a length less than 100, pad with 0. For those exceeding 100, truncate them, and denote the processed result as S T 。
[0057] Step S15: The obtained S T , process it in the way of Step S9 to obtain the 128-dimensional feature M 4 。
[0058] Step S16: The feature M obtained in Step S8 1 , the feature M obtained in Step S9 2 , the feature M obtained in Step S12 3 and the feature M obtained in Step S15 4 are concatenated to obtain the binding feature M of the drug molecule and the protein target.
[0059] Step S17: Use the KAN neural network (Kolmogorov - Arnold Network) as the prediction model. Its core comes from the Kolmogorov - Arnold representation theorem. Take the binding feature M obtained in Step S16 as the input in the KAN neural network for transformation. The transformation formula is as follows:
[0060]
[0061] where, φ q,p (xp) belongs to the internal function, Φ q (x q ) belongs to the external function
[0062] Step S18: Denote Steps S1 - S17 as EGK - DTA, implement it using Python's Pytorch and DGL frameworks, set batch_size to 8, the loss function is the mean square error loss (MSE), the number of training times is 100, and optimize it using the Adam function with a learning rate of 0.0005.
[0063] Step S19: Input the protein sequence information T and the drug SMILES sequence information S into EGK - DTA to obtain the sequence representation and structure representation of the protein and the drug, fuse them together and input them into the KAN neural network, and then output to obtain the binding affinity magnitude between the protein and the drug small molecule.
[0064] This embodiment takes the drug target binding affinity of drug CHEMBL373751 and protein P53350 as an example, and the drug target binding affinity of drug CHEMBL373751 and protein P53350 is calculated through the above steps S1 - S19 as Figure 2 shown.
[0065] Although the embodiments of the present invention are described with reference to the accompanying drawings, the calculation results obtained by taking the prediction of the drug target binding affinity of drug CHEMBL373751 and protein P53350 as an example do not limit the scope of implementation of the present invention. The patent owner can make various deformations or modifications within the scope of the appended claims. As long as it does not exceed the protection scope described in the claims of the present invention, it should be within the protection scope of the present invention.
Claims
1. A drug-target binding affinity prediction method based on multimodal graph learning driven by a large model, characterized in that: The following steps are involved: S1, input a protein sequence information T with a residue number of L1 and a length of L2 to be used for drug-target binding prediction and a drug SMILES sequence information S; S2, using the ESM-GearNet pre-trained large model to represent the protein sequence information T. ESM-GearNet uses the protein language model ESM combined with the GearNet graph neural network to adopt a feature serial fusion strategy for multi-view pre-training to generate a feature matrix F1; S3, download the pdb file corresponding to the protein sequence information T in the protein prediction structure database, and map the protein structure to G T =(V T , E T ) to indicate that Represents the set of all residue nodes in the protein sequence, The set of edges representing the interacting residues; S4, the protein sequence information T to be used for drug-target binding prediction with a residue number of L1 is represented by one-hot encoding, and each residue Generate the corresponding 20-dimensional vector representation, the category to which the vector belongs is 1, and the rest are 0, and add 3-dimensional physicochemical properties, 3-dimensional secondary structure, 1-dimensional B factor, 1-dimensional sequence position and 1-dimensional ASA feature. Finally, each residue The feature dimension is 29, and the feature matrix F2 size is L1×29; S5, in the protein graph obtained in step S3, residue nodes with close spatial distances are defined as edges of the protein graph; S6. Concatenate the feature matrix F1 and the feature matrix F2 obtained in step S2 and step S4 to obtain the final residue representation matrix The size is L1×4381, each residue node are all composed of a vector f of length 4381 i express; S7, truncate and fill the protein sequence information T in step S1 to obtain the processed sequence T T ; S8, the graph G constructed in step S3 T =(V T , E T ) is input into the network GinConvNet composed of three layers of GinCon, and the residue representation matrix F is transformed as follows: In the formula, represents the feature vector of node v in the kth layer, represents the feature vector of node v in the k-1th layer, represents the set of neighbor nodes of node v, 1+∈ represents the importance of node v, MLP represents multi-layer perceptron, and the output of the last layer is processed by global_max_pool to obtain the feature M1 of 4381 dimensions; S9, the sequence T processed in step S7 T Feature extraction is performed through a neural network consisting of one Embedding layer and three CNN layers. After processing, a vector of size 1000×128 is generated. The vector is max_pooled to obtain a 128-dimensional feature M2. S10. Use the chemical informatics tool library RDKit in Python to transform the drug SMILES sequence information S in step S1 to obtain the molecular graph structure G D =(V D , E D ), where V D Represents the set of all atomic nodes in the molecular graph, E D Represents the set of all connected edges of the molecular graph; S11, one-hot encode each atom of the drug SMILES sequence information S, calculate the connectivity of the atom, calculate the total number of hydrogen atoms on the atom, calculate the number of hidden valence electrons of the atom, and determine whether it is part of the aromatic ring, to obtain a feature matrix of size L2×78, which is used as the atomic representation matrix E; S12. Molecular graph structure G D =(V D , E D ), and process according to step S8 to obtain a 78-dimensional feature M3; S13. Use the custom function get_selfies to transform the drug SMILES sequence information S and generate a molecular representation SELFIES string S F ; S14. Use the tokenizer to tokenize the SELFIES string S F The word is segmented and mapped using a dictionary. The mapping range is between 1 and 25, and the generated digital sequence is controlled to 100. For those with a length less than 100, 0 is used for padding, and for those with a length greater than 100, truncation is performed and recorded as S after processing. T ; S15, obtained S T , use the method of step S9 to obtain the 128-dimensional feature M4 S16, combining the feature M1 obtained in step S8, the feature M2 obtained in step S9, the feature M3 obtained in step S12, and the feature M4 obtained in step S15 to obtain a binding feature M between the drug molecule and the protein target; S17, using the KAN neural network as a prediction model, and transforming the combined feature M obtained in step S16 as an input in the KAN neural network; S18, steps S1-S17 are recorded as EGK-DTA, implemented using Python's Pytorch and DGL framework, batch_size is set to 8, the loss function is the mean square error loss, the number of training times is 100, and the Adam function with a learning rate of 0.0005 is used for optimization; S19. Input the protein sequence information T and the drug SMILES sequence information S into EGK-DTA to obtain the sequence representation and structural representation of the protein and the drug, and fuse them together and input them into the KAN neural network, and then output the binding affinity between the protein and the drug small molecule.
2. The drug-target binding affinity prediction method based on large model-driven multimodal graph learning according to claim 1, characterized in that: The size of the feature matrix F1 in step S2 is L1×4352, and the feature matrix F1 is stored using the library function numpy in python.
3. The drug-target binding affinity prediction method based on large model-driven multimodal graph learning according to claim 1, characterized in that: In the step S4, encoding is performed according to the residue category.
4. The drug-target binding affinity prediction method based on large model-driven multimodal graph learning according to claim 1, characterized in that: The step S5 specifically comprises converting the C α The Euclidean distance between atoms is less than Defined as touching.
5. The drug-target binding affinity prediction method based on large model-driven multimodal graph learning according to claim 1, characterized in that: The transformation formula of step S17 is: Among them, φ q,p (x p ) belongs to the internal function, φ q (x q ) is an external function.