A method for predicting drug-target affinity based on multi-shell and extended connectivity fingerprints
Through the method of binding of multi-shell and extended connectivity fingerprints, the 3D structure of protein-ligand complexes is learned using the Transformer model, solving the accuracy of drug target binding affinity prediction and improving the efficiency and accuracy of drug discovery process.
Patent Information
- Application Number
- CN202310223854.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-03-09
AI Technical Summary
The prior art is difficult to accurately predict drug-target binding affinity, resulting in a complex and time-consuming process for drug discovery and lack of an effective scoring function to distinguish binding agents from non-binding agents.
Using a method based on multi-shell layer and extended connectivity fingerprint, the shell modeling of the protein-ligand complex is used to calculate the extended connectivity atom pair characteristics, and the shell atomic features are learned using the Transformer model, combining 3D slices and fully connected layers for affinity prediction.
It improves the accuracy of drug target binding affinity prediction, can better capture long-range interactions, improves the characterization performance of the model, and reduces experimental needs.
Smart Images

Figure CN116206678B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of biological informatics and computer application technologies, and more specifically, to a method for predicting drug-target affinity based on multi-shell and extended connectivity fingerprints. Background Art
[0002] Since the birth of the first drug to date, academic researchers and industrial workers have been dedicated to the discovery and development of drugs to combat various diseases. The process of drug discovery and development usually starts with target identification and ends with clinical trials. Due to the need for a large amount of analysis and testing and the high risk of failure, the entire process of developing a new drug usually takes 10 - 20 years and requires a large amount of investment, ranging from $500 million to $2.6 billion. The prediction of drug-target affinity is one of the most critical steps in discovering key and lead compounds in drug discovery.
[0003] The discovery of small molecule drugs usually identifies the protein targets against which hit compounds are directed through high-throughput screening (HTS). Subsequently, the hit compounds are optimized to obtain lead compounds with good pharmacodynamic and pharmacokinetic properties. With the development of computer technology and the emergence of a large amount of protein-ligand structure data, computer-aided drug discovery has played an important role in the development of new small molecule drugs in the past few decades, and the accurate prediction of drug-target binding affinity (DTA) is a key link in computer-aided drug design. One of the main goals of DTA prediction is to design a suitable scoring function (SF) to calculate the relative or absolute binding free energy to distinguish strong binders from weak binders (or non-binders) against a specific target. Without a reliable scoring function, it is difficult to ensure the performance of various tasks involved in computer-aided drug design. The rapid and accurate prediction of DTA will avoid many time-consuming and complex experiments.
[0004] How to accurately predict DTA remains a key challenge in the fields of computational biology and computational chemistry. In the past few decades, extensive efforts have been made to develop new SFs or improve existing SFs to enhance the performance of DTA prediction. However, there are still great challenges in further developing DTA prediction methods that can best simulate real physiological scenarios. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for predicting drug-target affinity based on multi-shell and extended connectivity fingerprints to overcome the defects of the prior art.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A method for predicting drug-target affinity based on multi-shell and extended connectivity fingerprints, comprising the following steps:
[0008] S1. Initialize parameters, including setting the number of shells num_shells, the chunk size patch_size, the number of multi-head attentions in the Transformer num_heads, the dimension of the final feature vector projection_dim, and the number of layers of the Transformer n;
[0009] S2. Perform shell modeling on the protein-ligand complex;
[0010] S3. Count all atom types in the PDBbind v2019 dataset to obtain the number of protein atom types protein_atom_types and the number of ligand atom types ligand_atom_types;
[0011] S4. For all shells, calculate the number of specific protein atom-ligand atom pairs located in the shell to obtain the extended connectivity atom pair feature of the complex in a single shell;
[0012] S5. Determine whether the extended connectivity atom pair features of all shells have been calculated. If so, stack the feature vectors of all shells to obtain the feature vector of the entire complex, and go to step S6; otherwise, go to step S4;
[0013] S6. Slice the feature vector of the complex to obtain num_patch slices, where a single slice is a feature vector of length projection_dim;
[0014] S7. Perform positional encoding on each slice, and sum the positional information of the slice after passing through the embedding layer and the feature vector of the slice to obtain the slice feature vector with positional information added;
[0015] S8. Use the obtained slices as sequences to input into n identical Transformer models to learn the representation vectors of the slices. Each Transformer model consists of a multi-head attention mechanism, layer normalization, and two fully connected networks;
[0016] S9. Determine whether the maximum number of iterations n has been reached. If so, output the feature vector as the final representation and go to step S10; otherwise, use the feature vector as a slice sequence and go to step S8;
[0017] S10. Perform layer normalization, global average pooling on the final representation vector and input it into the final fully connected layer to obtain the predicted value of the affinity strength.
[0018] Furthermore, step S2 specifically includes:
[0019] S21. Calculate the geometric center point O of the small molecule ligand according to the three-dimensional coordinates of each atom of the small molecule ligand;
[0020] S22. Divide a series of shells centered on the geometric center point O, where the distance between the first shell and the geometric center point O is d0, and the distance between the Kth shell and the (K - 1)th shell is d s , where K = 2, 3, 4, …, num_shells;
[0021] S23. For all protein atoms, calculate the distance d between it and the geometric center point O. If d < d0, mark it as the first-layer atom. If d > d0 + (num_shells - 1)*d s , then ignore this atom. Otherwise, mark it as the ((d - d0) / d s + 1)-layer atom.
[0022] Further, the step S3 specifically includes:
[0023] S31. According to the interpretation of the Python open-source cheminformatics software RDKit, obtain the ligand atom types from the standard data format file. Define the ligand atoms by considering six characteristics of the atoms according to the atomic connectivity proposed by the ECIF method, including atomic symbol, explicit valence, number of heavy atoms connected, number of hydrogens connected, atomic aromaticity, and ring structure. Finally, 70 ligand atom types are obtained;
[0024] S32., In the dictionary-based mapping, manually assign the atom types to 22 different atom types according to the residues and atom labels in the PDB structure file.
[0025] Further, the step S4 specifically includes:
[0026] S41. Count the types and quantities of all ligand atoms and all protein atoms within a single shell;
[0027] S42. Combine the protein atoms and ligand atoms in pairs to form protein-ligand atom pairs. Form a specific atom pair matrix with protein_atom_types rows and ligand_atom_types columns from protein_atom_types types of protein atoms and ligand_atom_types types of ligand atoms. Each element in the matrix is the quantity of the corresponding specific atom pair combination;
[0028] S43. The specific atom pair matrix is the extended connectivity atom pair feature of a single shell.
[0029] Further, the step S6 specifically includes:
[0030] S61. The eigenvector of the entire complex is a three-dimensional array composed of the shell layer, protein atom types, and ligand atom types. Using 3D-CNN, slice it according to the predefined patch_size in these three dimensions to obtain num_patch slices;
[0031] S62. Flatten the features of each slice after convolution to obtain the eigenvector of each slice. The length of the eigenvector is the number of convolution filters projection_dim.
[0032] Further, the step S8 specifically includes:
[0033] S81. The eigenvector v passes through LayerNormalization and the multi-head attention mechanism to obtain the first intermediate temporary result x1 (for residual connection in the neural network). Perform a residual connection between the eigenvector v and the first intermediate temporary result x1 to obtain the second intermediate temporary result x2;
[0034] S82. Pass the second intermediate temporary result x2 through LayerNormalization and two fully connected layers to obtain the third intermediate temporary result x3. Perform a residual connection between x2 and x3 to obtain the final eigenvector v.
[0035] Compared with the prior art, the advantages of the present invention are as follows: A drug target affinity prediction method based on multi-shell layer and extended connectivity fingerprint provided by the present invention combines the idea of multi-shell layer with the extended connectivity fingerprint. That is, taking the ligand as the center, the 3D structure of the protein-ligand complex is artificially divided into multiple shell layers, and the extended connectivity fingerprint is calculated for each shell layer respectively. The finally generated descriptor can solve the problem that the existing atomic extended connectivity fingerprint cannot characterize long-range interactions while maintaining the spatial rotation consistency of the complex; and through 3D slicing of the complex features, Transformer can well learn the shell atom pair features, thereby changing the performance of the representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 is a flowchart of the drug target affinity prediction method based on multi-shell layer and extended connectivity fingerprint of the present invention.
[0038] Figure 2It is a line graph of the performance using different numbers of shells in the present invention.
[0039] Figure 3 It is a line graph of the performance using different numbers of shells in the present invention. Detailed implementation manners
[0040] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0041] Refer to Figure 1 As shown, this embodiment discloses a method for predicting drug-target affinity based on multi-shell and extended connectivity fingerprints, including the following steps:
[0042] Step S1, initialize parameters, including setting the number of shells num_shells, the chunk size patch_size, the number of multi-head attentions of the Transformer num_heads, the dimension of the final feature vector projection_dim, and the number of layers n of the Transformer.
[0043] Step S2, perform shell modeling on the protein-ligand complex.
[0044] Step S3, count all atomic species in the PDBbind v2019 dataset to obtain the number of protein atom types protein_atom_types and the number of ligand atom types ligand_atom_types.
[0045] Step S4, for all shells, calculate the number of specific protein atom-ligand atom pairs located in this shell to obtain the extended connectivity atom pair feature of the complex in a single shell;
[0046] Step S5, determine whether the extended connectivity atom pair features of all shells have been calculated. If so, stack the feature vectors of all shells to obtain the feature vector of the entire complex, and go to step S6; otherwise, go to step S4;
[0047] Step S6, slice the feature vector of the complex to obtain num_patch slices, where a single slice is a feature vector with a length of projection_dim;
[0048] Step S7, perform positional encoding on each slice, and sum the positional information of the slice after passing through the embedding layer and the feature vector of the slice to obtain the slice feature vector with positional information added;
[0049] Step S8: Use the obtained slices as a sequence and input them into n identical Transformer models to learn the representation vectors of the slices. Each Transformer model consists of a multi-head attention mechanism, layer normalization, and two fully connected networks;
[0050] Step S9: Determine whether the maximum number of iterations n has been reached. If so, use the feature vector as the final representation output and go to Step S10; otherwise, use the feature vector as the slice sequence and go to Step S8;
[0051] Step S10: Perform layer normalization and global average pooling on the final representation vector and input it into the final fully connected layer to obtain the predicted value of the affinity strength.
[0052] The specific steps of Step S2 include:
[0053] Step S21: Calculate the geometric center point O of the small molecule ligand according to the three-dimensional coordinates of each atom of the small molecule ligand.
[0054] Step S22: With the geometric center point O as the center, divide a series of shell layers. The distance between the first shell layer and the geometric center point O is d0, and the distance between the Kth shell layer and the (K - 1)th shell layer is d s , where K = 2, 3, 4, …, num_shells.
[0055] Step S23: For all protein atoms, calculate their distance d from the geometric center point O. If d < d0, mark them as the first layer of atoms; if d > d0 + (num_shells - 1) * d s , then ignore the atom; otherwise, mark it as the ((d - d0) / d s +1)th layer of atoms.
[0056] The specific steps of Step S3 include:
[0057] Step S31: For ligand atoms, according to the interpretation of the Python open-source cheminformatics software RDKit (2020.03.1), obtain the ligand atom types from the standard data format (SDF) file. Define the ligand atoms by considering six characteristics of the atoms according to the atomic connectivity proposed by the ECIF method, including atomic symbol, explicit valence, number of connected heavy atoms, number of connected hydrogens, atomic aromaticity, and ring structure. Finally, 70 ligand atom types are obtained as ligand_atom_types.
[0058] Step S32: In the dictionary-based mapping, manually assign atom types to 22 different atom types of protein_atom_types according to the residues and atom labels in the PDB structure file.
[0059] The specific steps of step S4 include:
[0060] Step S41: Count the types and quantities of all ligand atoms and all protein atoms within a single shell.
[0061] Step S42: Combine protein atoms and ligand atoms pairwise to form protein-ligand atom pairs, and form a specific atom pair matrix with protein_atom_types rows and ligand_atom_types columns from protein_atom_types types of protein atoms and ligand_atom_types types of ligand atoms. Each element in the matrix is the quantity of the corresponding specific atom pair combination.
[0062] Step S43: The specific atom pair matrix is the extended connectivity atom pair feature of a single shell.
[0063] The specific steps of step S6 include:
[0064] Step S61: The feature vector of the entire complex is a three-dimensional array composed of the shell, protein atom type, and ligand atom type. Slice it using 3D-CNN with a predefined patch_size in these three dimensions to obtain num_patch slices.
[0065] Step S62: Flatten the features of each slice after convolution to obtain the feature vector of each slice. The length of the feature vector is the number of convolution filters projection_dim.
[0066] The specific steps of step S8 include:
[0067] Step S81: The feature vector v passes through LayerNormalization and the multi-head attention mechanism to obtain the first intermediate temporary result x1 (for residual connection in the neural network), and perform a residual connection between the feature vector v and the first intermediate temporary result x1 to obtain the second intermediate temporary result x2.
[0068] Step S82: Pass the second intermediate temporary result x2 through LayerNormalization and two fully connected layers to obtain the third intermediate temporary result x3, and perform a residual connection between x2 and x3 to obtain the final feature vector v.
[0069] There are n identical Transformer layers in total. The output of each layer is a v, and the output of each layer is used as the input of the next layer. If it is the nth layer currently, then the newly generated feature vector v is used as the input of step S9.
[0070] Combine Figure 2 and Figure 3As shown, the cut-off distance for statistically counting atom pairs in the existing method is fixed, and all atom pairs beyond this cut-off distance will be ignored, which may not well describe the binding process between proteins and ligands. As is well known, the main influencing factors for the overall binding affinity of protein-ligand include electrostatic interaction, van der Waals interaction, hydrogen bond, hydration / dehydration during the binding process, etc. Among them, in addition to short-range interactions, there is also this important long-range interaction of electrostatic interaction. If the cut-off distance is directly extended to capture long-range interactions, as the cut-off distance increases, the number of extended connectivity features will increase sharply, and the newly added long-range features will overwhelm the key short-range features, making it difficult to capture short-range interactions and resulting in a decline in the performance of the model.
[0071] The present invention combines the idea of multi-shells with extended connectivity fingerprints, that is, centered on the ligand, the 3D structure of the protein-ligand complex is artificially divided into multiple shells, and extended connectivity fingerprints are calculated for each shell respectively. The finally generated descriptors can, while maintaining the spatial rotation consistency of the complex, solve the problem that the existing atom extended connectivity fingerprints cannot characterize long-range interactions; secondly, through 3D slicing of the complex features, the Transformer can well learn the shell atom pair features, thus changing the performance of the characterization.
[0072] Although the embodiments of the present invention have been described with reference to the accompanying drawings, the patent owner can make various deformations or modifications within the scope of the appended claims. As long as it does not exceed the protection scope described in the claims of the present invention, it should be within the protection scope of the present invention.
Claims
1. A method for predicting drug target affinity based on multi-shell and extended connectivity fingerprints, characterized in that It includes the following steps: S1. Initialize parameters, including setting the number of shells num_shells, the chunk size patch_size, the number of multi-head attentions of the Transformer num_heads, the dimension of the final feature vector projection_dim, and the number of layers of the Transformer n; S2. Perform shell modeling on the protein-ligand complex; S3. Count all atom types in the PDBbind v2019 dataset to obtain the number of protein atom types protein_atom_types and the number of ligand atom types ligand_atom_types; S4. For all shells, calculate the number of specific protein atom-ligand atom pairs located in the shell to obtain the extended connectivity atom pair feature of the complex in a single shell; S5. Determine whether the extended connectivity atom pair features of all shells have been calculated. If so, stack the feature vectors of all shells to obtain the feature vector of the entire complex, and go to step S6; otherwise, go to step S4; S6. Slice the feature vector of the complex to obtain num_patch slices, where a single slice is a feature vector with a length of projection_dim; S7. Perform positional encoding on each slice, and sum the positional information of the slice after passing through the embedding layer and the feature vector of the slice to obtain the slice feature vector with positional information added; S8. Use the obtained slices as a sequence to input into the same n-layer Transformer model to learn the representation vectors of the slices. Each layer of the Transformer model consists of a multi-head attention mechanism, layer normalization, and two fully connected networks; S9. Determine whether the maximum number of iterations n has been reached. If so, output the feature vector as the final representation and go to step S10; otherwise, use the feature vector as a slice sequence and go to step S8; S10. Perform layer normalization, global average pooling on the final representation vector and input it into the final fully connected layer to obtain the predicted value of the affinity strength.
2. The method for predicting drug target affinity based on multi-shell and extended connectivity fingerprints according to claim 1, wherein The specific steps of step S2 include: S21. Calculate the geometric center point O of the small molecule ligand according to the three-dimensional coordinates of each atom of the small molecule ligand; S22. With the geometric center point O as the center, a series of shells are divided, where the distance between the first shell and the geometric center point O is d0, and the distance between the Kth shell and the (K - 1)th shell is d s , where K = 2, 3, 4, …, num_shells; S23. For all protein atoms, calculate their distance d from the geometric center point O. If d < d0, mark them as the first-layer atoms. If d > d0 + (num_shells - 1) * d s , then ignore the atom. Otherwise, mark it as the ((d - d0) / d) s +1 layer atom.
3. The method for predicting drug target affinity based on multi-shell and extended connectivity fingerprints according to claim 1, characterized in that The specific steps of step S3 include: S31. According to the interpretation of the Python open-source chemoinformatics software RDKit, obtain the ligand atom types from the standard data format file, and define the ligand atoms by considering six features of the atom connectivity according to the ECIF method, including atomic symbol, explicit valence, number of connected heavy atoms, number of connected hydrogens, atomic aromaticity, and ring structure. Finally, 70 ligand atom types are obtained; S32. In the dictionary-based mapping, manually assign atom types to 22 different atom types according to the residues and atom labels in the PDB structure file.
4. The method for predicting drug target affinity based on multi-shell and extended connectivity fingerprints according to claim 1, characterized in that The specific steps of step S4 include: S41. Count the types and numbers of all ligand atoms and all protein atoms in a single shell; S42. Pair the protein atoms and ligand atoms pairwise to form protein-ligand atom pairs, and form a specific atom pair matrix with protein_atom_types rows and ligand_atom_types columns from protein_atom_types kinds of protein atoms and ligand_atom_types kinds of ligand atoms. Each element in the matrix is the number of combinations of the corresponding specific atom pairs. S43. The specific atom pair matrix is the extended connectivity atom pair feature of a single shell.
5. The method for predicting drug target affinity based on multi-shell and extended connectivity fingerprints according to claim 1, wherein The specific steps of step S6 are as follows: S61. The feature vector of the whole complex is a three-dimensional array composed of the shell, protein atom type, and ligand atom type. Slice it using 3D-CNN with a predefined patch_size in these three dimensions to obtain num_patch slices. S62. Flatten the features of each slice after convolution to obtain the feature vector of each slice. The length of the feature vector is the number of convolution filters projection_dim.
6. The method for predicting drug target affinity based on multi-shell and extended connectivity fingerprints according to claim 1, wherein The specific steps of step S8 are as follows: S81. The feature vector v passes through LayerNormalization and the multi-head attention mechanism to obtain the first intermediate temporary result x1, and the feature vector v and the first intermediate temporary result x1 are subjected to residual connection to obtain the second intermediate temporary result x2. S82. The second intermediate temporary result x2 passes through LayerNormalization and two fully connected layers to obtain the third intermediate temporary result x3, and a residual connection is made between x2 and x3 to obtain the final feature vector v.
Citation Information
Patent Citations
Hydrophobic fluorescent particle and photosensitizer-carrying polymersome complex
KR101323848B1
Methods and apparatuses for using artificial intelligence trained to generate candidate drug compounds based on dialects
WO2022250726A1