Method and system for predicting drug target affinity based on three-dimensional molecular fragmentation
By employing a three-dimensional molecular fragmentation-based approach, and utilizing the SchNet network and Gaussian mixture model to encode and splice features of drug and protein molecules, the problems of spatial information loss and insufficient flexibility in drug-target affinity prediction are solved, thus achieving more accurate drug target affinity prediction.
Patent Information
- Application Number
- CN202511350397.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Existing technologies struggle to fully preserve and utilize the three-dimensional spatial geometric information of drug and protein molecules in drug-target affinity prediction, and they are unable to adaptively identify structurally significant atomic clusters, thus lacking flexibility.
A three-dimensional molecular fragmentation-based approach is adopted to obtain structural feature vectors through drug molecular graph structure encoding. Combined with protein molecular graph structure encoding and fragment feature extraction, the features are spliced using a SchNet network and a Gaussian mixture model and input into a pre-trained drug target affinity prediction model to achieve drug target affinity prediction.
By effectively preserving and utilizing three-dimensional spatial geometric information, we can identify more chemically significant functional fragments of drug molecules, thereby improving the accuracy and efficiency of drug target affinity prediction.
Smart Images

Figure CN120877844B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics and computer-aided drug design, and specifically relates to a method and system for predicting drug target affinity based on three-dimensional molecular fragmentation. Background Technology
[0002] Drug-target affinity (DTA) is a key indicator for measuring the strength of interactions between small molecules and proteins, and its accurate prediction is crucial for accelerating drug screening processes and reducing R&D costs. While traditional wet experimental methods are accurate, their long cycles and high costs make them unsuitable for the demands of large-scale drug screening. Therefore, developing efficient and accurate computational prediction methods has become an urgent need in this field.
[0003] Currently, existing technologies still face the following challenges in the field of drug-target affinity prediction:
[0004] 1. When identifying the three-dimensional molecular structure of drug molecules, it is inevitable to lose three-dimensional spatial information; for example, the precise spatial position and relative orientation of atoms are lost, and the changes in bond length and bond angle under different conformations cannot be intuitively reflected.
[0005] 2. It cannot adaptively discover structurally significant atomic clusters based on the actual conformation of molecules in a specific three-dimensional environment, lacking flexibility. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a method and system for predicting drug target affinity based on three-dimensional molecular fragmentation.
[0007] In a first aspect, embodiments of the present invention provide a method for predicting drug target affinity based on three-dimensional molecular fragmentation, the method comprising the following steps:
[0008] Obtain the molecular structure of the drug to be predicted and the molecular structure of the protein to be predicted;
[0009] The structure of a drug molecule is encoded to obtain a drug structure feature vector; the structure of a protein molecule is encoded to obtain a protein structure feature vector; and fragment features are extracted from the structure of the drug molecule to obtain a fragmented structure feature vector.
[0010] The drug structure feature vector, protein structure feature vector, and drug fragmentation structure feature vector are concatenated into a fusion feature vector, which is then input into a pre-trained drug target affinity prediction model to obtain the drug target affinity prediction result.
[0011] Secondly, embodiments of the present invention provide a drug target affinity prediction system based on three-dimensional molecular fragmentation, the system comprising:
[0012] The data acquisition module is used to acquire drug molecule diagrams and protein molecule diagrams.
[0013] The feature encoding module is used to encode the structure of drug molecules to obtain drug structure feature vectors; to encode the structure of protein molecules to obtain protein structure feature vectors; and to extract fragment features from the structure of drug molecules to obtain fragmented structure feature vectors.
[0014] The drug target affinity prediction module is used to concatenate the drug structure feature vector, protein structure feature vector, and drug fragmentation structure feature vector into a fused feature vector, which is then input into the drug target affinity prediction model to obtain the drug target affinity prediction result.
[0015] Thirdly, embodiments of the present invention provide an electronic device, including:
[0016] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the above-described method for predicting drug target affinity based on three-dimensional molecular fragmentation.
[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for predicting drug target affinity based on three-dimensional molecular fragmentation.
[0018] Fifthly, embodiments of the present invention provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-described method for predicting drug target affinity based on three-dimensional molecular fragmentation.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0020] This invention provides a drug target affinity prediction method based on three-dimensional molecular fragmentation. By encoding the structure of the drug molecule map, a drug structure feature vector is obtained; by encoding the structure of the protein molecule map, a protein structure feature vector is obtained. This method can completely preserve and utilize the three-dimensional spatial geometric information of drug molecules and protein molecules. By extracting fragment features from the drug molecule map structure, a fragmented structure feature vector of the drug is obtained. This invention, through drug molecule fragmentation processing, transcends the atomic level and identifies chemically more meaningful functional fragments of drug molecules in a data-driven and adaptive manner, overcoming the problem of insufficient characterization of key chemical features such as drug molecule functional fragments. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of the drug target affinity prediction method based on three-dimensional molecular fragmentation provided in the embodiments of the present invention;
[0023] Figure 2 This is a schematic diagram of the SchNet protein encoder network provided in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of the interaction module architecture provided in an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of a continuous filter convolutional layer provided in an embodiment of the present invention;
[0026] Figure 5 This is the result of the GMM algorithm provided in this embodiment of the invention classifying drug molecules;
[0027] Figure 6 This is a schematic diagram of the drug target affinity prediction model provided in an embodiment of the present invention;
[0028] Figure 7 This is a schematic diagram of a drug target affinity prediction system based on three-dimensional molecular fragmentation provided in an embodiment of the present invention;
[0029] Figure 8 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0031] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0032] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0033] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.
[0034] like Figure 1 As shown, this embodiment of the invention provides a method for predicting drug target affinity based on three-dimensional molecular fragmentation, the method comprising the following steps:
[0035] Step S1: Obtain the molecular structure of the drug to be predicted and the molecular structure of the protein to be predicted.
[0036] The expressions for drug molecule diagram structures and protein molecule diagram structures are as follows:
[0037]
[0038]
[0039] in, Representing the molecular structure of a drug. This is the set of three-dimensional coordinates of all atoms in a drug molecule. for The set of element types corresponding to atoms in the middle; Representing the structure of protein molecules, All of the proteins The set of three-dimensional coordinates of atoms. for Chinese correspondence The set of amino acid types of atoms.
[0040] For ease of description, it is assumed in this embodiment that the drug molecule contains Atoms, protein molecules contain indivual atom. This is the set of three-dimensional coordinates of all atoms in a drug molecule, for molecules containing... Drug molecules of 10 atoms, It can be written as:
[0041]
[0042] in, Let be a vector, representing the first... The three-dimensional coordinates of each atom.
[0043] for The set of element types corresponding to the atoms in the vector is in the form of a one-hot encoded vector, written as:
[0044]
[0045] in, As a scalar, used to characterize the first The element type of each atom.
[0046] Similarly, All of the proteins The set of three-dimensional coordinates of atoms, for atoms containing indivual Atom-based protein molecules, It can be written as:
[0047]
[0048] in, Let be a vector, representing the first... indivual The three-dimensional coordinates of an atom.
[0049] for Chinese correspondence The set of amino acid types of an atom, which is in the form of a one-hot encoded vector, is written as:
[0050]
[0051] in, As a scalar, used to characterize the first indivual The amino acid type of the atom.
[0052] Therefore, drug molecule diagram structure Depend on Data sequence composition, protein molecular map structure Depend on It consists of a data sequence.
[0053] Step S2, the molecular structure of the protein to be predicted The protein structure feature vector is obtained by encoding with the SchNet protein encoder. .
[0054] The structure diagram of the protein encoder SchNet is shown below. Figure 2 As shown, the protein encoder SchNet includes an interaction module, atomic-level layers, and shifted softplus. The specific processing procedure of the protein encoder SchNet on protein molecules is described below.
[0055] Step S21: The dimension of the new feature vector mapped by each atom through the neural network layer is 64, that is, each atom is mapped to a vector containing 64 elements. After embedding, the one-hot node feature vector of each atom is transformed into a continuous feature vector of fixed size. In this embodiment, the dimension of the new feature vector mapped by the neural network layer for each atom is 128, that is, the element type information of each atom is mapped into a vector containing 128 elements.
[0056] After passing through the embedding layer, the input Row 1 column vector The output after transformation is a The eigenma matrix with 128 rows and 128 columns is denoted as ,in accession to the throne The 128-dimensional feature vector mapped from the element type information of each atom.
[0057] Step S22 involves sequentially passing the protein molecule data through multiple interaction modules. Specifically, for the first... Each interacting module has the following input: and The output is the updated atomic feature representation. In particular, when At that time, the input to the interaction module is obtained after passing through the embedding layer. and . and They represent the first The atom passes through the first The and the first The 128-dimensional feature vectors obtained after updating the interaction modules can be represented as follows:
[0058]
[0059] in, For the first The atom passes through the first The residuals obtained after the operation calculation of each interacting module.
[0060] The operation and calculation process of each interaction module is as follows: Figure 3 As shown, its input passes through an atomic-wise layer, a continuous filter convolutional layer (cfconv), an atomic-wise layer, a nonlinear activation function (shiftedsoftplus), and another atomic-wise layer in sequence, finally yielding the corresponding residual.
[0061] Among them, the input of the atom-wise layer is A 128-column feature matrix is used, where each row represents the feature vector of each atom, and a linear transformation is performed on each row vector. In this embodiment, the atomic level layer specifically refers to a 128-dimensional fully connected layer. The output is the updated... A feature matrix with 128 rows and 128 columns.
[0062] The input of the continuous filter convolutional layer (cfconv) is The eigenma matrix with 128 rows and columns The output is the updated version. The feature matrix has 128 rows and 128 columns. For ease of understanding, assume that the first row is... The input of each atom is The output after passing through cfconv is Then we have:
[0063]
[0064] In this context, the dot product symbol · indicates that vector elements are multiplied together, and the summation is performed on the nth element. Adjacent atoms of each atom It was carried out. The filter weights are calculated as follows: Figure 4 As shown, the process includes the following steps:
[0065] Step S221, firstly based on all of the proteins The set of three-dimensional coordinates of an atom The formula for calculating interatomic distance is as follows:
[0066]
[0067] in, Indicates the first The atom and the first The distance between atoms.
[0068] Step S222: Calculate the interatomic distances. Extended using radial basis functions (RBF):
[0069]
[0070] in, The center of the radial basis functions, This is a parameter, and in this embodiment, it is set to 300.
[0071] Step S223: The expanded distance vector is input into two fully connected layers with the non-linear activation function Shiftedsoftplus. The Shiftedsoftplus function is:
[0072]
[0073] After steps S221-S223, the network output is the filter weights. .
[0074] In summary, after step S22, the final updated result is obtained. The feature matrix has 128 rows and 128 columns. In this embodiment, there are 3 interaction modules, so the final feature matrix is denoted as... ,in That is, the first The element type information of each atom is finally updated into a 128-dimensional feature vector.
[0075] Step S23, will The input is passed sequentially through an atomic-level layer, a shifted softplus layer, another atomic-level layer, and a pooling layer to finally obtain a 128-dimensional protein structure feature vector. .
[0076] Step S3, the molecular structure of the drug to be predicted The drug structure feature vector is obtained by encoding with the protein encoder SchNet. .
[0077] The following describes the process by which the protein encoder SchNet processes drug molecules.
[0078] Step S31: The dimension of the new feature vector mapped by each atom through the neural network layer is 64, that is, each atom is mapped to a vector containing 64 elements. After embedding, the one-hot node feature vector of each atom is transformed into a continuous feature vector of fixed size. In this embodiment, the dimension of the new feature vector mapped by the neural network layer for each atom is 128, that is, the element type information of each atom is mapped into a vector containing 128 elements.
[0079] After passing through the embedding layer, the input Row 1 column vector The output after transformation is a The eigenma matrix with 128 rows and 128 columns is denoted as ,in accession to the throne The 128-dimensional feature vector mapped from the element type information of each atom.
[0080] Step S32 involves sequentially passing the protein molecule data through multiple interaction modules. Specifically, for the first... Each interacting module has the following input: and The output is the updated atomic feature representation. In particular, when At that time, the module input is obtained after passing through the embedding layer. and . and They represent the first The atom passes through the first The and the first The 128-dimensional feature vectors obtained after updating the interaction modules can be represented as follows:
[0081]
[0082] in For the first The atom passes through the first The residuals obtained after the operation calculation of each interacting module.
[0083] The operation and calculation process of each interaction module is as follows: Figure 3As shown, its input passes through an atomic-wise layer, a continuous filter convolutional layer (cfconv), an atomic-wise layer, a nonlinear activation function (shiftedsoftplus), and another atomic-wise layer in sequence, finally yielding the corresponding residual.
[0084] Among them, the input of the atom-wise layer is A 128-column feature matrix is used, where each row represents the feature vector of each atom, and a linear transformation is performed on each row vector. In this embodiment, the atomic level layer specifically refers to a 128-dimensional fully connected layer. The output is the updated... A feature matrix with 128 rows and 128 columns.
[0085] The input of the continuous filter convolutional layer (cfconv) is The eigenma matrix with 128 rows and columns The output is the updated version. The feature matrix has 128 rows and 128 columns. For ease of understanding, assume that the first row is... The input of each atom is The output after passing through cfconv is Then we have:
[0086]
[0087] The dot product symbol This represents the element-wise multiplication of a vector, and the summation is performed on the element-wise multiplication. Adjacent atoms of each atom It was carried out. The filter weights are calculated as follows: Figure 4 As shown, the process includes the following steps:
[0088] Step S321: First, based on the set of three-dimensional coordinates of all atoms in the drug molecule... The formula for calculating interatomic distance is as follows:
[0089]
[0090] in Indicates the first The atom and the first The distance between atoms.
[0091] Step S222: Calculate the interatomic distances. Extended using radial basis functions (RBF):
[0092]
[0093] in The center of the radial basis functions, This is a parameter, and in this embodiment, it is set to 300.
[0094] Step S323: Input the expanded distance vector into two fully connected layers with the non-linear activation function Shiftedsoftplus. The Shiftedsoftplus function is:
[0095]
[0096] After steps S321-S323, the network output is the filter weights. .
[0097] In summary, after step S32, the final updated result is obtained. The feature matrix has 128 rows and 128 columns. In this embodiment, there are 3 interaction modules, so the final feature matrix is denoted as... ,in That is, the first The element type information of each atom is finally updated into a 128-dimensional feature vector.
[0098] Step S33, will The input is passed sequentially through an atomic-level layer, a shifted softplus layer, another atomic-level layer, and a pooling layer to finally obtain a 128-dimensional protein structure feature vector. .
[0099] Step S4, the molecular structure of the drug to be predicted Fragment feature extraction is performed to obtain the drug fragmentation structural feature vector. .
[0100] Specifically, step S4 includes the following sub-steps:
[0101] Step S41: Use a spatial clustering algorithm to analyze the drug molecule diagram structure. Fragmentation is performed. In this embodiment, a Gaussian Mixture Model (GMM) algorithm is used, which fragments the drug molecules. Three-dimensional geometric distribution of atoms As input, it is decomposed into several physically or chemically meaningful clusters of atoms (i.e., "fragments") under unsupervised conditions. In this embodiment, the spatial clustering algorithm used is the Gaussian Mixture Model (GMM) algorithm. It should be noted that other spatial clustering algorithms can also be applied to this sub-step.
[0102] When using the GMM method, drug molecules As input, the algorithm is based on Three-dimensional geometric distribution of atoms Spatial clustering is performed to divide all atoms in the drug molecule into Within a cluster of atoms. Therefore, the drug molecule is re-expressed as follows:
[0103]
[0104] in, It is a collection of atomic clusters of drug molecules; They represent the 1st to the 1st. Atom cluster. With the first Atom clusters For example, its form is:
[0105]
[0106] in, Represents the i-th atomic cluster The three-dimensional coordinates of the atoms contained therein The element type corresponding to the atom. Figure 5 The example shown is the division of drug molecules into atomic clusters based on the GMM algorithm.
[0107] It should be noted that the spatial clustering algorithm used in this embodiment is the Gaussian Mixture Model (GMM) algorithm. It should also be pointed out that other spatial clustering algorithms can also be applied to this sub-step. This method, through data-driven adaptive clustering, enables fragment segmentation to dynamically adapt to the conformation of molecules in a specific three-dimensional environment, fully preserving the three-dimensional spatial position information of atoms and capturing the overall structural features of functional groups composed of multiple atoms, without relying on a predefined chemical rule library.
[0108] Step S42: The intra-cluster atomic positions and element types corresponding to each atomic cluster are encoded using the protein encoder SchNet and subjected to nonlinear transformation to obtain the intra-cluster feature vector corresponding to each atomic cluster. .
[0109] Specifically, each atom cluster obtained in step S41 The data are then fed into the protein encoder SchNet for processing. (The text abruptly ends here, likely due to an incomplete sentence or a formatting error.) Atom clusters For example, the protein encoder SchNet will analyze the atomic positions within the cluster ( ) and element type ( The data is encoded to generate a high-dimensional vector. This high-dimensional vector is then subjected to a nonlinear transformation using a multilayer perceptron (MLP) to obtain the corresponding intra-cluster feature vector. The purpose of this step is to understand and learn each atomic cluster, which can be summarized mathematically as follows:
[0110]
[0111] It should be noted that this example uses the SchNet model to learn and represent the geometric structural features inside each atomic cluster, which makes up for the shortcomings of traditional network representations that ignore the synergistic effects inside functional groups. It can accurately capture the three-dimensional geometric information such as bond length and bond angle of atoms in the fragment, as well as the correlation between atom type and spatial position, providing more chemically meaningful local features for subsequent affinity prediction.
[0112] Step S43: Calculate the geometric center corresponding to each atom cluster. The geometric center corresponding to each atomic cluster and intra-cluster feature vectors As an inter-cluster graph .
[0113] This step will further study the spatial arrangement of each segment, thereby understanding the global morphology of the entire molecule. Specifically, it can be broken down into the following processes:
[0114] Calculate the geometric center of each atomic cluster. (The sentence is incomplete and requires more context to translate accurately.) Atom clusters For example, the formula for calculating its geometric center is:
[0115]
[0116] in, Atomic clusters The geometric center, for The number of atoms contained for The Middle The three-dimensional coordinates of each atom. The result obtained in step two Summarize and construct a new inter-cluster graph. :
[0117]
[0118] It should be noted that this example calculates the geometric center corresponding to each atomic cluster, and further learns the spatial arrangement of each segment, thereby understanding the global morphology of the entire molecule. It realizes hierarchical modeling from "local segment features" to "global molecular structure", which not only preserves the detailed information inside the segment, but also integrates the global relationships such as spatial distance and relative orientation between segments, and fully restores the three-dimensional topological structure of the molecule.
[0119] Step S44: Develop the inter-cluster graph constructed in step S43. The drug fragmentation structure feature vector is obtained by encoding with the protein encoder SchNet. .
[0120] It should be noted that this example employs a clustering algorithm to fragment drug molecules and identify fragment features separately, aiming to simulate the cognitive process of molecules and their combinations in chemical research. Traditional atomic-level modeling struggles to capture the overall effect of functional groups composed of multiple atoms, while fragmented feature encoding, through clustering to form chemically meaningful atomic clusters, can reflect the structural characteristics of key functional fragments (such as pharmacophores) in drug molecules, more closely aligning with the actual mechanism of drug-target binding. By modeling features within clusters and aggregating relationships between clusters, both the spatial interactions of atoms within fragments are preserved, and the global arrangement relationships between fragments are captured, fully utilizing the three-dimensional structural information of the molecule and solving the problem of lost spatial information in traditional sequence or topological graph representations. Furthermore, fragmentation reduces the combinatorial complexity of the molecular representation space, thus enabling more efficient learning of common features of different molecular structures.
[0121] Step S5: The drug structure feature vector, protein structure feature vector, and drug fragmentation structure feature vector are concatenated into a fused feature vector, which is then input into a pre-trained drug target affinity prediction model to obtain the drug target affinity prediction result.
[0122] Specifically, the protein structure feature vector obtained in step S2 The drug structure feature vector obtained in step S3 The drug fragmentation structural feature vector obtained in step S4 The features are concatenated into a 384-row, 1-column fusion feature vector, which is then input into a pre-trained drug target affinity prediction model to obtain the drug target affinity prediction result.
[0123] The structural diagram of the drug target affinity prediction model is shown below. Figure 6 As shown, the network consists of multiple hidden layers (each hidden layer includes a fully connected layer and a nonlinear activation function). The nonlinear activation function used is the PRELU function, whose expression is:
[0124]
[0125] in, For function input, These are trainable parameters.
[0126] The training process for the drug target affinity prediction model includes:
[0127] Step S100: Clean and filter the bioinformatics database to obtain drug molecule diagram structures, protein molecule diagram structures, and their corresponding drug target affinity values (in this example, the drug target affinity values include inhibition constants). dissociation constant and half-maximal inhibitory concentration ).
[0128] Specifically, in this example, publicly available bioinformatics databases (such as PDBbind (v2019), BindingMOAD, and CASF-2016) are cleaned and screened to obtain protein molecule three-dimensional structure data and drug molecule structure data. In this embodiment, PDBbind (v2019) and BindingMOAD data are used as the training set for the drug target affinity prediction model, and CASF-2016 data is used as the test set.
[0129] Furthermore, the bioinformatics database was cleaned and screened, with the following screening criteria:
[0130] (1) Only retain measurements with clear affinity (inhibition constant) dissociation constant and half-maximal inhibitory concentration Samples were discarded from the sample, excluding those with only approximate values or those that only showed the affinity range.
[0131] (2) Drug molecule data must have standardized identifiers (such as PDB IDs) and processable structural representations (such as SMILES strings). Samples with ambiguous names or missing structural information should be removed.
[0132] (3) For protein data, only retain data where each amino acid residue contains exactly one Atoms of protein structure
[0133] (4) During training and testing, all data entries in the test set data will be removed from the training set data to ensure the validity of the verification.
[0134] Furthermore, in this embodiment, the protein molecular structure is shown in the diagram. and drug molecular structure As input, the numerical values of the drug target affinity between the drug molecule and the protein molecule, i.e., the inhibition constant. dissociation constant and half-maximal inhibitory concentration This serves as the training label. To ensure consistency during training and prediction, the drug target affinity values are all expressed as logarithmic values in molar concentration units as training and testing labels. The conversion formula is as follows:
[0135]
[0136]
[0137]
[0138] Therefore, the training label sequence can be represented by a 3x1 vector. The drug target affinity prediction model is trained to learn the mapping relationship between the input data sequence and the training labels.
[0139] Step S200: Encode the drug molecule graph structure to obtain the drug structure feature vector; encode the protein molecule graph structure to obtain the protein structure feature vector; extract fragment features from the drug molecule graph structure to obtain the drug fragmentation structure feature vector.
[0140] Step S300: The drug structure feature vector, protein structure feature vector, and drug fragmentation structure feature vector are concatenated into a fused feature vector. The fused feature vector is used as input, and the drug target affinity values corresponding to the drug molecular graph structure and the protein molecular graph structure are used as labels to train the drug target affinity prediction model.
[0141] After training, the feasibility of the model was verified using the CASF-2016 test set. Specifically, in this embodiment, to demonstrate the accuracy of the system's predictions, the model was compared with other existing prediction models, and the results are shown in Table 1 below:
[0142] Table 1: Accuracy Comparison of This Example Model with Existing Prediction Models
[0143] method RMSE DeepDTA 1.476 Co-VAE 1.430 Attention DTA 1.513 TransVAE-DTA 1.624 MMSG-DTA 1.490 This example model 1.336
[0144] The formula for the evaluation index RMSE is as follows:
[0145]
[0146] in, This represents the total number of samples. To test the drug target affinity labeling results for the test set samples, This shows the predicted drug target affinity labels from the system. It can be seen that the model proposed in this example demonstrates excellent accuracy in prediction.
[0147] To demonstrate the improvement in model prediction brought about by the molecular fragmentation method proposed in this invention, an ablation experiment was conducted to compare the molecular fragmentation operation (corresponding to step S4 in this example). The system was trained using the training sets PDBbind (v2019) and BindingMOAD respectively, and the drug target affinity prediction results were obtained: inhibition constant. dissociation constant and half-maximal inhibitory concentration The drug target affinity prediction model is iteratively trained using the inhibition constant, dissociation constant, and half-inhibition concentration corresponding to the training samples as training labels.
[0148] After training, the feasibility of the model was verified using the CASF-2016 test set, and the results are shown in Table 2 below:
[0149] Table 2: Comparison of Ablation Experiment Data
[0150] method RMSE Exclude step S4 1.582 This example model 1.336
[0151] like Figure 7 As shown, this invention provides a drug target affinity prediction system based on three-dimensional molecular fragmentation, used to implement the above-mentioned drug target affinity prediction method based on three-dimensional molecular fragmentation. The system includes:
[0152] The data acquisition module is used to acquire drug molecule diagrams and protein molecule diagrams.
[0153] The feature encoding module is used to encode the structure of drug molecules to obtain drug structure feature vectors; to encode the structure of protein molecules to obtain protein structure feature vectors; and to extract fragment features from the structure of drug molecules to obtain fragmented structure feature vectors.
[0154] The drug target affinity prediction module is used to concatenate the drug structure feature vector, protein structure feature vector, and drug fragmentation structure feature vector into a fused feature vector, which is then input into the drug target affinity prediction model to obtain the drug target affinity prediction result.
[0155] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0156] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0157] like Figure 8As shown, this application provides an electronic device including a memory 101 for storing one or more programs and a processor 102. When the one or more programs are executed by the processor 102, they implement the method as described in any of the first aspects above.
[0158] The system also includes a communication interface 103. The memory 101, processor 102, and communication interface 103 are electrically connected directly or indirectly to each other to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, and the processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.
[0159] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0160] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor 102, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0161] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can also be implemented in other ways. The method and system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0162] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0163] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by processor 102, the computer program implements the methods described in any of the first aspects above. If the functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0164] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only.
[0165] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for predicting drug target affinity based on three-dimensional molecular fragmentation, characterized by, The method comprises the following steps: Obtaining a drug molecule graph structure to be predicted and a protein molecule graph structure to be predicted; the drug molecule graph structure includes a three-dimensional coordinate set of all atoms in the drug molecule and an element type set, and the protein molecule graph structure includes a three-dimensional coordinate set of all C α atoms in the protein and an amino acid type set; The drug molecule graph structure is encoded to obtain a drug structure feature vector; the protein molecule graph structure is encoded to obtain a protein structure feature vector; and fragment feature extraction is performed on the drug molecule graph structure to obtain a drug fragmented structure feature vector; The drug structure feature vector, the protein structure feature vector and the drug fragmented structure feature vector are spliced into a fusion feature vector, which is input into a pre-trained drug target affinity prediction model to obtain a drug target affinity prediction result; The process of performing fragment feature extraction on the drug molecule graph structure to obtain the drug fragmented structure feature vector comprises: The atoms in the drug molecule are divided into a plurality of atomic clusters according to the three-dimensional geometric distribution of the atoms in the drug molecule; the intra-cluster atomic positions and element types corresponding to each atomic cluster are encoded and nonlinearly transformed to obtain an intra-cluster feature vector corresponding to each atomic cluster; the geometric center corresponding to each atomic cluster is calculated, and the geometric center and the intra-cluster feature vector corresponding to each atomic cluster are taken as an inter-cluster graph; and the inter-cluster graph is encoded to obtain the drug fragmented structure feature vector.
2. The method of predicting drug target affinity based on three-dimensional molecular fragmentation according to claim 1, characterized in that, The process of encoding the drug molecule graph structure, the protein molecule graph structure and the inter-cluster graph comprises: encoding the drug molecule graph structure, the protein molecule graph structure and the inter-cluster graph by a protein encoder SchNet.
3. The method of predicting drug target affinity based on three-dimensional molecular fragmentation according to claim 1, characterized in that, The training process of the drug target affinity prediction model comprises: The bioinformatics database is cleaned and screened to obtain drug molecule graph structures, protein molecule graph structures and corresponding drug target affinity values; The drug molecule graph structure is encoded to obtain a drug structure feature vector; the protein molecule graph structure is encoded to obtain a protein structure feature vector; and fragment feature extraction is performed on the drug molecule graph structure to obtain a drug fragmented structure feature vector; The drug structure feature vector, the protein structure feature vector and the drug fragmented structure feature vector are spliced into a fusion feature vector, which is input into a pre-trained drug target affinity prediction model to obtain a drug target affinity prediction result; 4. The method of predicting drug target affinity based on three-dimensional molecular fragmentation according to claim 1, wherein, The drug target affinity prediction result comprises an inhibition constant, a dissociation constant and a half maximal inhibitory concentration.
5. A drug target affinity prediction system based on three-dimensional molecular fragmentation, characterized by, The system comprises: The data acquisition module is used for acquiring a drug molecule graph structure and a protein molecule graph structure; the drug molecule graph structure comprises a three-dimensional coordinate set and an element type set of all atoms in a drug molecule, and the protein molecule graph structure comprises a three-dimensional coordinate set and an amino acid type set of all C α atoms in a protein. The feature encoding module is configured to encode the drug molecule graph structure to obtain a drug structure feature vector, encode the protein molecule graph structure to obtain a protein structure feature vector, and extract fragment features of the drug molecule graph structure to obtain a drug fragmented structure feature vector. The process of extracting the fragment features of the drug molecule graph structure to obtain the drug fragmented structure feature vector includes: performing spatial clustering according to the three-dimensional geometric distribution of atoms in the drug molecule, thereby dividing all atoms in the drug molecule into a plurality of atomic clusters; encoding and nonlinearly transforming the intra-cluster atomic positions and element types corresponding to each atomic cluster to obtain an intra-cluster feature vector corresponding to each atomic cluster; calculating a geometric center corresponding to each atomic cluster, and taking the geometric center corresponding to each atomic cluster and the intra-cluster feature vector as an inter-cluster graph; and encoding the inter-cluster graph to obtain the drug fragmented structure feature vector. The drug target affinity prediction module is configured to splice the drug structure feature vector, the protein structure feature vector, and the drug fragmented structure feature vector into a fusion feature vector, input the fusion feature vector into a drug target affinity prediction model, and obtain a drug target affinity prediction result.
6. An electronic device, comprising: The computer program product comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the drug target affinity prediction method based on three-dimensional molecular fragmentation according to any one of claims 1-4.
7. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program product comprises:
8. A computer program product comprising computer programs / instructions, characterized in that, The computer program product comprises: The computer program product comprises:
Citation Information
Patent Citations
Drug target interaction relationship prediction method and system
CN119068972A
Application of multi-modal feature fusion model in drug target binding affinity prediction
CN119479783A
Drug-target interaction prediction method and device based on multi-view feature fusion and medium
CN120072036A