Method and device for constructing neural network model for target design drug molecules

By constructing a neural network model for target-oriented design of drug molecules, combining chemical structure inverse synthesis splitting algorithm and self-attention mechanism, the problem of time-consuming and costly design of traditional drug is solved, and the generated drug molecules have higher drug potential and reasonable topological structure.

CN120260730APending Publication Date: 2025-07-04SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510369903.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

Smart Images

  • Figure CN120260730A_ABST
    Figure CN120260730A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method and device of a neural network model for target design drug molecules. According to the method, target pockets are characterized in three levels of atoms, residues and among the residues, so that the information of the target is more comprehensively considered; a molecular fragment obtained by a chemical structure inverse synthesis splitting algorithm is taken as a unit, the drug molecules are characterized and generated, and the generated molecules have higher drug forming potential; the model designed by the invention is composed of two neural network encoders and a gated self-attention decoder, the encoders are used for extracting features of a target and ligand molecules, the decoder realizes association of the two features and generates a molecular fragment probability, and the model can effectively design drug molecules for a specific target. Molecules designed by the model constructed by the invention have an average synthesis accessibility score of 0.769 and an average drug-likeness score of 0.607, which are higher than those of a baseline model PMDM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application of deep generative models in the field of drug design, and to a construction method and device for a neural network model for designing drug molecules targeting a target. Specifically, the present invention relates to using technologies such as graph neural networks and chemical structure retrosynthetic cleavage algorithms to characterize protein targets and drug small molecules, and constructing a corresponding drug design model through a gated recurrent unit and a self-attention mechanism. Background Art

[0002] Designing new drug molecules targeting specific targets is a challenging task. In the past, drug design based on traditional methods such as high-throughput screening required high professional knowledge and skills of personnel. From design, synthesis to clinical testing, it often took a long cycle and consumed high costs.

[0003] In recent years, the successful application of artificial intelligence technologies such as deep learning, especially deep generative models, has greatly accelerated the process of drug design. The main idea is to continuously optimize the generative model to learn the characteristics of existing molecules in chemical space and generate new molecules with ideal properties based on this. Drug design methods based on deep learning represent an emerging trend in the field of drug design. These methods are efficient, low-cost, and easy to use. However, the previously proposed drug design models based on deep learning have some deficiencies, such as ignoring the fine structural features of protein pockets and generating molecules with unreasonable topological structures. In summary, developing a neural network model for designing drug molecules targeting a target has important application potential. Summary of the Invention

[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a construction method, device, computer device and medium for a neural network model for designing drug molecules targeting a target.

[0005] Technical Solution: To solve the above technical problem, the present invention provides a construction method for a neural network model for designing drug molecules targeting a target, including the following steps:

[0006] (1) Construction of a target feature encoding module for extracting the features of the target target pocket and obtaining the latent vector of the target pocket through reparameterization; the features of the target pocket include atomic-level features, residue features and inter-residue features, the atomic-level features include atomic type and relative coordinates, the residue features include residue type and residue backbone dihedral angles, and the inter-residue features include the distance between residues and relative position encoding;

[0007] (2) Construction of the ligand molecule feature encoding module, which is used to map the fragment sequence of the ligand molecule into the potential vector representation of the fragment sequence of the ligand molecule; specifically, the chemical structure retrosynthesis splitting algorithm is used to break the ligand molecule into molecular fragments, and thus obtain the fragment sequence of each ligand molecule; the skip-gram model is used to convert the fragment sequence of each ligand molecule into a feature embedding;

[0008] (3) Construction of the autoregressive decoding module, which consists of two gated recurrent unit modules and a self-attention module; add the potential vector of the target pocket and the potential vector of the ligand molecule fragment sequence to obtain a composite potential vector as the decoder input, and gradually generate the fragment sequence in an autoregressive manner. When the termination marker is generated or the generated fragment sequence reaches the specified maximum length, the generation terminates;

[0009] (4) Model training: Use the preprocessed dataset to train the model. The preprocessed dataset includes a training set, a test set, and a validation set. The loss function during the model training process consists of a reconstruction loss and a relative entropy. A learning rate scheduler is introduced during the training process to automatically adjust the learning rate;

[0010] (5) Use the protein target in the test set as the input to generate the ligand molecule.

[0011] In the present invention, the target pocket is also referred to as the protein pocket, and the two have the same meaning. The ligand molecule described in the present invention is essentially a small drug molecule in the actual design result of the present invention.

[0012] Among them, after obtaining various features in step (1), the atomic-level features are embedded through linear transformation and pooling. The residue features are embedded after normalization and linear transformation, and added to the corresponding atomic-level feature embeddings to obtain a composite embedding as the point feature embedding in the protein target graph. The residue-interaction features are embedded after normalization and linear transformation as the edge feature embedding in the protein target graph;

[0013] Among them, preferably, the atomic-level feature embeddings are added to obtain a 256-dimensional composite embedding, and the residue-interaction features are converted into 256-dimensional residue-interaction feature embeddings;

[0014] Among them, step (1) also includes representing the protein pocket as a k-nearest neighbor graph after obtaining the point feature embedding and the edge feature embedding. Each node in the graph represents each amino acid in the protein pocket, and each edge represents the connection between residues;

[0015] Among them, in step (1), it also includes mapping the point features and edge features in the protein graph to the hidden space through a linear layer. The target pocket encoder takes the mapped node features and edge features as inputs, performs information passing and aggregation to optimize the feature representation of each node, and then outputs the feature representation of each protein graph through global average aggregation as the latent representation p of the protein. For the latent representation p of each protein, calculate its mean μ(p) and variance σ 2 (p), and use the normal distribution z p as the distribution of the protein pocket latent vector:

[0016] z p = μ(p) + σ 2 (p)·∈

[0017] Among them, the construction steps of the molecular fragment encoder in step (2) include: using the chemical structure retrosynthetic cleavage algorithm to break the drug small molecule into molecular fragments, and thus obtaining the fragment sequence of each molecule. After obtaining the fragment sequence of the drug small molecule, regarding the molecular sequence as a sentence and each fragment decomposed from the molecular sequence as a vocabulary, convert different fragments into 256-dimensional word embeddings through the skip-gram model with negative sampling. The molecular fragment encoder takes each molecular fragment sequence embedding x = (x1, x2, …, x i ) as an input. For each fragment embedding x i , convert it into a hidden state h i through two layers of gated recurrent units, and use the hidden state h Final of the last fragment embedding obtained through forward propagation as the latent representation of the entire fragment sequence, and calculate the latent vector z x distribution of the molecular fragment sequence in the same way as the protein pocket encoder:

[0018] h i = GRU(x i .h i-1 )

[0019] z x ~N(μ(x), σ 2 (x))

[0020] Among them, the reason for using the chemical structure retrosynthetic cleavage algorithm to break the drug molecule into molecular fragments and construct the fragment dataset for the design of step (2) is as follows:

[0021] Since the model of the present invention conducts drug design in units of fragments, after obtaining the pocket-ligand pair data, it is necessary to further construct a fragment dataset for the embedding process of the fragments and the reconstruction of ligand molecules. Currently, the mainstream molecular fragmentation methods in the chemical field mainly include Retro Synthetic Combinatorial Analysis Procedure (RECAP), Chemical Structure Inverse Synthetic Disconnection Algorithm, and sequence- or graph-based fragmentation methods. Among them, RECAP breaks 11 common bonds in chemical reactions, with simple rules but a low-quality fragment set and prone to redundant fragments; sequence- or graph-based fragmentation methods are highly flexible but tend to generate chemically unreasonable or low-drugability fragments; while the advantage of the Chemical Structure Inverse Synthetic Disconnection Algorithm lies in that it incorporates more detailed medicinal chemistry concepts to redefine 16 disconnection rules, making the fragments after disconnection have higher drug-likeness and synthetic accessibility, and is more suitable for the field of drug design compared to other fragmentation methods. Therefore, the present invention uses the Chemical Structure Inverse Synthetic Disconnection Algorithm to break drug molecules into molecular fragments and construct a fragment dataset. Specifically, for a ligand molecule, first traverse all breakable chemical bonds in the molecule from left to right in the order corresponding to the Simplified Linear Molecular Input Specification Sequence, return the indices of the chemical bonds that conform to the disconnection rules of the Chemical Structure Inverse Synthetic Disconnection Algorithm, and then iteratively break the molecule through the returned bond indices in the same order, introducing virtual atom markers at the break points. The fragments after disconnection can be inversely synthesized into the molecule before disconnection through the virtual atom markers. Since the disconnection strategy of the Chemical Structure Inverse Synthetic Disconnection Algorithm is fixed, to avoid over-disconnection, the present invention defines the minimum fragment size for disconnection as 3. After each disconnection to obtain a fragment, calculate the number of atoms in the fragment. If the number of atoms in the fragment is less than 3, discard this disconnection and re-disconnect at the next breakable chemical bond. After breaking 172,216 ligand molecules in the dataset and filtering out duplicate fragments, a total of 10,475 unique fragments are obtained. Convert them into Simplified Linear Molecular Input Specification Sequences and define indices in sequence. Finally, a fragment dataset of the Chemical Structure Inverse Synthetic Disconnection Algorithm with a size of 10,475 is constructed. The process of constructing the fragment dataset and molecular fragmentation is as Figure 1 shown. After obtaining the fragment dataset of the Chemical Structure Inverse Synthetic Disconnection Algorithm for all ligands, it is also necessary to convert the fragment data into fragment embeddings that can be input into the model. The present invention regards the complete Simplified Linear Molecular Input Specification Sequence of the molecule as a sentence, and the Simplified Linear Molecular Input Specification Sequences of the decomposed individual fragments as words, and converts each fragment into a word embedding of a fixed dimension through the negative sampling skip-gram model. This embedding method can better learn rare fragments and retain the correlation between fragments during the training process.

[0022] Among them, the construction steps of the autoregressive decoding module in step (3) include: first, the protein pocket latent vector z p reparameterized and the molecular fragment sequence latent vector zx Add them to obtain the composite latent vector c z , and convert c z to an initial hidden representation through a linear layer as the input to the gated recurrent unit. The gated recurrent unit outputs the final hidden state hc Final . After that, perform self-attention calculation on it through a self-attention module with 8 heads:

[0023] Q = W Q ·hc Final ; K = W K ·hc Final ; V = W V ·hc Final

[0024]

[0025] Q, K, and V are the query, key, and value in the basic principle of the self-attention mechanism. W Q , W K , and W V are the product matrices for forming Q, K, and V respectively; T represents transpose, is the scaling factor;

[0026] After the multi-head self-attention scores are concatenated, they are used to calculate the probability distribution Probs(I) of the molecular fragment:

[0027] Multihead(Q, K, V) = Concat(head1, head2,..., head n )·W out

[0028] Probs(I) = softmax(Multihead(Q, K, V))

[0029] where Multihead(Q, K, V) represents multi-head self-attention, and head1, head2,..., head n refer to the 1st, 2nd,..., nth attention heads, and W out refers to the output matrix. The result obtained by multiplying the concatenated result of the attention heads and the output matrix is used as the final output of the multi-head self-attention;

[0030] Preferably, an autoregressive method is adopted to generate each fragment. First, a zero vector is initialized as the first Token (fragment) of the starting sequence. After the two encoders process the starting sequence and the corresponding target pocket, the hidden state is obtained. Then, the decoder uses the starting sequence and the hidden state to calculate the probability distribution of the molecular fragment, samples the next Token from this distribution in a polynomial sampling manner, and adjusts the sampling process using the temperature parameter:

[0031]

[0032] Next_Token = multinomial(Adjust_Probs(I), 1)

[0033] where I represents the molecular fragment, Adjust_Probs(I) represents the probability distribution of the molecular fragments with different indices after being adjusted by the temperature parameter, multinomial is the polynomial sampling function in pytorch. The meaning of this function is that after the model outputs the probability distribution, the next Token is sampled from the probability distribution Adjust_Probs(I) in a polynomial sampling manner. This sampling method is used in this model to increase the randomness of the sampling process and improve the robustness of the model.

[0034] Finally, using the sequence embedding of the current fragment and the updated hidden state, input them into the encoder again to generate the next Token until the decoder generates the end-of-sequence marker EOS or the generated fragment sequence reaches the specified maximum length (8 fragments), then the generation process stops. After the generation is completed, connect them in sequence at the virtual atoms of the fragments according to the generated order, so as to form a simplified linear molecular input specification representation of a complete molecule.

[0035] Among them, the construction method further includes the evaluation of the model generation result: the model is trained through the training set, and the model is applied to the protein pockets in the test set, calculate the general indicators of the designed drug molecules, and compare them with two advanced baseline models.

[0036] The present invention also includes a device comprising a neural network model for designing drug molecules targeting a target, the device includes:

[0037] A target feature encoding module, configured to extract the three-dimensional structure features and the protein pocket latent vector of the target protein pocket; the three-dimensional structure features of the protein pocket include atomic-level features, residue features, and inter-residue features, the atomic-level features include atomic types and relative coordinates, the residue features include residue types and residue backbone dihedral angles, and the inter-residue features include the distance between residues and relative position encoding;

[0038] Construction of a ligand molecule feature encoding module for mapping the input molecular fragment sequence into a potential vector representation of the molecular fragment sequence. Specifically, the chemical structure retrosynthesis splitting algorithm is used to break down small drug molecules into molecular fragments, and thus the fragment sequence of each molecule is obtained. The skip-gram model is used to transform the fragments in the sequence into feature embeddings.

[0039] An autoregressive decoding module, which consists of two gated recurrent unit modules and a self-attention module with 8 heads.

[0040] The present invention also includes a computer device, comprising a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, the steps of the method for constructing the neural network model for designing drug molecules targeting a target or the steps of the prediction method are implemented.

[0041] The present invention also includes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps of constructing the neural network model for designing drug molecules targeting a target or the steps of the prediction method are implemented.

[0042] Based on the method described above, it further includes a model generation result evaluation step: training the model with a training set, applying the model to the protein pockets in the test set, calculating the general indicators of the designed drug molecules, and comparing them with two advanced baseline models, FLAG and PMDM. Specifically, 100 protein pockets are randomly selected from the test set, 100 drug molecules are designed for each pocket, and the indicators widely used by mainstream models to evaluate the properties of candidate drugs are selected to evaluate various properties of the generated molecules: the synthetic accessibility score (SA) is the normalized synthetic accessibility score, which is used to measure the difficulty of compound synthesis, and the higher the score, the easier it is to synthesize; the drug-likeness score (QED) is used to evaluate the similarity between the molecules generated by the model and drugs, and the higher the score, the more likely the molecule is to become a drug; the proportion of high-affinity molecules (High Affinity) represents the proportion of molecules in the molecules predicted by the model whose docking scores are better than the docking scores of the real molecules; the octanol-water partition coefficient (LogP), for suitable drug molecules, this coefficient is generally between 0.4 and 5.6; Lipinski's rule of five (Lipinski) is the comprehensive feature of drug-like molecules; the docking score (Vina Score) is calculated by the AutoDock Vina software and is used to evaluate the binding affinity between the drug molecule and its protein target, and the lower the value, the better the match between the drug molecule and the binding target; diversity is used to evaluate the diversity of the generated molecules, and the higher the value, the more diverse the model generation results are. The calculation method is 1 minus the average Tanimoto similarity of pairwise molecules.

[0043] The present invention also includes a method for predicting drug molecules based on the model constructed by the above method. The method for predicting drug molecules includes inputting the target protein to be designed into the model to obtain candidate drug molecules, and selecting the drug molecule with the highest synthetic accessibility and druggability scores calculated from the docking scores of each molecule.

[0044] Advantages: The present invention characterizes the target pocket at three levels of atoms, residues, and between residues, thus considering the target information more comprehensively; uses molecular fragments obtained by chemical structure retrosynthetic splitting algorithm as units to characterize and generate drug molecules, and the generated molecules have higher drug development potential; the drug design model involved in the present invention consists of two encoders and a gated self-attention decoder. The encoders are used to extract the features of the target and ligand molecules, and the decoder realizes the association of the two features and generates the probability of molecular fragments. The model can effectively design drug molecules for specific targets. Through the fine protein pocket characterization method, the molecular fragment characterization method based on chemical structure retrosynthetic splitting algorithm, and the model architecture based on gated recurrent unit, two encoders of graph neural network and self-attention mechanism, the present invention can predict structurally reasonable drug molecules for specific protein pocket targets. The innovation points of the model of the present invention are: 1. Establish a more refined target characterization method, thus considering the structural information of the target more comprehensively; 2. Characterize and generate drug molecules with molecular fragments obtained by chemical structure retrosynthetic splitting algorithm as units, and the generated molecules have higher drug development potential; 3. Design an encoding-decoding model architecture based on gated recurrent unit and self-attention module, and the model can effectively design drug molecules for specific targets. The average synthetic accessibility score of the molecules designed by the model constructed in the present invention reaches 0.769, and the average druggability score reaches 0.607, higher than the baseline model PMDM. Description of the Drawings

[0045] Figure 1 The process of molecular fragmentation and construction of fragment data set;

[0046] Figure 2 It is a schematic diagram of the model process;

[0047] Figure 3 It is a graph of the change of model loss during training, the abscissa represents the training rounds, and the ordinate represents the corresponding training loss of each round;

[0048] Figure 4 It is a molecular generation case graph for two specific targets, ubiquitin-conjugating enzyme Ube2T and p21-activated kinase 4, using different methods. Calculate the synthetic accessibility, druggability, and docking score with the target of the generated molecules, and the best results among the three methods are shown in bold; Detailed Embodiments

[0049] Example 1 Construction of a Neural Network Model for Designing Drug Molecules Targeting Targets

[0050] The construction method of the neural network model for designing drug molecules targeting targets of the present invention includes the following steps:

[0051] (1) Construction of the target feature encoding module: A more refined method for characterizing protein pockets is designed, incorporating richer three-dimensional feature information. Compared with simply considering residue sequence information, it can represent the characteristics of protein pockets more accurately. Specifically, for each protein pocket, consider its characteristics at three levels: atomic, residue, and inter-residue:

[0052] (1) Atomic-level features: Include atomic type and relative coordinates. First, read the atomic type and coordinate information of the four backbone atoms (N, C α , C, O) of each amino acid residue in the pocket. Then, calculate the four-dimensional one-hot encoded vector of the atomic type corresponding to each atom, and calculate the relative coordinates (Δx α , Δy i , Δz i ) of each atom relative to the C i atom within the same residue (Eq.1):

[0053]

[0054] where (x i , y i , z i ) represents the coordinates of a certain backbone atom, and represents the coordinates of the C α atom of the residue where the atom is located.

[0055] (2) Residue features: Include residue type and residue backbone dihedral angle. First, calculate the one-hot encoded vector of 21 residue types corresponding to each residue in the pocket (including the 20 common amino acid residue types and the residue type X that does not belong to the common residue types). Then, calculate the dihedral angles of the target backbone according to the coordinates of the N, C α , C three backbone atoms of all residues. Specifically as follows: Calculate the displacement vector u between adjacent backbone atoms (Eq.2-3):

[0056] dC i = (x i , y i , z i ) - (x i-1 , y i-1 , z i-1 ) Eq.2

[0057]

[0058] Among them, dC i represents the coordinate difference between adjacent backbone atoms, and u i is the normalized dC i .

[0059] After obtaining the displacement vector u, calculate the normal vector n between adjacent displacement vectors. n represents the normal vector of the plane formed by every three adjacent atoms (Eq. 4):

[0060]

[0061] Calculate the angle D between two planes with respect to the common edge according to two adjacent normal vectors (Eq. 5):

[0062] cos D = n i × n i-1 ; D = arccos(cosD) Eq. 5

[0063] Assign the angle to the atom before the common edge, the first atom and the last two atoms on the residue backbone that have not obtained the angle, and the angle is replaced by zero. Each residue obtains three angles in total to represent the dihedral angle information of its backbone.

[0064] After that, calculate the sine value and cosine value for the three dihedral angles of each residue. The sine values and cosine values of the three dihedral angles are concatenated to obtain a six-dimensional dihedral angle feature vector (Eq. 6):

[0065] dihedrals = [sinD, cosD] Eq. 6

[0066] (3) Inter-residue features: including inter-residue distance and inter-residue relative position encoding. Calculate the coordinates of the residue center according to the coordinates of the four backbone atoms (Eq. 7):

[0067]

[0068] C i represents the residue center coordinate of the i-th residue, and X ij is the coordinate of the j-th atom in the i-th residue. The distance between the centers of two residues is represented by the Euclidean distance (Eq. 8):

[0069]

[0070] The relative position encoding is calculated using the position encoding formula in the Transformer model based on the residue index difference, aiming to represent the relative position relationship between residues. Specifically: calculate the index difference between each residue and its neighboring residues (Eq. 9):

[0071] dij = e ij -i Eq.9

[0072] where i represents the index value of the i-th residue in the pocket residue sequence, and E ij represents the index of the j-th neighbor of the i-th residue. After obtaining the index difference, multiply it by the frequency f k to get the product (Eq. 10 - 11):

[0073]

[0074] where N emb represents the dimension of the position encoding, and k represents the dimension index of different position encodings. Finally, generate the relative position encoding through the sine and cosine functions (Eq. 12):

[0075]

[0076] The relative position encoding calculated in this way can help the model learn the relative position relationships of each residue in the pocket residue sequence.

[0077] Furthermore, after obtaining various features, the atomic-level features of the four backbone atoms in each residue are linearly transformed and pooled to obtain the atomic-level feature embedding of the residue. After the residue features are converted into residue feature embeddings through normalization and linear transformation, they are added to the corresponding atomic-level feature embeddings to obtain a composite embedding, which serves as the feature of the nodes in the subsequent protein graph. The inter-residue features are converted into inter-residue feature embeddings through normalization and linear transformation, which serve as the features of each edge in the subsequent protein graph. Among them, the sum of the atomic-level feature embeddings obtains a 256-dimensional composite embedding, and the inter-residue features are converted into 256-dimensional inter-residue feature embeddings;

[0078] Furthermore, after obtaining the point feature embedding and the edge feature embedding, the protein pocket is represented as a k-nearest neighbor graph (k = 30), where each node in the graph represents each amino acid in the protein pocket, and each edge represents the connection between residues.

[0079] Furthermore, map the point features and edge features in the protein graph to the hidden space through a linear layer. The message-passing graph neural network takes the mapped node features and edge features as inputs, performs information passing and aggregation to optimize the feature representation of each node, and then outputs the feature representation of each protein graph through global average aggregation as the latent representation p of the protein. For the latent representation p of each protein, calculate its mean μ(p) and variance σ 2 (p), and use the normal distribution z p (Eq. 13) as the distribution of the protein pocket latent vector.

[0080] z p= μ(p) + σ 2 (p)·∈ Eq.13

[0081] (2) Construction of the molecular fragment sequence feature encoding module: The construction of the molecular fragment sequence feature encoding module includes: Using the chemical structure reverse synthesis splitting algorithm, the drug small molecule is sequentially broken into molecular fragments from left to right according to the simplified linear molecule input specification, and the fragment sequence of each molecule is obtained. After obtaining the fragment sequence of the drug small molecule, the molecular sequence is regarded as a sentence, and each fragment decomposed from the molecular sequence is regarded as a word. Different fragments are converted into 256-dimensional word embeddings through the skip-gram model with negative sampling. The molecular fragment encoder accepts each molecular fragment sequence embedding x = (x1, x2, …, x i ) as input. For each fragment embedding x i , it is converted into a hidden state h i (Eq.14) through two layers of gated recurrent units. The hidden state h Final of the last fragment embedding obtained through forward propagation is used as the latent representation of the entire fragment sequence, and the latent vector z x (Eq.15) distribution of the molecular fragment sequence is calculated in the same way as the protein pocket encoder:

[0082] h i = GRU(x i .h i-1 ) Eq.14

[0083] z x ~N(μ(x), σ 2 (x)) Eq.15

[0084] where μ represents the mean, σ represents the variance, and N represents the normal distribution calculated with this mean and method;

[0085] (3) Autoregressive decoding module: The decoding module consists of two layers of gated recurrent unit modules and an 8-head self-attention module. First, the protein pocket latent vector z p and the molecular fragment sequence latent vector z x obtained by reparameterization are added to obtain a composite latent vector c z , which is linearly transformed and used as the input of the gated recurrent unit. After the gated recurrent unit outputs the final hidden state hc Final , self-attention calculation (Eq.16 - 17) is performed on it through an 8-head self-attention module:

[0086] Q = W Q ·hc Final ; K = W K ·hc Final ; V = W V ·hcFinal Eq.16

[0087]

[0088] Q, K, and V are the query, key, and value in the basic principle of the self-attention mechanism, and W Q , W K , W V are the product matrices for forming Q, K, and V respectively; in this embodiment, d k is 128.

[0089] After the multi-head self-attention scores are concatenated, they are used to calculate the probability distribution Probs(I) of the molecular fragments (Eq. 18 - 19):

[0090] Multihead(Q, K, V) = Concat(head1, head2, …, head n ) · W out Eq. 18

[0091] Probs(I) = softmax(Multihead(Q, K, V)) Eq. 19

[0092] where n refers to the nth attention head, and W out refers to the output matrix, and the result obtained by multiplying the concatenated result of the attention heads and the output matrix is used as the final output of the multi-head self-attention.

[0093] Furthermore, an autoregressive method is used to generate each fragment one by one. First, a zero vector is initialized as the first Token (fragment) of the starting sequence. After two encoders process the starting sequence and the corresponding target pocket, the hidden states are obtained. Then, the decoder uses the starting sequence and the hidden states to calculate the probability distribution of the molecular fragments, samples the next Token from this distribution in a multinomial sampling manner, and adjusts the randomness of the sampling process using the temperature parameter (introducing the temperature parameter is a method to adjust the sampling randomness. Setting the temperature coefficient greater than 1 will increase the diversity and randomness of the sampling, and less than 1 will increase the certainty of the sampling and reduce the randomness. As long as it is set greater than 1, the randomness can be increased. In this experiment, it is set to 1.2, and users can adjust it according to their own situations) (Eq. 20 - 21):

[0094]

[0095] Next_Token = multinomial(Adjust_Probs(I), 1) Eq. 21

[0096] Among them, $I$ represents the molecular fragment, which represents the probability distribution of molecular fragments with different indices after temperature parameter adjustment. Multinomial is a polynomial sampling function in PyTorch. The meaning of this function is that after the model outputs the probability distribution, the next token is sampled from the probability distribution by polynomial sampling. This sampling method is used in this model to increase the randomness of the sampling process and improve the robustness of the model.

[0097] The current fragment sequence embedding and the updated hidden state are input into the encoder again to generate the next token. The generation process stops until the decoder generates the end token or the generated fragment sequence reaches the specified maximum length (8 fragments).

[0098] (4) Loss function and training method: The CrossDocked standard dataset (FRANCOEUR P G, MASUDAT, SUNSERI J, et al. Three-Dimensional Convolutional Neural Networks and a Cross-Docked Data Set for Structure-Based Drug Design[J]. J Chem Inf Model, 2020, 60(9): 4200 - 15. doi: 10.1021 / acs.jcim.0c00411.) is used to train the model. The loss function consists of the reconstruction loss and the relative entropy (Eq. 22):

[0099] Loss = CE(X′, X) + KL(N(μ(x), σ 2 (x)), N(0, 1)) Eq. 22

[0100] Among them, the reconstruction loss term CE calculates the cross-entropy loss for the fragment index values at each position of the output sequence and the true sequence, aiming to make the sequence generated by the decoder as close as possible to the true sequence; the purpose of the KL divergence loss term is to make the latent variables generated by the molecular fragment encoder conform as much as possible to the standard normal distribution;

[0101] The Adam optimizer is used in the model training process, and a learning rate scheduler is introduced to automatically adjust the learning rate.

[0102] (5) Evaluation of Model Generation Results: The model is trained using the training set and applied to the protein pockets in the test set. General metrics are calculated for the designed drug molecules and compared with two advanced baseline models, FLAG (ZHANG Z, MIN Y, ZHENG S, et al. Molecule Generation For Target Protein Binding with Structural Motifs[J]. International Conference on Learning Representations, 2023. doi: 10.48550 / arXiv.2305.13997.) and PMDM (HUANG L, XU T, YU Y, et al. A dual diffusion model enables 3D molecule generation and lead optimization based on target pockets[J]. Nature Communications, 2024, 15. doi: 10.1038 / s41467-024-46569-1.). Specifically, 100 protein pockets are randomly selected from the test set, and 100 drug molecules are designed for each pocket. Metrics widely used by mainstream models to evaluate the properties of candidate drugs are selected to evaluate various properties of the generated molecules: The synthetic accessibility score (SA) is the normalized synthetic accessibility score, which is used to measure the difficulty of compound synthesis. The higher the score, the easier the synthesis; The drug-likeness score (QED) is used to evaluate the similarity between the molecules generated by the model and drugs. The higher the score, the more likely the molecule is to become a drug; The proportion of high-affinity molecules (High Affinity) represents the proportion of molecules in the molecules predicted by the model whose docking scores are better than the docking scores of the real molecules; The octanol-water partition coefficient (LogP). For suitable drug molecules, this coefficient is generally between -0.4 and 5.6; The Lipinski's rule of five (Lipinski) is the comprehensive characteristic of drug-like molecules; The docking score (Vina Score) is calculated by the AutoDock Vina software and is used to evaluate the binding affinity between the drug molecule and its protein target. The lower the value, the better the match between the drug molecule and the binding target; Diversity is used to evaluate the diversity of the generated molecules. The higher the value, the more diverse the model generation results. The calculation method is 1 minus the average Tanimoto similarity of pairwise molecules.

[0103] Example 2 Ablation Experiment of Different Protein Pocket Features

[0104] In theory, an effective pocket representation method enables the model to better learn the differences between different pockets, thereby generating ligand molecules more specifically. Ablation experiments were conducted on 100 randomly selected target pockets, and 100 drug molecules were designed for each target pocket. After various features were gradually incorporated into the model, the average molecular docking score (Vina Score) of all generated drug molecules with the corresponding targets was calculated.

[0105] The table shows the results of the ablation experiments. The results indicate that, based on only considering residue types, the docking score of the model's generated results was improved to a certain extent with the incorporation of each new type of feature information. After all feature information was incorporated, the model achieved the best result. From the change in the docking score, the features of residue backbone dihedral angles, residue atom types, and inter-residue distances contributed significantly to the model. This result shows that compared with the representation method that only considers residue types, the new target pocket representation method designed in the present invention can effectively enhance the learning ability of the model.

[0106] Table 1 Ablation Experiments on Different Features of Protein Pockets

[0107]

[0108] Note: "√" represents considering this feature, and "-" represents not considering this feature

[0109] Example 3 Training of the Model and Evaluation of the Results

[0110] First step, prepare the data for model testing. The model constructed in Example 1 of the present invention was trained using the CrossDocked standard dataset, which contains 183,461 pairs of target pocket-ligand pairs. The present invention removed the data that could not be converted into canonical SMILES representation by RDkit, as well as the data with incomplete coordinate information of backbone atoms (N, C α , C, O) in the target pocket files. Finally, 172,216 pairs of target pocket-ligand pair data were retained and randomly divided into a training set, a test set, and a validation set according to the ratio of 0.7:0.15:0.15.

[0111] Second step, model training. The loss function consists of a reconstruction loss and a relative entropy. The Adam optimizer was used during the training process, the initial learning rate was set to 0.001, and a ReduceLROnPlateau learning rate scheduler was introduced to automatically adjust the learning rate. The batch size was set to 32. The training device was a GeForce RTX 4090 GPU, and the code was implemented using python3.9 and pytorch2.2.1. The dataset was randomly split according to the ratio of the training set, the test set, and the validation set of 0.7:0.15:0.15. The loss curve during the model training process is as Figure 2As shown, the results indicate that when the model was only trained for 30 epochs, it had successfully approached convergence.

[0112] Step 3: Generate tests. Use 100 protein targets randomly selected from the test set as inputs, and generate 100 candidate drug molecules for each test target through the trained model. After training two baseline models, FLAG and PMDM, generate candidate drug molecules in the same way.

[0113] Step 4: Generate result evaluation and comparison. Evaluate and compare all the molecules generated by this model and the two baseline models, as well as the molecules in the test set, using general metrics. Among them, the synthetic accessibility score SA is calculated as the weighted calculation score of fingerprint frequencies and the proportion of non-standard structural features; the drug-likeness score QED, the octanol-water partition coefficient LogP, and the Lipinski's five rules are calculated by the RDkit package; the docking score Vina Score is calculated by the AutoDock Vina software; the high-affinity molecule ratio High Affinity represents the proportion of molecules in the molecules predicted by the model whose docking scores are better than the docking scores of the real molecules; the diversity score Diversity is calculated as 1 minus the average Tanimoto similarity of pairwise molecules, and the Tanimoto similarity is calculated by the RDkit package. The evaluation results are shown in Table 2. Table 2 shows the general metric scores of the test set and the molecule sets generated using different methods. The meanings of each metric are as described in the previous text. Among them, the up arrow in parentheses indicates that the higher the score of this item, the better, and the down arrow indicates that the lower the score of this item, the better. The best results in each method are shown in bold; the results indicate that the average synthetic accessibility score SA of the molecules designed by this model reaches 0.769, and the average drug-likeness score QED reaches 0.607, higher than the baseline model PMDM (0.611 and 0.594); in terms of Lipinski and docking scores, this model also achieved scores better than those of the molecules in the test set (real molecules). The above evaluation results show that the molecules designed by this model have a more reasonable structure.

[0114] Table 2 Evaluation Results of Test Set Molecules and Different Models

[0115]

[0116] Among them, the bold is the best among different models.

[0117] Step 5: Case generation study. Two specific targets were selected for case studies, namely ubiquitin-conjugating enzyme (Ube2T, PDB id: 5NGZ) and p21-activated kinase 4 (PAK4, PDB id: 5I0B). Input the targets into the three models to design 100 candidate drug molecules respectively, and calculate and compare the synthetic accessibility and drug-likeness scores of the drug molecules with the best docking scores for each, and the comparison results are asFigure 4 As shown. The results indicate that in target 1, the drug-likeness score QED of the drug molecules designed by this model is 0.88, which is 23.9% higher than the best result in the baseline model; in target 2, the synthetic accessibility score SA and the drug-likeness score QED of the drug molecules designed by this model are 0.63 and 0.85 respectively, which are 8.6% and 16.4% higher than the best results in the baseline model respectively. Overall, the performance of this model is excellent.

Claims

1. A method for constructing a neural network model for designing drug molecules targeting specific sites, characterized in that, It includes the following steps: (1) Construction of a target feature encoding module for extracting the features of the target target pocket and obtaining the latent vector of the target pocket through reparameterization; the features of the target pocket include atomic-level features, residue features, and inter-residue features. The atomic-level features include atomic type and relative coordinates, the residue features include residue type and residue backbone dihedral angles, and the inter-residue features include the distance between residues and relative position encoding; (2) Construction of a ligand molecule feature encoding module for mapping the fragment sequence of the ligand molecule into a latent vector representation of the fragment sequence of the ligand molecule; specifically, the chemical structure retrosynthetic cleavage algorithm is used to break the ligand molecule into molecular fragments, and thus the fragment sequence of each ligand molecule is obtained; the skip-gram model is used to transform the fragment sequence of each ligand molecule into a feature embedding; (3) Construction of an autoregressive decoding module, which consists of two gated recurrent unit modules and a self-attention module; the latent vector of the target pocket and the latent vector of the ligand molecule fragment sequence are added together to obtain a composite latent vector as the decoder input, and the fragment sequence is gradually generated in an autoregressive manner. When the termination marker is generated or the generated fragment sequence reaches the specified maximum length, the generation terminates; (4) Model training: The model is trained using the preprocessed dataset; (5) Using the targets in the test set as inputs to generate ligand molecules.

2. The method for constructing a neural network model for designing drug molecules targeting a target according to claim 1, wherein After obtaining various features in step (1), the atomic-level features are embedded through linear transformation and pooling, the residue features are embedded through normalization and linear transformation, and added to the corresponding atomic-level feature embeddings to obtain a composite embedding as the point feature embedding in the target pocket graph. The inter-residue features are embedded through normalization and linear transformation as the edge feature embedding in the target pocket graph.

3. The method for constructing a neural network model for designing drug molecules targeting a target according to claim 1, characterized in that, Step (1) also includes representing the target pocket as a k-nearest neighbor graph after obtaining the point feature embedding and the edge feature embedding. Each node in this graph represents each residue in the target pocket, and each edge represents the connection between residues.

4. The method for constructing a neural network model for designing drug molecules targeting a target according to claim 1, characterized in that, In step (1), it also includes mapping the point features and edge features in the target pocket graph to the hidden space through a linear layer. The target pocket encoder takes the mapped node features and edge features as inputs, performs information passing and aggregation to optimize the feature representation of each node, and then outputs the feature representation of each target pocket graph through global average aggregation as the latent representation p of the target pocket. For the latent representation p of each target pocket, ∈ is the noise sampled from the standard normal distribution, and its mean μ(p) and variance σ 2 (p) are calculated, and the normal distribution z p is used as the distribution of the target pocket latent vector: z p = μ(p) + σ 2 (p)·∈ 5. The method for constructing a neural network model for designing drug molecules targeting a target according to claim 1, characterized in that The construction steps of the molecular fragment sequence feature encoding module described in step (2) include: using the chemical structure retrosynthetic cleavage algorithm to break the ligand molecule into molecular fragments, and thus obtaining the fragment sequence of each molecule. After obtaining the fragment sequence of the ligand molecule, regarding the ligand molecule sequence as a sentence and each fragment decomposed from the ligand molecule sequence as a vocabulary, converting different fragments into word embeddings through the skip-gram model with negative sampling. The molecular fragment encoder accepts the embedding x = (x1, x2, …, x i ) of each ligand molecule fragment sequence as input. For each molecular fragment embedding x i , converting it into a hidden state h i through two gated recurrent units, and taking the hidden state h Final of the last fragment embedding obtained through forward propagation as the latent representation of the entire fragment sequence, and calculating the latent vector z x distribution of the fragment sequence of the ligand molecule in a similar way to the target pocket encoder: h i = GRU(x i .h i-1 ) z x ~ N(μ(x), σ 2 (x)), GRU represents the gated recurrent unit, and N represents the normal distribution.

6. The method for constructing a neural network model for designing drug molecules targeting a target according to claim 1, wherein The construction steps of the autoregressive decoding module described in step (3) include: First, add the target pocket latent vector z p reparameterized and the latent vector z x of the fragment sequence of the ligand molecule to obtain a composite latent vector c z . Then, convert c z into an initial hidden representation through a linear layer as the input of the gated recurrent unit. After the gated recurrent unit outputs the final hidden state hc Final , perform self-attention calculation on it through the self-attention module; The autoregressive method is used to generate the molecule fragment by fragment. First, a zero vector is initialized as the first Token (fragment) of the starting sequence. After the molecule fragment encoder and the target pocket encoder process the starting sequence and the corresponding target pocket, the hidden state is obtained. Then, the decoder uses the starting sequence and the hidden state to calculate the probability distribution of the molecule fragment, samples the next Token from this distribution in a multinomial sampling manner, and adjusts the sampling process using the temperature parameter: Next_Token = multinomial(Adjust_Probs(I), 1) where I represents the molecule fragment, Adjust_Probs(I) represents the probability distribution of molecule fragments with different indices after adjustment by the temperature parameter, and multinomial is the multinomial sampling function in pytorch; Finally, the sequence embedding and updated hidden state of the current fragment are used to input the encoder again to generate the next Token. The generation process stops when the decoder generates the termination mark EOS or the generated fragment sequence reaches the specified maximum length. After the generation is completed, the virtual atoms of the fragments are connected in sequence in the order of generation to form a simplified linear molecular input specification representation of a complete molecule.

7. A method for predicting a drug molecule by using a model constructed by the method according to any one of claims 1 to 6, characterized in that, The method for predicting drug molecules comprises inputting the target protein to be designed into the model to obtain candidate drug molecules, and taking the drug molecule with the highest docking score to calculate the drug molecule with the highest synthetic accessibility and drug-likeness score.

8. A device comprising a neural network model for target-oriented drug molecule design, the device comprising: A target feature encoding module, used to extract the features of the target pocket and calculate the target pocket potential vector; the features of the target pocket include atomic-level features, residue features and inter-residue features, the atomic-level features include atomic types and relative coordinates, the residue features include residue types and residue backbone dihedrals, and the inter-residue features include distances between residues and relative position encoding; The ligand molecule feature encoding module is used to map the input fragment sequence of the ligand molecule into a latent vector representation of the fragment sequence of the ligand molecule. Specifically, the chemical structure retrosynthetic splitting algorithm is used to split the ligand molecule into molecular fragments, and thus the fragment sequence of each ligand molecule is obtained; the skip-gram model is used to convert the fragment sequence of each ligand molecule into feature embedding; Autoregressive decoding module,The decoding module consists of two gated recurrent unit modules and a self-attention module.

9. A computer device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the computer program, the steps of the method for constructing a neural network model for target-oriented design of drug molecules according to any one of claims 1 to 6 or the steps of the prediction method according to claim 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of constructing a neural network model for target-oriented drug molecule design according to any one of claims 1 to 6 or the steps of the prediction method according to claim 7 are implemented.

Citation Information

Cited By

  • Small molecule generation method based on target protein under deep neural network

    CN121215021A