Protein pocket hypergraph constraint-based drug molecule generation method and system
By constructing protein pocket hypergraphs and diagrams, extracting and fusion feature embedding representations, as constraints of the conditional molecular generation model, the problems of long-term, high cost and low binding specificity of traditional drug design methods are solved, and more efficient and safer drug molecule generation is achieved.
Patent Information
- Application Number
- CN202510177390.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional drug design methods are time-consuming, costly and accompanied by high risks, and ignoring the structural information of protein pockets may reduce the binding specificity of the drug to the target and increase the risk of side effects.
Using a drug molecule generation method based on protein pocket hypergraph constraints, a protein pocket hypergraph and graph are constructed, feature embedding representations at amino acid and atom level are extracted, and feature fusion is performed to optimize drug molecule generation as a constraint on the conditional molecular generation model.
It improves the binding affinity and specificity of drug molecules and protein pockets, reduces the risk of side effects, and improves the effectiveness, safety and efficiency of drug development.
Smart Images

Figure CN120126549A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics medicine, and particularly relates to a method and system for generating drug molecules based on protein pocket hypergraph constraints. Background Art
[0002] The main purpose of drug design is to identify and develop active small molecule compounds closely related to disease development. By designing and optimizing these small molecules, they can precisely bind to the target protein pocket, thereby regulating the activity of the protein and achieving the purpose of treating diseases. Traditional drug design methods are time-consuming, costly, and accompanied by high risks. Integrating data-driven deep learning technology with molecular generation technology is expected to significantly accelerate the process of drug design and generate molecules that bind tightly to the target protein. In addition, by generating new molecules from scratch, problems such as drug resistance can be alleviated, which has important practical significance.
[0003] De novo drug design based on deep learning is to learn the representation and probability distribution of molecules, extract representative features, generate low-dimensional continuous representations, and finally sample and generate new data from the learned data distribution. Currently, a widely adopted strategy is based on the assumption that "similar structures may have similar properties". By deeply studying the molecular structure characteristics of specific disease targets, a series of similar molecules with high affinity and specific biological activity for the target are designed and generated. Although this method can utilize known information to accelerate the discovery of new drugs, it may lead to inaccurate property prediction due to minor structural differences, and over-reliance on this method may limit the breadth and depth of drug innovation. In addition, due to ignoring the structural information of the protein pocket, the binding specificity of the drug to the target may be reduced, the risk of side effects may be increased, and the R & D efficiency may be reduced. Summary of the Invention
[0004] Aiming at the above problems, the purpose of the present invention is to provide a method and system for generating drug molecules based on protein pocket hypergraph constraints, and specifically proposes a strategy that integrates hypergraph feature embedding representation and graph feature embedding representation to constrain the generation of drug molecules.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] The first aspect of the present invention proposes a method for generating drug molecules based on protein pocket hypergraph constraints, including the following steps:
[0007] S1, constructing a drug small molecule dataset and a protein pocket-ligand dataset, where the protein pocket-ligand dataset includes protein pocket data and corresponding ligand data information;
[0008] S2. Construct a protein pocket hypergraph and a protein pocket graph using the protein pocket data described in step S1. Among them, the protein pocket hypergraph includes an amino acid hypergraph representing the protein pocket with amino acids as nodes and an atomic hypergraph with atoms as nodes, and the protein pocket graph is an atomic graph representing the protein pocket with atoms as nodes;
[0009] S3. Construct a feature extraction module, map the nodes and edges in the amino acid hypergraph and atomic hypergraph of the protein pocket in step S2 to a low-dimensional vector space, and respectively obtain a hypergraph feature embedding representation F HG at the amino acid level and a hypergraph feature embedding representation F' HG . At the same time, generate an atomic-level graph feature embedding representation F G for the nodes in the protein pocket atomic graph;
[0010] S4. Construct a feature fusion module, fuse the hypergraph feature embedding representations F HG and F' HG obtained in step S3, and the graph feature embedding representation F G to obtain a fused feature embedding representation F fus of the protein pocket;
[0011] S5. Construct a conditional molecule generation model based on a gated recurrent unit, and pre-train the model using the drug small molecule dataset described in step S1 so that it learns the rules corresponding to drug molecules and actual chemical structures;
[0012] S6. Use the fused feature embedding representation F fus of the protein pocket obtained in step S4 as a constraint for molecule generation of the pre-trained model obtained in step S5, and combine with the ligand data described in step S1 to retrain the pre-trained model to optimize the conditional molecule generation model;
[0013] S7. Use the conditional molecule generation model optimized in step S6 to generate drug molecules for specific protein targets.
[0014] Furthermore, in step S1, the construction of the drug small molecule dataset includes:
[0015] (1) Collect drug small molecule data in the form of SMILES sequences from the ChEMBL database, select drug small molecules with SMILES sequence lengths between 20 characters and 100 characters, and clean the SMILES sequences;
[0016] (2) Convert the cleaned SMILES sequences into SELFIES sequences, and split the SELFIES sequences into token sequences through a predefined tokenization table;
[0017] (3) Generate a vocabulary of small molecule drugs based on all tokens, and map the token sequence to the corresponding integer index to form the final small molecule drug dataset;
[0018] The construction of the protein pocket-ligand dataset includes:
[0019] (1) Collect protein pocket source-ligand source data from the PDBbind database, where the protein pocket source data is in PDB format and the ligand source data is in SDF and Mol2 formats;
[0020] (2) Convert the ligand source data in Mol2 format to a SMILES sequence and clean it, convert the cleaned SMILES sequence to a SELFIES sequence; split the SELFIES sequence into a token sequence through a predefined tokenization table; generate a ligand vocabulary based on all tokens, and map the token sequence to the corresponding integer index;
[0021] (3) Obtain the pocket information of the protein corresponding to the ligand from the protein pocket source data in PDB format; integrate the ligand sequence information and the corresponding protein pocket information to form a protein pocket-ligand dataset.
[0022] Furthermore, in step S2, representing the protein pocket as an amino acid hypergraph with amino acids as nodes includes:
[0023] Obtain the amino acid sequence of the protein pocket, use the central carbon atom of a single amino acid as the current node, calculate the distance between the central carbon atoms of the remaining amino acids and the current node, and construct a spatial structure hyperedge with all the central carbon atoms whose distance is less than a predetermined threshold and the current node; traverse all the amino acids of the protein pocket in turn to obtain the spatial structure hyperedge corresponding to each amino acid;
[0024] Representing the protein pocket as an atom hypergraph with atoms as nodes includes:
[0025] Obtain the amino acid sequence of the protein pocket, use each heavy atom in the protein pocket as a node, for each amino acid, save all its heavy atoms in a set, and each set constructs a hyperedge;
[0026] Representing the protein pocket as an atom graph with atoms as nodes includes:
[0027] Obtain the amino acid sequence of the protein pocket, use each heavy atom in the protein pocket as a node, traverse each heavy atom in turn, calculate the distance between the current heavy atom and the remaining heavy atoms, and define an edge between two heavy atoms whose distance is less than a predetermined threshold; if there is a chemical bond between two heavy atoms, also define it as an edge.
[0028] Further, in step S3, a hypergraph neural network is used to extract the hypergraph feature embedding representation F at the amino acid level HG and the hypergraph feature embedding representation F' at the atomic level HG , and the hypergraph neural network is composed of at least two layers of hypergraph convolutional layers. The node feature update of each layer is shown in formula (1):
[0029]
[0030] where X (l) represents the node feature matrix of the l-th layer of the protein pocket, W is the hyperedge weight, H is the incidence matrix of the protein pocket, Θ is used to extract useful node features in the protein pocket, D v represents the degree matrix of the hyperedges, D e represents the degree matrix of the vertices, and σ represents the activation function
[0031] Further, in step S3, a graph attention neural network is used to generate the atomic-level graph feature embedding representation F for the nodes in the protein pocket atomic graph G , and the graph attention neural network is composed of at least two layers of graph attention layers. The node feature update of each layer is shown in formula (2):
[0032]
[0033] where is the initial node feature is the updated node feature, σ represents the activation function, K represents K attentions, and α ij k represents the corresponding attention coefficient, W K is the corresponding weight matrix, and α ij k and W K are both learnable parameters
[0034] Further, in step S4, obtaining the fused feature embedding representation F of the protein pocket fus , includes
[0035] First, the hypergraph feature embedding representation F' at the atomic level HG and the graph feature embedding representation F G are concatenated, and an average pooling strategy is used to reduce the dimension and extract features from the concatenated features to achieve preliminary feature fusion; then it is further fused with the hypergraph feature embedding representation F at the amino acid level HG to obtain the fused feature embedding representation F of the protein pocket through concatenation fus .
[0036] Further, in step S5, the conditional molecule generation model based on the gated recurrent unit includes an embedding layer, at least three gated recurrent unit layers, and an activation layer;
[0037] During pre-training, the token sequence mapped to the corresponding integer index in the small molecule drug dataset is used as the input of the model embedding layer. The probability of the next possible output token in the predicted sequence by the conditional molecule generation model is shown in formula (3):
[0038]
[0039] where P(y i ) represents the probability of predicting the i-th possible token, and m represents the total number of tokens in the small molecule drug vocabulary.
[0040] Further, in step S6, the optimization of the conditional molecule generation model includes:
[0041] Embedding the fusion feature representation F of the protein pocket fus as the initial input of the hidden state of the gated recurrent unit to constrain molecule generation. The token sequence of the corresponding ligand mapped to the corresponding integer index is used as the input of the model embedding layer. A conditional probability model is constructed for the probability of the molecule output token to learn the molecular probability distribution conditional on the protein pocket. The conditional molecule generation probability is shown in formula (4):
[0042]
[0043] where P(molecule|pocket) represents the molecule generation probability conditional on the protein pocket, n is the number of tokens after word segmentation of the corresponding ligand sequence, and x t is the token of the molecular sequence at the t-th iteration.
[0044] Further, the conditional molecule generation model uses the cross-entropy between the predicted value and the true value as the loss function, and the loss function is shown in formula (5):
[0045]
[0046] where Loss is the value of the loss function, n is the number of tokens after word segmentation of a sequence, y i represents the true value of the i-th token, and y′ i represents the predicted value of the i-th token.
[0047] In a second aspect of the present invention, a drug molecule generation system based on protein pocket hypergraph constraints is proposed for implementing the drug molecule generation method based on protein pocket hypergraph constraints as described above, including:
[0048] A dataset construction module configured to construct a small molecule drug dataset and a protein pocket-ligand dataset, where the protein pocket-ligand dataset includes protein pocket data and corresponding ligand data information;
[0049] A hypergraph and graph construction module configured to construct a protein pocket hypergraph and a protein pocket graph. Among them, the protein pocket hypergraph includes an amino acid hypergraph representing the protein pocket with amino acids as nodes and an atomic hypergraph with atoms as nodes, and the protein pocket graph is an atomic graph representing the protein pocket with atoms as nodes;
[0050] A feature extraction module configured to map the nodes and edges in the two protein pocket hypergraphs to a low-dimensional vector space to obtain a hypergraph feature embedding representation F HG at the amino acid level and a hypergraph feature embedding representation F' HG at the atomic level. At the same time, generate an atomic-level graph feature embedding representation F G for the nodes in the protein pocket graph;
[0051] A feature fusion module configured to fuse the hypergraph feature embedding representations F HG and F' HG , as well as the graph feature embedding representation F G to obtain a fused feature embedding representation F fus of the protein pocket;
[0052] A model construction and pre-training module configured to construct a conditional molecule generation model based on a gated recurrent unit and pre-train the model using the small molecule drug dataset so that it learns the rules corresponding to the actual chemical structure of the drug molecule;
[0053] A model optimization module configured to use the fused feature embedding representation F fus of the protein pocket as a constraint condition for molecule generation of the pre-trained model, and combine with the ligand data to retrain the pre-trained model to optimize the conditional molecule generation model;
[0054] A drug molecule generation module configured to use the optimized conditional molecule generation model to generate drug molecules for specific protein targets.
[0055] Compared with the prior art, the present invention has the following beneficial effects:
[0056] (1) The present invention represents protein pockets through hypergraphs with amino acids as nodes, which can simplify the structural complexity and highlight the functional regions, facilitating the analysis of the interactions between amino acids and the binding characteristics with ligands. At the same time, by representing protein pockets through hypergraphs with atoms as nodes, detailed structural information can be provided to accurately capture the complex interaction details between atoms.
[0057] (2) The present invention further represents protein pockets through graphs with atoms as nodes, which can more accurately represent the atomic arrangement and three-dimensional structure inside the pocket, providing richer geometric information. It can not only intuitively display the interactions between atoms, such as hydrogen bonds and hydrophobic interactions, but also help analyze the binding mode and affinity between the ligand and the pocket. By analyzing the atomic graph, hidden structural features and rules can be discovered.
[0058] (3) The present invention uses a method that fuses hypergraph feature embedding representation and graph feature embedding representation to comprehensively characterize the structural features of protein pockets from both the amino acid and atomic levels, and uses this as a conditional constraint for the generation of drug molecules. It can more accurately capture the molecular structure that matches the protein pocket, improve the binding affinity between the generated molecule and the protein pocket, which not only improves the quality of the molecule, but also significantly enhances its binding tightness and specificity with the protein target, helping to improve the effectiveness, safety, and efficiency of drug research and development. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is the basic flowchart of an embodiment of the present invention.
[0060] Figure 2 is Figure 1 a more intuitive flow schematic diagram of the shown basic flowchart.
[0061] Figure 3 is the structural diagram of the protein pocket feature fusion module in an embodiment of the present invention.
[0062] Figure 4 is the structural diagram of the conditional molecule generation model in an embodiment of the present invention.
[0063] Figure 5 is the schematic diagram of the synthetic accessibility score (SA) of the generated molecule in an embodiment of the present invention.
[0064] Figure 6 is the schematic diagram of the quantitative estimation of drug likeness (QED) of the generated molecule in an embodiment of the present invention.
[0065] Figure 7 is the schematic diagram of the docking binding energy of the generated molecule in an embodiment of the present invention.
[0066] Figure 8It is a docking schematic diagram of the generated molecule and protease 3Clpro in an embodiment of the present invention.
[0067] Figure 9 It is a binding schematic diagram of the generated molecule and the pocket of protease 3Clpro in an embodiment of the present invention. Detailed implementation manners
[0068] In order to make the technical means, creative features, achieved purposes and functions of the present invention easy to understand, the preferred implementation schemes of the present invention are described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.
[0069] Term explanations:
[0070] Protein pocket: refers to a cavity or depression on the surface or inside of a protein that is suitable for binding to a ligand;
[0071] Hypergraph: an important concept in graph theory, which expands the expressive power of the traditional graph (Graph). A hypergraph allows an edge to connect multiple vertices, unlike a simple graph that only allows an edge to connect two vertices;
[0072] SMILES (Simplified Molecular-Input Line-Entry System): a text string encoding system for representing molecular structures. A SMILES string can describe the atomic connection order and topological structure of a molecule in a linear manner, including ring, branch and stereochemical information.
[0073] SELFIES (SELF-referencIng Embedded Strings): a string encoding system for representing molecular structures, which is an improved form of SMILES and aims to solve the limitations of SMILES in representing some complex molecular structures.
[0074] Heavy atoms: refer to all atoms in a protein molecule except hydrogen atoms. Unless otherwise specified, the "atoms" in this article all refer to "heavy atoms");
[0075] Hypergraph Neural Networks (HGNN): a deep learning model for processing hypergraph data. Different from traditional Graph Neural Networks (GNN), it can process more complex relationships and structures;
[0076] Graph Attention Network (GAT): A neural network architecture specifically designed to process graph-structured data. It effectively captures complex relationships and feature information between nodes by introducing an attention mechanism to perform weighted aggregation on the neighbors of nodes;
[0077] Gate Recurrent Unit (GRU): A variant of the Recurrent Neural Network (RNN) used for sequence data processing. By introducing a gating mechanism to control the flow of information, it improves the performance and stability of the model;
[0078] Token: In the context of this invention, it refers to a series of basic units or segments into which a SELFIES string is split. These tokens are the result of tokenization of the SELFIES string and can be individual characters, atomic symbols, chemical bonds, bracket pairs, or more complex molecular structure units;
[0079] Tokenization Dictionary: A tokenization dictionary is a rule or dictionary used to split a string into tokens, which defines which character sequences should be regarded as a single token;
[0080] Vocabulary: During the process of string tokenization, a vocabulary is constructed, which contains all possible tokens and their corresponding integer indices;
[0081] Pocket2Drug: An encoder-decoder deep neural network for target-based drug design;
[0082] HGAF (Hypergraph Graph Attention Fusion): A multimodal learning framework that combines hypergraph embedding representation and attention mechanism;
[0083] Synthetic Accessibility score (SA): A method used to evaluate the ease of compound synthesis. It evaluates the ease of small molecule synthesis using a numerical value in the range of 1 to 10. The closer to 1, the easier the synthesis, and the closer to 10, the more difficult the synthesis;
[0084] Quantitative Estimate Of Drug-Likeness (QED): An indicator for evaluating the drug-likeness of compounds, which quantifies drug-likeness as a value between 0 and 1.
[0085] Database:
[0086] ChEMBL database: A target and bioactivity database dedicated to storing small molecule drug data related to bioactivity. Its official website is https: / / www.ebi.ac.uk / chembl / ;
[0087] PDBbind database: A database that specifically collects and organizes the three-dimensional structure information of various complexes in the Protein Data Bank (PDB) and their corresponding affinity experimental data. The website is http: / / www.pdbbind-cn.org / ;
[0088] RCSB PDB database: A global scientific database mainly used to store and share the three-dimensional structure information of biological macromolecules (mainly proteins, nucleic acids, and some polysaccharides). The website is https: / / www.rcsb.org / .
[0089] Refer to Figures 1 to 4 , this embodiment provides a method for generating drug molecules based on protein pocket hypergraph constraints. The specific steps of this method are as follows:
[0090] S1. Construct a small molecule drug dataset and a protein pocket-ligand dataset, where the protein pocket-ligand dataset includes protein pocket data and corresponding ligand data information.
[0091] Specifically, step S1 includes:
[0092] S101. Data acquisition.
[0093] (1) Obtain small molecule drug data: Download small molecule drug data from the ChEMBL database. The small molecule drug data is represented in the form of SMILES sequences.
[0094] (2) Obtain protein pocket source-ligand source data: Download the general dataset of protein-ligand complexes from the PDBbind database, where the protein pocket source data is in PDB format and the ligand source data is in SDF and Mol2 formats.
[0095] S102. Data preprocessing. It includes:
[0096] (1) Preprocess the small molecule drug data: Screen the small molecule drugs in the form of SMILES sequences, obtain molecules with a length range of 20 to 100 characters, and clean the selected small molecule drugs, including removing salts, stereochemical information, and molecules that cannot be converted, etc. Then further convert the cleaned SMILES sequences into SELFIES sequences;
[0097] The SELFIES sequence is segmented character - by - character using the tokenization dictionary obtained from the ChEMBL database to convert it into a token sequence composed of multiple tokens. For example, a SELFIES sequence X can be represented as X = x 1 x 2 ···x n , and each item x in the sequence i is a token. Since neural network models can only process numerical values and cannot understand characters, a vocabulary for small - molecule drugs is generated based on all tokens, and the token sequence is mapped to the corresponding integer indices to form the final small - molecule drug dataset for subsequent pre - training of the model.
[0098] Since the purpose of pre - training is to enable the model to learn the rules corresponding to drug molecules and actual chemical structures, all the pre - processed small - molecule drug datasets are used as the training set.
[0099] (2) Pre - process the general dataset of protein - ligand complexes: Use the RDKit toolkit to convert the ligand source data in Mol2 format into SMILES sequences, and clean them, including removing salts, stereochemical information, and molecules that cannot be converted in the SMILES sequences. Convert the obtained SMILES sequences into SELFIES sequences, and further remove molecules that cannot be converted;
[0100] According to the above method, the string of the SELFIES sequence is segmented by word - splitting to convert it into a token sequence. A ligand vocabulary is generated based on all tokens, and the token sequence is mapped to the corresponding integer indices; at the same time, information such as atoms, coordinates, and relative accessible surface area of the protein pocket is obtained from the PDB - format file of the protein pocket.
[0101] After the above pre - processing, the protein pocket information and the corresponding ligand sequence information are integrated into a protein pocket - ligand dataset. To test the targeting of the molecules generated by the optimized model, it is divided into a training set and a test set in a ratio of 9:1 for further optimization of the pre - trained model.
[0102] It should be noted that in this embodiment, initially, 2,474,590 small - molecule drug data were extracted from the ChEMBL database, and finally 1,127,080 SMILES data available for modeling were obtained, all of which were used for training. At the same time, the initial number of protein pocket - ligand complexes was 14,127, and after data pre - processing, 11,840 were obtained, including 10,656 in the training set and 1,184 in the test set.
[0103] S2. Construct a protein pocket hypergraph and a protein pocket graph using the protein pocket data in the protein pocket-ligand dataset in step S1. Among them, the protein pocket hypergraph includes an amino acid hypergraph representing the protein pocket with amino acids as nodes and an atomic hypergraph with atoms as nodes, and the protein pocket graph is an atomic graph representing the protein pocket with atoms as nodes.
[0104] Specifically, step S2 includes:
[0105] S201. Represent the protein pocket as a hypergraph with amino acids as nodes and construct the amino acid hypergraph of the protein pocket.
[0106] Obtain the amino acid sequence of the protein pocket. Using the central carbon atom of a single amino acid as the current node, calculate the Euclidean distance between the central carbon atoms of the remaining amino acids and the current node. All central carbon atoms with a Euclidean distance less than the threshold and the current node are constructed into a hyperedge. Traverse all amino acids in the protein pocket in turn to obtain the hyperedge corresponding to each amino acid, thus completing the construction of the hypergraph with amino acids as nodes.
[0107] S202. Represent the protein pocket as a hypergraph with atoms as nodes and construct the atomic hypergraph of the protein pocket.
[0108] Based on the amino acid sequence of the protein pocket, obtain all the heavy atoms of the protein pocket through the Biopython tool. Use each heavy atom of the protein pocket as a node. For each amino acid in the protein pocket, save all its heavy atoms in a set, and each set is constructed into a hyperedge.
[0109] S203. Represent the protein pocket as a graph with atoms as nodes and construct the atomic graph of the protein pocket. It represents the heavy atoms in the protein pocket as nodes in the graph, and the edges represent the interactions or spatial relationships between atoms. The characteristics of the edges include the distance between atoms, the type and strength of the bond, etc.
[0110] Specifically, based on the amino acid sequence of the protein pocket, using each heavy atom in the protein pocket as a node, traverse each heavy atom in turn, calculate the Euclidean distance between the current heavy atom and the remaining heavy atoms. If the Euclidean distance is less than the threshold then define that there is an edge between the two heavy atoms; in addition, if there is a chemical bond between the two heavy atoms, it is also defined as an edge.
[0111] S3. Construct a feature extraction module, map the nodes and edges in the amino acid hypergraph and atomic hypergraph of the protein pocket in step S2 to a low-dimensional vector space, and respectively obtain the hypergraph feature embedding representation F HG at the amino acid level and the hypergraph feature embedding representation F' HG; Meanwhile, generate an atomic-level graph feature embedding representation F for the nodes in the protein pocket atomic graph G , and this representation contains information about the nodes and their neighboring nodes. Specifically, step S3 includes:
[0112] S301, construct a first feature extraction module to extract and represent the protein pocket hypergraph features F at the amino acid level HG .
[0113] In this embodiment, the first feature extraction module uses a hypergraph neural network (HGNN) to obtain the feature embedding representation F of the protein pocket hypergraph at the amino acid level HG , and the hypergraph convolution operator is designed based on the hypergraph Laplacian matrix for extracting the global features of the protein pocket. Preferably, this hypergraph neural network consists of two layers of hypergraph convolution layers, and the output dimensions are 256 and 128 in sequence. For the input hypergraph at the amino acid level, the node feature update of each layer is shown in formula (1):
[0114]
[0115] where X (l) represents the node feature matrix of the protein pocket at the l-th layer, W is the hyperedge weight, H is the incidence matrix of the protein pocket, Θ is used to extract the useful node features in the protein pocket, D v represents the degree matrix of the hyperedges, D e represents the degree matrix of the vertices, and σ represents the activation function. It should be noted that this formula is divided into two parts:
[0116] 1) Aggregate the information of the hyperedges (the global information of the amino acid groups) to the nodes (amino acids), realizing the propagation of features from point to edge to point, and being able to better learn the global information of the protein pocket;
[0117] 2) Aggregate the features of the nodes (amino acids) in the protein pocket to the hyperedges (amino acid sets).
[0118] S302, construct a second feature extraction module to extract and represent the protein pocket hypergraph features F' at the atomic level HG .
[0119] In this embodiment, the second feature extraction module also uses a hypergraph neural network to obtain the feature embedding representation F' of the protein pocket hypergraph at the atomic level HG , and it adopts the same feature extraction module structure as in step S301.
[0120] S303. Construct a third feature extraction module to extract and represent the protein pocket graph features F at the atomic level G .
[0121] In this embodiment, the third feature extraction module uses the multi-head attention neural network (MHA) in the graph attention neural network (Graph Attention Network, GAT) to obtain the feature embedding representation F of the protein pocket graph at the atomic level G . Preferably, the network model consists of two layers of graph attention layers. The first layer contains 8 attention heads. Using multi-head attention can obtain non-single relationships between neighbor nodes; the second layer contains 1 attention head, which aggregates the various node features and neighbor relationships obtained in the previous layer and obtains the final features. The output dimensions are 256 and 128 in sequence. For the input atomic graph of the protein pocket, the node feature update of each layer is shown in formula (2):
[0122]
[0123] Among them, is the initial node feature, is the updated node feature, σ represents the activation function, K represents K attentions, α ij k represents the corresponding attention coefficient, W K is the corresponding weight matrix, α ij k and W K are both learnable parameters.
[0124] S4. Construct a feature fusion module to fuse the hypergraph feature embedding representation F HG and F' HG , as well as the graph feature embedding representation F G to obtain the fused feature embedding representation F fus of the protein pocket.
[0125] Specifically, in this embodiment, the feature fusion module includes a graph pooling module with a Set2Set network layer. It first concatenates the hypergraph feature embedding representation F' HG at the atomic level and the graph feature embedding representation F G , adopts the average pooling strategy to reduce the dimension and extract features of the concatenated features, retains all the information in the two representations, realizes the preliminary fusion of features, and then fuses it with the hypergraph feature embedding representation F HG at the amino acid level, and obtains the fused feature embedding representation F fus of the protein pocket through concatenation.
[0126] S5. Construct a conditional molecular generation model based on the Gate Recurrent Unit (GRU), and use the small molecule drug dataset obtained in step S1 to pre-train the model so that it learns the rules corresponding to drug molecules and actual chemical structures.
[0127] Combine Figure 2 and Figure 4 , the conditional molecular generation model in this embodiment is divided into two parts: one part is used to obtain the feature embedding representation of the protein pocket to constrain the generation of molecules; the other part is the molecular generator. Specifically, the conditional molecular generation model consists of an Embedding layer, three Gate Recurrent Unit (GRU) layers, and an activation (Softmax) layer. Among them, each GRU layer has a hidden state vector of size 256.
[0128] Furthermore, use this model to train the small molecule drugs obtained from the ChEMBL database to learn the chemical structures and generation rules of small molecule drugs. The pre-training process of the molecular generation model is as follows: Take the token sequence mapped to the corresponding integer index in the small molecule drug dataset in step S1 as the input of the embedding layer of the model, and obtain the vector representation of the token after being processed by the embedding layer; then, learn through the GRU layer to capture the long-term dependencies in the sequence; finally, predict the next possible output token in the sequence through the activation layer. The probability of the predicted token is shown in formula (3):
[0129]
[0130] where P(y i ) represents the probability of predicting the i-th possible token, and m represents the total number of tokens in the small molecule drug vocabulary.
[0131] Furthermore, take the cross-entropy between the predicted value and the true value as the loss function of the model. The loss function is shown in formula (5):
[0132]
[0133] where Loss is the value of the loss function, n is the number of tokens after segmenting a sequence, y i represents the true value of the i-th token, and y′ i represents the predicted value of the i-th token.
[0134] S6. The fused feature embedding representation F of the protein pocket obtained in step S4 fusAs a constraint condition for the molecular generation of the pre-trained model obtained in step S5, and in combination with using the ligand data in the protein pocket-ligand dataset in step S1 to retrain the pre-trained model, so as to optimize the conditional molecular generation model.
[0135] Specifically, the process of retraining is as follows: Use the embedded representation F of the protein pocket fus As the initial input of the hidden state of the gated recurrent unit, use the token sequence of the corresponding ligand that has been mapped to the corresponding integer index in step S1 as the input of the model embedding layer. The settings and processes of the remaining parameters are the same as those in step S5. Use the fused feature embedded representation F of the protein pocket fus As the condition for molecular generation, construct a conditional probability model for the probability of the molecular output token, and learn the molecular probability distribution conditional on the protein pocket. The conditional molecular generation probability is shown in formula (4):
[0136]
[0137] Among them, P(molecule|pocket) represents the molecular generation probability conditional on the protein pocket, n is the number of tokens after segmenting the corresponding ligand sequence, and x t Is the token of the molecular sequence at the t-th iteration.
[0138] It should be noted that the loss function in the model optimization process is the same as that in step S5, and will not be elaborated here.
[0139] S7. Use the conditional molecular generation model optimized in step S6 to generate drug molecules for specific protein targets, and evaluate the generated drug molecules.
[0140] In this embodiment, the main protease 3Clpro (3C-like protease) of the novel coronavirus is selected as the target protein target. It specifically includes the following steps:
[0141] S701. Download the PDB file (PDBID: 7DPU) of the main protease 3Clpro of the novel coronavirus from the RCSB PDB database, and remove water molecules and ligands.
[0142] S702. Use the conditional molecular generation model optimized in step S6, use the pocket feature of the main protease 3Clpro of the novel coronavirus as the initial input of the hidden state of the model gated recurrent unit, and generate a batch of potential drug molecules for the pocket feature of this protein.
[0143] S703. Conduct a comprehensive quantitative evaluation of the generated molecules. Figure 5The synthetic accessibility score (SA) is shown. As can be seen from the figure, the SA values obtained by the method of this embodiment are mainly concentrated between 0.6 and 1.0, and there is a significant peak when the SA value is about 0.9. This indicates that the method has good performance in synthetic accessibility and can generate molecules with higher SA values. Compared with other methods such as Poket2Drug and HGAF, the method of this embodiment has a more obvious distribution in the high-end region of the SA value, which means that when generating compounds, the method can more effectively use the fusion features to improve the synthetic accessibility of molecules, especially in the region where the SA value is close to 1, the method shows obvious advantages.
[0144] Figure 6 It is the effect of the generated molecules on the quantitative estimation of drug likeness (QED). As can be seen from the figure, the QED values obtained by the method of this embodiment are widely distributed, covering the range from close to 0 to close to 1, showing significant diversity. The density peak of the method appears near the QED value of about 0.45, indicating that more molecules are generated in this range and have good drug likeness. Compared with other methods such as Poket2Drug and HGAF, its QED distribution curve is smoother, which means that the distribution of the generated molecules on QED is more uniform and has a wider range of possibilities, making the generated molecules perform better in terms of drug likeness.
[0145] Figure 7 Intuitively presents the docking binding energy of the generated molecules under different features. It is observed that the median line of the fusion features is at a lower position. This feature indicates that the molecules generated by the method of this embodiment generally show lower docking binding energies, reflecting a stronger interaction between the generated molecules and the target. In addition, its average binding energy also reaches the lowest value, which strongly proves that when the molecules generated by the method of this embodiment bind to the target, they can form a more stable interaction relationship.
[0146] S704, using molecular docking technology, docks the generated molecules with the main protease 3Clpro to measure the interaction between the generated molecules and the active site. As Figure 8 shown, more hydrogen bonds are formed between the generated molecules and the active pocket of the main protease, and the bond lengths are mostly between 2.2 and 2.8, both of which are less than These hydrogen bonds provide stronger stability for the binding state between the molecule and the protease. In addition, Figure 9 shows the complementarity and adaptability of the shape of the generated molecules combined with the pocket shape. As can be seen from it, the generated molecules can closely fit the active pocket of the main protease, effectively filling the space of the pocket, thereby significantly enhancing the interaction strength between the molecule and the main protease.
[0147] It should be understood that the above embodiments are only a specific embodiment of the present invention. For those of ordinary skill in the art, various modifications and deformations made on the basis of the above description should be considered within the protection scope of the present invention.
Claims
1. A method for generating drug molecules based on protein pocket hypergraph constraints, characterized in that: The steps include: S1, constructing a drug small molecule data set and a protein pocket-ligand data set, wherein the protein pocket-ligand data set includes protein pocket data and corresponding ligand data information; S2, constructing a protein pocket hypergraph and a protein pocket graph using the protein pocket data described in step S1, wherein the protein pocket hypergraph includes an amino acid hypergraph representing protein pockets as amino acids as nodes and an atom hypergraph representing protein pockets as atoms as nodes, and the protein pocket graph represents protein pockets as an atom graph representing protein pockets as atoms as nodes; S3, construct a feature extraction module to map the nodes and edges in the amino acid hypergraph and atom hypergraph of the protein pocket in step S2 to a low-dimensional vector space, and obtain the amino acid-level hypergraph feature embedding representation F HG and atomic-level hypergraph feature embedding representation F′ HG At the same time, the nodes in the protein pocket atomic graph are generated into atomic-level graph feature embedding representation F G ; S4, construct a feature fusion module to embed the hypergraph features obtained in step S3 into representation F HG and F′ HG , and graph feature embedding representation F G Fusion is performed to obtain the fusion feature embedding representation F of the protein pocket fus ; S5, constructing a conditional molecule generation model based on a gated recurrent unit, and pre-training the model using the drug small molecule dataset described in step S1, so that it learns the rules corresponding to the actual chemical structure of the drug molecule; S6, embed the fusion feature of the protein pocket obtained in step S4 into the representation F fus As the constraint conditions for the molecule generation of the pre-trained model obtained in step S5, and re-training the pre-trained model in combination with the ligand data described in step S1 to optimize the conditional molecule generation model; S7, using the conditional molecule generation model optimized in step S6 to generate drug molecules for specific protein targets.
2. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 1, characterized in that: In step S1, constructing a drug small molecule data set includes: (1) Collect drug small molecule data in the form of SMILES sequences from the ChEMBL database, select drug small molecules with SMILES sequence lengths ranging from 20 to 100 characters, and clean the SMILES sequences; (2) Convert the cleaned SMILES sequence into a SELFIES sequence, and segment the SELFIES sequence into a token sequence using a predefined word segmentation table; (3) Generate a drug small molecule vocabulary based on all tokens, and map the token sequence to the corresponding integer index to form the final drug small molecule dataset; The construction of the protein pocket-ligand data set includes: (1) Collect protein pocket source-ligand source data from the PDBbind database, where the protein pocket source data is in PDB format and the ligand source data is in SDF and Mol2 formats; (2) Convert the ligand source data in Mol2 format into a SMILES sequence and clean it, and then convert the cleaned SMILES sequence into a SELFIES sequence; segment the SELFIES sequence into a token sequence using a predefined word segmentation table; generate a ligand vocabulary based on all tokens, and map the token sequence to a corresponding integer index; (3) Obtain the pocket information of the protein corresponding to the ligand from the protein pocket source data in PDB format; integrate the ligand sequence information and the corresponding protein pocket information to form a protein pocket-ligand dataset.
3. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 1, characterized in that: In step S2, the protein pocket is represented as an amino acid hypergraph with amino acids as nodes, including: Obtain the amino acid sequence of the protein pocket, take the central carbon atom of a single amino acid as the current node, calculate the distance between the central carbon atoms of the remaining amino acids and the current node, and construct all central carbon atoms and the current node with a distance less than a predetermined threshold into a spatial structure hyperedge; traverse all amino acids in the protein pocket in turn to obtain the spatial structure hyperedge corresponding to each amino acid; The method of representing the protein pocket as an atomic hypergraph with atoms as nodes includes: Get the amino acid sequence of the protein pocket, take each heavy atom in the protein pocket as a node, and for each amino acid, save all its heavy atoms in a set, and each set is constructed into a hyperedge; The method of representing the protein pocket as an atomic graph with atoms as nodes comprises: Obtain the amino acid sequence of the protein pocket, take each heavy atom in the protein pocket as a node, traverse each heavy atom in turn, calculate the distance between the current heavy atom and the remaining heavy atoms, and define the distance between two heavy atoms whose distance is less than a predetermined threshold as an edge; if there is a chemical bond between the two heavy atoms, it is also defined as an edge.
4. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 1, characterized in that: In step S3, a hypergraph neural network is used to extract the amino acid level hypergraph feature embedding representation F HG and atomic-level hypergraph feature embedding representation F′ HG , the hypergraph neural network consists of at least two layers of hypergraph convolutional layers, and the node feature update of each layer is shown in formula (1): Among them, X (l) represents the node feature matrix of the l-th layer protein pocket, W is the hyperedge weight, H is the incident matrix of the protein pocket, Θ is used to extract useful node features in the protein pocket, and D v represents the degree matrix of the hyperedge, D e represents the degree matrix of the vertex, and σ represents the activation function.
5. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 1, characterized in that: In step S3, a graph attention neural network is used to generate an atomic-level graph feature embedding representation F for the nodes in the protein pocket atom graph. G , the graph attention neural network consists of at least two graph attention layers, and the node feature update of each layer is shown in formula (2): in, is the initial node feature, is the updated node feature, σ represents the activation function, K represents K attentions, α ij k Represents the corresponding attention coefficient, W K is the corresponding weight matrix, α ij k and W K These are all learnable parameters.
6. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 1, characterized in that: In step S4, the fusion feature embedding representation F of the protein pocket is obtained. fus ,include: First, embed the atomic-level hypergraph features into F′ HG And graph feature embedding representation F G The average pooling strategy is used to reduce the dimension and extract features of the spliced features to achieve the initial fusion of features; then it is embedded with the amino acid level hypergraph feature representation F HG Further fusion is performed to obtain the fusion feature embedding representation F of the protein pocket through splicing fus .
7. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 2, characterized in that: In step S5, the conditional molecule generation model based on the gated recurrent unit includes an embedding layer, at least three gated recurrent unit layers and an activation layer; During pre-training, the token sequence mapped to the corresponding integer index in the drug small molecule dataset is used as the input of the model embedding layer. The conditional molecule generation model predicts the probability of the next possible output token in the sequence as shown in formula (3): Among them, P(y i ) represents the probability of predicting the i-th possible token, and m represents the total number of tokens in the drug small molecule vocabulary.
8. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 7, characterized in that: In step S6, the optimization condition molecule generation model includes: The fusion features of the protein pocket are embedded into the representation F fus As the initial input of the hidden state of the gated recurrent unit, it is used to constrain the generation of molecules. The token sequence of the corresponding ligand mapped to the corresponding integer index is used as the input of the model embedding layer. A conditional probability model is constructed for the probability of the molecular output token to learn the molecular probability distribution conditioned on the protein pocket. The conditional molecule generation probability is shown in formula (4): Among them, P (molecule|pocket) represents the probability of molecule generation with the protein pocket as the constraint condition, n is the number of tokens after the segmentation of the corresponding ligand sequence, and x t It is the token of the molecular sequence of the tth iteration.
9. The method for generating drug molecules based on protein pocket hypergraph constraints according to claim 7 or 8, characterized in that: The conditional molecule generation model uses the cross entropy between the predicted value and the true value as the loss function, and the loss function is shown in formula (5): Among them, Loss is the loss function value, n is the number of tokens after a sequence segmentation, and y i Represents the true value of the i-th token, y′ i Represents the predicted value of the i-th token.
10. A drug molecule generation system based on protein pocket hypergraph constraints, used to implement the drug molecule generation method based on protein pocket hypergraph constraints according to any one of claims 1 to 9, characterized in that: include: A data set construction module is configured to construct a drug small molecule data set and a protein pocket-ligand data set, wherein the protein pocket-ligand data set includes protein pocket data and corresponding ligand data information; A hypergraph and graph construction module, configured to construct a protein pocket hypergraph and a protein pocket graph, wherein the protein pocket hypergraph includes an amino acid hypergraph representing the protein pocket as a node and an atom hypergraph representing the protein pocket as a node, and the protein pocket graph is an atom graph representing the protein pocket as a node; The feature extraction module is configured to map the nodes and edges in the two protein pocket hypergraphs to low-dimensional vector spaces, and obtain amino acid-level hypergraph feature embedding representations F and HG and atomic-level hypergraph feature embedding representation F′ HG At the same time, the nodes in the protein pocket graph are generated into atomic-level graph feature embedding representation F G ; The feature fusion module is configured to embed the hypergraph features into the representation F HG and F′ HG , and graph feature embedding representation F G Fusion is performed to obtain the fusion feature embedding representation F of the protein pocket fus ; A model building and pre-training module is configured to build a conditional molecule generation model based on a gated recurrent unit, and to pre-train the model using a drug small molecule dataset so that it learns the rules corresponding to the actual chemical structure of the drug molecule; The model optimization module is configured to embed the fusion features of the protein pocket into the representation F fus As a constraint condition for the molecule generation of the pre-trained model, the pre-trained model is retrained in combination with the ligand data to optimize the conditional molecule generation model; The drug molecule generation module is configured to generate drug molecules for specific protein targets using an optimized conditional molecule generation model.
Citation Information
Cited By
Small molecule generation method based on target protein under deep neural network
CN121215021A