A New Method for Designing Novel Bioactive Small Molecules with Controllable Properties Based on Protein Structure

By constructing the CProMG model, combining protein structure and small molecule generation algorithm, the problem of difficult to control the biochemical and physical and chemical properties of the generated small molecules in the existing technology is solved, and small molecule generation with high binding force and controllable attributes is achieved, which promotes the progress of drug discovery.

CN116758978BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310707583.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2025-07-11
Estimated Expiration
2043-06-15

AI Technical Summary

Technical Problem

Existing deep learning-based molecular generation methods are difficult to control biochemical and physical and chemical properties on the basis of satisfying high binding forces.

Method used

A new active small molecule design method based on protein structure is designed. By constructing a small molecule generation model CProMG, a protein embedding module, a two-view encoder module, a small molecule embedding module and a decoder module are used to combine the beam search algorithm to generate small molecules with high binding force and controllable attributes.

Benefits of technology

The resulting small molecules have high binding force and controllable biochemical and physical and chemical properties, which significantly improve the quality and efficiency of drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758978B_ABST
    Figure CN116758978B_ABST
Patent Text Reader

Abstract

The present invention discloses a novel method for designing active small molecules with controllable properties based on protein structures, and proposes a small molecule generation model based on Transformer, namely CproMG. Based on the hierarchical view of the fusion protein, it significantly enhances the expression of the protein binding pocket by associating amino acid residues with their constituent atoms. By jointly embedding the molecular sequence, its drug-like properties and the binding affinity with the protein, it automatically regresses to generate new molecules with desired properties in a controllable manner by measuring the proximity of molecular markers to protein residues and atoms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer-aided drug R & D, and particularly relates to a method for designing a new active small molecule with controllable properties based on protein structure. Background Art

[0002] In the process of drug design, screening or designing candidate compounds that bind to protein targets is crucial. However, the chemical space of small molecules is very large, estimated to include 10 23 -10 60 compounds. Therefore, it is extremely difficult to find suitable small molecules in such a space.

[0003] In the development of computer-aided drug design, high-throughput screening and virtual screening technologies were first proposed, and target molecules were obtained by filtering molecules in a large compound library. High-throughput screening is computer-aided and can detect tens of millions of samples in one experiment. Molecular docking technology and quantitative structure-activity relationship method (QSAR) based on machine learning are applied to virtual screening, which are two virtual screening methods based on small molecule structure screening and drug action mechanism screening respectively. With the development of artificial intelligence, molecular biochemical property prediction models based on deep learning are also applied to virtual screening, bringing hope for the discovery of lead compounds. However, the above methods are all based on known databases for screening, which greatly limits the search scope in chemical space, and the molecules obtained by screening lack originality.

[0004] De novo design of drug small molecules is essentially searching for small molecules in chemical space, but it is not restricted by existing databases and can explore the entire chemical space more fully. With the development of artificial intelligence, many deep generative models have emerged and have been successfully applied to the fields of natural language processing and images. Inspired by this, currently, generative models are applied to small molecule generation, learning the physical and chemical properties and structural features of small molecule data, and finally generating ideal small molecules that meet specific conditions.

[0005] Currently, deep learning-based molecular generation methods can be roughly divided into ligand structure-based generation methods and receptor structure-based generation methods. Ligand structure-based generation methods do not consider target information or are restricted by target-specific ligand datasets, and it is difficult to meet the requirement of high binding affinity with new targets. Although receptor structure-based generation methods can solve the above problems, the biochemical and physical and chemical properties of the generated molecules are difficult to control.

[0006] In view of this, it is necessary to design a new generation method to enable the generated molecules to have controllable properties on the basis of high binding affinity. Summary of the Invention

[0007] The object of the present invention is to solve the technical problem that the molecular generation method based on deep learning cannot control its biochemical and physical properties while satisfying high binding affinity when designing molecules, and thus provides a novel method for designing active small molecules with controllable properties based on protein structure.

[0008] To achieve the above object, the technical solution provided by the present invention is as follows:

[0009] A novel method for designing active small molecules with controllable properties based on protein structure, characterized in that it includes the following steps:

[0010] 1) Construct a small molecule generation model CProMG:

[0011] The small molecule generation model CProMG includes a protein embedding module, a dual-view encoder module, a small molecule embedding module, and a decoder module, and uses the beam search algorithm to gradually generate a complete SMILES sequence;

[0012] The protein embedding module is used to obtain the amino acid graph feature and atomic graph feature of the protein (i.e., its input is the protein 3D structure, and the output is the amino acid graph feature and atomic graph feature), and includes an amino acid graph embedding unit and an atomic graph embedding unit;

[0013] The dual-view encoder module is used to fuse the amino acid graph feature and atomic graph feature of the protein to obtain the fused protein feature (i.e., the input is the feature representation of the amino acid graph and atomic graph, and the output is the fused protein feature), and includes a multi-head attention network, a feed-forward neural network, and an information cross-fusion unit;

[0014] The small molecule embedding module is used to obtain the initial feature of the small molecule (i.e., the input is the small molecule sequence, and the output is the small molecule embedding feature), and includes an embedding unit for small molecule SMILES and attributes, a segment encoding unit, and a position encoding unit; wherein, the segment encoding unit is used to distinguish the molecular sequence and molecular attributes, and the position encoding unit is used to obtain the position information;

[0015] The decoder module is used to generate the small molecule sequence (i.e., the input is the small molecule embedding feature and the protein feature, and the output is the generated small molecule sequence), and includes a masked multi-head attention network, an interactive multi-head attention network, and a feed-forward neural network;

[0016] 2) Obtain sample data, train the small molecule generation model CProMG constructed in step 1) to obtain a trained small molecule generation model; the specific training process is as follows:

[0017] 2.1) Collect sample data and construct a training data set and a test data set

[0018] The sample data is a protein-small molecule pair with a combined root mean square deviation of pose less than (screened from existing datasets, generating small molecules based on protein structures so that the generated small molecules have target affinity), which includes the three-dimensional structure information of the protein and the SMILES sequence information of the small molecule;

[0019] 2.2) Obtain protein features and initial small molecule features

[0020] The protein features are obtained through the following methods:

[0021] A1. Characterize the three-dimensional structure of the protein in step 2.1), and use the K-nearest neighbor algorithm (KNN) to construct a protein amino acid graph and a protein atom graph Encode the node information through one-hot encoding, and then add Laplacian positional encoding to obtain the initial features of the nodes. Use the Gaussian kernel function to convert the edge lengths into edge features;

[0022] A2. Use a dual-view encoder module to fuse and train the initial features of the protein amino acid graph and the protein atom graph obtained in A1 to obtain protein features;

[0023] The dual-view encoder module includes parallel amino acid view encoders En r and atom view encoders En a , each encoder includes t encoding layers. Each encoding layer first uses edge features to enhance the information of each node, then uses the multi-head attention mechanism to calculate the attention scores of each node and its adjacent nodes, and uses them as weights to aggregate the adjacent nodes and update the node information. Finally, it is passed into a feed-forward neural network; the information cross-fusion unit fuses the information of the two views (that is, aggregates the information of the atom view into the amino acid view through attention calculation to update the node features of the amino acid view); finally, the outputs of En r and En a are concatenated to obtain the final protein feature representation;

[0024] The initial small molecule features are obtained through the following methods:

[0025] Characterize the SMILES sequence information of the small molecule in step 2.1), use RDKit to obtain the physicochemical properties of the small molecule, and concatenate them in front of the small molecule SMILES sequence as generation conditions. Encode the entire sequence through one-hot encoding to obtain the initial small molecule features;

[0026] 2.3) Use the decoder module to train the initial features of the small molecule obtained in step 2.2). The decoder is similar to the decoder of the original Transformer and contains t decoding layers. Each decoding layer first passes through a masked multi-head attention network to learn the features of the molecule itself, and then uses an interactive multi-head attention network to calculate the proximity between the molecule tokens and the protein features obtained in step 2.2) to update the molecule features. Finally, it is passed into a feed-forward neural network to predict the complete molecular output;

[0027] 2.4) Use the molecule predicted in step 2.3), calculate the model loss using the cross-entropy loss function, and adjust the model parameters through negative feedback according to the loss. After training is completed, obtain the small molecule generation model CProMG;

[0028] 3) Use the CProMG model trained in step 2) and combine it with the beam search algorithm to gradually generate the complete small molecule SMILES sequence. The beam search algorithm is a strategy for searching the space used when generating molecules after the model is trained. When the model runs once, it can only predict the next character based on the known sequence. Therefore, a complete SMILES sequence needs to be generated step by step through multiple loop runs.

[0029] Furthermore, in step 2.2), the three-dimensional structure of the protein is represented as an amino acid graph atomic graph where, is the node set, v i represents the feature of node i, represents the three-dimensional coordinates of the node, ε = {e ij , i, j = 1, 2,..., n & i ≠ j} represents the edge feature;

[0030] For the amino acid graph, the node feature v i is the one-hot encoding of the residue type of the i-th residue; based on the three-dimensional coordinates of the amino acid, use the K-nearest neighbor algorithm to construct the protein amino acid graph; use multiple Gaussian kernel functions to represent the edge length as an n-dimensional vector as the edge feature ε;

[0031] For the atomic graph, the node feature v i is the one-hot encoding including information such as atomic type, the amino acid it belongs to, and whether it is a backbone; based on the three-dimensional coordinates of the atom, use the K-nearest neighbor algorithm to construct the protein atomic graph; use multiple Gaussian kernel functions to represent the edge length as an n-dimensional vector as the edge feature ε.

[0032] Furthermore, Laplacian positional encoding is a generalization of the positional encoding used in the original Transformer in the graph, which can better help encode distance perception information, that is, nearby nodes have similar positional features, and distant nodes have different positional features. Therefore, in step 2.2), the Laplacian eigenvector is used as the positional encoding in CProMG, where the eigenvector is defined by the factorization of the Laplacian matrix of the graph, and the formula is as follows:

[0033]

[0034] where is the identity matrix, the n×n diagonal matrix D is the degree matrix of the graph and A represents the adjacency matrix of contains a set of eigenvectors which correspond to a set of eigenvalues {λ k}; Adding the positional encoding to the embedding features of the protein graph nodes gives the initial features of the protein graph nodes with global spatial features (that is, adding positional confidence to the graph nodes to facilitate subsequent model modules to consider the positional information of the graph nodes):

[0035]

[0036] where and are weight matrices.

[0037] Furthermore, the architecture of the dual-view encoder module in step 2.2) is specifically as follows:

[0038] The embeddings of the amino acid graph and the atom graph (that is, the initial features of the protein amino acid graph and the protein atom graph obtained by A1) are respectively input into the En r and En a in the dual-view encoder to obtain the final representation of the protein binding pocket;

[0039] Each encoder consists of t cascaded encoding units, and each encoding unit contains an edge enhancement encoding block and a multi-head attention block The first block enhances the node representation, and the attention block further updates the node representation through the self-attention mechanism;

[0040] The edge-enhanced q, k, v are defined as follows:

[0041]

[0042]

[0043] Among them, {W} is a learnable weight matrix, ⊙ represents element-wise multiplication, and e ij represents the edge feature between node i and node j;

[0044] Update the node features through the multi-head attention block:

[0045]

[0046] represents the number of nodes in the graph, and d k is a hyperparameter representing the feature dimension; the node representation and are designed with a residual connection, that is, η(·) represents the regularization function; then the node features are input into the FNN, and there is also a residual connection, that is,

[0047] Define the outputs of the encoders En r and En a as H (t) and Z (t) respectively, and concatenate them to obtain the final representation feature H P of the protein structure = [H (t) ; Z (t) .

[0048] Furthermore, the information cross-fusion unit in step 2.2) is specifically:

[0049] The information cross-fusion unit is implemented through multi-head attention. The node features output by the multi-head attention of the atom view encoder are regarded as Keys and Values, and the node features output by the multi-head attention of the amino acid view encoder are regarded as Queries:

[0050]

[0051] Among them, the three W matrices represent linear layers;

[0052] The node feature of the r-th attention head of the amino acid view at the i-th node is updated through the following formula:

[0053]

[0054] n represents the number of nodes in the atomic graph; d k is a hyperparameter representing the feature dimension; after concatenating the node features of multiple attention heads, pass through a linear layer Then update the nodes using residual connections

[0055] Further, the specific method for obtaining the initial features of small molecules in step 2.2) is as follows:

[0056] Given the SMILES sequence of a small molecule, use the open-source chemistry toolkit RDKit to calculate its physicochemical properties, including the water-octanol partition coefficient (LogP), topological polar surface area (TPSA), drug-likeness (QED), and synthetic accessibility (SA); splice the four property values with the docking score of the protein-small molecule pair (obtained through the docking software Autodock vina) to form a generation condition vector y; obtain the molecular representation as:

[0057] h m =[yW p ; SW s

[0058] where ';' is the stacking operation of matrices; S represents the one-hot encoding of the SMILES sequence, and represent two linear layers respectively;

[0059] The position encoding of the sequence is defined as follows:

[0060]

[0061] where j = 1, 2,..., N. If d is even, N = d / 2; if d is odd, N = (d + 1) / 2;

[0062] The position representation of the molecule

[0063] The molecular segment encoding h token =[t1; t0;...; t0], and the final embedding of the small molecule is defined as:

[0064] h 0 =h m +h token +H pos

[0065] Further, the interactive multi-head attention network in step 2.3) is specifically as follows:

[0066] The interactive attention of the decoder uses the attention mechanism to learn the key dependencies between the small molecule substructure and the protein, and the protein feature H P ​Using small molecule features as queries for attention calculation with values and keys, the updated feature of the i-th token of the small molecule is Denoting the r-th attention head of the l-th layer decoder, the formula is as follows:

[0067]

[0068] Where, Denotes the number of nodes of the protein feature H P , d k Is a hyperparameter representing The feature dimension of.

[0069] Furthermore, in step 2.4), the cross-entropy loss function is used to calculate the loss, specifically as follows:

[0070]

[0071] Where, x0 = [p, b], p and b respectively represent the attribute condition and the start symbol, x i Denotes the token in the SMILES sequence, and P represents the probability of generating x i .

[0072] Furthermore, step 3) generating the molecular SMILES sequence based on the beam search algorithm is specifically as follows:

[0073] The beam search contains a hyperparameter, the beam width k, which represents the width of the search; at time step 1, given the desired molecular properties and the start symbol '$' as the first two tokens of k candidate output sequences; at each subsequent time step, based on the k candidate output sequences of the previous time step, continue to pick out k candidate output sequences with the highest conditional probability from several possible choices; repeat the above steps until the end symbol '&' is searched, and the search ends. When k = 1, the beam search degenerates into a greedy search.

[0074] Meanwhile, the present invention also provides an electronic device and a computer-readable storage medium, on which a computer program is stored, and the special feature is that: when the computer program is executed by a processor, it implements the steps of the above method.

[0075] The principle of the present invention:

[0076] Based on the graph Transformer, the present invention designs a molecular generation model CProMG, which can be used to generate molecules with high binding affinity to proteins and controllable properties. This is mainly because the present invention generates small molecules based on protein structures, so it can be regarded as a "translation" process from protein structures to small molecule sequences, and the generated small molecules have high target affinity. During the decoder decoding process, the desired properties are generated as conditions, ensuring that the generated molecules have specific properties, that is, property controllability. The generation model of the present invention includes a protein embedding module, a dual-view encoder module, a small molecule embedding module, and a decoder module. First, CProMG enhances the features of the protein binding pocket by fusing protein hierarchical structure information. Secondly, the protein interactive multi-head attention module in the decoder calculates the attention scores between small molecules and protein residues and atoms, so as to capture the key interactions between the protein pocket and small molecules. Finally, by jointly embedding small molecules and their drug-like properties, new molecules with the desired properties are automatically regressed in a controllable manner.

[0077] The advantages of the present invention are as follows:

[0078] By learning the distribution law of small molecules in the known chemical space and exploring small molecules in the unknown space, the present invention proposes a new method for designing active small molecules with controllable properties based on protein structures, namely CProMG. This model learns the amino acid view and atomic view features of proteins through the multi-head attention mechanism in step 2.2) and effectively fuses them, significantly enhancing the features of the protein pocket. Secondly, the decoder of this model can effectively control the properties of the generated molecules by embedding the drug-like properties of small molecules. In addition, the interactive attention module of the decoder in step 2.3) calculates the attention scores between small molecule features and protein features, learns the key interactions between proteins and small molecules, and enables the generated molecules to have high binding affinity to proteins. The evaluation of CProMG on the dataset shows that the small molecules generated by CProMG have good binding affinity and drug-like properties. The present invention can provide a small molecule generation tool to promote drug discovery and development. The present invention not only improves the generation quality but also provides a certain degree of interpretability. Description of the Drawings

[0079] Figure 1 It is the overall architecture of the method CProMG proposed by the present invention. Detailed Embodiments

[0080] The following further describes in detail the content of the present invention in conjunction with the drawings and specific embodiments:

[0081] An embodiment is implemented according to the new method for designing active small molecules with controllable properties based on protein structures proposed by the present invention, wherein:

[0082] This example uses a protein-ligand pair dataset: This dataset contains approximately 180,000 protein-ligand docking pairs. Each pair of data has a docking score calculated by Autodock Vina. 100,000 pairs of data are selected to train the model, 1,000 pairs of data are randomly selected from them as the validation set, and the remaining are used as the training set; 100 pairs of data are selected to test the model. The sequence similarity of the data for training the model and testing the model is less than 30%.

[0083] Regarding the three-dimensional structure information of proteins, the K-nearest neighbor algorithm (KNN) is used to construct a protein amino acid graph and a protein atom graph, and the initial features of the graph nodes and edges are obtained through one-hot encoding. Then, Laplacian positional encoding is used to obtain the positional information between the graph nodes;

[0084] For the initial features of the obtained protein atom graph and amino acid graph, a sampled dual-view encoder is used for processing. Edge features are used to enhance the information of each node, and then the multi-head attention mechanism is used to calculate the attention scores of each node with its adjacent nodes, aggregate the information on the nodes, and finally input it into a feed-forward neural network. The information fusion module fuses the information of these two views to obtain the features of each protein.

[0085] Regarding the small molecule SMILES sequence information, with the help of the RDKit tool, the physicochemical properties of small molecules are obtained, and together with the SMILES sequence, the initial features of small molecules are obtained through one-hot encoding.

[0086] For the obtained initial features of small molecules, during the training process, the decoder predicts the complete small molecule SMILES sequence through a multi-head attention module, an interactive attention module, and a feed-forward neural network.

[0087] For the predicted generated small molecule SMILES sequence, the cross-entropy loss function is used to calculate the model loss.

[0088] After training, a small molecule generation model is obtained, and beam search algorithm is combined for generation. Each time, the next character is generated based on the current sequence, and the complete small molecule SMILES sequence is generated through multiple loops.

[0089] To evaluate the model generation quality, the present invention selects docking score (VS), drug-likeness (QED), synthetic accessibility (SA), and diversity (Diversity) as the basic evaluation indicators. The smaller VS and SA are, the better; the larger QED and Diversity are, the better. The calculation of VS is achieved through Autodock Vina, and the calculation of the other three indicators is achieved through the RDKit tool.

[0090] The trained generative model is tested using the test set data. Meanwhile, in the unified dataset, the present invention is compared with other baseline basic methods, and the test results are shown in Table 1 as follows:

[0091] Table 1 Performance demonstration of CProMG

[0092]

[0093] [1] Luo, S., et al. (2021) A 3D Generative Model for Structure-Based Drug Design. Advances in Neural Information Processing Systems, 34.

[0094] [2] Skalic, M., et al. (2019) From Target to Drug: Generative Modeling for the Multimodal Structure-Based Ligand Design. Mol. Pharm., 16, 4282–4291.

[0095] As can be seen from Table 1, the docking scores (VS), drug-likeness (QED), synthetic accessibility (SA), and diversity of the molecules generated by the method of the present invention are significantly higher than those of other baseline basic methods, showing remarkable effects.

[0096] In summary, the present invention can be used to generate small molecules with high binding strength for specific properties. The well-known implementation methods and common knowledge of the above-described solutions are not described in detail here. It should be noted that for those skilled in the art, several improvements can be made without departing from the present invention, and these should also be regarded as the protection scope of the present invention, which will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of the claims, and the specific implementation manners described in the specification are used to explain the content of the claims.

Claims

1. A novel method for designing active small molecules with controllable properties based on protein structure, characterized in that, Including the following steps: 1) Construct a small molecule generation model CProMG: The small molecule generation model CProMG includes a protein embedding module, a dual-view encoder module, a small molecule embedding module, and a decoder module, and uses a beam search algorithm to gradually generate a complete SMILES sequence; The protein embedding module is used to obtain the amino acid graph features and atomic graph features of the protein, including an amino acid graph embedding unit and an atomic graph embedding unit; The dual-view encoder module is used to fuse the amino acid graph features and atomic graph features of the protein to obtain the fused protein features, including a multi-head attention network, a feed-forward neural network, and an information cross-fusion unit; The small molecule embedding module is used to obtain the initial features of the small molecule, including an embedding unit for the small molecule SMILES and properties, a segment encoding unit, and a position encoding unit; The decoder module is used to generate a small molecule sequence, including a masked multi-head attention network, an interactive multi-head attention network, and a feed-forward neural network; 2) Obtain sample data, train the small molecule generation model CProMG constructed in step 1) to obtain a trained small molecule generation model; the specific training process is as follows: 2.1) Collect sample data and construct a training data set and a test data set The sample data is a protein-small molecule pair with a root mean square deviation of the combined pose less than , which includes the three-dimensional structure information of the protein and the SMILES sequence information of the small molecule; 2.2) Obtain protein features and initial small molecule features The protein features are obtained in the following way: A1. Characterize the three-dimensional structure of the protein in step 2.1), and construct a protein amino acid graph using the K-nearest neighbor algorithm and a protein atom graph Encode the node information through one-hot encoding, and then add Laplacian positional encoding to obtain the initial features of the nodes. Use the Gaussian kernel function to convert the edge lengths into edge features; A2. Use the dual-view encoder module to perform fusion training on the initial features of the protein amino acid graph and the protein atomic graph obtained in A1 to obtain protein features; The dual-view encoder module contains parallel amino acid view encoder En r and atomic view encoder En a , each encoder contains t encoding layers respectively. Each encoding layer first uses edge features to enhance the information of each node, then uses the multi-head attention mechanism to calculate the attention scores of each node and its adjacent nodes, and uses them as weights to aggregate the adjacent nodes and update the node information. Finally, it is passed into the feed-forward neural network; the information cross-fusion unit fuses the information of the two views; finally, the outputs of En r and En a are concatenated to obtain the final protein feature representation; The initial small molecule features are obtained in the following way: Characterize the small molecule SMILES sequence information in step 2.1), use RDKit to obtain the physicochemical properties of the small molecule, and splice them in front of the small molecule SMILES sequence as generation conditions, and obtain the initial small molecule features by one-hot encoding the entire sequence; 2.3) Use the decoder module to train the initial small molecule features obtained in step 2.2). The decoder is similar to the decoder of the original Transformer and contains t decoding layers. Each decoding layer first passes through a masked multi-head attention network to learn the molecular self-features, and then uses the interactive multi-head attention network to calculate the proximity of the molecular token to the protein features obtained in step 2.2) to update the molecular features, and finally passes them into the feed-forward neural network to predict the complete molecular output; 2.4) Use the molecule predicted in step 2.3), calculate the model loss using the cross-entropy loss function, and adjust the model parameters through negative feedback according to the loss. After training is completed, the small molecule generation model CProMG is obtained; 3) Use the CProMG model trained in step 2) combined with the beam search algorithm to gradually generate a complete small molecule SMILES sequence.

2. The method for designing a novel active small molecule with controllable properties based on protein structure according to claim 1, characterized in that: In step 2.2), the three-dimensional structure of the protein is represented as an amino acid graph Atom graph Among them, is a set of nodes, and v i represents the feature of node i, represents the three-dimensional coordinates of the node, and ε = {e ij , i, j = 1, 2,..., n & i ≠ j} represents the edge feature; For the amino acid graph, the node feature v i is the one-hot encoding of the residue type of the i-th residue; based on the three-dimensional coordinates of the amino acids, the protein amino acid graph is constructed using the K-nearest neighbor algorithm; the side length is represented as an n-dimensional vector using multiple Gaussian kernel functions as the edge feature ε; For the atomic graph, the node feature v i is a one-hot encoding including information such as atomic type, the amino acid it belongs to, and whether it is a backbone; based on the three-dimensional coordinates of the atoms, the protein atomic graph is constructed using the K-nearest neighbor algorithm; the side length is represented as an n-dimensional vector using multiple Gaussian kernel functions as the edge feature ε.

3. The novel active small molecule design method based on controllable properties of protein structures according to claim 1, characterized in that, In step 2.2), the Laplace eigenvector is used as the position encoding in CProMG, where the eigenvector is defined by the factorization of the Laplace matrix of the graph, and the formula is as follows: Among them, is the identity matrix, the n×n diagonal matrix D is the degree matrix of graph , and A represents the adjacency matrix; contains a set of eigenvectors corresponding to a set of eigenvalues {λ k}; adding the positional encoding to the embedded features of the protein graph nodes, the initial features of the protein graph nodes with global spatial features are obtained: Among them, and are weight matrices.

4. The method for designing a novel active small molecule with controllable properties based on protein structure according to claim 1, characterized in that, The architecture of the dual-view encoder module in step 2.2) is specifically as follows: Amino acid map and atomic map embeddings are respectively input into En r and En a to obtain the final representation of the protein binding pocket; Each encoder is composed of t serially-connected encoding units, and each encoding unit contains an edge enhancement encoding block and a multi-head attention block The first block Enhanced node representation, attention block Further update the node representation through the self-attention mechanism; The edge-enhanced q, k, and v are defined as follows: Among them, {W} is a learnable weight matrix, ⊙ represents element-wise multiplication, and e ij represents the edge feature between node i and node j; Update the node features through the multi-head attention block: Represents the number of nodes in the graph, d k is a hyperparameter representing the dimensionality of the features; the node representation and is designed with a residual connection between them, that is η(·) represents the regularization function; then the node features are input into the FNN, and there is also a residual connection, that is Define the encoder En r and the outputs of En a are H (t) and Z (t) respectively. Concatenate them to obtain the final representation feature H P of the protein structure = [H (t) ; Z (t) .

5. The novel active small molecule design method based on controllable properties of protein structure according to claim 1, characterized in that, The information cross-fusion unit in step 2.2) is specifically as follows: The information cross - fusion unit is implemented through multi - head attention, taking the node features output by the multi - head attention of the atomic view encoder as Keys and Values, and taking the node features output by the multi - head attention of the amino - acid view encoder as Queries: Among them, the three W matrices represent linear layers; The i-th node feature of the r-th attention head in the amino acid view is updated through the following formula: n represents the number of nodes in the atomic graph; d k is a hyperparameter, representing the feature dimension of; after concatenating the node features of multiple attention heads, pass through a linear layer and then use residual connection to update the nodes 6. The method for designing a novel active small molecule with controllable properties based on protein structure according to claim 1, characterized in that, Obtaining the initial features of small molecules in step 2.2) is specifically as follows: Given a small molecule SMILES sequence, use the open-source chemistry toolkit RDKit to calculate its physicochemical properties, including water-octanol partition coefficient (LogP), topological polar surface area (TPSA), drug-likeness (QED), and synthetic accessibility (SA); splice the four property values with the docking score of the protein-small molecule pair to form a generated conditional vector y; obtain the molecular representation as: Among them, ';' is the stacking operation of the matrix; S represents the one-hot encoding of the SMILES sequence, and respectively represent two linear layers; The position encoding of the sequence is defined as follows: Among them, j = 1, 2,..., N. If d is even, N = d / 2. If d is odd, N = (d + 1) / 2; Position representation of molecules Molecular segment encoding h token = [t1; t0;...; t0], and the embedding of the final small molecule is defined as: h 0 = h m + h token + H pos 。 7. The novel active small molecule design method based on the controllable properties of protein structure according to claim 1, characterized in that: The interactive multi-head attention network in step 2.3) is specifically as follows: The interactive attention of the decoder uses the attention mechanism to learn the key dependencies between small molecule substructures and proteins, taking the protein feature H P as values and keys, taking the small molecule feature as queries for attention calculation, and the updated feature of the i-th token of the small molecule is denoting the r-th attention head of the l-th layer of the decoder, and the formula is as follows: Among them, represents the number of nodes of protein feature H P , d k is a hyperparameter representing the feature dimension.

8. The novel active small molecule design method based on controllable properties of protein structures according to claim 1, characterized in that: In step 2.4), calculate the loss using the cross-entropy loss function, specifically as follows: Among them, x0 = [p, b], where p and b represent the attribute condition and the start symbol respectively, and x i represents a token in the SMILES sequence, and P represents the probability of generating x i of.

9. The method for designing a novel active small molecule with controllable properties based on protein structure according to claim 1, characterized in that, Generating the molecular SMILES sequence based on the beam search algorithm in step 3) is specifically as follows: The beam search contains a hyperparameter, beam width k, which represents the width of the search; at time step 1, given the desired molecular properties and the start symbol '$' as the first two tokens of k candidate output sequences; at each subsequent time step, based on the k candidate output sequences of the previous time step, continue to select the k candidate output sequences with the highest conditional probability from several possible choices; repeat the above steps until the end symbol '&' is searched, and the search ends.

10. An electronic device and a computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Compound-protein interaction prediction method fusing multi-view information

    CN116230113A

  • Systems and methods for predicting potential inhibitors of target protein

    US20220084627A1