Method and apparatus for providing molecular structures of drugs
By generating and screening candidate molecular structures through diffusion models and determining pharmacophores based on binding energy and interactions, the problem of low efficiency in drug design in traditional methods is solved, and efficient generation of candidate molecular structures that interact with target protein structures is achieved.
Patent Information
- Application Number
- CN202510926092.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional virtual screening methods based on drug molecule databases are unable to design new molecular structures outside the drug molecule database. Moreover, as the database size increases, computing power and time requirements increase, resulting in low drug design efficiency.
The diffusion model is used to generate candidate molecular structures, and the complex conformation is determined by binding energy and interaction. The pharmacophore is extracted and saved as conditional database data. The candidate molecular structures are screened and updated, and the diffusion model is trained using the pharmacophore conditions.
It improves the efficiency and quality of molecular structure design, can generate candidate molecular structures that interact stably with target protein structures, reduces dependence on large-scale data, and improves the training efficiency of diffusion models.
Smart Images

Figure CN120748544A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of drug design, and more particularly to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for providing the molecular structure of a drug. Background Art
[0002] Currently, structure-based drug design (SBDD) is commonly used to propose new drug molecular structures. The goal of SBDD is to design effective active molecules or small molecule structures based on the known target protein structure, so that they can form stable complexes with the target protein structure to activate or inhibit the corresponding biological function. Traditional SBDD is generally implemented through virtual screening of drug molecule databases. Specifically, by traversing each molecular structure in the drug molecule database, the molecular structure is docked with the target protein structure, and the docking score is obtained as the screening basis. However, this virtual screening method based on drug molecule databases cannot design new molecular structures outside of the drug molecule database. As the size of the drug molecule database increases, it will require a large amount of computing power and time for screening, making drug design very inefficient. Summary of the Invention
[0003] According to some embodiments of the present disclosure, there is provided a method for providing a molecular structure of a drug, comprising:
[0004] Inputting a first conditional feature and a first random noise into a diffusion model to obtain a candidate molecular structure, wherein the first conditional feature is generated according to a pocket position of the target protein structure;
[0005] determining a conformation of a complex formed by the candidate molecular structure and the target protein structure based on the binding energy between the candidate molecular structure and the target protein structure;
[0006] determining a first pharmacophore in the candidate molecular structure based on the interaction between the candidate molecular structure and the target protein structure in the complex conformation;
[0007] determining a conformational score of the complex conformation based on the binding energy and the interaction;
[0008] storing the candidate molecular structure, the first pharmacophore, and the conformational score in association with each other as data in a condition database;
[0009] Filtering a first number of data from the condition database according to the conformation scores of all data in the condition database, and generating a second condition feature according to one or more first pharmacophores in the first number of data; and
[0010] The second conditional feature and the second random noise are input into the diffusion model to obtain an updated candidate molecular structure.
[0011] According to other embodiments of the present disclosure, there is provided a device for providing a molecular structure of a drug, comprising:
[0012] a molecular structure acquisition module configured to input a first conditional feature and a first random noise into a diffusion model to acquire a candidate molecular structure, wherein the first conditional feature is generated according to a pocket position of a target protein structure;
[0013] a complex conformation determination module, configured to determine the conformation of the complex formed by the candidate molecular structure and the target protein structure based on the binding energy between the candidate molecular structure and the target protein structure;
[0014] a pharmacophore determination module configured to determine a first pharmacophore in the candidate molecular structure based on the interaction between the candidate molecular structure and the target protein structure in the complex conformation;
[0015] a conformation score determination module configured to determine a conformation score of the complex conformation based on the binding energy and the interaction;
[0016] a condition database configured to store the candidate molecular structure, the first pharmacophore, and the conformational score in association with each other as data in the condition database;
[0017] a condition generation module configured to screen a first number of data from the condition database based on conformational scores of all data in the condition database, and generate a second condition feature based on one or more first pharmacophores in the first number of data;
[0018] The molecular structure acquisition module is further configured to input the second conditional feature and the second random noise into the diffusion model to obtain an updated candidate molecular structure.
[0019] According to some embodiments of the present disclosure, an electronic device is provided, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to execute a method of any embodiment described in the present disclosure based on instructions stored in the at least one memory.
[0020] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the instructions are executed by a processor, the method of any embodiment described in the present disclosure is implemented.
[0021] According to some embodiments of the present disclosure, a computer program product is provided. When the computer program product is run on a computer, the computer is caused to implement the method of any embodiment described in the present disclosure.
[0022] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The following describes embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:
[0024] Figure 1 A schematic flow chart showing a method for providing the molecular structure of a drug according to some embodiments of the present disclosure;
[0025] Figure 2 A schematic diagram illustrating a process for generating a first conditional feature in a method for providing a molecular structure of a drug according to some embodiments of the present disclosure;
[0026] Figure 3 A schematic diagram illustrating an operation for generating a first conditional feature in a method for providing a molecular structure of a drug according to some embodiments of the present disclosure;
[0027] Figure 4 A schematic flow chart illustrating step S104 in the method for providing the molecular structure of a drug according to some embodiments of the present disclosure;
[0028] Figure 5 A schematic diagram illustrating a process for generating a second conditional feature in a method for providing a molecular structure of a drug according to some embodiments of the present disclosure;
[0029] Figure 6 A schematic diagram illustrating an operation for generating a second conditional feature in a method for providing a molecular structure of a drug according to some embodiments of the present disclosure;
[0030] Figure 7 A schematic flow chart illustrating step S107 in the method for providing the molecular structure of a drug according to some embodiments of the present disclosure;
[0031] Figure 8 A schematic flow chart showing a method for providing the molecular structure of a drug according to other embodiments of the present disclosure;
[0032] Figure 9 A schematic diagram illustrating an iterative process of a method for providing a molecular structure of a drug according to some embodiments of the present disclosure;
[0033] Figure 10A block diagram showing an apparatus for providing the molecular structure of a drug according to some embodiments of the present disclosure;
[0034] Figure 11 A block diagram illustrating an electronic device according to some embodiments of the present disclosure;
[0035] Figure 12 A block diagram of an electronic device showing some other embodiments of the present disclosure.
[0036] It should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to scale. The same or similar reference numerals are used throughout the drawings to indicate the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. DETAILED DESCRIPTION
[0037] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described here.
[0038] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values of the parts and steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.
[0039] The term “including” and its variations used in the present disclosure are open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least partially based on.”
[0040] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.
[0041] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0042] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0043] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0044] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0045] In order to break away from the limitations of drug molecule databases on molecular structure design and improve the efficiency of molecular structure design, in some embodiments of the present disclosure, a diffusion model can be used to provide candidate molecular structures. The diffusion model is a generative model that can generate new molecular structure data by simulating the diffusion process of molecular structure data gradually from an ordered state to a disordered state (e.g., noise) and learning the reverse process. During the training process of the diffusion model, various conditions can also be provided, such as pharmacophore conditions associated with the molecular structure, target protein conditions associated with the environment in which the molecular structure is located, etc., so as to utilize the attention mechanism to improve the training effect of the diffusion model, especially its denoising network, so that the diffusion model can provide better candidate molecular structures. However, due to the small amount of data related to both the molecular structure and the target protein structure, the training of the diffusion model has certain difficulties.
[0046] In some embodiments of the present disclosure, a method for providing a molecular structure of a drug is proposed, which generates a candidate molecular structure through a diffusion model, and then "evolves" guided by the conformational score of the complex conformation formed by the candidate molecular structure and the target protein structure, thereby generating an updated candidate molecular structure. The method in the embodiments of the present disclosure can generate pharmacophore conditions associated with the molecular structure based on limited target protein conditions, thereby helping to improve the design of the molecular structure. It should be noted that, in this article, the pharmacophore condition can uniquely indicate a clear molecular fragment in the molecular structure (that is, the pharmacophore condition indicates the specific atomic type and its arrangement contained in the molecular fragment, etc.), or it can indicate a certain type of structure in the molecular structure, and this type of structure may include multiple specific molecular fragment structures (for example, the pharmacophore condition can indicate the presence of an aromatic ring in the molecular structure, etc.).
[0047] In some embodiments of the present disclosure, Figure 1 As shown, the method 100 for providing the molecular structure of a drug may include: step S101, inputting a first conditional feature and a first random noise into a diffusion model to obtain a candidate molecular structure; step S102, determining a complex conformation formed by the candidate molecular structure and the target protein structure based on the binding energy between the candidate molecular structure and the target protein structure; step S103, determining a first pharmacophore in the candidate molecular structure based on the interaction between the candidate molecular structure and the target protein structure in the complex conformation; step S104, determining a conformation score of the complex conformation based on the binding energy and the interaction; step S105, storing the candidate molecular structure, the first pharmacophore and the conformation score in association as data in a conditional database; step S106, screening a first number of data from the conditional database based on the conformation scores of all data in the conditional database, and generating a second conditional feature based on one or more first pharmacophores in the first number of data; and step S107, inputting the second conditional feature and the second random noise into the diffusion model to obtain an updated candidate molecular structure.
[0048] In step S101, a first conditional feature can be generated based on the pocket position of the target protein structure. The pocket position refers to a position on the surface or inside of a protein that can specifically bind to a drug molecule, which plays a key role in the function of the protein and is therefore crucial for the development of drug molecular structures. By inputting the first conditional feature and the first random noise into the diffusion model together, the diffusion model can take into account the target protein conditions and generate candidate molecular structures. In some embodiments, the vector or matrix representing the first conditional feature and the vector or matrix representing the first random noise can be concatenated into a higher-dimensional vector or matrix and input into the diffusion model to obtain a candidate molecular structure. In some embodiments, the first random noise can be Gaussian noise.
[0049] In order to generate the first conditional feature according to the pocket position of the target protein structure, in some embodiments, as Figure 2 and Figure 3 As shown, the method for providing the molecular structure of a drug may also include: step S201, determining the amino acid type data of the second number of amino acids at the pocket position, and the atomic number data and atomic coordinate data of the non-hydrogen atoms contained in each amino acid; for each amino acid, performing the following operations: step S202, using the first embedding layer 301 to convert the amino acid type data into an amino acid type feature, step S203, using the second embedding layer 302 to convert the atomic number data of each non-hydrogen atom contained in the amino acid into an atomic number feature, using the first linear layer 303 to convert the atomic coordinate data of each non-hydrogen atom contained in the amino acid into an atomic coordinate feature, and generating the atomic feature of the amino acid based on the atomic number feature and atomic coordinate feature of all non-hydrogen atoms in the amino acid, and step S204, generating the amino acid feature of the amino acid based on the amino acid type feature and the atomic feature of the amino acid; and step S205, splicing the amino acid features of the second number of amino acids to generate the first conditional feature.
[0050] In some embodiments, discrete data can be used to represent amino acid types and atomic numbers (or atomic species), while continuous data can be used to represent atomic coordinates. For example, different integers can be used to represent the corresponding amino acid type and the atomic number or atomic species (e.g., carbon, nitrogen, oxygen, etc.) of the non-hydrogen atoms contained in the amino acid, respectively, and continuous numerical values can be used to represent the atomic coordinates of the non-hydrogen atoms contained in the amino acid. Furthermore, an autoencoding mechanism can be used to uniformly map the aforementioned heterogeneous characteristics of amino acids into a shared latent space, thereby forming amino acid features for input into the diffusion model. The encoder used to perform this autoencoding can be pre-trained and may include a first embedding layer 301 for converting amino acid type data into amino acid type features, a second embedding layer 302 for converting atomic number data into atomic number features, and a first linear layer 303 for converting atomic coordinate data into atomic coordinate features. It should be noted that the first embedding layer 301 or the second embedding layer 302 can each include one or more embedding layers, organized together in a certain manner. Similarly, the first linear layer 303 can include one or more linear layers, organized together in a certain manner to improve the encoding effect, which is not limited here. In some embodiments, the amino acid type characteristics of the amino acid, the atomic number characteristics of each non-hydrogen atom, and the atomic coordinate characteristics can have the same dimensions. Therefore, the matrix or vector addition of the atomic number characteristics and atomic coordinate characteristics of all non-hydrogen atoms can be performed to obtain the atomic characteristics of the amino acid. Furthermore, the matrix or vector addition of the amino acid type characteristics and atomic characteristics of the amino acid can be performed to obtain the amino acid characteristics of the amino acid. In addition, a decoder corresponding to the encoder can also be provided to decode the corresponding amino acid structure or protein structure based on the amino acid characteristics, wherein the decoder can be obtained by training together with the encoder.
[0051] In some embodiments, step S102 may include: performing force field relaxation on the candidate molecular structure and the target protein structure, and determining the complex conformation with the local minimum energy or the global minimum energy as the complex conformation formed by the candidate molecular structure and the target protein structure. Specifically, during the force field relaxation process, energy parameters and force parameters can be generated based on the initial coordinates of the candidate molecular structure and the target protein structure, where the force parameters can indicate the direction of the force; then, the relative position between the candidate molecular structure and the target protein structure can be changed stepwise according to the direction of the force, and energy parameters and force parameters can be generated again; this iteration is carried out until a complex conformation that converges to the local minimum energy or the global minimum energy is determined. It will be understood that in some other embodiments, other methods can also be used to optimize the complex conformation formed by the candidate molecular structure and the target protein structure, which is not limited here.
[0052] Furthermore, after determining the conformation of the complex formed by the candidate molecular structure and the target protein structure, step S103 can be performed to determine the first pharmacophore in the candidate molecular structure based on the interaction between the candidate molecular structure and the target protein structure in the complex conformation. In some embodiments, the type of interaction between the candidate molecular structure and the target protein structure can be detected based on a set of rules. This set of rules can be determined by users such as pharmaceutical experts based on previous knowledge. When the first pharmacophore in the candidate molecular structure and the second pharmacophore in the target protein structure meet a certain distance and angle in space, it can be considered that a certain interaction can be formed between them. Some interaction types and related pharmacophore types are given in the following table:
[0053]
[0054] Furthermore, since the target protein structure can be unchanged in the embodiments disclosed herein, the second pharmacophore in the target protein structure can be ignored, focusing on the first pharmacophore in the candidate molecular structure. This first pharmacophore is then extracted as a condition for the subsequent generation of the diffusion model. This allows the diffusion model, trained on a large amount of molecular or small molecule data, to be used to generate candidate molecular structures, without relying on training data related to both the molecular structure and the target protein structure. This helps the diffusion model provide candidate molecular structures of higher quality. It is understood that in other embodiments, other methods can also be used to analyze the complex conformation and extract the corresponding first pharmacophore, which is not limited here.
[0055] In some embodiments, as Figure 4As shown, step S104 may include: step S1041, determining a binding energy score according to the relaxation energy of the target protein structure alone in solution, the relaxation energy of the candidate molecular structure alone in solution, and the relaxation energy of the complex conformation formed by the candidate molecular structure and the target protein structure in solution; step S1042, determining an interaction score according to the number of pharmacophore pairs in the complex conformation corresponding to the designated pharmacophore pairs, which are formed by the first pharmacophore from the candidate molecular structure and the second pharmacophore from the target protein structure, and the number of designated pharmacophore pairs; and step S1043, determining a conformational score of the complex conformation according to the binding energy score and the interaction score. In some embodiments, the binding energy score may be negatively correlated with the energy of relaxation of the complex conformation in solution, that is, the greater the binding energy between the candidate molecular structure and the target protein structure, the lower the binding energy score will be; and the interaction score may be positively correlated with the number of pharmacophore pairs corresponding to the specified pharmacophore pair in the complex conformation, that is, the more pharmacophore pairs in the complex conformation that hit the specified pharmacophore pair, the higher the interaction score will be.
[0056] In some embodiments, the binding energy score can be the sum of the energy of the candidate molecular structure and the target protein structure when they are relaxed in solution separately minus the difference in the energy of the complex conformation formed by the candidate molecular structure and the target protein structure when they are relaxed in solution. In particular, since the target protein structure can be unchanged in the embodiments of the present disclosure, the energy of the target protein structure when it is relaxed in solution alone can be calculated only once to participate in subsequent calculations. In addition, in order to calculate the energy of the candidate molecular structure when it is relaxed in solution alone, a certain number of molecular conformations can be randomly generated for the same candidate molecular structure, and each molecular conformation can be relaxed separately to obtain multiple local minimum energies. The minimum value of the multiple local minimum energies is then taken to approximate the global minimum energy of the candidate molecular structure (or the energy of the candidate molecular structure when it is relaxed in solution alone), so as to participate in subsequent calculations. It can be seen that this form of binding energy score reflects the static energy difference between the candidate molecular structure and the target protein structure in the unbound state and the bound state. The higher the binding energy score, the greater the possibility of the candidate molecular structure and the target protein structure binding. It is understandable that in some other embodiments, other methods may be used to calculate the binding energy score, which is not limited here.
[0057] In some embodiments, the interaction score can be the ratio of the number of pharmacophore pairs corresponding to a specified pharmacophore pair in the complex conformation, formed by a first pharmacophore from the candidate molecular structure and a second pharmacophore from the target protein structure, to the number of specified pharmacophore pairs. For example, a set of pharmacophore pairs (or interactions) can be pre-specified, and the interaction score calculated based on the first pharmacophore (or interaction) extracted in step S103. For example, a user can specify the following two pharmacophore pairs: [(ASP12-C, HBAcceptor), (TYR13-C, PiStacking)]. If the pharmacophore pair (ASP12-C, HBAcceptor) is actually extracted in step S103, but the pharmacophore pair (TYR13-C, PiStacking) is not, the interaction score will be 1 / 2 = 0.5. It is understood that in other embodiments, other methods can be used to calculate the interaction score, which is not limited here.
[0058] In some embodiments, the product of the binding energy score and the interaction score can be used as a conformational score to evaluate the corresponding complex conformation. In this case, it is not only helpful to screen out molecular structures that can minimize the binding energy, but also to screen out molecular structures that can bind to the desired position of the target protein structure. In some embodiments, after the molecular structure is screened, it can be further evaluated by methods such as free energy perturbation (FEP) or experiments, and the specified pharmacophore pairs or interactions can be continuously updated based on the accumulated knowledge, thereby improving the candidate molecular structures that can be provided by the diffusion model. It is understood that in other embodiments, other methods can also be used to calculate the conformational score, which is not limited here.
[0059] In step S105, for each candidate molecular structure, the candidate molecular structure, the first pharmacophore of the candidate molecular structure, and the conformation score of the complex conformation formed by the candidate molecular structure and the target protein structure can be associated and saved as a data item in the condition database for subsequent operations. For example, the candidate molecular structure, the corresponding first pharmacophore, and the conformation score can be saved as data in the condition database in JSON format. It is understood that in other embodiments, the candidate molecular structure, the corresponding first pharmacophore, and the conformation score can also be saved in other formats, which are not limited here.
[0060] In step S106, one or more data with the best conformation score can be screened out from the condition database, and a second conditional feature can be generated based on one or more first pharmacophores in these data. In some embodiments, the screened out one or more first pharmacophores can be represented as the second conditional feature in a manner similar to representing multiple amino acids at the pocket position of the target protein structure as the first conditional feature. For example, Figure 5 and Figure 6 As shown, generating a second conditional feature based on one or more first pharmacophores in the first number of data may include: step S1061, determining pharmacophore type data and pharmacophore coordinate data of the first pharmacophore for each of the one or more first pharmacophores; step S1062, for each first pharmacophore, converting the pharmacophore type data of the first pharmacophore into a pharmacophore type feature using the third embedding layer 601, converting the pharmacophore coordinate data of the first pharmacophore into a pharmacophore coordinate feature using the second linear layer 602, and generating a fourth conditional feature corresponding to the first pharmacophore based on the pharmacophore type feature and the pharmacophore coordinate feature; step S1063, combining the fourth conditional features of the one or more first pharmacophores into a second conditional feature.
[0061] In some embodiments, discrete data can be used to represent pharmacophore types, and continuous data can be used to represent pharmacophore coordinates. For example, different integers can be used to represent the corresponding pharmacophore types, and continuous values can be used to represent the pharmacophore coordinates. Furthermore, an auto-encoding mechanism can be used to uniformly map the heterogeneous features of the pharmacophores into a shared latent space, thereby forming pharmacophore features for input into the diffusion model. Similar to the generation of the first conditional features described above, the encoder used to perform this auto-encoding can be pre-trained and may include a third embedding layer 601 for converting pharmacophore type data into pharmacophore type features and a second linear layer 602 for converting pharmacophore coordinate data into pharmacophore coordinate features. It should be noted that the third embedding layer 601 may include one or more embedding layers organized in a specific manner. Similarly, the second linear layer 602 may include one or more linear layers organized in a specific manner to improve encoding performance, without limitation. In some embodiments, the pharmacophore type feature and the pharmacophore coordinate feature of a pharmacophore can have the same dimensions, so a matrix or vector addition of the pharmacophore type feature and the pharmacophore coordinate feature can be performed to obtain the pharmacophore feature. Furthermore, a decoder corresponding to the encoder can be provided to decode the pharmacophore structure from the pharmacophore feature, wherein the decoder can be trained together with the encoder. In some embodiments, the fourth conditional feature of one or more first pharmacophores can be concatenated to obtain the second conditional feature.
[0062] In some embodiments, in step S107, only the second conditional features and the second random noise may be input into the diffusion model to obtain an updated candidate molecular structure. In other words, in this case, the conditions used to input the diffusion model may only include the second conditional features associated with the extracted first pharmacophore. Such input features can have a lower dimensionality, thereby helping to improve the efficiency of training the diffusion model and the efficiency of the diffusion model in generating candidate molecular structures. In some embodiments, the second random noise may also be Gaussian noise, which, because it is randomly generated, is typically different from the first random noise.
[0063] In other embodiments, Figure 7 As shown, step S107 may include: step S1071, combining the second conditional feature with the first conditional feature to generate a third conditional feature; and step S1072, inputting the third conditional feature and the second random noise into the diffusion model to obtain an updated candidate molecular structure. In other words, the extracted first pharmacophore and target protein structure can be input together as conditions into the diffusion model to generate an updated candidate molecular structure. In such an embodiment, the diffusion model will take into account the characteristics of both the first pharmacophore and the target protein structure, thereby potentially generating a better candidate molecular structure. Furthermore, in some embodiments, a user, such as a pharmaceutical chemist, may specify a fifth conditional feature. This fifth conditional feature can be used to characterize the pharmacophore (including molecular fragments) contained in the desired candidate molecular structure. By combining the fifth conditional feature with the second conditional feature, or by combining the fifth conditional feature with both the first and second conditional features, and inputting them into the diffusion model, a better candidate molecular structure is likely to be generated. Here, combining conditional features may specifically refer to concatenating the conditional features. Alternatively, in other embodiments, conditional features may be combined in other ways, which are not limiting here.
[0064] In some embodiments, as Figure 8 As shown, the method for providing the molecular structure of a drug may further include: step S108, determining whether the conformation score satisfies a conformation score condition, wherein, in response to the conformation score satisfying the conformation score condition, step S106 is executed, and a first number of data are screened out from the condition database according to the conformation scores of all data in the condition database; and in response to the conformation score not satisfying the conformation score condition, step S109 is executed, and the first condition feature and the third random noise are input into the diffusion model to obtain an updated candidate molecular structure.
[0065] In some cases, the conformational score condition may include whether the conformational score is greater than or equal to a score threshold, which may be zero or another score threshold. If the conformational score associated with the candidate molecular structure generated during a certain iteration of the diffusion model is very low, for example, the conformational score is zero due to the lack of a specified pharmacophore pair, then in order to improve the effect of the subsequent diffusion model in generating candidate molecular structures, the target protein structure can be re-started and the candidate molecular structure can be generated by the diffusion model, that is, returning to the step of inputting the first conditional feature and random noise into the diffusion model to obtain the candidate molecular structure. It should be noted that since the noise is randomly generated, the new third random noise is generally different from the first random noise during the process of returning to perform the above steps, so that the diffusion model is expected to generate a new candidate molecular structure. In addition, if the conformational score meets the conformational score condition, for example, the conformational score is greater than zero, then step S106 can be continued, that is, a second conditional feature is generated based on the data screened from the conditional database, and the candidate molecular structure is updated at least based on the second conditional feature. It is understood that if the conformational score satisfies the conformational score condition, then step S105 generally needs to be executed, that is, the candidate molecular structure, the first pharmacophore, and the conformational score are associated and saved as data in the conditional database; however, if the conformational score does not satisfy the conformational score condition, then step S105 may be executed to ensure the integrity of the data in the conditional database, or step S105 may be omitted to reduce the amount of data to be saved, thereby saving storage space in the conditional database. This is not limited here.
[0066] In some embodiments, as Figure 9 As shown, the candidate molecular structures generated by the diffusion model can be continuously iterated and updated based on the conformational scores. Specifically, starting from the first conditional feature, step S101 can be executed to generate an initial candidate molecular structure. Then, corresponding data can be screened out from the conditional database based on the conformational score, and a second conditional feature can be generated. Then, the second conditional feature, or a combination of the first and second conditional features, can be input into the diffusion model again to generate an updated candidate molecular structure, and the process returns to screen out corresponding data from the conditional database based on the conformational score, and the candidate molecular structures are continuously iterated until a certain stopping condition is reached. For example, the stopping condition can be that a sufficient number of candidate molecular structures have been obtained, or an ideal candidate molecular structure has been obtained, or a certain number of iterations has been reached, or a certain iteration time has been reached, etc.
[0067] In the method disclosed herein, by extracting the first pharmacophore from the effective candidate molecular structure as a condition, the updated candidate molecular structure can stably maintain the interaction with the target protein structure, thereby improving the iterative activity. Moreover, the diffusion model trained based on the pharmacophore condition can be independent of the target protein structure-molecular structure data, so it is possible to pre-train the diffusion model using larger-scale data, thereby ensuring the generation effect of the diffusion model. In addition, in the method disclosed herein, the conformational score of the complex conformation can be determined based on energy and interaction, so that the calculation can be completed quickly and efficiently, while providing the direction of evolution.
[0068] The present disclosure also provides a device 900 for providing the molecular structure of a drug, such as Figure 10 As shown, the apparatus 900 may include: a molecular structure acquisition module 901, a complex conformation determination module 902, a pharmacophore determination module 903, a conformation score determination module 904, a condition database 905, and a condition generation module 906, wherein the molecular structure acquisition module 901 may be configured to input a first condition feature and a first random noise into a diffusion model to obtain a candidate molecular structure, wherein the first condition feature may be generated according to a pocket position of a target protein structure; the complex conformation determination module 902 may be configured to determine a complex conformation formed by the candidate molecular structure and the target protein structure according to the binding energy between the candidate molecular structure and the target protein structure; the pharmacophore determination module 903 may be configured to determine a conformation score of the complex formed by the candidate molecular structure and the target protein structure according to the binding energy between the candidate molecular structure and the target protein structure; The first pharmacophore in the candidate molecular structure is determined based on the interaction between the structures; the conformation score determination module 904 can be configured to determine the conformation score of the complex conformation according to the binding energy and the interaction; the condition database 905 can be configured to store the candidate molecular structure, the first pharmacophore and the conformation score as data in the condition database in an associated manner; the condition generation module 906 can be configured to screen out a first number of data from the condition database 905 according to the conformation scores of all data in the condition database 905, and generate a second condition feature according to one or more first pharmacophores in the first number of data; in addition, the molecular structure acquisition module 901 can also be configured to input the second condition feature and the second random noise into the diffusion model to obtain an updated candidate molecular structure.
[0069] In some embodiments, the conformation score determination module 904 can also be configured to determine whether the conformation score satisfies the conformation score condition, wherein, in response to the conformation score satisfying the conformation score condition, the condition generation module 906 can be configured to filter out a first number of data from the condition database based on the conformation scores of all data in the condition database, and in response to the conformation score not satisfying the conformation score condition, the molecular structure acquisition module 901 can be configured to input the first condition feature and the third random noise into the diffusion model to obtain an updated candidate molecular structure.
[0070] In some embodiments, the condition generation module 906 can be configured to combine the second condition feature with the first condition feature to generate a third condition feature; the molecular structure acquisition module 901 can be configured to input the third condition feature and the second random noise into the diffusion model to obtain an updated candidate molecular structure.
[0071] In some embodiments, the condition generation module 906 can be configured to generate a first conditional feature based on the pocket position of the target protein structure. For example, the condition generation module 906 can be configured to: determine the amino acid type data of the second number of amino acids at the pocket position, and the atomic number data and atomic coordinate data of the non-hydrogen atoms contained in each amino acid; for each amino acid, perform the following operations: using a first embedding layer to convert the amino acid type data into an amino acid type feature, using a second embedding layer to convert the atomic number data of each non-hydrogen atom contained in the amino acid into an atomic number feature, using a first linear layer to convert the atomic coordinate data of each non-hydrogen atom contained in the amino acid into an atomic coordinate feature, and generate the atomic feature of the amino acid based on the atomic number features and atomic coordinate features of all non-hydrogen atoms in the amino acid, and generate the amino acid feature of the amino acid based on the amino acid type feature and atomic feature of the amino acid; and splicing the amino acid features of the second number of amino acids to generate the first conditional feature.
[0072] In some embodiments, the condition generation module 906 can be configured to: determine the pharmacophore type data and pharmacophore coordinate data of each first pharmacophore in one or more first pharmacophores; for each first pharmacophore, use the third embedding layer to convert the pharmacophore type data of the first pharmacophore into a pharmacophore type feature, use the second linear layer to convert the pharmacophore coordinate data of the first pharmacophore into a pharmacophore coordinate feature, and generate a fourth conditional feature corresponding to the first pharmacophore based on the pharmacophore type feature and the pharmacophore coordinate feature; and combine the fourth conditional features of one or more first pharmacophores into a second conditional feature.
[0073] In some embodiments, the complex conformation determination module 902 can be configured to perform force field relaxation on the candidate molecular structure and the target protein structure, and determine the complex conformation with the local lowest energy or the global lowest energy as the complex conformation formed by the candidate molecular structure and the target protein structure.
[0074] In some embodiments, the conformation score determination module 904 can be configured to: determine the binding energy score based on the relaxation energy of the target protein structure alone in solution, the relaxation energy of the candidate molecular structure alone in solution, and the relaxation energy of the complex conformation formed by the candidate molecular structure and the target protein structure in solution; determine the interaction score based on the number of pharmacophore pairs in the complex conformation corresponding to the specified pharmacophore pairs, which are formed by the first pharmacophore from the candidate molecular structure and the second pharmacophore from the target protein structure, and the number of specified pharmacophore pairs; and determine the conformation score of the complex conformation based on the binding energy score and the interaction score.
[0075] In the device disclosed herein, by extracting the first pharmacophore from the effective candidate molecular structure as a condition, the updated candidate molecular structure can stably maintain the interaction with the target protein structure, thereby improving the iterative activity. Moreover, the diffusion model trained based on the pharmacophore condition can be independent of the target protein structure-molecular structure data, so it is possible to pre-train the diffusion model using larger-scale data, thereby ensuring the generation effect of the diffusion model. In addition, in the device disclosed herein, the conformational score of the complex conformation can be determined based on energy and interaction, so that the calculation can be completed quickly and efficiently, while providing the direction of evolution.
[0076] Figure 11 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0077] Memory 111 is used to store one or more computer-readable instructions. Memory 111 may include any combination of various forms of computer-readable storage media, such as volatile and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 111 may store, for example, an operating system, applications, a boot loader, databases, and other programs, as well as various applications and data.
[0078] The processor 112 is configured to execute computer-readable instructions to implement the method for providing the molecular structure of a drug described in any of the aforementioned embodiments or the method described in any of the aforementioned embodiments. The specific implementation of each step of the method can be found in the aforementioned embodiments, and any repetitive details are omitted here.
[0079] The processor 112 may be configured to execute Figures 1 to 9 The processor 112 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) may be an X86 or ARM architecture, etc.
[0080] The processor 112 and the memory 111 can communicate with each other directly or indirectly. For example, the processor 112 and the memory 111 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 112 and the memory 111 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0081] It should be noted that Figure 11 The components of the electronic device 11 shown are merely exemplary and non-limiting. The electronic device 11 may further include other components according to actual application requirements. The processor 112 may control other components in the electronic device 11 to perform desired functions.
[0082] The electronic device 11 may be implemented by software, firmware and / or hardware, and may be integrated into a device installed with relevant application programs.
[0083] Figure 12 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.
[0084] Figure 12 The electronic device 12 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.
[0085] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital televisions and desktop computers, etc.
[0086] like Figure 12As shown, the central processing unit (CPU) 121 executes various processes according to the program stored in the read-only memory (ROM) 122 or the program loaded from the storage part 128 to the random access memory (RAM) 123. In the RAM 123, data required when the CPU 121 executes various processes is stored as needed. The central processing unit is only an example, and it can also be other types of processors, such as the various processors described above. The ROM 122, RAM 123 and the storage part 128 can be various forms of computer-readable storage media. It should be noted that although Figure 12 ROM 122, RAM 123 and storage portion 128 are shown separately in FIG, but one or more of them may be combined or located in the same or different memory or storage modules.
[0087] The CPU 121, the ROM 122, and the RAM 123 are connected to one another via a bus 124. An input / output interface 125 is also connected to the bus 124.
[0088] The following components are connected to the input / output interface 125: an input portion 126 such as a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output portion 127 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage portion 128 including a hard disk, a magnetic tape, etc.; and a communication portion 129 including a network interface card such as a LAN card, a modem, etc. The communication portion 129 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 12 Some of the electronic devices 12 are shown to communicate via the bus 124, but they can also communicate via a network or other means, where the network can include a wireless network, a wired network, and / or any combination of wireless networks and wired networks.
[0089] A drive 1210 is also connected to the input / output interface 125 as needed. A removable medium 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1210 as needed so that a computer program read therefrom is installed into the storage section 128 as needed.
[0090] When the above-described series of processing is implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1211 .
[0091] In the electronic device disclosed herein, by extracting the first pharmacophore from the valid candidate molecular structure as a condition, the updated candidate molecular structure can stably maintain the interaction with the target protein structure, thereby improving the iterative activity. Moreover, the diffusion model trained based on the pharmacophore condition can be independent of the target protein structure-molecular structure data, so it is possible to pre-train the diffusion model using larger-scale data, thereby ensuring the generation effect of the diffusion model. In addition, in the electronic device disclosed herein, the conformational score of the complex conformation can be determined based on energy and interaction, so that the calculation can be completed quickly and efficiently, while providing the direction of evolution.
[0092] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when the computer program product is run on a computer, causes the computer to implement the method described in any of the aforementioned embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for executing the method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from the network through the communication part 129, or installed from the storage part 128, or installed from the ROM 122. When the computer program is executed by the CPU 121, the method of the embodiment of the present disclosure is executed.
[0093] In the computer program product disclosed herein, by extracting the first pharmacophore from the valid candidate molecular structure as a condition, the updated candidate molecular structure can stably maintain the interaction with the target protein structure, thereby improving the iterative activity. Moreover, the diffusion model trained based on the pharmacophore condition can be independent of the target protein structure-molecular structure data, so it is possible to pre-train the diffusion model using larger-scale data, thereby ensuring the generation effect of the diffusion model. In addition, in the computer program product disclosed herein, the conformational score of the complex conformation can be determined based on energy and interaction, so that the calculation can be completed quickly and efficiently, while providing the direction of evolution.
[0094] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.
[0095] The computer readable medium may be a computer readable storage medium, or a computer readable signal medium, or any combination of the two.
[0096] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. Computer instructions are stored on the computer-readable storage medium, and when the instructions are executed by the processor, the method described in any of the aforementioned embodiments is implemented.
[0097] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0098] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0099] In the computer-readable medium disclosed herein, by extracting the first pharmacophore from the effective candidate molecular structure as a condition, the updated candidate molecular structure can stably maintain the interaction with the target protein structure, thereby improving the iterative activity. Moreover, the diffusion model trained based on the pharmacophore condition can be independent of the target protein structure-molecular structure data, so it is possible to pre-train the diffusion model using larger-scale data, thereby ensuring the generation effect of the diffusion model. In addition, in the computer-readable medium disclosed herein, the conformational score of the complex conformation can be determined based on energy and interaction, so that the calculation can be completed quickly and efficiently, while providing the direction of evolution.
[0100] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the method described in any of the aforementioned embodiments. For example, the instructions may be embodied as computer program codes.
[0101] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In situations involving a remote computer, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).
[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0103] The functions described above may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0104] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for providing the molecular structure of a drug, comprising: Inputting a first conditional feature and a first random noise into a diffusion model to obtain a candidate molecular structure, wherein the first conditional feature is generated according to a pocket position of the target protein structure; determining a conformation of a complex formed by the candidate molecular structure and the target protein structure based on the binding energy between the candidate molecular structure and the target protein structure; determining a first pharmacophore in the candidate molecular structure based on the interaction between the candidate molecular structure and the target protein structure in the complex conformation; determining a conformational score of the complex conformation based on the binding energy and the interaction; storing the candidate molecular structure, the first pharmacophore, and the conformational score in association with each other as data in a condition database; Filtering a first number of data from the condition database according to the conformation scores of all data in the condition database, and generating a second condition feature according to one or more first pharmacophores in the first number of data; and The second conditional feature and the second random noise are input into the diffusion model to obtain an updated candidate molecular structure.
2. The method according to claim 1, further comprising: determining whether the conformational score satisfies a conformational score condition, wherein, in response to the conformation score satisfying the conformation score condition, performing an operation of filtering out the first number of data from the condition database according to the conformation scores of all the data in the condition database, In response to the conformational score not satisfying the conformational score condition, the first conditional feature and third random noise are input into the diffusion model to obtain an updated candidate molecular structure.
3. The method according to claim 1, wherein Inputting the second conditional feature and the second random noise into the diffusion model to obtain the updated candidate molecular structure includes: combining the second conditional feature with the first conditional feature to produce a third conditional feature; and The third conditional feature and the second random noise are input into the diffusion model to obtain the updated candidate molecular structure.
4. The method according to claim 1, further comprising: determining amino acid type data of a second number of amino acids at the pocket position, and atomic number data and atomic coordinate data of non-hydrogen atoms contained in each amino acid; For each amino acid, perform the following: The amino acid type data is converted into amino acid type features using a first embedding layer, converting the atomic number data of each non-hydrogen atom contained in the amino acid into an atomic number feature using a second embedding layer, converting the atomic coordinate data of each non-hydrogen atom contained in the amino acid into an atomic coordinate feature using a first linear layer, and generating an atomic feature of the amino acid based on the atomic number feature and the atomic coordinate feature of all non-hydrogen atoms in the amino acid, and generating an amino acid signature of the amino acid based on the amino acid type signature and the atomic signature of the amino acid; and The amino acid signatures of the second number of amino acids are concatenated to generate the first conditional signature.
5. The method according to claim 1, generating a second conditional feature according to one or more first pharmacophores in the first number of data comprises: determining, for each first pharmacophore of the one or more first pharmacophores, pharmacophore type data and pharmacophore coordinate data of the first pharmacophore; For each first pharmacophore, converting the pharmacophore type data of the first pharmacophore into a pharmacophore type feature using a third embedding layer, converting the pharmacophore coordinate data of the first pharmacophore into a pharmacophore coordinate feature using a second linear layer, and generating a fourth conditional feature corresponding to the first pharmacophore based on the pharmacophore type feature and the pharmacophore coordinate feature; The fourth conditional features of the one or more first pharmacophores are combined into the second conditional feature.
6. The method according to claim 1, wherein Determining, based on the binding energy between the candidate molecular structure and the target protein structure, a conformation of a complex formed by the candidate molecular structure and the target protein structure includes: Force field relaxation is performed on the candidate molecular structure and the target protein structure, and the complex conformation with the local minimum energy or the global minimum energy is determined as the complex conformation formed by the candidate molecular structure and the target protein structure.
7. The method according to claim 1, wherein Determining a conformational score of the complex conformation based on the binding energy and the interaction comprises: Determining the binding energy score according to the relaxation energy of the target protein structure alone in the solution, the relaxation energy of the candidate molecular structure alone in the solution, and the relaxation energy of the complex conformation formed by the candidate molecular structure and the target protein structure in the solution; determining the interaction score according to the number of pharmacophore pairs corresponding to the designated pharmacophore pairs in the conformation of the complex, which are formed by the first pharmacophore from the candidate molecule structure and the second pharmacophore from the target protein structure, and the number of the designated pharmacophore pairs; A conformational score of the complex conformation is determined based on the binding energy score and the interaction score.
8. The method according to claim 7, wherein: The binding energy score is negatively correlated with the energy of relaxation of the complex conformation in the solution; and The interaction score is positively correlated with the number of the pharmacophore pairs corresponding to the specified pharmacophore pairs in the conformation of the complex.
9. A device for providing the molecular structure of a drug, comprising: a molecular structure acquisition module configured to input a first conditional feature and a first random noise into a diffusion model to acquire a candidate molecular structure, wherein the first conditional feature is generated according to a pocket position of a target protein structure; a complex conformation determination module, configured to determine the conformation of the complex formed by the candidate molecular structure and the target protein structure based on the binding energy between the candidate molecular structure and the target protein structure; a pharmacophore determination module configured to determine a first pharmacophore in the candidate molecular structure based on the interaction between the candidate molecular structure and the target protein structure in the complex conformation; a conformation score determination module configured to determine a conformation score of the complex conformation based on the binding energy and the interaction; a condition database configured to store the candidate molecular structure, the first pharmacophore, and the conformational score in association with each other as data in the condition database; a condition generation module configured to screen a first number of data from the condition database based on conformational scores of all data in the condition database, and generate a second condition feature based on one or more first pharmacophores in the first number of data; The molecular structure acquisition module is further configured to input the second conditional feature and the second random noise into the diffusion model to obtain an updated candidate molecular structure.
10. An electronic device comprising: at least one memory; as well as At least one processor coupled to the at least one memory, the at least one processor being configured to execute the method for providing a molecular structure of a drug according to any one of claims 1 to 8 based on instructions stored in the at least one memory.
11. A computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the method for providing the molecular structure of a drug according to any one of claims 1 to 8 is implemented. 12 . A computer program product, which, when run on a computer, causes the computer to implement the method for providing the molecular structure of a drug according to claim 1 .
Citation Information
Cited By
Model training method, diffusion model-based molecule generation method and data processing device
CN120804707A