An automated molecular design method and apparatus
By obtaining the docking of the molecules to be screened with protein targets and performing 3D automated screening, the problem of automated screening of molecular design in new drug development has been solved, and efficient molecular optimization and discovery of potential active compounds have been achieved.
Patent Information
- Application Number
- CN202310975218.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-03
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-08-03
AI Technical Summary
Existing technologies struggle to quickly and automatically screen for potentially effective molecules in new drug development, especially in the molecular design process where it is difficult to efficiently traverse the chemical space and combine the experience of medicinal chemists for molecular optimization.
By docking the molecule to be screened with the protein target, the conformation of the complex is obtained, and 3D automatic screening is performed. Candidate molecules are screened out using molecular generation algorithms and 3D automatic screening technology.
It enables efficient molecular design and optimization, provides potential active compounds, and improves the efficiency and accuracy of new drug development.
Smart Images

Figure CN117133378B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of drug design technology, and in particular to an automated molecular design method and apparatus. Background Technology
[0002] Drug discovery is expensive, lengthy, and challenging, with a stark contrast between costs and benefits. With the rapid development of artificial intelligence (AI) in the pharmaceutical field, this new paradigm of drug discovery is considered one of the most promising directions. AI technology can significantly reduce the time and cost of drug discovery, especially AI molecular generation technologies represented by deep learning, such as backbone transition technology, fragment generation technology, and de novo design technology, which are currently the core methods of AI-driven drug development.
[0003] The chemical space is vast; therefore, when researchers generate molecules in a targeted manner, they aim to traverse this space more efficiently, limiting the number of molecules to be considered. Molecular design utilizes AI generative models, including recurrent neural networks (RNNs), variational autoencoders (VAEs), generative adversarial networks (GANs), generative glow-based models (FLOW), and generative pre-training (GPT), to explore molecules in a continuous latent space at the atomic, fragment, or reaction dimensions. Compared to humans, AI molecular generation models possess the ability to learn from vast amounts of data and the potential for molecular design beyond chemical intuition. They can explore a broader chemical space and optimize molecules for specific properties, and have been widely applied in many important molecular design and optimization projects.
[0004] With the continuous development of computer technology, especially cloud computing, in recent years, AI molecular generation models can generate a large number of molecules for given targets and property constraints. Further improving R&D efficiency can be achieved by rapidly and automatically screening potentially effective molecules from these based on druggability, 3D molecular features, key interactions between molecules and targets, and the experience of medicinal chemists. Summary of the Invention
[0005] The technical problem to be solved by the embodiments of this application is to provide an automated molecular design method and apparatus for efficiently designing potential active compounds for selected targets.
[0006] In a first aspect, embodiments of this application provide an automated molecular design method, the method comprising:
[0007] Obtain the molecules to be screened;
[0008] The molecules to be screened are docked with protein targets to obtain a complex conformation;
[0009] The conformation of the complex was subjected to 3D automated screening to obtain candidate molecules.
[0010] Optionally, the molecules to be screened are obtained by performing a molecular generation algorithm based on a molecular generation scheme.
[0011] Optionally, the molecular generation scheme includes at least one of: R-base generation, side chain generation, linker generation, backbone transition, and de novo generation;
[0012] The R-base generation, side chain generation, linker generation, and backbone transition are molecular generation based on specific substructures. The R-base generation specifies the growth site, while the side chain generation does not specify a specific growth site.
[0013] Optionally, the molecular generation algorithm is obtained through real-time fine-tuning training based on binding pocket information and / or active molecule data in protein crystallization data.
[0014] Optionally, the 3D automated screening process includes: determining whether the molecule to be screened in the complex conformation forms a specific interaction with a specific amino acid residue of the protein target.
[0015] Optionally, the 3D automatic screening process includes: using a known active molecule of the protein target as a reference molecule, and calculating the degree of spatial overlap between the specific substructure of the reference molecule and the specific substructure of the molecule to be screened when they bind to the protein target.
[0016] Optionally, the 3D automatic screening process includes: determining whether a specific type of atom exists in a specific region of the complex conformation.
[0017] Optionally, before docking the molecule to be screened with the protein target, the molecule to be screened is subjected to automated screening for unreasonable substructures; the unreasonable substructures include at least one of the following: substructures with low stability, rare substructures, and substructures with high synthesis difficulty.
[0018] Optionally, before docking the molecule to be screened with the protein target, the molecule to be screened is subjected to automated property screening; the property screening conditions include at least one of molecular weight, topological polar surface area, lipid solubility, water solubility, PAINS, Alerts, and syntheticity.
[0019] Secondly, embodiments of this application provide an automated molecular design device, the device comprising:
[0020] The module for acquiring molecules to be screened is used to acquire molecules to be screened.
[0021] The complex conformation acquisition module is used to dock the molecule to be screened with the protein target to obtain the complex conformation.
[0022] The candidate molecule acquisition module is used to perform 3D automatic screening of the complex conformation to obtain candidate molecules.
[0023] Optionally, the molecules to be screened are obtained by performing a molecular generation algorithm based on a molecular generation scheme.
[0024] Optionally, the molecular generation scheme includes at least one of: R-base generation, side chain generation, linker generation, backbone transition, and de novo generation;
[0025] The R-base generation, side chain generation, linker generation, and backbone transition are molecular generation based on specific substructures. The R-base generation specifies the growth site, while the side chain generation does not specify a specific growth site.
[0026] Optionally, the molecular generation algorithm is obtained through real-time fine-tuning training based on binding pocket information and / or active molecule data in protein crystallization data.
[0027] Optionally, the 3D automated screening process includes: determining whether the molecule to be screened in the complex conformation forms a specific interaction with a specific amino acid residue of the protein target.
[0028] Optionally, the 3D automatic screening process includes: using a known active molecule of the protein target as a reference molecule, and calculating the degree of spatial overlap between the specific substructure of the reference molecule and the specific substructure of the molecule to be screened when they bind to the protein target.
[0029] Optionally, the 3D automatic screening process includes: determining whether a specific type of atom exists in a specific region of the complex conformation.
[0030] Optionally, the device further includes:
[0031] The molecule screening module is used to automatically screen the molecule for unreasonable substructures before docking the molecule for screening with the protein target; the unreasonable substructures include at least one of the following: substructures with low stability, rare substructures, and substructures with high synthesis difficulty.
[0032] Optionally, the device further includes:
[0033] The molecular attribute screening module is used to automatically screen the attributes of the molecule to be screened before docking it with the protein target. The attribute screening conditions include at least one of the following: molecular weight, topological polar surface area, lipid solubility, water solubility, PAINS, Alerts, and syntheticity.
[0034] Thirdly, embodiments of this application provide an electronic device, including:
[0035] A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the automated molecular design method described in any of the preceding claims.
[0036] Fourthly, embodiments of this application provide a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the automated molecular design method described in any of the preceding claims.
[0037] Compared with the prior art, the embodiments of this application have the following advantages:
[0038] In this embodiment, the molecule to be screened is obtained, and then docked with a protein target to obtain a complex conformation. The complex conformation is then subjected to 3D automated screening to obtain candidate molecules. This embodiment establishes a complete molecular design workflow based on 3D automated molecular screening technology, enabling efficient molecular design and optimization, and providing potential active molecules.
[0039] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0040] Figure 1 A flowchart illustrating the steps of an automated molecular design method provided in this application embodiment;
[0041] Figure 2 A flowchart illustrating the steps of an automated molecular design method provided in this application embodiment;
[0042] Figure 3A schematic diagram of a molecular property screening process provided in an embodiment of this application;
[0043] Figure 4 A schematic diagram of the structure of an automated molecular design device provided in an embodiment of this application;
[0044] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0045] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0047] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes said element.
[0048] Reference Figure 1 This document illustrates a flowchart of the steps involved in an automated molecular design method provided in an embodiment of this application. Figure 1 As shown, this automated molecular design method may include the following steps:
[0049] Step 101: Obtain the molecules to be screened.
[0050] In this embodiment, the molecules to be screened can be obtained by executing a molecular generation algorithm based on a molecular generation scheme. The molecular generation algorithm can be obtained through real-time fine-tuning training based on binding pocket information and / or active molecule data from protein crystallization data.
[0051] In practical implementation, when designing potential drug compounds targeting protein targets, target data information can be obtained. Specifically, this involves in-depth research of patent literature and biochemical databases to analyze the biological mechanisms and metabolic pathways of the targets, and collecting and organizing publicly reported active molecules and bioactivity data, protein crystallization data, etc. After obtaining the target data information, the corresponding molecular generation scheme and molecular screening conditions can be determined based on the target data information.
[0052] In this embodiment, the molecular generation scheme may include at least one of the following: R-base generation, side chain generation, linker generation, backbone transition, and de novo generation.
[0053] Among them, R-base generation, side chain generation, linker generation, and backbone transition are molecular generation based on specific substructures. R-base generation specifies the growth site, while side chain generation does not specify a specific growth site.
[0054] After determining the molecular generation scheme and molecular screening conditions corresponding to the protein target based on the target data, the molecular generation scheme can be executed to obtain the molecules to be screened corresponding to the protein target. Specifically, based on the molecular generation scheme, the molecular generation algorithm can be used to generate the molecules to be screened corresponding to the protein target using the binding pocket information and / or active molecule data in the protein crystallization data.
[0055] In this embodiment, after obtaining the target data information of the protein target, a molecular generation scheme corresponding to the protein target can be designed based on the active molecule data in the target data information. Specifically, a molecular generation scheme can be designed based on the structure-activity relationship (SAR). The retained region is analyzed according to the molecular generation scheme. The retained region can be one or more, and can be a single atom or a group; this embodiment does not impose any limitations on this. The molecular generation scheme can include fragment generation, backbone transition, side chain generation, linker generation, and de novo generation schemes, etc.; this embodiment also does not impose any limitations on this.
[0056] After obtaining the molecules to be screened, proceed to step 102.
[0057] Step 102: The molecule to be screened is docked with the protein target to obtain the complex conformation.
[0058] After obtaining the molecule to be screened, it can be docked with the protein target to obtain the complex conformation.
[0059] Before docking the molecule to be screened with the protein target, the molecule to be screened can be automatically screened for unreasonable substructures. The unreasonable substructures can include at least one of the following: substructures with low stability, rare substructures, and substructures with high synthesis difficulty.
[0060] Before docking the molecules to be screened with protein targets, automated property screening can be performed on the molecules. The property screening conditions can include at least one of the following: molecular weight, topological polar surface area, lipid solubility, water solubility, PAINS, Alerts, and syntheticity.
[0061] In practice, after obtaining the molecules to be screened, attribute screening can be performed on the molecules to be screened based on attribute screening conditions.
[0062] In this example, the attribute screening criteria may include at least one of the following: molecular weight, topological polar surface area, lipophilicity, water solubility, PAINS, Alerts, irrational substructure, quantitative assessment of drug-likeness, number of hydrogen bond donors, number of hydrogen bond acceptors, syntheticity, number of rotatable bonds, number of rings, number of spirocyclic rings, number of bridged rings, number of macrocycles, and similarity to the active molecule.
[0063] Wherein, molecular weight (MW) is the relative molecular weight of a molecule.
[0064] Quantitative Estimate of Drug-Likeness (QED) quantifies drug-likeness into a value between 0 and 1 by combining multiple molecular descriptors. The closer the value is to 1, the greater the potential drug-likeness of the molecule.
[0065] Topological Polar Surface Area (TPSA) is used to predict drug flow and can represent the correlation between human intestinal absorption and blood-brain barrier penetration.
[0066] Hydrogen Bond Donors Count (HDB) is the number of hydrogen bond donors in a molecule.
[0067] Hydrogen bond acceptor count (HDA) is the number of hydrogen bond acceptors in a molecule.
[0068] The calculated logarithm lipophilicity (ClogP) is the calculated logarithm octanol–water coefficient; ClogP is a lipophilicity index that has a significant impact on membrane permeability and hydrophobic binding with macromolecules.
[0069] The calculated logarithm solubility (ClogS) is an indicator of molecular water solubility; low solubility is detrimental to oral absorption.
[0070] The Synthetic Accessibility Score (SA_score) is used to quantitatively assess the syntheticity of a molecule. It is calculated by assigning a fractional number of molecular fragments and applying a complexity penalty, resulting in a value between 0 and 10. A SA_score closer to 0 indicates that the molecule is easier to synthesize, while a SA_score closer to 10 indicates that the molecule is more difficult to synthesize.
[0071] The number of rotational bonds (NRB) is used to calculate the number of rotational bonds in a molecule.
[0072] The number of rings (NR) is used to calculate the number of rings in a molecular ring.
[0073] The number of spirocyclic rings (NSC) is used to calculate the number of spirocyclic rings in a molecule.
[0074] The number of spiro atoms (NS) is used to calculate the number of spiro atoms in a molecule.
[0075] The number of bridge heads (NBH) is the number of bridge atoms in a molecule.
[0076] The number of macrocycles (NMC) is used to calculate the number of multiple rings in a molecule. The presence of multiple triple rings in a molecule is detrimental to its synthesis.
[0077] PAINS (Pan-Assay Interference Compounds) may cause false positives due to their corresponding structures. It should be noted that the filtering effectiveness of PAINS has been repeatedly questioned in the industry, with a tendency to reject too many false positives; therefore, PAINS-filtered results should only be considered as a reference.
[0078] Alerts substructures refer to harmful structures, including toxic groups or groups that are unstable and easily degraded in the body, or that may undergo specific chemical reactions. Early detection and screening during the initial stages of drug development can reduce unnecessary experimental costs.
[0079] Unreasonable substructures, in addition to the molecular structures listed above, are supplemented by a compilation of irrational chemical substructures identified through accumulated drug experimental data. These irrational chemical substructures can be uncommon in drug-like molecules or arbitrary substructures specified by experts. In this example, irrational substructures can include at least one of the following: substructures with low stability, rare substructures, and substructures with high synthetic difficulty. In practical implementations, corresponding thresholds, such as stability thresholds and synthetic difficulty thresholds, can be pre-set to determine whether a substructure is irrational.
[0080] Tanimoto Similarity refers to the two-dimensional topological similarity between molecules to be screened for attributes and known active molecules, based on extended connectivity fingerprints.
[0081] The molecular property screening process can be as follows: Figure 3 As shown, for the input .sdf.mol.smi.csv file, the system can read the legal molecules and convert them into mol objects. It can then calculate the drug-like properties of the molecules and query for special substructures, such as SMILES, MW, QED, TPSA, HBD, HAD, etc. Finally, molecules can be filtered according to predetermined rules to obtain the final .sdf file, which represents the molecules to be docked.
[0082] After the above processing, the molecules to be screened can be docked with protein targets using molecular docking algorithms to obtain the complex conformation.
[0083] After docking the molecule to be screened with the protein target to obtain the complex conformation, step 103 is performed.
[0084] Step 103: Perform 3D automatic screening on the conformation of the complex to obtain candidate molecules.
[0085] After obtaining the complex conformation, the complex conformation can be screened using a 3D automatic screening method to obtain candidate molecules.
[0086] In this example, the 3D automated screening process may include determining whether the molecule to be screened in the complex conformation forms a specific interaction with a specific amino acid residue of the protein target.
[0087] 3D automated screening can also include: using known active molecules of protein targets as reference molecules, and calculating the degree of spatial overlap between the specific substructure of the reference molecule and the specific substructure of the molecule to be screened when they bind to the protein target.
[0088] 3D automatic screening can also include: determining whether a specific type of atom exists in a specific region of the complex conformation.
[0089] Specifically, based on the conformation of the complex after molecular docking, the root mean square deviation (RMSD) of the selected retention region in the molecular generation scheme can be calculated, the interaction between the compound conformation and the target site can be calculated, and molecules can be filtered based on the RMSD and the interaction. The specific screening process is as follows:
[0090] 1. Candidate molecules are screened based on whether the complex conformation contains a specific substructure.
[0091] 2. Candidate molecules are screened based on the degree of spatial overlap between specific substructures in the complex conformation and specific substructures of known active molecules, wherein the degree of spatial overlap is quantified by RMSD.
[0092] 3. Candidate molecules are screened based on whether a specific atom appears in a specific region of the complex conformation.
[0093] 4. Candidate molecules are screened based on whether the complex conformation forms a specific key interaction with a specific amino acid residue of the protein.
[0094] 5. Candidate molecules were screened based on molecular docking sequencing.
[0095] After candidate molecules are obtained by screening the complex conformation using 3D automated screening technology, business personnel or CADD (Carbon Derivative Development and Analysis) can conduct a comprehensive evaluation of the candidate molecules. Specifically, retrosynthetic planning is carried out on the molecules retained after 3D automated molecular screening. Molecules that fail to complete retrosynthetic planning or whose time exceeds a threshold are screened out, while those that meet the requirements are retained. For the remaining molecules, drug experts comprehensively evaluate and select potential active molecules based on expert experience, patent data, druggability prediction data, post-dock molecular conformation, and affinity scores. The potential active molecules selected and retained through comprehensive evaluation can be used as training data to fine-tune the molecular generation model.
[0096] Among them, CADD (Computer-Aided Drug Design) is a method for designing and optimizing lead compounds. Based on research findings in life sciences such as biochemistry, enzymology, molecular biology, and genetics, it targets potential drug targets revealed in these fundamental studies, including enzymes, receptors, ion channels, and nucleic acids. It also references the chemical structural characteristics of other ligands or natural products, using computer-based chemistry to simulate, calculate, and estimate the interactions between drugs and receptor biomolecules. This involves examining the structural and property complementarity between the drug and the target to design rational drug molecules.
[0097] After the business personnel conduct a comprehensive evaluation and selection of candidate molecules, the selected candidate molecules can be ranked based on free energy calculation methods, and potential active molecules that meet the criteria can be screened according to the ranking results. Specifically, through molecular dynamics simulations, the binding free energies of the target and the compound are calculated, such as molecular mechanics / Poisson Boltzmann (Generalized Born) Surface Area (MM / PB(GB)SA) and free energy perturbation (FEP), to further evaluate the binding free energies of the target and the compound more accurately and establish a ranking of compound affinities.
[0098] In practical applications, 3D automatic screening includes one or more of the following steps:
[0099] 1. Determine whether the complex conformation contains a specific substructure. If the complex conformation does not contain a specific substructure, the complex conformation should be filtered out; otherwise, it should be retained. The specific substructure can be a key substructure known to be the binding of an active molecule to a protein and the resulting interaction, or it can be any substructure specified by an expert. The specific substructure can be represented in the form of SMILES, or in chemical formats such as Mol or SDF. The method for determining whether the complex conformation contains a specific substructure can be to use the GetSubstructMatch function of RDKit to identify whether the complex conformation contains a specific substructure, or it can be other methods. This embodiment does not limit this method.
[0100] 2. Measure the degree of spatial overlap between a specific substructure of the complex conformation and a specific substructure of a known active molecule. The degree of spatial overlap can be quantified by RMSD. If the degree of spatial overlap is higher than the spatial overlap threshold, the complex conformation should be filtered out; otherwise, it should be retained. The specific substructure of the complex conformation has the same constituent atom types and inter-atomic connections as the specific substructure of the known active molecule.
[0101] 3. Determine whether a specific atom appears in a specific region of the complex conformation and the protein complex conformation. The specific atom can be a key atom known to bind and interact with the known active molecule and protein, or any atom specified by an expert. The process for determining the specific region can be as follows: determine the spatial position of the specific atom based on the known complex conformation of the active molecule and the protein; the specific region is the area within a sphere with a radius of 1.5 angstroms centered on the spatial position of the specific atom. Specifically, determine whether a specific atom appears in the specific region of the complex conformation and the protein complex conformation: locate the atom located in the sphere region within the complex conformation; determine whether the atom in the sphere region contains the specific atom type. If not, the complex conformation should be filtered out; otherwise, it should be retained.
[0102] 4. Determine whether the complex conformation forms a specific key interaction with a specific amino acid residue of the protein. If no specific key interaction is formed, the complex conformation should be filtered out; otherwise, it should be retained. The specific amino acid residue can be an amino acid residue in the protein that forms a key interaction with a known active molecule, or an amino acid residue specified by experts that the complex conformation is expected to form a key interaction with. The specific key interaction can be one or more of the following interaction types: salt bridge, hydrogen bond, pistacking, pication, hydrophobic interaction, halogen bond, metal chelation, and family interaction. The key interaction can be analyzed and labeled using PLIP (Protein-Ligand Interaction Profiler).
[0103] Based on molecular generation schemes and automated molecular screening technology, this application establishes a complete molecular design process, enabling efficient molecular design and optimization, and providing potential active molecules.
[0104] The above molecular design process can be combined with Figure 2 The following is a detailed description.
[0105] Reference Figure 2 The diagram shows a flowchart of the steps of a molecular design method provided in an embodiment of this application.
[0106] like Figure 2 As shown, this complete molecular design method may include the following steps:
[0107] 1. Obtain the target as input, i.e., the protein target.
[0108] 2. Target analysis, which is conducted in two ways: one is biochemical database retrieval, and the other is literature retrieval and structure extraction. The analysis results are target data information, namely bioactivity data and protein crystallization data.
[0109] 3. Screening condition analysis, including: 1) docking algorithm analysis; 2) active molecule property analysis; 3) QSAR model establishment; 4) determination of key interactions. The analysis results are: molecular screening and docking conditions.
[0110] 4. The molecular generation scheme is based on information about active molecules to formulate a generation plan, while also defining 3D screening conditions. Its output is the molecular generation task.
[0111] 5. Molecular Generation: The molecular generation task is performed using at least one of the following models: skeletal transition model, fragment generation model, linker generation model, side chain generation model, and de novo generation model. The output is a target-specific molecule.
[0112] 6. Molecular attribute screening filters molecules based on a given range of molecular attributes, and outputs attribute-specific molecules.
[0113] 7. Molecular docking: Molecular docking is performed based on the selected docking algorithm, and the output result is the docking result.
[0114] 8. 3D Molecular Screening: By calculating the RMSD of the retained region and the key interactions after docking, 3D specific molecules are output.
[0115] 9. CADD / Medicinal Chemistry Expert Analysis: Analyze the results based on CADD and / or the experience of medicinal chemistry experts to identify potential active molecules.
[0116] 10. Molecular ranking: MM / PB(GB)SA, FEP and other free energy calculations are used to rank molecules by affinity, and the final potential active molecules are screened from the potential active molecules based on the ranking results.
[0117] The automated molecular design method provided in this application involves obtaining a molecule to be screened, docking the molecule with a protein target to obtain a complex conformation, and then performing 3D automated screening on the complex conformation to obtain candidate molecules. Based on 3D automated molecular screening technology, this application establishes a complete molecular design workflow, enabling efficient molecular design and optimization, and providing potential active molecules.
[0118] Reference Figure 4 The diagram shows a schematic representation of an automated molecular design device provided in an embodiment of this application. Figure 4 As shown, the automated molecular design device 400 may include the following modules:
[0119] The molecule acquisition module 410 is used to acquire molecules to be screened.
[0120] The complex conformation acquisition module 420 is used to dock the molecule to be screened with the protein target to obtain the complex conformation.
[0121] The candidate molecule acquisition module 430 is used to perform 3D automatic screening of the complex conformation to obtain candidate molecules.
[0122] Optionally, the molecules to be screened are obtained by performing a molecular generation algorithm based on a molecular generation scheme.
[0123] Optionally, the molecular generation scheme includes at least one of: R-base generation, side chain generation, linker generation, backbone transition, and de novo generation;
[0124] The R-base generation, side chain generation, linker generation, and backbone transition are molecular generation based on specific substructures. The R-base generation specifies the growth site, while the side chain generation does not specify a specific growth site.
[0125] Optionally, the molecular generation algorithm is obtained through real-time fine-tuning training based on binding pocket information and / or active molecule data in protein crystallization data.
[0126] Optionally, the 3D automated screening process includes: determining whether the molecule to be screened in the complex conformation forms a specific interaction with a specific amino acid residue of the protein target.
[0127] Optionally, the 3D automatic screening process includes: using a known active molecule of the protein target as a reference molecule, and calculating the degree of spatial overlap between the specific substructure of the reference molecule and the specific substructure of the molecule to be screened when they bind to the protein target.
[0128] Optionally, the 3D automatic screening process includes: determining whether a specific type of atom exists in a specific region of the complex conformation.
[0129] Optionally, the device further includes:
[0130] The molecule screening module is used to automatically screen the molecule for unreasonable substructures before docking the molecule for screening with the protein target; the unreasonable substructures include at least one of the following: substructures with low stability, rare substructures, and substructures with high synthesis difficulty.
[0131] Optionally, the device further includes:
[0132] The molecular attribute screening module is used to automatically screen the attributes of the molecule to be screened before docking it with the protein target. The attribute screening conditions include at least one of the following: molecular weight, topological polar surface area, lipid solubility, water solubility, PAINS, Alerts, and syntheticity.
[0133] The full-process molecular design apparatus provided in this application involves acquiring molecules to be screened, docking these molecules with protein targets to obtain complex conformations, and then performing 3D automated screening on the complex conformations to obtain candidate molecules. Based on 3D automated molecular screening technology, this application establishes a complete molecular design workflow, enabling efficient molecular design and optimization, and providing potential active molecules.
[0134] This application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the above-described automated molecular design method.
[0135] Figure 5 A schematic diagram of the structure of an electronic device 500 according to an embodiment of the present invention is shown. Figure 5 As shown, the electronic device 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 502 or loaded from storage unit 508 into random access memory (RAM) 503. The RAM 503 can also store various programs and data required for the operation of the electronic device 500. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0136] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, microphone, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0137] The various processes and handling described above can be executed by processing unit 501. For example, the methods of any of the above embodiments can be implemented as computer software programs tangibly contained in a computer-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by CPU 501, one or more actions of the methods described above can be performed.
[0138] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned automated molecular design method.
[0139] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0140] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminals (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal, generate instructions for implementing the flowchart illustrations. Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0142] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal, causing a series of operational steps to be executed on the computer or other programmable terminal to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0144] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0145] The above provides a detailed description of an automated molecular design method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An automated molecular design method, characterized in that, The method includes: Obtain the molecules to be screened; The molecules to be screened are obtained by executing a molecular generation algorithm based on a molecular generation scheme. The molecular generation scheme includes at least one of R-base generation, side chain generation, linker generation, skeletal transition, and de novo generation; the R-base generation, side chain generation, linker generation, and skeletal transition are molecular generation based on specific substructures, the R-base generation specifies a growth site, and the side chain generation does not specify a specific growth site. The molecular generation algorithm is obtained through real-time fine-tuning training based on binding pocket information and / or active molecule data in protein crystallization data. The molecules to be screened are docked with protein targets to obtain a complex conformation; The conformation of the complex was subjected to 3D automated screening to obtain candidate molecules; The 3D automatic screening process includes: using a known active molecule of the protein target as a reference molecule, and calculating the degree of spatial overlap between the specific substructure of the reference molecule and the specific substructure of the molecule to be screened when they bind to the protein target. Before docking the molecule to be screened with the protein target, the molecule to be screened is subjected to automated property screening; the property screening conditions include at least one of the following: molecular weight, topological polar surface area, lipid solubility, water solubility, PAINS, Alerts and syntheticity.
2. The method according to claim 1, characterized in that, The 3D automated screening process includes determining whether the molecule to be screened in the complex conformation forms a specific interaction with a specific amino acid residue of the protein target.
3. The method according to claim 1, characterized in that, The 3D automatic screening process includes: determining whether a specific type of atom exists in a specific region of the complex conformation.
4. The method according to claim 1, characterized in that, Before docking the molecule to be screened with the protein target, the molecule to be screened is subjected to automated screening for unreasonable substructures. The unreasonable substructures include at least one of the following: substructures with low stability, rare substructures, and substructures that are difficult to synthesize.
5. An automated molecular design device, characterized in that, The device includes: The module for acquiring molecules to be screened is used to acquire molecules to be screened. The complex conformation acquisition module is used to dock the molecule to be screened with the protein target to obtain the complex conformation. The candidate molecule acquisition module is used to perform 3D automatic screening of the complex conformation to obtain candidate molecules. The molecular attribute screening module is used to automatically screen the attributes of the molecule to be screened before docking it with the protein target. The attribute screening conditions include at least one of the following: molecular weight, topological polar surface area, lipid solubility, water solubility, PAINS, Alerts, and syntheticity. The molecules to be screened are obtained by executing a molecular generation algorithm based on a molecular generation scheme. The molecular generation scheme includes at least one of R-base generation, side chain generation, linker generation, skeletal transition, and de novo generation; the R-base generation, side chain generation, linker generation, and skeletal transition are molecular generation based on specific substructures, the R-base generation specifies a growth site, and the side chain generation does not specify a specific growth site. The molecular generation algorithm is obtained through real-time fine-tuning training based on binding pocket information and / or active molecule data in protein crystallization data. The 3D automatic screening process includes: using a known active molecule of the protein target as a reference molecule, and calculating the degree of spatial overlap between the specific substructure of the reference molecule and the specific substructure of the molecule to be screened when they bind to the protein target.
Citation Information
Patent Citations
General molecular library construction platform for screening small molecular drugs
CN113096723A
Compound library construction method and device based on artificial intelligence, equipment and storage medium
CN113436686A