Method for constructing the molecular structure of a compound
A method using a personal computer to construct compound molecular structures addresses the challenges of resource-intensive and expert-dependent drug design, allowing professionals to generate novel compounds efficiently.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- 岩瀬 一彦
- Filing Date
- 2025-09-12
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for designing compound molecular structures for drug discovery are time-consuming, costly, and require significant computational resources and expertise in advanced computer science or artificial intelligence, making them inaccessible to many drug discovery professionals.
A method involving the use of a personal computer to construct compound molecular structures by dropping spheres onto the molecular surface of a three-dimensional structural model, forming a carbon graph, and converting atoms to interact with organic polymers, using tools like PyMOL, Blender, and AutoDock Vina to visualize and stabilize the structure.
Enables drug discovery professionals to easily design compounds without extensive computing resources or advanced technical expertise, facilitating the generation of novel compound ideas based on their experience.
Smart Images

Figure 0007854675000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a compound molecular structure that can be used for the molecular design of compounds having biological activity. The present invention also relates to a program for executing the method.
Background Art
[0002] It is said that it takes a long time and a large amount of money to develop a single new drug. Even with such an investment, the success rate of developing new drugs is extremely low. So far, compounds have been widely collected by synthesis, culture, extraction, etc., and a library consisting of a compound group exceeding hundreds of thousands to millions has been created, and high-throughput screening (HTS) has been performed to search for candidate compounds for new drugs. Alternatively, virtual screening (VS) has been performed by inputting the atomic coordinates of proteins and compounds into a computer, specifying the compound binding site of the protein, and narrowing down the compounds that may bind thereto to search for candidate compounds for new drugs. However, HTS has problems such as high cost and low hit rate of compounds, and although VS has a higher hit rate than HTS, it also has problems such as the need for an enormous amount of calculation time to significantly improve the accuracy. Therefore, VS has been used for compound library design to improve the hit rate of HTS, and a supercomputer has been used to improve the accuracy of VS. Recently, in addition to VS using a supercomputer, artificial intelligence technology that deeply learns a huge amount of data has also been utilized. For such large-scale compound design methods, there is a need for a method that allows drug discovery researchers to easily design compounds.
[0003] Regarding compound design methods, Kuntz et al. generate spheres along the surface normal at each surface point across the entire surface of the protein three-dimensional structure to fill in depressions, and search for compounds so that the distance between the sphere pairs and the compound atom pairs matches (Non-Patent Literature 1). Itai et al. also set dummy atoms at the positions of heteroatoms that can become hydrogen bonding partners of hydrogen-bonding functional groups of protein amino acid residues located in depressions on the molecular surface of the protein, and search for protein compound complex structures so that the distance between the dummy atom pairs and the hydrogen-bonding heteroatom pairs of compound atoms matches (Patent Literature 1). However, these methods cannot generate compounds.
[0004] Itai et al. have proposed a method for designing the molecular structure of a compound by determining the type and position of each atom in the cavities where compounds bind in a three-dimensional protein structure model using random numbers and electrostatic potential to determine the dihedral angle, evaluating the interaction energy, and arranging the atoms accordingly (Non-Patent Literature 2). Fujitani et al. have also proposed a method for designing the compound structure by evaluating the interaction energy and arranging substructures of a compound in the cavities where compounds bind in a three-dimensional protein structure model, and then chemically linking the substructures in a way that is possible (Non-Patent Literature 3). However, the former method has difficulty in generating appropriate compounds by adding atoms one by one, and the accuracy of the energy calculations is insufficient, while the latter method requires running numerous independent molecular dynamics simulations in a large-scale parallel environment to perform energy calculations with high accuracy.
[0005] In recent years, a method has been proposed (Non-Patent Document 4) that uses a conditional variational autoencoder to learn the atomic density within the cavities where compounds bind to a three-dimensional protein structure model, thereby generating the molecular structure of a compound. However, this method, which uses artificial intelligence technology that requires a large amount of data and long learning times, is a concern because it is highly dependent on the quality and quantity of existing data, and furthermore, it is not something that can be easily implemented by drug discovery personnel who are not computer science experts. [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] International Publication No. 93 / 20525 [Non-patent literature]
[0007] [Non-Patent Document 1] Kuntz, Irwin D. et al., A geometric approach to macromolecule-ligand interactions, Journal of Molecular Biology, vol.161, no.2, pp.269-288, 1982. [Non-Patent Document 2] Nishibata, Y. et al., Automatic creation of drug candidate structures based on receptor structure. Starting point for artificial lead generation, Tetrahedron, vol.47, no.43, pp.8985-8990, 1991. [Non-Patent Document 3] Yamashita, T. et al., Molecular Dynamics Simulation-Based Evaluation of the Binding Free Energies of Computationally Designed Drug Candidates: Importance of the Dynamical Effects, Chemical and Pharmaceutical Bulletin, vol.62, no.7, pp.661-667, 2014. [Non-Patent Document 4] Ragoza, M. et al., Generating 3D molecules conditional on receptor binding sites with deep generative models, Chemical Science, vol.13, no.9, pp.2701-2713, 2022. [Overview of the Initiative] [Problems that the invention aims to solve]
[0008] There is a need for a method that allows drug discovery professionals to easily construct compound molecular structures that may bind to organic polymers (proteins, DNA / RNA, etc.) using a personal computer, without requiring large computing resources, computation time, complex operating procedures, or expertise in cutting-edge computer science or artificial intelligence. Furthermore, there is a need for a method of constructing compound molecular structures that is useful for drug discovery professionals to generate novel compound ideas and is based on their experience, rather than a one-stop compound generation method like conditional variational autoencoders or diffusion models of artificial intelligence. [Means for solving the problem]
[0009] This invention involves loosely randomly filling the molecular surface of a three-dimensional structural model of an organic polymer by dropping spheres of a specified radius onto the surface of the cavity where the compound bonds. Next, the radius of the filling spheres is increased, and spheres that extend beyond the molecular surface of the organic polymer cavity are removed. The centers of the remaining filling spheres are then connected with a three-dimensional minimal spanning tree to construct a carbon graph. The ring structure is then reconstructed visually, or using genetic algorithms or reinforcement learning. Finally, carbon atoms within interacting distance of the main atoms of the organic polymer are converted to atoms (oxygen, nitrogen, etc.) or bonds (aromatic rings, double bonds, etc.) suitable for interaction to construct the compound molecular structure, then strain is reduced, and the stabilized organic polymer-compound complex structure is calculated using a docking tool.
[0010] In other words, the present invention is as follows. (1) This method involves using a computer to perform the following steps 1 through 5 to construct a compound molecular structure on the surface of a depression in an organic polymer to which the compound is bound. 1st step Download the PDB file of the target organic polymer's three-dimensional structural model from the Protein Databank. Open the PDB file in a molecular graphics tool to display the three-dimensional structural model of the organic polymer. Then, specify the recessed surface of the compound that binds to the organic polymer and the organic polymer anchor (the main atom of the organic polymer that interacts with the compound) and output it as a VRML (X3D) file. 2nd process The VRML(X3D) file created in the first step is opened in a 3D-CG toolkit. The surface of the organic polymer depression and the organic polymer anchors are treated as rigid bodies (passive objects / collision shape mesh). Then, spheres with radii of 0.3 to 1.2 Å (filling spheres) are treated as rigid bodies (active objects / collision shape convex hulls) and placed on top of the organic polymer depressions. After dropping and filling the depressions, the top of the organic polymer depressions is blocked with several plates treated as rigid bodies (passive objects / collision shape meshes). Rigid body simulations are performed one to several times with the ratio of the convex hulls of the collision shapes of the filling spheres ranging from 10% to 100%, and the results are output as a VRML(X3D) file. 3rd process In the second step, the VRML(X3D) file created is opened in the 3D-CG toolkit. The radius of the packed spheres is slightly increased to 0.4-1.5 Å to remove spheres that protrude from the surface of the organic polymer depressions. Then, the radius of the packed spheres is further increased to 0.9-3.0 Å to remove spheres that protrude from the surface of the organic polymer depressions, except near the organic polymer anchors, to narrow down the packed spheres to be used for compound construction. The result is output as a VRML(X3D) file, and then the VRML(X3D) file is converted to a PDB file with the packed spheres as carbon atoms. 4th step In the third step, the PDB file created is loaded into a three-dimensional minimal spanning tree tool, and a carbon graph is constructed by connecting the centers of carbon atoms. Next, a ring structure is formed from the carbon graph, and the side chains of the ring structure are specified to construct the carbon skeleton structure. Finally, the carbon skeleton structure and organic polymer anchors are output as a PDB file. The formation of the ring structure can be done visually, or using genetic algorithms or reinforcement learning. 5th step In the fourth step, the PDB file created is opened with a molecular graphics tool. Carbon atoms that are within interacting distance of the organic polymer anchor are converted into interacting atoms (such as oxygen and nitrogen) and bonds (such as aromatic rings and double bonds) to construct the compound molecular structure. After reducing the strain, the stabilized organic polymer-compound complex structure is calculated using a docking tool.
[0011] (2) A program that causes an information processing device to perform a process for constructing a compound molecular structure on the surface of a depression in an organic polymer to which the compound is bound, and which causes the device to perform the following procedure. Step 1 Download the PDB file of the target organic polymer's three-dimensional structural model from the Protein Databank. Open the PDB file in a molecular graphics tool to display the three-dimensional structural model of the organic polymer. Then, specify the recessed surface of the compound that binds to the organic polymer and the organic polymer anchor (the main atom of the organic polymer that interacts with the compound) and output it as a VRML (X3D) file. Step 2 The VRML(X3D) file created in the first step is opened in the 3D-CG toolkit. The surface of the organic polymer depression and the organic polymer anchors are set as rigid bodies (passive objects / collision shape mesh). Then, spheres with radii of 0.3 to 1.2 Å (filling spheres) are set as rigid bodies (active objects / collision shape convex hulls) and placed on top of the organic polymer depressions. After dropping and filling the depressions, the top of the organic polymer depressions is blocked with several plates set as rigid bodies (passive objects / collision shape meshes). Rigid body simulations are performed one to several times with the ratio of the convex hulls of the collision shapes of the filling spheres set to 10-100%, and the results are output as a VRML(X3D) file. Step 3 In the second step, open the VRML(X3D) file created in the 3D-CG toolkit, slightly increase the radius of the packed spheres to 0.4-1.5 Å and remove the spheres that protrude from the surface of the organic polymer depressions, then further increase the radius of the packed spheres to 0.9-3.0 Å and remove any spheres that protrude from the surface of the organic polymer depressions except near the organic polymer anchors to narrow down the packed spheres to be used for compound construction, output the result as a VRML(X3D) file, and then convert the VRML(X3D) file to a PDB file with the packed spheres as carbon atoms. Step 4 The PDB file created in the third step is loaded into a three-dimensional minimal spanning tree tool, and a carbon graph is constructed by connecting the centers of carbon atoms. Then, a ring structure is formed from the carbon graph, and the side chains of the ring structure are specified to construct the carbon skeleton structure. Finally, the carbon skeleton structure and organic polymer anchors are output as a PDB file. The formation of the ring structure can be done visually, or using genetic algorithms or reinforcement learning. Step 5 In the fourth step, the PDB file created is opened with a molecular graphics tool. Carbon atoms that are within interacting distance of the organic polymer anchor are converted into interacting atoms (such as oxygen and nitrogen) or bonds (such as aromatic rings and double bonds) to construct the compound molecular structure. Strain reduction is then performed, and the stabilized organic polymer-compound complex structure is calculated using a docking tool. [Effects of the Invention]
[0012] Virtual screening (VS) has been used to efficiently search for new drug candidates, but significantly improving accuracy requires enormous computation time. In recent years, VS using supercomputers and artificial intelligence technologies that perform deep learning on vast amounts of data have also been utilized. Against the backdrop of such large-scale compound design methods, there is a need for a method that allows drug discovery professionals to design compounds simply. It is extremely beneficial for drug discovery professionals who are not computer science experts to design compounds using their own experience in order to generate new compound ideas. The present invention provides a method for constructing compound molecular structures that allows drug discovery professionals to construct compound molecular structures themselves using a personal computer.
Brief Description of Drawings
[0013] [Figure 1] It is a conceptual diagram of a method for constructing the molecular structure of a compound of the present invention. [Figure 2] It is a flowchart of a method for constructing the molecular structure of a compound of the present invention. [Figure 3] It is a schematic diagram for narrowing down the packing spheres of a method for constructing the molecular structure of a compound of the present invention. [Figure 4] It is a diagram showing the procedure for constructing the skeletal structure from the side chain of the influenza drug Relenza that binds to protein neuraminidase. [Figure 5] It is a diagram of the molecular surface of the protein within 5 Å around chain A of Relenza in PDB 3B7E, the protein anchor interacting with Relenza, and 176 ICO spheres stacked in a grid pattern at intervals of 1.6 Å vertically, horizontally, and diagonally above the depression. [Figure 6] It is a diagram in which the upper part of the depression on the molecular surface of the protein is covered with 8 planes for performing rigid body simulation of 119 falling-packed ICO spheres. [Figure 7] It is a diagram showing 48 spheres remaining after deleting the spheres protruding from the molecular surface of the protein by setting the radius of the packing spheres to 0.9 Å after rigid body simulation. [Figure 8] It is a diagram showing 30 spheres remaining after further deleting the spheres protruding from the molecular surface of the protein except near the protein anchor by setting the radius of the 48 spheres in the above figure to 2.1 Å. [Figure 9] It is a diagram showing the procedure for constructing a carbon graph by connecting 30 spheres with a minimum spanning tree, then visually reconstructing it into a 6-membered ring, and selecting the longest span for the 6-membered ring side chain to construct the carbon skeletal structure of the compound molecule. [Figure 10] It is a diagram showing the procedure for constructing the compound molecular structure by converting some carbon atoms of the carbon skeleton created in the above figure into atoms and bonds that interact with the protein anchor. [Figure 11]This diagram illustrates the procedure for constructing a carbon skeleton structure containing bicyclooctane using a genetic algorithm after connecting 37 spheres with a minimal spanning tree and then narrowing down the side chains. [Figure 12] This diagram illustrates the procedure for identifying the bicyclooctane structure using the multi-bandit problem, based on eight carbon spheres and nine CC bonds. [Figure 13] This figure shows the three-dimensional structural model of the complex of the PDB compound and the three-dimensional structural model of the complex of the compound constructed using the proposed method. [Figure 14] This diagram shows the procedure for constructing an N-butanoylglucoamine derivative that binds to the protein β-1,4-galactosyltransferase 1. [Figure 15] This diagram illustrates the procedure for constructing a carbon skeleton structure containing bicycloheptane by connecting 18 spheres using a minimal spanning tree and then utilizing a genetic algorithm. [Figure 16] This figure shows the procedure for constructing a benzamide derivative that binds to the bromodomain protein of the human p300 / CBP-associating factor. [Figure 17] This diagram shows the procedure for constructing a shikimic acid derivative that binds to the shikimic acid kinase protein of Mycobacterium tuberculosis. [Figure 18] This diagram shows 46 spheres with a radius of 1.8 Å (left) and 25 spheres with a radius of 2.1 Å (right) remaining in the cavities of Mycobacterium tuberculosis DNA gyrase and gyrase protein. [Figure 19] This figure shows a carbon graph created by connecting 25 spheres with a minimal spanning tree, and a carbon graph reconstructed from interactions with Mg2+ ions and DNA bases. [Figure 20] This figure shows the constructed compound structure (left) and the gatifloxacin GFN structure (right). [Figure 21] This diagram shows a carbon map created by connecting 26 spheres (left) and 20 spheres with a radius of 1.8 Å, which remained within the ribosomal RNA depressions of Staphylococcus aureus, using a minimal spanning tree. [Figure 22] This figure shows a carbon graph connecting 20 spheres using a minimal spanning tree (top left), a carbon graph connecting 23 spheres using a minimal spanning tree (right), and a carbon model of linezolid (LZD) (bottom left). [Figure 23] This figure shows the constructed 6-membered ring compound structure (top left), 5-membered ring compound structure (bottom left), and linezolid (LZD) compound structure (bottom right). [Modes for carrying out the invention]
[0014] Polymers are defined as "structures composed of many repetitions of units that are substantially or conceptually obtained from molecules with smaller molecular weights, using molecules with larger molecular weights." They are classified into organic polymers and inorganic polymers depending on the molecules that make them up. Organic polymers include natural polymers, which are products of nature, synthetic polymers, which are artificially synthesized, and semi-synthetic polymers, which are chemically derived from natural polymers. Therefore, organic polymers include polypeptides, proteins, DNA, RNA, lipids, polyamines, cellulose, amylose, starch, chitin, and the like. The explanation will be based on the conceptual diagram in Figure 1 and the flowchart in Figure 2. In the first step, the surface of the cavity where the compound binds is created in the three-dimensional structural model of the organic polymer. This can be done by estimating the binding sites and organic polymer anchors from a three-dimensional structural model of the organic polymer alone, but the following explanation will use a three-dimensional structural model of the organic polymer ligand complex. The three-dimensional structure of the organic polymer ligand complex is obtained from the Protein Databank (PDB, https: / / www.rcsb.org / ) and displayed using, for example, PyMOL™, an open-source molecular graphics tool from Schrodinger. After removing water molecules, metals, etc., and generating hydrogen atoms, the molecular surface of the organic polymer is created in 5 angstroms (Å, in the range of 2 to 12 Å) around the ligand. This molecular surface of the organic polymer, as well as the major atoms of the organic polymer interacting with the ligand (organic polymer anchors; atoms and functional groups involved in hydrogen bonding and charge interactions, and aromatic rings involved in π-π and CH-π interactions, etc.), are output as a VRLM or X3D file.
[0015] In the second step, for example, the VRLM or X3D file is opened in Blender®, an integrated 3D-CG toolkit developed as open source by the Blender Foundation, and a rigid body simulation is performed by dropping and filling ICO spheres (polyhedra made up of triangular faces) with a radius of 0.75 Å (Blender®'s scale is meters (m), but atoms are measured in angstroms (Å), so it is written in Å). This allows the phenomenon of spheres bouncing off the molecular surface of organic polymers or other spheres can be calculated.
[0016] To explain in more detail, the molecular surface of the organic polymer, as well as the organic polymer anchors that interact with the ligand, have their surfaces smoothed (adjacent faces are smoothly connected to appear curved), their rigid bodies are passive objects (they do not move themselves but influence the motion of other rigid bodies), and their collision shape is a mesh consisting only of triangular faces. The ICO sphere has its surface smoothed, its rigid body is an active object (it moves under the influence of gravity and collisions with other rigid bodies; Blender's mass scale is in kilograms (Kg), so it is considered a 1Kg rigid body), and its collision shape is a convex hull (a polyhedron covering the mesh). Then, for depressions on the molecular surface of the organic polymer, ICO spheres are stacked in a grid pattern on top of the depressions at intervals of 1.6 Å in all directions until the depressions are filled, and then the collision speed is -9.81 m / s 2 (-2~19.62m / sec 2The ICO spheres are dropped and filled within the specified range. Any ICO spheres that spill out of the depressions are removed. After dropping and filling, several planes are created to cover the top of the depressions on the molecular surface of the organic polymer. The surfaces of the planes are given a smooth shade, the rigid bodies are set to passive objects, and the collision shapes are set to meshes. Then, the collision shapes of the ICO spheres are randomly divided into convex hulls and meshes (with an appropriate convex hull ratio of 50 percent), and a rigid body simulation is performed. This is similar to the operation of filling a box with spheres and then shaking the box to fill the gaps, and it utilizes the phenomenon in Blender™ where rigid bodies that overlap will explode against each other. This simulation can be performed several times as needed.
[0017] In the third step, referring to the schematic in Figure 3, the radius of the packed spheres (ICO spheres) is changed to 0.9 Å to remove spheres that extend beyond the molecular surface of the organic polymer. Next, the radius of the packed spheres is changed to 1.8 Å or 2.1 Å to further remove spheres that extend beyond the molecular surface of the organic polymer, except near the organic polymer anchors, and then output as a VRLM or X3D file. The VRLM or X3D file then outputs the coordinates of the packed spheres as carbon atoms in a PDB file.
[0018] In the fourth step, the PDB file created in the third step is loaded using a three-dimensional minimum spanning tree (Prim's method) tool (for example, https: / / yatt.hatenablog.jp / entry / 20090525 / 1243179000) to construct a carbon graph. Note that the minimum spanning tree does not create a ring structure, so the ring structure is reconstructed by visually inspecting the carbon graph or by using a genetic algorithm or reinforcement learning.
[0019] In the fifth step, carbon atoms within interacting distance of the organic polymer anchor are converted into interacting atoms and bonds (atoms and functional groups involved in hydrogen bonds and charge interactions, as well as aromatic rings involved in π-π and CH-π interactions, and double bonds, etc.) to construct the compound molecular structure, and then strain is reduced. Subsequently, the stabilized organic polymer compound complex structure is calculated using, for example, AutoDock Vina, an open-source docking tool from Scripps Research Institute. By comparing the docking scores at this stage, an indication of bond strength can be obtained.
[0020] For a novel compound to become a potential drug, it is important to pre-determine that it is synthesizable (low difficulty of synthesis), possesses appropriate physical properties, has appropriate absorption / distribution / metabolism / excretion characteristics, and is non-toxic. For these purposes, tools such as RDkit, an open-source toolkit for cheminformatics, DruMAP (Kawashima, H. et al., DruMAP: A Novel Drug Metabolism and Pharmacokinetics Analysis Platform, Journal of Medicinal Chemistry, vol.66, no.14, pp.9697-9709, 2023.), a drug metabolism and pharmacokinetic analysis software, and Cardiotoxicity Database (Sato, T. et al., Construction of an integrated database for hERG blocking small molecules, PLoS ONE, vol.13, no.7, e0199348, 2018.), an hERG analysis software, can be used. [Examples]
[0021] Figure 4 shows a diagram illustrating the procedure for constructing the skeletal structure of the anti-influenza drug Relenza by retaining only the side chains that bind to the protein neuraminidase, as one embodiment of the present invention, and linking them together. (PDB 3B7E(Xu, X. et al., Structural Characterization of the 1918 Influenza Virus H1N1 Neuraminidase, Journal of Virology, vol.82, no.21, pp.10493-10501) The image shows the molecular surface of the protein 5 Å around the A-chain relenza of the 2008 protein, along with protein anchors (guazinino group of arginine 118, carboxylate and main-chain carbonyl group of aspartic acid 151, guazinino group of arginine 152, main-chain carbonyl group of tryptophan 178, carboxylate of glutamic acid 227, guazinino group of arginine 292, and guazinino group of arginine 371), and ligand anchors (relenza side chain; carboxylate at position 2, guazinino group at position 4, acetylamide group at position 5, and (1,2,3)-trihydroxypropyl group at position 6, containing some polar hydrogens).
[0022] The ligand anchor was made into an ICO sphere with a radius of 0.75 Å (0.35 Å per hydrogen) (rigid body as a passive object and collision shape as a mesh), and its surface was smoothed. Then, 176 packed spheres (ICO spheres) were stacked in a grid pattern with 1.6 Å spacing on all sides above, below, and above on the protein anchor and on the top of the protein depression. When they were then dropped and filled using normal gravity, 76 spheres remained in the depression. After dropping and filling, eight planes were created to cover the top of the depression on the molecular surface of the protein. The collision shapes of the packed spheres were then randomly divided into convex hulls and meshes (convex hull ratio 50 percent), and a rigid body simulation was performed once. In the final 250 frames (24 frames per second), the radius of the packed spheres was made 2.3 Å (larger because it is the internal skeletal structure that connects the relenza side chains), packed spheres that protruded from the surface of the protein depression were removed, and 6 spheres that could bind to the ligand anchor were connected from the remaining 12 spheres. As a result, the six-membered carbon skeleton of Relenza was reproduced with the correct stereochemistry (white-filled triangles indicate side chains pointing over the ring, and open-circle triangles indicate side chains pointing below the ring). Incidentally, performing one rigid-body simulation resulted in a cleaner construction of the six-membered ring compared to the case of drop-filling alone. [Examples]
[0023] As one embodiment of the present invention, a derivative of the anti-influenza drug Relenza that binds to the protein neuraminidase was constructed. Figure 5 shows the molecular surface of the protein and protein anchors 5 Å around the A-chain Relenza of PDB 3B7E, and 176 packing spheres (ICO spheres) stacked in a grid pattern at 1.6 Å intervals on all sides above and below the protein depression. When the ICO spheres were dropped and filled using normal gravity, 119 spheres remained in the depression. After dropping and filling, eight planes were created to cover the top of the depression on the molecular surface of the protein (Figure 6). The collision shape of the packing spheres was then randomly divided into convex hulls and mesh (convex hull ratio 50 percent), and a rigid-body simulation was performed once. In the final 250 frames, when the radius of the packing spheres was set to 0.9 Å and the packing spheres that protruded from the surface of the protein depression were removed, 48 spheres remained in the depression (Figure 7).
[0024] Next, the radius of the packed spheres was changed to 2.1 Å, and spheres that protruded from the protein molecular surface, except near the protein anchor, were further removed, leaving 30 spheres in the depressions (Figure 8). These 30 spheres were output as an X3D file, then converted to a PDB file with the packed spheres as carbon atoms, and a carbon graph was constructed using a three-dimensional minimal spanning tree. Then, one CC bond was uniquely added by visual inspection to construct a six-membered ring, and then the side chains at each position of the six-membered ring were removed, leaving only the longest span. As a result, the carbon skeleton structure of the stereochemistry of Relenza (filled white triangles indicate side chains pointing over the ring, and open white triangles indicate side chains pointing below the ring) was constructed (Figure 9).
[0025] The constructed carbon skeleton structure is surrounded by a positively charged (δ+) region composed of arginine 118, 292, and 371, and a negatively charged (δ-) region composed of aspartic acid 151, the carbonyl group of tryptophan 178, and glutamic acid 227. For the substructure of the compound molecule, negatively charged carboxylates are preferred in the former region, while positively charged amino groups and guanidino groups are preferred in the latter region. The compound molecular structure was constructed by converting the substructure of the carbon skeleton to these functional groups (Figure 10). While there may be cases where no carbon spheres are available to connect when converting to carboxylates or guanidino groups, since the packing spheres near the protein anchors have already been removed, there is no particular problem in generating and connecting new carbon spheres. Furthermore, since the distance between carbon atoms is approximately 1.5 Å, intramolecular steric hindrance (dotted line in the structural formula of Figure 10) may occur during reconstruction, and it may be necessary to remove that portion. As a result, a compound similar to Relenza was constructed from the smallest spanning tree with 30 spheres (radius 2.1 Å), and a compound similar to Tamiflu was constructed from the smallest spanning tree with 37 spheres (radius 1.8 Å).
[0026] From a minimal spanning tree of 37 spheres (radius 1.8 Å), visually adding one CC bond constructs a 7-membered ring, and adding another constructs a bicyclononane consisting of two 7-membered rings and one 6-membered ring. Adding yet another constructs a carbon skeleton structure consisting of bicyclooctane (three 6-membered rings) and one 3-membered ring. Cleavaging this 3-membered ring constructs bicyclooctane. Alternatively, it is possible to construct a carbon graph with the longest span of the side chain narrowed down prior to ring formation, select carbon atoms that have the potential to form a ring structure, and then use a genetic algorithm (e.g., https: / / qiita.com / spwimdar / items / e46487e25f35953c7cfe) to search for a path from their coordinates and form a ring. However, the genetic algorithm used in this study does not close the path, and there are cases where the start and goal of the carbon spheres are far apart, so caution is required when closing the path to form a ring. Specifically, the three-dimensional coordinates of the carbon spheres enclosed by the dotted line in the upper left of Figure 11 were input. The gene length was set to 8 carbon spheres, 20 gene individuals, 20,000 evolutionary generations, the number of elites was increased to 10 to enhance fitness, and the mutation (translocation) probability was increased to 0.3 to ensure diversity. A pathway with a short sum of CC bonds was then searched for. As a result, 7 CC bonds were selected for 8 carbon spheres (211 generations, distance 10.25 Å), and connecting the start and goal of the carbon spheres formed an 8-membered ring. However, an 8-membered ring is undesirable for compound synthesis, so two CC bonds were added to form bicyclooctane consisting of three 6-membered rings. While it is easy to form a 6-membered ring structure without using a genetic algorithm, using a genetic algorithm allowed for diversity in the ring structure.
[0027] In addition to the genetic algorithm, the multi-armed bandit problem (e.g., https: / / qiita.com / tsugar / items / b809f8d6399cc988aa69), a reinforcement learning problem, was used to identify ring structures by rewarding bonds involved in ring structure formation. The genetic algorithm found that, in addition to the eight carbon spheres and six CC bonds enclosed by the dotted line in the upper left of Figure 12, three CC bonds were involved in the formation of six-membered rings. Therefore, the problem was set up as a Bernoulli trial where acquiring an edge resulted in a reward of 1, and failing to acquire an edge resulted in a reward of 0. The success probability p for the six bonds that already had bonds was set to 0.5 each, and the p for the three bonds where six-membered ring formation was estimated was set to 0.6 each. The ring CC bonds were then identified by sequentially finding the bandits that acquired rewards. The first to acquire a reward was CC bond number 6, followed by CC bond number 8, and one six-membered ring was identified. Furthermore, CC bond number 2 acquired a reward, and bicyclooctane consisting of three six-membered rings was identified.
[0028] The constructed compound molecular structure was subjected to strain reduction, and then the stabilized protein-compound complex structure was calculated using the docking tool AutoDock Vina. The calculation conditions were as follows: the center coordinates X, Y, and Z of the receptor search volume were specified, with size: X, Y, Z = 20, number of binding modes: 9, search coverage: 8, and maximum energy difference (Kcal / mol): 3. Figure 13 shows the three-dimensional structural models of the protein-ligand complexes stabilized in AutoDock Vina (3B7E and 2HU4 in PDB) for Relenza and Tamiflu, as well as the stabilized complex models of the construct compounds. The construct compounds are shown with a dihydropyran ring (Relenza) instead of cyclohexane or cyclohexene rings (Tamiflu). The binding energy dG = -8.028 kcal / mol for the construct compound from 37 spheres (1.8 Å) and -7.434 kcal / mol for the bicyclooctane compound with the reconstituted ring, and dG = -7.649 kcal / mol for the construct compound from 30 spheres (2.1 Å), are equivalent to or better than Relenza's dG = -7.943 kcal / mol and Tamiflu's -6.431 kcal / mol, suggesting that the binding strength is equivalent to that of Relenza and Tamiflu. Furthermore, the physical properties and synthesis difficulty calculated using RDkit, the drug metabolism and pharmacokinetics calculated using DruMAP, and the hERG toxicity calculated using the Cardiotoxicity Database were estimated to be equivalent to those of Relenza. [Examples]
[0029] Examples 3 to 5 describe the construction of compound molecular structures targeting three proteins (Non-Patent Document 4) evaluated using a state-of-the-art artificial intelligence conditional variational autoencoder. Figure 14 shows the procedure for constructing an N-butanoylglucoamine derivative that binds to the protein β-1,4-galactosyltransferase 1. In PDB's 1NWG (Ramakrishnan, B. et. al., α-Lactalbumin (LA) Stimulates Milk β-1,4-Galactosyltransferase I(β4Gal-T1) to Transfer Glucose from UDP-glucose to N-Acetylglucosamine, The Journal of Biological Chemistry, vol.276, no.40, pp.37665-37671, 2001), 122 packed spheres (ICO spheres) were stacked in a lattice pattern at 1.6 Å intervals in all directions on the 5 Å molecular surface of the protein surrounding N-butanoylglucoamine in the A and B chains, along with protein anchors (main chain nitrogen of glycine 316 in the B chain, carboxylates of aspartic acid 318 and 319, and the guanidino group of arginine 359) and above the protein depressions. Then, when ICO spheres were dropped and filled into the depressions under normal gravity, 87 spheres remained in the depressions. After dropping and filling, 10 planes were created to cover the top of the depressions on the molecular surface of the protein. Then, the collision shape of the filled spheres was randomly divided into convex hulls and meshes, and rigid body simulations were performed three times (convex hull ratios of 50%, 100%, and 50 percent).
[0030] In the final 250 frames, when the radius of the packing spheres was set to 0.9 Å and the packing spheres that protruded from the surface of the protein depression were removed, 31 spheres remained within the depression. Next, the radius of the packing spheres was changed to 2.1 Å, and when spheres that protruded from the molecular surface of the protein, except near the protein anchor, were further removed, 18 spheres remained within the depression. A PDB file was created for these 18 spheres, and a carbon graph was constructed using a three-dimensional minimal spanning tree. Then, four CC bonds were added visually to construct a six-membered ring, and all bonds except the longest side chain were removed. The carbon graphs in the first and second rigid body simulations were such that the six-membered ring structure was easily identifiable, but the carbon graph in the third simulation was more appropriate for side chain construction, so the six-membered ring structure common to all structures was used. As a result, we were able to construct a carbon skeleton structure of N-butanoylglucoamine that maintained the stereochemical configuration of the region that interacts with the protein anchor (filled white triangles indicate that the side chains are facing above the ring, and open white triangles indicate that the side chains are facing below the ring). The partial structure of the carbon skeleton was converted to oxygen and nitrogen atoms that can interact with the protein anchor to construct the compound molecular structure.
[0031] The reconstruction of the ring structure can also be done by selecting carbon atoms that may form a ring structure from the carbon graph and using a genetic algorithm to find a path from their coordinates. As shown in Figure 15, a carbon skeleton structure with bicycloheptane consisting of two five-membered rings and one six-membered ring was constructed. [Examples]
[0032] Figure 16 shows the procedure for constructing a benzamide derivative that binds to the bromodomain protein of the human p300 / CBP-associating factor. 109 packed spheres (ICO spheres) were stacked in a lattice pattern at 1.6 Å intervals on the top of the protein's depression, around the 5 Å molecular surface of the A-chain benzamide of PDB (Navratilova, I. et al., Discovery of New Bromodomain Scaffolds by Biosensor Fragment Screening, ACS Medicinal Chemistry Letters, vol.7, no.12, pp.1213-1218, 2016.), along with protein anchors (amide group of asparagine 803, phenol of tyrosine 809), and above the protein's depression. When the ICO spheres were dropped and filled using normal gravity, all 109 spheres remained in the depression. After the dropping and filling, eight planes were created to cover the top of the depression on the protein's molecular surface. Then, the collision shape of the packed spheres was randomly divided into a convex hull and a mesh, and rigid-body simulations were performed three times.
[0033] In the final 250 frames, when the radius of the packing spheres was set to 0.9 Å and the packing spheres that protruded from the surface of the protein depression were removed, 38 spheres remained within the depression. Next, the radius of the packing spheres was changed to 1.8 Å, and spheres that protruded from the molecular surface of the protein, except near the protein anchor, were further removed, leaving 26 spheres within the depression. A PDB file was created for these 26 spheres, and a carbon graph was constructed using a three-dimensional minimal spanning tree. Then, two unique CC bonds were added visually to construct two 6-membered rings (a planar ring and a thick ring), and the longest side chain was rearranged based on its interaction with the protein anchor, while the other side chains were removed. As a result, the carbon skeleton structure of a benzamide derivative that interacts with the protein anchor was constructed. The partial structure of the carbon skeleton was then converted to oxygen, nitrogen, and aromatic rings that can interact with the protein anchor to construct the compound molecular structure. Incidentally, the distance between carbon atoms is approximately 1.5 Å, and steric hindrance occurred when the side chains were rearranged (triangle-bordered area in the upper right of Figure 16), so it is also necessary to remove that part. [Examples]
[0034] Figure 17 shows the procedure for constructing a shikimic acid derivative that binds to the Mycobacterium tuberculosis shikimate kinase protein. 120 filling spheres (ICO spheres) were stacked in a lattice pattern at 1.6 Å intervals on the top of the depressions in the protein's molecular surface and protein anchors (aspartic acid 34 carboxylate, arginine 58 guazinyl group, glycine 81 main chain nitrogen, arginine 136 guazinyl group) within 5 Å of the shikimic acid of PDB's 1ZYU (Gan, J. et al., Crystal Structure of Mycobacterium tuberculosis Shikimate Kinase in Complex with Shikimic Acid and an ATP Analogue, Biochemistry, vol.45, no.28, pp.8539-8545, 2006.), with 120 filling spheres at 1.6 Å intervals on all sides. When the ICO spheres were dropped and filled using normal gravity, 71 spheres remained in the depressions. After the dropping and filling, four planes were created to cover the top of the depressions on the molecular surface of the protein. Then, the collision shape of the packed spheres was randomly divided into a convex hull and a mesh, and rigid-body simulations were performed three times.
[0035] In the final 250 frames, when the radius of the filling spheres was set to 0.9 Å and the filling spheres that protruded from the surface of the protein depression were removed, 23 spheres remained within the depression. Next, when the radius of the filling spheres was changed to 2.1 Å and the spheres that protruded from the molecular surface of the protein, except near the protein anchor, were further removed, 12 spheres remained within the depression. A PDB file was created for the 12 spheres, and a carbon graph was constructed using a three-dimensional minimal spanning tree. Then, one CC bond was uniquely added by visual inspection to construct a 6-membered ring, and the side chains were rearranged based on their interaction with the protein anchor, allowing the carbon skeleton structure of the shikimic acid derivative interacting with the protein anchor (filled triangles indicate the side chains are facing above the ring, and open triangles indicate the side chains are facing below the ring) to be constructed. Finally, the partial structure of the carbon skeleton was converted to carboxylates and oxygen that can interact with the protein anchor to construct the compound molecular structure (3,4,5,6-tetrahydroxy compound). [Examples]
[0036] Examples 1 to 5 targeted proteins, while Example 6 describes examples of constructing compound molecular structures targeting DNA (and proteins), and Example 7 describes examples targeting RNA. In the PDB's 5BTD (Blower, Tim R. et al., Crystal structure and stability of gyrase–fluoroquinolone cleaved complexes from Mycobacterium tuberculosis, PNAS, vol.113, no.7, pp.1706-1713, 2016.), 8 Å of DNA surrounding the E-chain gatifloxacin GFN, the molecular surface of the Mycobacterium tuberculosis DNA gyrase protein, Mg2+ ions chelated with the GFN, 128 arginine residues of the C-chain protein with hydrogen bonds, 14 cytosine and 15 adenine bases of the stacked F-chain DNA, 10 thymine and 11 guanine bases of the E-chain, and 222 filling spheres (ICO spheres) were stacked on top of the DNA and protein cavities at 1.6 Å intervals in all directions. When the ICO spheres were dropped and filled into the cavities under normal gravity, 196 spheres remained in the cavities. After the drop-filling process, five planes were created to cover the tops of the depressions on the molecular surfaces of the DNA and proteins. Then, the collision shape of the packed spheres was randomly divided into convex hulls and meshes (convex hull ratio 50%), and a rigid-body simulation was performed once.
[0037] In the final 250 frames, when the radius of the packed spheres was set to 0.9 Å and the packed spheres that protruded from the surface of the depression were removed, 83 spheres remained within the depression. Subsequently, when the radius of the packed spheres was set to 1.8 Å and then 2.1 Å and the spheres that protruded from the curved surface except around the anchor were removed, and then the spheres that overlapped in two layers were removed, 25 spheres remained within the depression (Figure 18). A carbon graph was then constructed using a three-dimensional minimal spanning tree. Visually inspecting the graph of the minimal spanning tree, two CC bonds were uniquely added to construct the fused ring decalin, which can then be converted to a carbonyl compound with chelation ability to Mg2+ ions and a carboxylic acid or hydroxypyrazole. Furthermore, the 10-membered ring can be converted to 1,4-dihydro-4-oxoquinoline or 4-oxo-4H-quinolidine as an aromatic ring to stack with 10 thymine bases and 11 guanine bases of the E chain and 15 adenine bases of the F chain of DNA (Figure 19). The carboxylic acid has a negatively charged ionic structure, and by introducing a terminal amino group into a region without interactions, it was transformed into an amphoteric structure to ensure good tissue distribution. This structure was similar to that of gatifloxacin GFN (Figure 20). [Examples]
[0038] In the PDB's 4WFA (Eyal, Z. et al., Structural insights into species-specific features of the ribosome from the pathogen Staphylococcus aureus, PNAS, vol.112, no.43, pp.E5805-E5814, 2015), 175 packing spheres (ICO spheres) were stacked on the surface of Staphylococcus aureus ribosomal RNA 9 Å around the linezolid ZLD, along with 2478 adenine bases, 2532 ribose 2' hydroxyl group, and 2612 uracil bases that are hydrogen-bonded to the ZLD, as well as 2478 stacking adenine bases and 2479 CH-π interacting cytosine bases. 175 packing spheres were then stacked at 1.6 Å intervals in all directions above the RNA cavity. When the ICO spheres were dropped and filled using normal gravity, all 175 spheres remained within the cavity. After the drop-filling process, the top of the depressions on the RNA molecular surface was covered with a single plane. Then, the collision shape of the packed spheres was randomly divided into convex hulls and meshes (convex hull ratio 50%), and a rigid-body simulation was performed once.
[0039] In the final 250 frames, when the radius of the packed spheres was set to 0.9 Å and the packed spheres that protruded from the surface of the depression were removed, 52 spheres remained within the depression. Next, when the radius of the packed spheres was set to 1.8 Å and the spheres that protruded from the curved surface, excluding those around the RNA anchor, were removed, 26 spheres remained. Then, when the spheres overlapping in two layers were removed, 20 spheres remained within the depression (Figure 21). However, since stacking and hydrogen bond formation were possible, a carbon graph was constructed using a three-dimensional minimal spanning tree with the 3 removed spheres, and the ring structure was reconstructed to form a 6-membered ring. On the other hand, since there was no difference in CC bond length between 1.6 Å and 1.7 Å, a 5-membered ring was formed instead of a 6-membered ring, resulting in a structure similar to the carbon model of the X-ray crystal structure LZD (Figure 22). Finally, 6-membered ring compounds and 5-membered ring compounds were constructed by converting carbon to oxygen, nitrogen, and aromatic rings based on interactions with the RNA anchor. It was confirmed that the five-membered ring compound has the same stereochemistry as LZD in its X-ray crystal structure (filled white triangles indicate that the side chains are facing the top of the ring, and open white triangles indicate that the side chains are facing the bottom of the ring) (Figure 23).
[0040] The method for constructing the molecular structure of a compound according to the present invention is not limited to the target proteins, DNA, and RNA shown in Examples 1 to 7, but can be applied to organic polymers including polypeptides, lipids, polyamines, cellulose, amylose, starch, chitin, and the like. [Industrial applicability]
[0041] This invention relates to a method for constructing compound molecular structures that can be used in the molecular design of biologically active compounds, and also to a program for carrying out this method. This invention is useful for drug discovery professionals in generating ideas for novel compounds and can be used to construct compound molecular structures themselves.
Claims
1. A method for constructing a compound molecular structure on the surface of a cavity in an organic polymer to which a compound is bonded, by having a computer perform the following steps 1 through 5, The first step involves downloading a PDB file of the three-dimensional structural model of the target organic polymer from the Protein Databank, opening the downloaded PDB file using a molecular graphics tool to display the three-dimensional structural model of the organic polymer, specifying the recessed surface of the organic polymer to which the compound is bound and the organic polymer anchor, and outputting it as a VRML or X3D file. The second step involves opening the VRML or X3D file created in the first step using a 3D-CG toolkit, creating a mesh of collision shapes for the organic polymer depression surface and organic polymer anchors using rigidbody passive objects, arranging filling spheres with radii of 0.3 to 1.2 Å as rigidbody active objects with convex hulls as collision shapes on top of the organic polymer depressions, and then dropping and filling them. The top of the organic polymer depressions is then blocked with one to several plates whose collision shapes are meshed using rigidbody passive objects, and rigidbody simulations are performed one to several times with a ratio of convex hulls of the filling spheres' collision shapes ranging from 10% to 100%. As a result of the rigidbody simulations, the organic polymer depression surface, filling spheres, and organic polymer anchors are output as VRML or X3D files. The third step involves opening the VRML or X3D file created in the second step with a 3D-CG toolkit, slightly increasing the radius of the packed spheres to 0.4–1.5 Å to remove spheres that protrude from the surface of the organic polymer depressions, further increasing the radius of the packed spheres to 0.9–3.0 Å to remove spheres that protrude from the surface of the organic polymer depressions except near the organic polymer anchors, thereby narrowing down the packed spheres to be used for compound construction, outputting the narrowed-down packed spheres together with the organic polymer anchors as a VRML or X3D file, and then converting the packed spheres in the output file into a PDB file with the organic polymer anchors as carbon atoms. The fourth step involves reading only the carbon atoms that were packed spheres from the PDB file created in the third step using a three-dimensional minimal spanning tree tool, constructing a carbon graph by connecting the centers of the carbon atoms, and then forming a ring structure using a genetic algorithm that searches for a path from the coordinates of the carbon atoms selected from the constructed carbon graph, or by using reinforcement learning that selects bonds involved in the formation of a ring structure and earns a reward. After constructing the carbon skeleton structure by specifying the longest spanned side chain of the ring structure, the carbon skeleton structure and organic polymer anchors are output as a PDB file. Step 5 involves opening the PDB file created in Step 4 with a molecular graphics tool, converting carbon atoms that are within interacting distance of the organic polymer anchor into interacting atoms or bonds to construct the compound molecular structure, then reducing the strain, and finally calculating the stabilized organic polymer / compound complex structure with a docking tool. method.
2. A program that causes an information processing device to perform a process to construct a compound molecular structure on the surface of a depression in an organic polymer to which the compound is bound, by executing the following steps 1 through 5, The first step is to download a PDB file of the three-dimensional structural model of the target organic polymer from the Protein Databank, open the downloaded PDB file using a molecular graphics tool to display the three-dimensional structural model of the organic polymer, specify the recessed surface of the organic polymer to which the compound is bound and the organic polymer anchor, and output it as a VRML or X3D file. The second step involves opening the VRML or X3D file created in the first step using a 3D-CG toolkit, creating a mesh of collision shapes for the organic polymer depression surface and organic polymer anchors using rigidbody passive objects, arranging filling spheres with radii of 0.3 to 1.2 Å on top of the organic polymer depression using rigidbody active objects with convex hulls as collision shapes, and then dropping and filling them. The top of the organic polymer depression is then blocked with one to several plates with mesh collision shapes using rigidbody passive objects, and rigidbody simulations are performed one to several times with a ratio of convex hulls for the collision shapes of the filling spheres ranging from 10% to 100%. As a result of the rigidbody simulations, the organic polymer depression surface, filling spheres, and organic polymer anchors are output as VRML or X3D files. The third step involves opening the VRML or X3D file created in the second step using a 3D-CG toolkit, slightly increasing the radius of the packed spheres to 0.4–1.5 Å to remove spheres that protrude from the surface of the organic polymer depressions, further increasing the radius of the packed spheres to 0.9–3.0 Å to remove spheres that protrude from the surface of the organic polymer depressions except near the organic polymer anchors, thereby narrowing down the packed spheres to be used for compound construction, outputting the narrowed-down packed spheres together with the organic polymer anchors as a VRML or X3D file, and then converting the output file's packed spheres as carbon atoms together with the organic polymer anchors into a PDB file. The fourth step involves reading only the carbon atoms that were packed spheres from the PDB file created in the third step using a three-dimensional minimal spanning tree tool, constructing a carbon graph by connecting the centers of the carbon atoms, and then forming a ring structure using a genetic algorithm that searches for a path from the coordinates of the carbon atoms selected from the constructed carbon graph, or by using reinforcement learning that selects bonds involved in the formation of a ring structure to obtain a reward. After constructing the carbon skeleton structure by specifying the longest spanned side chain of the ring structure, the carbon skeleton structure and organic polymer anchors are output as a PDB file. The fifth step involves opening the PDB file created in the fourth step using a molecular graphics tool, converting carbon atoms that are within interacting distance of the organic polymer anchor into interacting atoms or bonds to construct the compound molecular structure, then reducing the strain, and finally calculating the stabilized organic polymer / compound complex structure using a docking tool. program.
Citation Information
Patent Citations
Systems and Methods for Structure-Based Drug Design Including Accurate Prediction of Binding Free Energy
JP2002533477A
Calculation of the induced fit effect
JP2021500661A
Compound searching device, compound searching method, and compound searching program
WO2024038700A1
Method of searching the structure of stable biopolymer-ligand molecule composite
WO1993020525A1