Method for generating compound derivatives for target protein to construct artificial intelligence-based new drug platform
By generating derivatives of effective compounds within a target protein's defined space, the method addresses the inefficiencies of traditional drug development and computational screening, enhancing binding affinity and accelerating drug discovery.
Patent Information
- Application Number
- PCT/KR2024/014763
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2024-09-27
- Publication Date
- 2025-09-25
AI Technical Summary
Traditional drug development processes are lengthy, costly, and have low success rates, while existing computational screening methods struggle with accuracy and scalability, particularly in predicting protein-ligand binding affinities for large-scale analysis.
A method for generating derivatives of effective compounds by selecting a substitutable portion, setting a target space within a target protein, and substituting with a suitable moiety to improve binding affinity, implemented in a computational system.
The method enhances the potential for derivatives to bind effectively to target proteins, improving binding affinity and activity, and optimizes molecular dynamics simulations for faster drug candidate discovery.
Smart Images

Figure KR2024014763_25092025_PF_FP_ABST
Abstract
Description
A method for generating compound derivatives of target proteins for building an artificial intelligence new drug platform.
[0001] The present disclosure relates to a technology that can be utilized in computer-aided drug discovery (CADD) for new drug development or in silico prescreening applied to the process of discovering new drug active substances for building an AI drug platform, and relates to a method for generating various forms of derivatives from existing effective compounds (Hit-compounds) for a specific target protein, an in silico derivative generated by the method, and an AI drug platform constructed by the method for generating a derivative of the present disclosure.
[0002] Traditional drug development involves identifying the target mechanism or protein for a disease, identifying and screening candidate compounds that act on it, and then proceeding through optimization, preclinical / toxicity testing, and clinical trials. This traditional drug development process takes an average of five years just to identify a candidate compound, and an additional two years to select a compound for clinical trials. This process incurs enormous R&D / clinical trial costs and a development period of an average of 15 years. However, only about 10% of these compounds ultimately receive final approval from regulatory authorities (such as the U.S. FDA).
[0003] Recently, computational analysis technologies (such as AI) are being actively utilized in the new drug development process to reduce the time and cost required in the existing new drug development process and increase the final success rate.
[0004] In the new drug development process, AI technologies can be utilized for a variety of purposes, including rapid data analysis, the search for optimal new drug candidates, and the analysis of suitable clinical subjects. In particular, various analytical tools are being used in the field of computational screening, which involves the exploration and analysis of new drug candidates. Representative analytical tools are listed in Table 1 below.
[0005] Program NameTypePrincipleAnalysis Time Per PoseAutoDock3D dockingGrid based semi-empherical scoring function (considering van der Waals forces, electrostatic interactions, and desolvation)Minutes to hoursGlide3D dockingEmpirical scoring function (GlideScore) (considering Coulomb, van der Waals, and solvation effects)0.2-2.4 minutesGOLD (Genetic Optimization for Ligand Docking)3D dockingChemScore based scoring function (multiple binding modes)Minutes to hoursFEP+Molecular DynamicsFree energy perturbation (depending on system size and settings)hours to daysPLUMED / GROMACSMolecular DynamicsMeta dynamicshours to daysGaussianQM / MMQuantum mechanics / molecular dynamics (depending on QM domain size)hours to daysAMBERMolecular DynamicsAlchemical Free Energy (e.g., thermodynamic integration)hours to days UnitsDeepChemDeep LearningNeural Networks for Molecular SystemsSeconds to minutes(once trained)gninaDeep Learning3D convolutional neural network based affinity prediction2.5 minutes(once trained)PhaseDiscovery StudioMOEPharmacophore modelingGenerate pharmacophore models from existing active compounds or protein structuresMinutes to hoursLigandScoutPharmacophore ensemble approachGenerate pharmacophore ensembles from multiple active ligands or protein conformationsTime unitsROCS3D Alignment2D FingerprintRapid overlay of chemical structures using shape and chemical features for force field virtual screeningMinutes to hours
[0006] Meanwhile, the process of virtually screening new drug candidates from a group of compounds is divided into (1) a pre-screening stage based on chemical properties or structural similarity, and (2) a deep screening stage that utilizes information on protein-ligand interactions through 3D docking. The pre-screening stage is generally used to reduce the number of candidate compounds from a large group of compounds, and thus uses simple, number-based discriminatory algorithms such as Lipinski's rule of 5 (a guideline for drug design). Because the information available for screening is limited, the reliability is typically maintained at around a 10% screening rate.
[0007] However, as computing technology has advanced, the demand for methods to compute billions of analysis targets has increased, creating a need to improve existing screening methods.
[0008] To this end, an advanced strategy for screening methods based on similarity between compounds was studied, such as comparing and analyzing the similarity of features including chemical properties between compounds or characterizing the two-dimensional pattern of molecular structures. However, the results of screening according to the strategy showed that the accuracy (T / P, true / positive) ratio was not high and there was a limitation that the structure of all substances could not be patterned.
[0009] Meanwhile, in the case of 3D and 4D-based protein-ligand binding affinity prediction technologies, which are known to have relatively high screening reliability, there was a problem that it was difficult to apply them to screening large-scale analysis targets because the analysis time required ranged from several minutes to several hours per structure.
[0010] In relation to,
[0011] (Patent Document 1) Republic of Korea Patent No. 10-2496208
[0012] (Patent Document 2) Republic of Korea Patent No. 10-0984735
[0013] (Patent Document 3) Republic of Korea Patent No. 10-2181058
[0014] (Non-patent Document 1) Garrett Morris et al. (2009) AutoDock4 and AutoDockTools4: Automated docking with selective receptor flexibility. Journal of Computational Chemistry 30(16): 2785-2791
[0015] (Non-patent Document 2) Richard Friesner et al. (2004) Glide: A New Approach for Rapid, Accurate Docking and Scoring. 1.Method. Journal of Medicinal Chemistry 47: 1739-1749
[0016] (Non-patent Document 3) G. Jones et al. (1995) Molecular recognition of receptor sites using a genetic algorithm with a description of desolvation. Journal of Molecular Biology, 245(1): 43-53
[0017] (Non-patent Document 4) Daniel Cappel et al. (2016) Relative Binding Free Energy Calculations Applied to Protein Homology Models. Journal of Chemical Information and Modeling 56(12): 2388-2400.
[0018] (Non-patent Document 5) Alessandro Laio et al. (2002) Escaping free-energy minima. Proceedings of the National Academy of Sciences 99 (20) 12562-12566
[0019] There are patent and non-patent literature.
[0020] The present disclosure was made to solve the above problems, and provides a method for generating derivatives capable of improving existing hit compounds whose efficacy has been confirmed or derived in terms of pharmacological efficacy or pharmacodynamics, and generating various derivatives for learning analysis AI algorithms.
[0021] This is accomplished by establishing a target space within the target protein centered around the region where the target protein and the effective compound bind, and generating derivatives of the effective compound that can be accommodated within that space. Furthermore, this method can be implemented in a computational system.
[0022] Through the derivative generation method of the present disclosure, it was confirmed that the generated derivative has an improved possibility of actually binding to a target protein, despite being an in silico derivative generation method, and that a derivative having improved binding affinity to a target protein or an improved effect on the activity of a target protein compared to an effective compound can be generated, thereby completing the present disclosure.
[0023] 1. One embodiment of the present disclosure is:
[0024] (A) A step of selecting a scaffold excluding a substitutable portion and a substitutable portion from the chemical structure of an effective compound;
[0025] (B) a step of setting a target space within a target protein centered on a region where a substitution target portion of a selected effective compound binds;
[0026] (C) a step of generating a derivative by selecting a substituent that replaces a substitutable portion of an effective compound acceptable within the target space of the set target protein;
[0027] It may relate to a method for producing a derivative of an effective compound from an effective compound for a target protein.
[0028] 2. In one embodiment, the method can be performed in an operating system.
[0029] 3. In one embodiment, the method may be for building an artificial intelligence new drug platform.
[0030] 4. In one embodiment, the method may be such that in step (A), the portion to be substituted in the chemical structure of the effective compound is selected based on an indicator that the binding interaction is lower than that of other chemical structures from the interaction profile of the effective compound and the target protein.
[0031] 5. In one embodiment, the binding interaction may be calculated by cleaving one bond among the bonds in the effective compound that interacts with the target protein, selecting a substitution candidate moiety that includes an atom of the cleaved portion for each cleaved bond, and calculating the average value of the binding energy of each atom constituting the substitution candidate moiety with the target protein.
[0032] 6. In one embodiment, the substitution candidate moiety may contain 1 to 12 atoms.
[0033] 7. In one embodiment, the method may further include a step of filtering the substitution candidate portions based on the number of constituent atoms.
[0034] 8. In one embodiment, the method may be such that the size of the target space within the target protein in step (B) is set to accommodate a substitutable portion of an effective compound selected from among chemical structures of effective compounds that bind to the target protein.
[0035] 9. In one embodiment, the method may be configured to include a step of stepwise dividing a region of the target protein constituting a target space within the target protein in accordance with the interaction energy between target protein atoms present in the region in step (B).
[0036] 10. In one embodiment, the interaction energy between target protein atoms for distinguishing regions of the target protein can be divided into three to five stages.
[0037] 11. In one embodiment, the method may include a step of clustering regions in which the interaction energy between target protein atoms is relatively low, and further selecting a clustering region of the target space to be used for generating a derivative of the effective compound based on the proximity to the scaffold of the effective compound and the size of the clustering region.
[0038] 12. In one embodiment, in step (B) of the method, the distinction of the interaction energy between target protein atoms may be performed by creating a region of the target protein constituting the target space as a spatial filter, setting dots arranged at equal intervals in the spatial filter, and distinguishing the dots by the interaction energy between the target protein atoms.
[0039] 13. In one embodiment, the spatial filter may be a spherical, rectangular, cylindrical, or amorphous filter.
[0040] 14. In one embodiment, the selection of a clustering region of the target space to be used for generating a derivative of the effective compound in step (B) is
[0041] (B1) A step of creating a region of a target protein constituting a target space in the form of a cylinder filter, setting up marking points (dots) arranged at equal intervals on the cylinder filter, and dividing the marking points step by step according to the interaction energy between target protein atoms.
[0042] (B2) It may include a step of approaching the scaffold of the effective compound to a cylinder filter, excluding from the cylinder filter a region where the interaction energy between target protein atoms is greater than a preset value, and then clustering the indicated points into spatial units.
[0043] 15. In one embodiment, the selection of a clustering region of a target space to be used for generating a derivative of an effective compound may include, in addition to steps (B1) and (B2) above, a step (B3) of selecting some clustering regions as a target space to be used for generating a derivative of an effective compound based on the proximity to the scaffold of the effective compound and the size of the clustering region in the clustered regions.
[0044] 16. In one embodiment, clustering may be performed based on the density of regions where the interaction energy between target protein atoms is relatively low, or may be Gaussian Mixture Model (GMM) clustering.
[0045] 17. In one embodiment, step (C) may be performed by selecting a compound in which the substitution target moiety binding to the scaffold of the effective compound selected in step (A) is substituted with a substituent selected from a database of substituent groups that are acceptable for the target space within the target protein established in step (B).
[0046] 18. In one embodiment, step (C) may be to generate a derivative having multiple different bonding structures for the same substituent by varying the bonding position within the substituent that binds to the scaffold of the effective compound.
[0047] 19. In one embodiment, step (C) may be to generate a derivative having multiple different bonding configurations for a substituent having the same bonding structure by varying the bonding angles of the substituent and the scaffold of the effective compound.
[0048] 20. In one embodiment, step (C) may be to generate a derivative having multiple different bonding forms (single bond, double bond, triple bond, etc.) for a substituent having the same bonding structure by changing the bonding of the substituent and the scaffold of the effective compound.
[0049] 21. In one embodiment, the method comprises:
[0050] In step (A), the portion to be substituted in the chemical structure of the effective compound is selected based on an indicator that the binding interaction is lower than that of other chemical structures from the interaction profile of the effective compound and the target protein.
[0051] In step (B), selection of target space to be used for generating derivatives of effective compounds among target proteins is performed.
[0052] (B1) A step of generating a region of a target protein constituting a target space in the form of a cylinder filter, setting up marking points (dots) arranged at equal intervals on the cylinder filter, and distinguishing the marking points by the interaction energy between target protein atoms.
[0053] (B2) A step of approaching the scaffold of the effective compound to the cylinder filter, excluding areas where the interaction energy is greater than a preset value from the cylinder filter, and then clustering the indicated points into spatial units, and
[0054] (B3) A method comprising a step of selecting some clustering regions based on the proximity to the scaffold of the effective compound in the clustered regions and the size of the clustering region,
[0055] The derivative generation in step (C) may be accomplished by selecting a substituent corresponding to the target space selected in step (B).
[0056] 22. In one embodiment, the method may further include a step of filtering the generated derivatives (D).
[0057] 23. In one embodiment, step (D) may be filtering the bond angle of the scaffold of the substituent and the effective compound by comparing it with a database of existing substances.
[0058] 24. In one embodiment, step (D) may filter the derivatives according to the amount of collisions between atoms of the derivative and atoms of the target protein generated within the target space set in step (B) according to the bonding form of the substituent of the derivative.
[0059] 25. One embodiment of the present disclosure may be directed to an in silico derivative generated by a method for generating a derivative according to one embodiment of the present disclosure.
[0060] 26. One embodiment of the present disclosure may relate to an artificial intelligence new drug platform constructed by a method for generating a derivative according to one embodiment of the present disclosure.
[0061] A method for generating a derivative of an effective compound from an effective compound for a target protein according to one embodiment of the present disclosure can improve an effective compound whose efficacy has been confirmed or derived previously in terms of pharmacological efficacy or pharmacodynamics, and generate various derivatives for learning an analysis AI algorithm.
[0062] Through a method for producing a derivative according to one embodiment of the present disclosure, a derivative having an improved actual binding potential to a target protein can be produced.
[0063] Through a method for producing a derivative according to one embodiment of the present disclosure, a derivative having improved binding affinity to a target protein or having improved effects on the activity of a target protein can be produced compared to an effective compound.
[0064] Through the method for generating a derivative according to one embodiment of the present disclosure, since the binding form (pose) of a compound that can be combined within a target space is derived together during the process of generating a derivative, when a molecular dynamics simulation is performed on an artificial intelligence new drug platform, there is an effect of improving the possibility of deriving an optimal binding form.
[0065] Figure 1 is a schematic diagram illustrating an exemplary overall configuration of an artificial intelligence drug platform (AI-drug platform) constructed by applying a method for producing compound derivatives for a target protein of the present disclosure.
[0066] Figure 2 is a conceptual diagram illustrating an exemplary cloud service structure of an artificial intelligence new drug platform constructed by applying a method for producing compound derivatives for a target protein of the present disclosure.
[0067] Figure 3 is a conceptual diagram illustrating the process of discovering effective substances for an artificial intelligence new drug platform constructed by applying a method for producing compound derivatives for a target protein of the present disclosure.
[0068] Figure 4 is a conceptual diagram illustrating the process of discovering a lead substance of an artificial intelligence new drug platform constructed by applying a method for producing compound derivatives for a target protein of the present disclosure.
[0069] Figure 5 is a flow chart illustrating a method for discovering a lead substance through derivative production according to a specific example of a method for producing compound derivatives for a target protein of the present disclosure.
[0070] FIG. 6 is a conceptual diagram illustrating a method for selecting a portion to be substituted in the process of a method for producing a compound derivative for a target protein of the present disclosure, and is an exemplary conceptual diagram including a process for selecting an anchor atom, which is an atom on a scaffold to which a portion to be substituted is bound.
[0071] Figures 7a to 9 are conceptual diagrams illustrating a process for calculating the size of a target space in a method for producing a compound derivative for a target protein of the present disclosure.
[0072] Figures 10 to 11b are conceptual diagrams illustrating a method for generating a derivative by selecting a substituent (R-group) in the process of a method for generating a compound derivative for a target protein of the present disclosure.
[0073] Figures 12a and 12b are exemplary diagrams illustrating a process of extracting a linker of an anchor portion of a target protein in the process of producing a compound derivative for a target protein of the present disclosure, and filtering the binding form of the linker and a substituent by comparing it with the binding form of an existing substance.
[0074] Figures 13a and 13b are exemplary diagrams illustrating a process of filtering according to whether or not there is an interatomic collision within the target space of the derivative generated in the process of the method for generating a compound derivative for the target protein of the present disclosure.
[0075] FIG. 14 shows a specific example of a method for producing a compound derivative for a target protein of the present disclosure, which calculates an interaction profile based on an effective compound (effective compound X) that is already known to bind to a specific target protein (protein A), and shows the steps of selecting a portion to be substituted and other scaffold portions.
[0076] Figure 15 shows a shape in which a region filter is extracted based on atoms (i.e., anchor atoms) that bind to the substitution target portion on the scaffold of an effective compound within the target region (aka pocket region) of a target protein, and display points are set at equal intervals (total cylinder in Figure 15). The region extraction proceeded in a cylinder shape. According to the r value calculated from the formula for calculating the van der Waals radius (vdw radius; r) of the atoms of the target protein within the region, the indicated points within about 0.7r were indicated as red points (red dot; red zone; collision zone), the indicated points within about 0.7r to about 1r were indicated as yellow points (yellow dot; yellow zone; buffer zone), the indicated points within about 1r to about 1.3r were indicated as green points (green dot; green zone; contact zone), and the indicated points above about 1.3r were indicated as gray points (gray dot; gray zone).
[0077] Figure 16 is a graph showing the process of a method for producing compound derivatives for a target protein of the present disclosure, in which a target space is extracted, a portion to be substituted is substituted with an acceptable substituent in the target space, and derivatives are selected in order of lowest binding energy (e.g., bottom 5%).
[0078] Figure 17 is a diagram modeling the spatial configuration of derivatives expected to have optimal substituents and the bonding forms of the derivatives.
[0079] The various embodiments or examples described in this disclosure are exemplified for the purpose of clearly explaining the technical idea of this disclosure and are not intended to be limited to specific embodiments. The technical idea of this disclosure includes various modifications, equivalents, alternatives, and optional combinations of all or part of each embodiment or example described in this document.
[0080] All technical and scientific terms used in this disclosure, unless otherwise defined, have the meaning commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0081] The singular expressions used in this disclosure may include the plural meaning unless the context clearly indicates otherwise, and the same applies to the singular expressions set forth in the claims.
[0082] The combination of each block of the block diagram and each step of the flowchart described in the present disclosure may be performed by computer program instructions (execution engine), and these computer program instructions may be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, so that the instructions executed through the processor of the computer or other programmable data processing equipment create a means for performing the functions described in each block of the block diagram or each step of the flowchart.
[0083] In addition, since the computer program instructions can be installed on a computer or other programmable data processing equipment, a series of operation steps are performed on the computer or other programmable data processing equipment to create a computer-executable process, and the instructions that perform the computer or other programmable data processing equipment can also provide steps for executing the functions described in each block of the block diagram and each step of the flowchart.
[0084] Additionally, each block or each step may represent a module, segment or portion of code that includes one or more executable instructions for performing specific logical functions, and in some alternative embodiments, the functions mentioned in the blocks or steps may occur out of order.
[0085] In addition, in the field of programs to which the technology according to the present disclosure is applied, most technical terms are not defined in Korean and English names are used as common names. Therefore, in the case of technical terms written in Korean, they should be interpreted as having the meaning of the English name commonly used as a common name in the technical field.
[0086] I. Definition
[0087] In this disclosure, the term "about" can indicate a typical error range for each value, as is well known to those skilled in the art. This can mean, in the context of a numerical value or range described in this disclosure, ±20%, ±15%, ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, or ±1% of the numerical value or range stated or claimed in one embodiment.
[0088] Expressions such as “including,” “comprising,” “having,” and the like, as used in this disclosure, should be understood as open-ended terms that imply the possibility of including other embodiments in a similar manner to “comprising,” unless otherwise stated in the phrase or sentence in which the expression is included.
[0089] The term "and / or" as used herein may mean any one or more of the items, any combination of the items, or all of the items associated with the term.
[0090] An embodiment of the present disclosure described in this disclosure is understood to include "comprising," "consisting of," and / or "consisting essentially of" an embodiment.
[0091] II. Compositions related to the method for producing derivatives of effective compounds in the present disclosure
[0092] Composition included in the technology of the present disclosure
[0093] The term "target protein" as used herein refers to a protein that binds to an effective compound of the present disclosure, and whose properties, such as activity or stability, may be affected by binding to the effective compound. In the present disclosure, derivatives are generated based on the properties of the binding between the target protein and the effective compound.
[0094] As used herein, the term "effective compound" may refer to a molecule, compound, or ligand that binds with high affinity to a target protein in the present disclosure. That is, it may refer to any molecule, ligand, or compound that binds covalently or non-covalently to a target protein in the present disclosure.
[0095] As used herein, the term "derivative of an effective compound" or "derivative" may refer to a compound produced by substituting, deleting, or adding a specific chemical group to a portion of the effective compound of the present disclosure. For example, a derivative of the effective compound may be produced by dividing the chemical structure of the effective compound of the present disclosure into a portion to be substituted and a scaffold portion, and substituting the portion to be substituted with a different substituent.
[0096] The term "substitution target moiety" as used herein may refer to a portion of a compound selected as a target for substitution from among the chemical structures of an effective compound that binds to a target protein in the present disclosure. This portion may be selected based on the binding interaction between the target protein and the effective compound. For example, a chemical structure within the compound with a lower binding interaction than other chemical structures may be selected based on the interaction profile of the target protein and the effective compound. The portion of the effective compound excluding the substitution target moiety is designated as a "scaffold."
[0097] The "interaction profile" used in this disclosure is an organized collection of information, including the types of atoms, physicochemical properties, and types of interactions. Specifically, the interaction profile can be obtained using commercially available software (e.g., FEP+, MMPBSA). In this context, there is the existing DSSP (Dictionary of Protein Secondary Structure) program, which is based on Coulomb's force and Coulomb's law, and the profile can be calculated using software based on various algorithms improved from the DSSP program.
[0098] The "binding interaction" used in the present disclosure can be measured by calculating the binding energy of atoms involved in the binding between a target protein and an effective compound. For example, the binding energy can be calculated for each atom involved in the binding between the target protein and the effective compound, and the median or average thereof can be calculated. Furthermore, in order to calculate the binding interaction value, one bond in the chemical structure of the effective compound that interacts with the target protein is cleaved, and for each cleaved bond, a candidate for substitution is selected, including an atom or a chemical structure portion of the cleaved portion that interacts with the target protein, and the binding energy of each atom constituting the candidate for substitution with the target protein can be calculated, and the average thereof can be calculated. The binding interaction in the present disclosure can mean the interaction efficiency commonly used in the art, and can be calculated as the average value of the binding energy of each atom that constitutes a specific chemical group or a specific portion of a compound.
[0099] Meanwhile, in the present disclosure, the “interaction relevance” can be calculated, for example, by the following equation.
[0100] [Calculation formula]
[0101]
[0102] The term "substitution candidate moiety" used herein refers to an atom or chemical group that may be a candidate for selection of the "substitution target moiety" within the effective compound. This can be achieved by cleaving bonds (e.g., single bonds, double bonds, etc.) one by one in a portion interacting with a target protein within the effective compound to create two atomic fragments, and then selecting the portion that interacts closer to or more strongly with the target protein as the "substitution candidate moiety." The substitution candidate moiety may comprise 1 to 12 atoms.
[0103] The term "target space within a target protein", "target region", "pocket space" or "target space" used in the present disclosure refers to a region to be analyzed for performing a method for producing a derivative in the present disclosure, and may be set with a focus on a region within the target protein where a substitution target portion of an effective compound binds. The size of the target space within the target protein may be set to accommodate a substitution target portion of an effective compound selected from among chemical structures of effective compounds that bind to the target protein. For example, the size of the target space within the target protein may be set to a size that accommodates (1) the substitution target portion and (2) a region excluding the substitution target portion (i.e., a portion belonging to the binding region among the scaffold portions of the effective compound) in the chemical structure of the effective compound that binds to the target protein. In the present disclosure, the target space may be set to a preset size, or may be set to a size that includes all of the substitution target portions selected based on the binding structure of the target protein and the effective compound and the atoms of the target protein that interact with the substitution target portion. The region of the target protein constituting the set target space may be three-dimensionally modeled. The target protein region constituting the target space within the set target protein can be divided stepwise according to the interaction energy of the target protein atoms present in the region. Meanwhile, in the present disclosure, the target protein region constituting the target space is extracted in the form of a cylinder filter, and after setting the marking points (dots) arranged at equal intervals on the cylinder filter, the marking points can be divided stepwise according to the interaction energy of the target protein atoms.
[0104] "Clustering" as used in the present disclosure refers to setting regions with identical or similar characteristics within a certain space or area as one region, and in the present disclosure, the region can be set using the interaction energy of target protein atoms as an indicator. For example, in the present disclosure, the region of the target protein constituting the target space within the target protein is divided into stages according to the interaction energy of the target protein atoms existing in the region, and then the region with a high interaction energy of the target protein atoms is excluded or only the region with a low interaction energy is left, and the region can be clustered. In the present disclosure, the clustering can be set by additionally considering the proximity to the scaffold of the effective compound of the region and / or the size of the clustering region. In one specific embodiment, the region of the target protein constituting the target space is extracted in the form of a cylinder filter, and the dot points arranged at equal intervals are set on the cylinder filter, and then the dot points are divided into stages according to the interaction energy of the target protein atoms, and the dot points with a relatively low interaction energy are left and clustered. In the present disclosure, clustering may be performed based on the density of regions with relatively low interaction energy or by Gaussian Mixture Model (GMM) clustering.
[0105] The "computational system" used in this disclosure may be a system that receives data as input, processes it, and outputs results through rules or algorithms. The computational system may have the function of automatically performing tasks such as calculations, processing, and / or conversions, and may include hardware and / or software. The computational system may include input, processing, and output stages as its components. The input stage is a stage where data is input to solve complex problems through multiple stages of calculations, and may be input from a user or automatically input by another system or sensor. The processing stage is a stage where calculations are performed through an algorithm or program, and may be a stage where input data is processed. The output stage is a stage where the processed result is output, and the output may be displayed on a screen, transmitted to another system, or stored. Examples of the computational system may include a computer, an artificial intelligence (AI) system, etc.
[0106] III. Aspects of a method for producing a derivative of an effective compound in the present disclosure
[0107] Before explaining the method for producing derivatives of effective compounds for target proteins according to the present disclosure, the entire artificial intelligence new drug platform to which the method for producing derivatives of effective compounds for target proteins according to the present disclosure is applied will be explained.
[0108] FIG. 1 is a diagram showing the overall configuration of an AI-drug platform to which the technology according to the present disclosure is applied, FIG. 2 is a conceptual diagram showing the cloud service structure of an AI-drug platform to which the technology according to the present disclosure is applied, FIG. 3 is a conceptual diagram showing an effective substance discovery process of an AI-drug platform to which the technology according to the present disclosure is applied, and FIG. 4 is a conceptual diagram showing a lead substance discovery process of an AI-drug platform to which the technology according to the present disclosure is applied.
[0109] The artificial intelligence new drug platform to which the technology of this disclosure is applied is a platform that basically performs the entire process of discovering new drug candidates at the preclinical stage, and can be serviced through the cloud (STB CLOUD).
[0110] At this time, the new drugs that can be serviced on the artificial intelligence new drug platform include new synthetic drugs (small molecules) and new antibody drugs, and the artificial intelligence new drug platform according to the present disclosure provides a discovery process for all of these.
[0111] Meanwhile, for this purpose, the artificial intelligence new drug platform according to the present disclosure is configured to include an effective (hit) substance automation discovery platform, a lead substance automation discovery platform, and a drug response (ADMET, Absorption, Distribution, Metabolism, Excretion & Toxicity) automation analysis platform, as illustrated in FIG. 1.
[0112] That is, the artificial intelligence new drug platform according to the present disclosure is an artificial intelligence platform configured to perform the entire process of new drug development, which includes selecting effective substances (e.g., effective compounds), discovering lead substances among them, and then selecting candidate substances through drug response analysis.
[0113] FIG. 2 illustrates the cloud service process of the artificial intelligence new drug platform according to the present disclosure. As illustrated therein, the technology according to the present disclosure discovers effective substances, produces lead substances, and provides all areas of the drug discovery and development process, including ADMET / PK, pharmacogenetics, and biomarkers.
[0114] In addition, in order to operate the platform for each discovery stage of these new drug development, the artificial intelligence new drug platform according to the present disclosure applies three individual artificial intelligence systems: a generative artificial intelligence system (GPT / BERT), a three-dimensional structural artificial intelligence system (ED-CNN), and a molecular dynamics analysis system (Auto-MD simulation).
[0115] And, using the artificial intelligence systems of the artificial intelligence new drug platform, specific methods for executing each effective substance automation discovery platform, leading substance automation discovery platform, and drug response automation analysis platform are as follows: effective substance discovery through 3D structural information between protein and ligand (hereinafter referred to as 'DMC-PRE', a coined technical name of the applicant), protein-ligand docking structure analysis based on central atom vector (hereinafter referred to as 'GAP-Dock', a coined technical name of the applicant), protein-compound optimal binding structure prediction using a 3D-CNN learning model (hereinafter referred to as 'DMC-SCR', a coined technical name of the applicant), generation of compound derivatives for target proteins (hereinafter referred to as 'LEAD-GEN', a coined technical name of the applicant), protein-compound mutual binding stability analysis through molecular dynamics simulation data (hereinafter referred to as 'DMC-MD', a coined technical name of the applicant), and a generated artificial intelligence model trained using 3D interaction data between protein and compound (hereinafter referred to as 'DMC-SCR', a coined technical name of the applicant). (called '3bmGPT') can be applied.
[0116] Here, the DMC-PRE and GAP-Dock are technologies that are applied prior to discovering effective substances through an effective substance automation platform, DMC-SCR is a technology that is applied after discovering effective substances through the effective substance automation platform by applying a molecular dynamics analysis system, and LEAD-GEN is a technology that is applied to discover leading substances through an automated leading substance automation platform.
[0117] And DMC-MD is a technology applied to the above molecular dynamics analysis system to verify the binding stability of the results derived from the effective substance automated discovery platform, the lead substance automated discovery platform, and the drug response automated analysis platform, and 3bmGPT is a technology applied to the generative artificial intelligence system to select the target substance for analysis when producing effective substances through the effective substance automated discovery platform.
[0118] Specifically, if we look at the process of discovering effective substances of the artificial intelligence new drug platform according to the present disclosure, as illustrated in FIG. 3, the target substance for analysis selected through the 3bmGPT is subjected to preliminary screening by applying the DMC-PRE and GAP-Dock, and then the DMC-SCR is applied to conduct in-depth screening, and then the DMC-MD is applied to verify binding stability to derive effective substances.
[0119] And, looking at the process of discovering a lead substance of the artificial intelligence new drug platform according to the present disclosure, as shown in Fig. 4, the lead substance is discovered by applying the above LEAD-GEN, and the binding stability is verified by applying DMC-MD to derive the lead substance.
[0120] The technology according to the present disclosure relates to a method for producing compound derivatives for a target protein (so-called LEAD-GEN) applied to the artificial intelligence new drug platform described above, and specific embodiments thereof will be described below with reference to drawings.
[0121] FIG. 5 is a flowchart illustrating a method for discovering a lead substance through the generation of derivatives of an effective compound for a target protein according to the present disclosure, FIG. 6 is a conceptual diagram illustrating a method for selecting a substitution target portion and a scaffold (anchor molecule; anchor atom) in the process of generating a derivative according to the present disclosure, FIGS. 7 to 9 are conceptual diagrams illustrating a process for calculating the size of a target space within a target protein in the process of generating a derivative according to the present disclosure, FIGS. 10 and 11 are conceptual diagrams illustrating a method for generating a derivative by selecting a substituent in the process of generating a derivative according to the present disclosure, and FIG. 12 is an exemplary diagram illustrating a process for filtering a substituent (filtering an existing material in the form of a bond between a substituent and a scaffold and a bonding group of the substituent by comparing the bonding form) in the process of generating a derivative according to the present disclosure.
[0122] First, the method for producing a derivative of an effective compound for a target protein according to the present disclosure can be broadly performed including the steps of (A) selecting a substitutable portion and a scaffold excluding the substitutable portion in the chemical structure of the effective compound; (B) setting a target space within the target protein centered on a region where the substitutable portion of the selected effective compound binds; and (C) selecting a substituent that replaces the substitutable portion of the effective compound that is acceptable within the set target space of the target protein to produce a derivative.
[0123] The method for producing a derivative of an effective compound for a target protein according to the present disclosure can be performed in an operating system.
[0124] In the above step (A), the portion to be substituted among the chemical structures of the effective compound may be selected based on an indicator that the binding interaction is lower than that of other chemical structures from the interaction profile of the effective compound and the target protein, and the binding interaction may be calculated by cutting one of the bonds in the effective compound that interacts with the target protein, selecting a candidate for substitution that includes an atom of the cut portion for each cut bond, and calculating the average value of the binding energy of each atom constituting the candidate for substitution with the target protein.
[0125] In the present disclosure, a candidate for substitution can be selected through a step of cleaving a bond, for example, a single bond, within an effective compound to generate atomic fragments on either side of the cleavage site. The candidate for substitution can include 1 to 12 atoms, and the method for generating a derivative according to the present disclosure can additionally include a step of filtering the candidate for substitution based on the number of constituent atoms.
[0126] Specifically, as shown in 'A' of Fig. 6, when an effective compound, which is an effective substance, is bound to a target protein, an interaction (binding information) profile between the effective compound and the target protein is calculated.
[0127] At this time, the above interaction profile can also be obtained through commercialized or separately developed software (FEP+, MMPBSA, etc.) that can extract the necessary information.
[0128] Thereafter, as shown in 'B' of Fig. 6, from the interaction profile of the effective compound and the target protein, one of the bonds constituting the effective compound that binds to the target protein is cleaved to generate atomic fragments on both sides of the cleaved portion.
[0129] In this way, cleavage and atomic fragmentation are performed across the bonds of the compound, producing atomic fragments twice the number of bonds.
[0130] And, among these produced atomic fragments, only the atomic fragments containing atoms within a preset range (e.g., 1 to 12) can be selected, and the remaining atomic fragments can be excluded from the analysis target.
[0131] The purpose of selecting the above atomic fragments is to derive atomic fragments that are valuable as replacement targets, with minimal impact on the bonding between the effective compound and the target protein. If the number of atoms constituting the selected atomic fragment is too small, the likelihood of deriving a new compound through substitution is low. Conversely, if the number of atoms constituting the atomic fragment is too large, there is a concern that the properties of the compound after substitution may significantly change.
[0132] For selected atomic fragments, interaction efficiency can be calculated to select a target substitution moiety. Atomic fragments with low interaction efficiency can be considered weaker in terms of bonding interactions and thus have a lower impact on bonding interactions, making them suitable for substitution. Within the chemical structure of a valid compound, the remaining portions, excluding the target substitution moiety, can be selected as the scaffold.
[0133] Meanwhile, in the chemical structure of an effective compound that binds to a target protein, atoms on the scaffold that bind to the target substitution site can be selected as anchor atoms. That is, for the selected atomic fragments, as illustrated in 'B' of Figure 6, the interaction efficiency can be calculated, and the cut portion of the atomic fragment with a low influence on the binding interaction can be selected as the anchor.
[0134] At this time, the interaction efficiency can be calculated as the average value of the binding energy of each atom constituting the atomic fragment.
[0135] The above step (B) may be to generate a target space, which is a space for subsequent derivative generation, including a region (also referred to as a pocket space) that interacts with an effective compound within the target protein.
[0136] Information on the binding form of these target proteins and effective compounds, and the form of the interacting region, can be obtained through RCSB PDB, etc.
[0137] In the above step (B), the size of the target space within the target protein may be set so that the area excluding the substitution target area of the selected effective compound in the effective compound portion belonging to the area binding to the target protein is acceptable. In addition, the setting of the target space in the above step (B) may be set including a step of dividing the area of the target protein constituting the target space within the target protein into stages according to the interaction energy between the target protein atoms present in the corresponding area. The interaction energy may be divided into 3 to 5 stages, and the area in which the interaction energy is relatively low among the divided stages may be clustered, and further, the clustering area of the target space to be used for generating a derivative of the effective compound may be selected according to the proximity to the scaffold of the effective compound and the size of the clustering area.
[0138] For example, the distinction of interaction energy between target protein atoms can be performed by extracting a region of the target protein constituting the target space in the form of a cylinder filter, setting dot points arranged at equal intervals on the cylinder filter, and then dividing the dot points step by step according to the interaction energy between the target protein atoms.
[0139] Meanwhile, the selection of the clustering region of the target space to be used for generating derivatives of the effective compound may be performed by including (B1) a step of extracting the region of the target protein constituting the target space in the form of a cylinder filter, setting dot points arranged at equal intervals on the cylinder filter, and then dividing the dot points in stages according to the interaction energy between the atoms of the target protein, and (B2) a step of approaching the scaffold of the effective compound to the cylinder filter, excluding from the cylinder filter the region in which the interaction energy is higher than a preset value, and then clustering the remaining dot points into spatial units, and additionally (B3) a step of selecting some clustering regions as the target space to be used for generating derivatives of the effective compound according to the proximity between the dot points in the clustered regions and the size of the clustering region.
[0140] The above clustering may be performed based on the density of regions where the interaction energy between target protein atoms is relatively low, or may be Gaussian Mixture Model (GMM) clustering, but is not limited thereto.
[0141] Specifically, (B) the step of calculating the target space (pocket space) within the target protein is to calculate the internal space that affects the interaction when the scaffold with the replacement target portion removed is bound within the interaction region between the target proteins, as illustrated in FIGS. 7 to 9.
[0142] To this end, as shown in Fig. 7, in a specific example, a cylindrical region of a preset size (length 10 Å, radius 10 Å) was extracted centered on the substitution target portion or anchor portion of the effective compound for the target protein, thereby generating a cylinder-shaped region filter.
[0143] The size of the above region filter can be changed depending on the structure of the scaffold and the resources of the computational system, but can be set so that the region interacting with the target protein of the scaffold is sufficiently covered.
[0144] Meanwhile, the cylinder region filter is set with equally spaced marker points (dots), and the marker points (dots) are distinguished by the interaction energy between target protein atoms. In 'A' of Fig. 7, the marker points included in the cylinder region filter are distinguished by color according to the interaction energy between target protein atoms.
[0145] As illustrated in 'A' and 'B' of Fig. 8, the anchor portion of the scaffold was then positioned close to the anchor portion of the cylinder region filter. At this time, the scaffold bonding axis direction can maintain the original bonding direction.
[0146] Afterwards, as shown in 'C' of Fig. 8, the areas (Red zone, Yellow zone, Green zone) with high interaction energy between target protein atoms among the indicated points were excluded from the cylinder area filter, and only the indicated points (dots) belonging to areas with low interaction energy (Gray zone, Dark Gray zone) among the indicated points were left.
[0147] The remaining marker points were clustered (e.g., GMM clustering) into spatial units, as illustrated in 'D' of Fig. 8. Among the clustered regions, only some clustering regions were derived based on their proximity to the scaffold anchors and their sizes. In a specific embodiment, the two largest clustering regions connected to the scaffold anchors were selected ('E' of Fig. 8).
[0148] As illustrated in 'F' of Fig. 8, the selected clustering region was derived as a target space (aka pocket space; target volume) within the target protein. The reason for deriving the target space within the target protein in this way is to select a substituent (R-group; Replace atom group) to replace the target portion based on the size and shape of the target space (target volume).
[0149] Figure 9 shows an example of deriving a target volume within an actual target protein using the method described above.
[0150] In the example illustrated in Fig. 9, for compound 6op0, nine anchors were derived, and as a result of analysis, the analysis results were obtained that matched the actual number of ligand atoms of the compound. Specifically, the number of dots included in the target space was divided by 300 to indicate the size, and when divided by 300, which is the average volume typically occupied by heavy atoms (C, N, O) included in the target space, it was confirmed that this matched the number of atoms included in the ligand of the actual compound.
[0151] The above step (C) is a step for generating a derivative, and can be performed by selecting a compound in which the substitution target portion that binds to the scaffold of the effective compound selected in the above step (A) is substituted with a substituent selected from a database of substituent groups that are acceptable for the target space within the target protein set in the above step (B).
[0152] That is, the step of generating a step (C) derivative is to select a substituent that matches the target space size within the target protein produced in step (B), as illustrated in FIG. 10.
[0153] Here, a substituent is a group of atoms that replaces the target moiety, i.e., the atom removed from the scaffold. The group is selected from a database containing various atomic configurations. The selected substituent is then attached to an anchor in the scaffold to generate a derivative.
[0154] At this time, by varying the bonding position within the substituent that binds to the scaffold of the effective compound, derivatives having multiple different bonding structures for the same substituent can be generated. That is, as illustrated in Fig. 10, by varying the bonding position within a selected substituent to the anchor atom of the scaffold, derivatives having various bonding structures can be generated.
[0155] In addition, even for substituents having the same bonding structure, derivatives having multiple different bonding configurations can be generated by varying the bonding angles between the substituent and the scaffold of the effective compound. That is, as illustrated in Fig. 12, even for derivatives having the same bonding structure, multiple derivatives can be generated by varying the bonding configurations (bonding angles) between the substituent and the scaffold.
[0156] In this case, by changing the bonding form of the substituent and the scaffold, derivatives with various bonding forms can be generated by extracting the bonding portion (aka linker) of the anchor portion of the target protein, as shown in Fig. 12.
[0157] Here, the linker is extracted only from the portion adjacent to the anchor, and can be generated by extracting only up to a preset number of linking atoms from the anchor. Fig. 12 illustrates an example in which up to three atomic bonds from the anchor atom are extracted as the linking portion.
[0158] In summary, by diversifying the bonding forms of the above bonding moieties and substituents, derivatives having various bonding forms can be produced.
[0159] The method for producing derivatives of effective compounds for the target protein of the present disclosure may additionally include a step of (D) filtering the produced derivatives.
[0160] Filtering criteria for generated derivatives can be selected in various ways. First, based on structural data on existing compounds, derivatives with a low probability of actual synthesis or existence can be excluded. Second, filtering can be done by considering the bonding patterns of substituents and the potential interatomic collisions that may occur within the target space. Here, interatomic collisions can refer to collisions between atoms in the derivative and atoms in the target protein.
[0161] Filtering of derivatives generated in the present disclosure can be largely accomplished by two criteria. The first criterion is to filter by comparing the bonding form of the bonding site and the substituent with a structural database of existing compounds (so-called linker filtering), and the second criterion is to filter derivatives by the amount of interatomic collisions generated within the target space after binding them to a cylinder area filter according to the bonding form of the substituent on the derivative (so-called bonding form filtering; shape filtering).
[0162] The above linker filtering can select various bonding forms of the linker and substituents by comparing them with the bonding structures of existing substances in an existing compound database (e.g., ChEMBL, etc.), as illustrated in FIG. 12.
[0163] Specifically, the bonding structures having the same configuration as the linker and substituent are selected from the compound database, and their bonding forms are analyzed to select only derivatives in which the bonding forms of the linker and substituent exist in actual substances. At this time, the process of calculating the bonding form of a specific bonding portion from actual compound data can be calculated according to the SMARTS method for describing molecular patterns, as illustrated in FIG. 12. In addition, as illustrated in FIG. 13, the derivatives generated in the process of the method for generating compound derivatives for the target protein of the present disclosure can be filtered according to whether or not there is an interatomic collision within the target space or the amount of collision. The interatomic collision may refer to a collision between an atom of the derivative and an atom of the target protein.
[0164] Hereinafter, the configuration and effects of the present disclosure will be described in more detail with examples. However, these examples are provided solely for illustrative purposes to aid understanding of the present disclosure, and the scope and range of the present disclosure are not limited by the examples below.
[0165] [Example 1] Analysis of effective compounds binding to target proteins
[0166] In order to verify the effectiveness of the method of making a derivative of an effective compound based on an effective compound that binds to a target protein of the present disclosure, an exemplary target protein and an effective compound that binds thereto were selected and the method of the present disclosure was performed.
[0167] In this example, the target protein is a protein with kinase activity (hereinafter referred to as protein A), and the portion to which the effective compound binds is the ATP binding pocket portion where the kinase and ATP bind. Specifically, protein A has a structure comprising five beta sheets at the top, a central cavity surrounding the top, and alpha helices at the left, right, and bottom, respectively, which corresponds to a common and similar structure in many kinase proteins having an ATP binding pocket of the kinase. The specific structure can be confirmed in the modeling schematic diagrams of FIGS. 15 and 17.
[0168] The compound C21H20N6O2S (SMILES: O=C(NC1=CC=CC(C2=NNC(NC(C3=NC(C=CC=N4)=C4S3)=O)=C2)=C1)C(C)(C)C), which the inventors had previously confirmed to bind to protein A, was set as the effective compound X and as the parent compound in the subsequent derivative production process.
[0169] [Effective compound X]
[0170]
[0171] The structural information of protein A and effective compound X was input into a binding interaction calculation program to determine the interaction profile, which includes information on whether there is a bond between two atoms and the strength of the bond.
[0172] Specifically, after calculating the interaction profile between the effective compound X and protein A using the program, single bonds within the effective compound X at the part interacting with protein A were individually cut to generate atomic fragments on both sides of the cut portion. Among these, atomic fragments close to protein A were selected as substitution candidate parts. A substitution target part can be selected from the substitution candidate parts based on the binding interaction between the target protein and the effective compound. In the present disclosure, the binding interaction may mean an interaction relevance commonly used in the present technical field.
[0173] First, the atomic fragments corresponding to the generated substitution candidate portion were filtered according to the number of constituent atoms of the atomic fragment. The atomic fragments containing 1 to 12 constituent atoms were filtered, and the substitution target portion was selected based on the low bonding interaction among the filtered atomic fragments. Specifically, the interaction efficiency of the filtered atomic fragments was calculated. The interaction efficiency can be calculated as the average value of the binding energy of each atom constituting the atomic fragment. At this time, the interaction efficiency was calculated using the following formula.
[0174] [Calculation formula]
[0175]
[0176] As a result, as can be seen in <2. Interaction profile of effective compound> of Fig. 14, the part interacting with protein A in the structure of effective compound X was identified, and the strength of the interaction was confirmed.
[0177] At this time, the part of the structure of the effective compound X that has a weak interaction with protein A may correspond to the atom side with a low calculated interaction correlation, which may also correspond to a lower bonding interaction compared to other chemical structures among the interaction profiles.
[0178] The atomic fragment with low interaction correlation was selected as the target for replacement.
[0179] In addition, the atoms on the scaffold to which the atomic fragments selected as the replacement target are combined were selected as anchors.
[0180] Accordingly, the two terminal parts of the compound corresponding to the weak interaction part were selected as the parts to be substituted (R1 and R2 in <3. Substitution Anchor> of Fig. 14), and the part excluding this was selected as the scaffold part (the part excluding R1 and R2 of 3 of Fig. 14), and the atoms on the scaffold that bind to R1 and R2 of 3 of Fig. 14, i.e., the N atom and the C atom, were selected as the so-called anchor atoms.
[0181] [Example 2] Setting the target space within the target protein
[0182] In Example 1, a target space within a target protein that interacts with a substitution target portion within the selected effective compound X is set and utilized for subsequent derivative generation. In this regard, the region interacting with the effective compound X within protein A was analyzed.
[0183] Specifically, a cylindrical region filter was created by extracting the region that acts on binding, centered on the anchor site within protein A. At this time, a cylindrical region with a radius of 10 Å and a height of 10 Å was extracted so as to sufficiently reflect the environment that affects the interaction between the effective compound X and protein A (Fig. 15).
[0184] Dots were set at equal intervals in the extracted region filter, and the dots were distinguished by the interaction energy between protein A atoms. Specifically, according to the r value calculated from the van der Waals radius (vdw radius; r) calculation formula shown in 'B' of Fig. 7, the dot points within 0.7r were indicated as red dots (red zone; collision zone), the dot points within about 0.7r to about 1r were indicated as yellow dots (yellow zone; buffer zone), the dot points within about 1r to about 1.3r were indicated as green dots (green zone; contact zone), and the dot points within about 1.3r to 3 Å were indicated as gray dots (gray zone; close zone).
[0185] At this time, it was judged that the gray point with low interaction energy between target protein atoms was likely to be applied as a pocket area with adequate free space where binding with the compound could be induced.
[0186] The anchor atom site of the scaffold from which the atomic fragment selected as the substitution target site in the effective compound X was removed was placed close to the anchor site of the extracted cylinder region filter, and the region with high interaction energy between target protein atoms among the dot sites was excluded from the cylinder region filter. Thereafter, the dot sites in the remaining filter were clustered into spatial units (GMM clustering), and the clustered regions were selected based on the proximity to the anchor atoms and the size of the clustered regions to derive only two clustered regions. This was set as the target space (pocket space; target volume) within the target protein, and the size of the target space was calculated.
[0187] [Example 3] Generation of derivatives through substituent selection
[0188] A derivative was generated by substituting the target portion of the effective compound selected in Example 1 with a substituent acceptable within the target space of the target protein established in Example 2. The substituent was selected by selecting a compound substituted with a substituent selected from a database of substituent groups.
[0189] Specifically, when one substituent was selected, the optimal bonding structure in which the substituent and the anchor atom are bonded was examined, and after selecting the bonding structure, various bonding forms were examined by changing the bonding angle, thereby generating a number of derivatives.
[0190] The binding energy of each generated derivative was calculated, and derivatives with stable bonds with binding energies in the lower 5% were selected (see Figure 16). Among these, compounds with high potential for existence and synthesis were selected through comparison with existing compounds, and actual compounds were prepared. As shown in Table 2 below, specific examples according to the present disclosure confirmed that derivatives generated in silico were synthesized into actual compounds.
[0191] [Example 4] Confirmation of the target protein activity inhibition effect of the manufactured effective compound derivative.
[0192] Through the above Example 3, it was confirmed whether the derivative of the effective compound actually manufactured could inhibit the activity of protein A, which is a kinase, and also whether it could have an effect similar to or superior to the effective compound.
[0193] Specifically, 20 uM protein A was incubated with 8 mM MOPS (pH 7.0), 0.2 mM EDTA, 10 mM MgAcetate, and 10 uM [gamma-33P]-ATP, and additionally 10 uM YRRAAVPPSPSLSRHSSPHQSEDEEE, the target peptide for phosphorylation. After incubation at room temperature for 40 min, the reaction was stopped by adding phosphoric acid to a concentration of 0.5%. An aliquot of the reaction was spotted onto a filter, which was washed four times with 0.425% phosphoric acid for 4 min each, and then washed once with methanol before drying and scintillation counting (i.e., measurement of the YRRAAVPPSPSLSRHSSPHQS(p)EDEEE peptide).
[0194] The results are summarized in Tables 2 and 3 below. Specifically, Table 2 shows the structure and IC of each compound after confirming the inhibitory effect of the derivative compound actually manufactured in Example 3 on target protein A. 50 and pIC 50 This is a table that summarizes the numerical values, and Table 3 shows the predicted binding energy and pIC for each compound with protein A. 50 This is a table summarizing the results.
[0195] Compound No. SMILES Chemical Formula IC 50 pIC 50D00651O=C(NC1=CC=CC(C2=NNC(NC(C3=NC(C=CC=N4)=C4S3)=O)=C2)=C1)C(C)(C)CC21H20N6O2S0.1196.924D00832O=C(NC1=CC(C2=CC(NS(N3CCOCC3)(=O)=O)=CC=C2)=NN1)C4=CC=CC5=C4SC=N5C21H20N6O4S20.1296.889D01353C#CC(NC1=CC=CC(C2=NNC(NC(C3=CC4=C(O3)C=C(O)C=C4)=O)=C2)=C1)=OC21H14N4O40.0637.201D00254O=C(NC1=CC(C2=CC(NS(N3CCOCC3)(=O)=O)=CC=C2)=NN1)C4=CC5=C(S4)N=CC=N5C20H19N7O4S20.0627.208D01375C#CC(NC1=CC=CC(C2=NNC(NC(C3=CC4=C(S3)N=C(O)C=N4)=O)=C2)=C1)=OC19H12N6O3S0.0407.398D10106CC#CC(NC1=CC(C2=NNC(NC(C3=CC4=C(S3)CCNC4)=O)=C2)=CC=C1)=OC21H19N5O2S0.0317.509D00697O=C(NC1=CC=CC(C2=NNC(NC(C3=CC4=C(C=CC(OC)=C4)O3)=O)=C2)=C1)C#CCC23H18N4O40.0307.523D00548C#CC(NC1=CC=CC(C2=NNC(NC(C3=CC4=NON=C4C=C3)=O)=C2)=C1)=OC19H12N6O30.0267.585D01399C#CC(NC1=CC=CC(C2=NNC(NC(C3=CC4=C(O3)C=C(O)C=C4)=O)=C2)=C1)=OC21H14N4O40.0227.658D100810CC#CC(NC1=CC(C2=NNC(NC(C3=CC4=C(S3)CCOC4)=O)=C2)=CC=C1)=OC21H18N4O3S0.0197.721D010111C#CC(NC1=CC=CC(C2=NNC(NC(C3=CC4=C(N=CC=N4)S3)=O)=C2)=C1)=O.ClC19H13ClN6O2S0.0137.886DM01012C#CC(=O)Nc4cccc(c3cc(NC(=O)c2cc1ncc(C(N)=O)nc1s2)[nH]n3)c4C20H13N7O3S0.0108.000D101113CC#CC(NC1=CC (C2=NNC(NC(C3=CC(CN4)=C(S3)CC4=O)=O)=C2)=CC=C1)=OC21H17N5O3S0.0088.097DM00914C#CC(=O)Nc4cccc(c3cc(NC(=O)c2cc1 ncc(C(=O)O)nc1s2)[nH]n3)c4C20H12N6O4S0.0078.155D001115O=C(C1=CC(N=CC=N2)=C2S1)NC3=CC(C4=CC(NC(C#C)=O)=CC=C4)= NN3C19H12N6O2S0.0137.886D015716C#CC(NC1=CC=CC(C2=NNC(NC(C3=CC4=C(CCOC4)S3)=O)=C2)=C1)=OC20H16N4O3S0.0048.398.
[0196]
[0197] As shown in Tables 2 and 3 above, the pIC of each derivative 50 , and showed similar or superior protein A inhibition potency (lower pIC ) than compound X. 50 It was confirmed that many derivatives having numerical values were generated. In particular, as can be confirmed in the case of D0157 in Fig. 17, it was at the level of log square (e) higher than the existing effective compounds. 2 ) It was confirmed that a derivative showing excellent inhibitory efficacy could be produced.
[0198] As a result of modeling the binding form of a derivative substituted with a substituent expected to be the optimal substituent and binding to protein A, it was expected to have a binding form as shown in Figure 17.
[0199] From the above description, those skilled in the art will understand that the present disclosure can be implemented in other specific forms without altering the technical spirit or essential characteristics of the present disclosure. In this regard, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. The scope of the present disclosure should be interpreted as encompassing all changes or modifications derived from the meaning and scope of the claims described below, and their equivalent concepts, rather than the detailed description above.
[0200] As described above, the relevant contents have been described in the best form for carrying out the invention.
[0201] It can be used in the process of discovering new drug substances for building an artificial intelligence new drug platform.
Claims
1. A method for producing a derivative of an effective compound from an effective compound for a target protein, (A) A step of selecting a scaffold excluding a substitutable portion and a substitutable portion from the chemical structure of an effective compound; (B) a step of setting a target space within a target protein centered on a region where a substitution target portion of a selected effective compound binds; (C) A method comprising the step of generating a derivative by selecting a substituent that replaces a substitutable portion of an effective compound acceptable within the target space of a set target protein.
2. A method according to paragraph 1, wherein the method is for building an artificial intelligence new drug platform.
3. In paragraph 1, A method in which, in the above step (A), the portion to be substituted in the chemical structure of the effective compound is selected based on an indicator that the binding interaction is lower than that of other chemical structures from the interaction profile of the effective compound and the target protein.
4. In paragraph 3, The above binding interaction is a method in which bonds in an effective compound that interacts with a target protein are cleaved one by one, a substitution candidate portion including an atom of the cleaved portion is selected for each cleaved bond, and the binding energy of each atom constituting the substitution candidate portion with the target protein is calculated as an average value.
5. A method according to claim 4, wherein the substitution candidate portion comprises 1 to 12 atoms.
6. A method according to claim 5, further comprising a step of filtering the substitution candidate portion according to the number of constituent atoms.
7. In paragraph 1, A method in which the size of the target space within the target protein in the above step (B) is set to accommodate a substitution target portion of an effective compound selected from among chemical structures of effective compounds that bind to the target protein.
8. A method of setting, in the first paragraph, including a step of dividing a region of a target protein constituting a target space within the target protein in stages according to the interaction energy between target protein atoms existing in the region in the step (B).
9. A method according to claim 8, wherein the interaction energy is divided into three to five stages.
10. A method for selecting a clustering region of a target space to be used for generating a derivative of an effective compound according to the size of the clustering region and the proximity to the scaffold of the effective compound in the 8th paragraph, wherein the region is clustered in a stage where the interaction energy is relatively low.
11. In the 8th paragraph, the method of classifying the interaction energy between target protein atoms in the step (B) is performed by extracting the region of the target protein constituting the target space using a spatial filter, setting the marking points (dots) arranged at equal intervals in the spatial filter, and then classifying the marking points step by step according to the interaction energy between the target protein atoms.
12. In the 11th paragraph, the spatial filter is a spherical, rectangular, cylindrical, or amorphous filter.
13. In the 10th paragraph, the selection of the clustering region of the target space to be used for generating derivatives of the effective compound in the step (B) is (B1) A step of extracting the area of the target protein constituting the target space in the form of a cylinder filter, setting the marking points (dots) arranged at equal intervals on the cylinder filter, and then dividing the marking points step by step according to the interaction energy between the target protein atoms, and (B2) A method comprising the step of approaching a scaffold of an effective compound to a cylinder filter, excluding from the cylinder filter an area where the interaction energy is greater than a preset value, and then clustering the remaining marker points into spatial units.
14. In the 13th paragraph, a method further comprising a step of selecting some clustering regions as target spaces to be used for generating derivatives of effective compounds according to the proximity between the marker points in the clustered regions and the size of the clustering regions (B3).
15. A method according to claim 10, claim 13 or claim 14, wherein the clustering is performed based on the density of an area where the interaction energy is relatively low, or is a Gaussian Mixture Model (GMM) clustering.
16. In the first paragraph, the step (C) is a method performed by selecting a compound in which the substitution target portion that binds to the scaffold of the effective compound selected in the step (A) is substituted with a substituent selected from a database of substituent groups that are acceptable for the target space within the target protein set in the step (B).
17. In paragraph 1, In the above step (A), the portion to be substituted in the chemical structure of the effective compound is selected based on an indicator that the binding interaction is lower than that of other chemical structures from the interaction profile of the effective compound and the target protein. In the above step (B), selection of the target space to be used for generating derivatives of effective compounds among target proteins is (B1) A step of extracting the area of the target protein constituting the target space in the form of a cylinder filter, setting the marking points (dots) arranged at equal intervals on the cylinder filter, and then dividing the marking points step by step according to the interaction energy between the target protein atoms. (B2) A step of approaching the scaffold of the effective compound to the cylinder filter, excluding the area where the interaction energy is higher than a preset value from the cylinder filter, and then clustering the remaining marker points into spatial units. (B3) It is performed including a step of selecting some clustering areas based on the proximity between the display points in the clustered areas and the size of the clustering area, A method wherein the derivative production in step (C) is performed by selecting a substituent acceptable to the target space selected in step (B).
18. A method for producing a derivative having multiple different bonding structures for the same substituent by changing the bonding position within the substituent that binds to the scaffold of the effective compound in the 17th paragraph.
19. A method for producing a derivative having multiple different bonding forms for a substituent having the same bonding structure by changing the bonding angle of the scaffold of the substituent and the effective compound in the 17th paragraph.
20. A method for producing a derivative having multiple different bonding forms for a substituent having the same bonding structure by changing the bonding between the substituent and the scaffold of the effective compound in the 17th paragraph.
21. In the 17th paragraph, the binding interaction is a method in which bonds in an effective compound that interacts with a target protein are cleaved one by one, a substitution candidate portion including an atom of the cleaved portion is selected for each cleaved bond, and the binding energy of each atom constituting the substitution candidate portion with the target protein is calculated as an average value.
22. A method according to claim 1, further comprising the step of filtering the generated derivatives (D).
23. In paragraph 19, Additionally, (D) includes a step of filtering the generated derivatives, Step (D) above, A method of filtering the bond angles of the scaffold of the above substituent and the effective compound by comparing them with a database of existing substances.
24. In paragraph 19, Additionally, (D) includes a step of filtering the generated derivatives, Step (D) above, A method of filtering derivatives according to the amount of collision between atoms of the derivative and atoms of the target protein, generated within the target space set in step (B), according to the bonding form of the substituent of the derivative.
Citation Information
Patent Citations
Method of using a water-based pharmacophore
KR1020160128288A
Molten Salt Reactor
KR1020240047020A
System for multi-channel fluorescence detection having improved rotatable cylindrical filter-wheel
KR102879947B1