Method and apparatus for discovering active substance through information about three-dimensional binding structure of protein and ligand

By converting three-dimensional protein-ligand binding information into one-dimensional data for analysis, the method enhances the speed and accuracy of drug discovery by selecting compounds and determining optimal binding affinities, addressing the inefficiencies of traditional methods.

WO2025198106A1PCT designated stage Publication Date: 2025-09-25SYNTEKABIO INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/014292
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2024-09-23
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing drug discovery methods are slow and costly, and there is a need for more accurate and efficient computational analysis techniques to identify effective substances through three-dimensional protein-ligand binding structures.

Method used

A method and device that convert three-dimensional protein-ligand binding structural information into one-dimensional data for analysis, using bond energy and atom accessibility information to select compounds, screen for effective substances, generate conformers, and determine optimal poses and binding affinities through deep learning.

Benefits of technology

Improves processing speed and accuracy in identifying effective substances, enabling efficient screening and derivation of optimal binding structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024014292_25092025_PF_FP_ABST
    Figure KR2024014292_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and an apparatus for discovering an active substance through information about the three-dimensional binding structure of a protein and a ligand. An embodiment of the present invention provides a method comprising: a step for determining selected compounds among compounds to be analyzed, the selected compounds being determined on the basis of bond energy atom property information, generated from the three-dimensional binding structure of a protein to be analyzed and a ligand, and atom accessibility atom property information, generated from the three-dimensional structures of the compounds to be analyzed; a step for identifying active substances by screening the selected compounds on the basis of the ratio of key interaction property information, generated through three-dimensional docking (3D docking) of the protein to be analyzed and a positive control, to interaction property information, generated through three-dimensional docking (3D docking) of the protein to be analyzed and the selected compounds; a step for generating conformers on the basis of compound poses generated by docking the active substances with ligand derivatives; and a step for inputting the conformers to a deep learning model to determine the optimal pose and binding affinity of the compound with respect to the protein to be analyzed, wherein the bond energy atom property information and the atom accessibility atom property information are one-dimensional data compressed from the three-dimensional binding structure and the three-dimensional structure, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for discovering effective substances using three-dimensional protein-ligand binding structure information

[0001] The present disclosure provides a method and device for discovering effective substances through three-dimensional protein-ligand binding structure information.

[0002] Typically, new drug development is carried out through a process of candidate discovery and screening, followed by optimization, non-clinical / toxicity testing, and clinical trials. However, recently, computational analysis technologies (such as AI) are being applied to reduce the time and cost required for the discovery of new drug candidates.

[0003] With the recent advancement of computing technology, the demand for computation on analysis targets in the thousands or hundreds of billions is increasing, and this has led to the need for advanced methods that can replace existing screening methods in terms of speed and accuracy.

[0004] The background technology described above is technical information that the inventor possessed for the purpose of deriving the present invention or acquired in the process of deriving the present invention, and cannot necessarily be considered as publicly known technology disclosed to the general public prior to the application for the present invention.

[0005] The purpose of the present disclosure is to provide a method and device for discovering effective substances using three-dimensional protein-ligand binding structural information. The problems addressed by the present disclosure are not limited to those mentioned above, and other problems and advantages of the present disclosure not mentioned herein can be understood through the following description and will be more clearly understood through the embodiments of the present disclosure. Furthermore, it will be appreciated that the problems and advantages addressed by the present disclosure can be realized by the means and combinations thereof set forth in the claims.

[0006] A first aspect of the present disclosure comprises: a step of determining selected compounds among compounds to be analyzed based on bond energy atomic characteristic information generated from a three-dimensional binding structure of a target protein and a ligand, and atom accessibility atomic characteristic information generated from a three-dimensional structure of the compounds to be analyzed; a step of discovering effective substances by screening the selected compounds based on a ratio of key interaction characteristic information generated through three-dimensional docking of the target protein and a positive control to interaction characteristic information generated through three-dimensional docking of the target protein and the selected compound; a step of generating conformers based on compound poses generated by docking the effective substances with a ligand derivative; And a step of inputting the conformers into a deep learning model to determine the optimal pose and binding affinity of the compound for the protein to be analyzed; wherein the binding energy atomic characteristic information and the compound atomic accessibility atomic characteristic information can provide a method for discovering effective substances through the three-dimensional binding structure and the three-dimensional binding structure information between the protein and the ligand, which is one-dimensional data in which the three-dimensional structure is compressed.

[0007] A second aspect of the present disclosure comprises a memory storing at least one program; And a processor for performing a calculation by executing the at least one program; wherein the processor determines selected compounds among the compounds to be analyzed based on bond energy atomic characteristic information generated from a three-dimensional binding structure of a target protein and a ligand, and compound atom accessibility atomic characteristic information generated from a three-dimensional structure of the compounds to be analyzed, and discovers effective substances by screening the selected compounds based on a ratio of key interaction characteristic information generated through three-dimensional docking of the target protein and a positive control to interaction characteristic information generated through three-dimensional docking of the target protein and the selected compound, and generates conformers based on compound poses generated by docking the effective substances with a ligand derivative, and inputs the conformers into a deep learning model to determine the optimal pose and binding affinity of the compound for the target protein, wherein the bond energy atomic characteristic information The atomic characteristic information and the above compound atomic accessibility atomic characteristic information can provide a device for discovering effective substances through three-dimensional protein-ligand binding structure information, which is one-dimensional data.

[0008] A third aspect of the present disclosure can provide a computer-readable recording medium having recorded thereon a program for executing the method of the first aspect on a computer.

[0009] In addition, other methods for implementing the present invention, other devices, and computer-readable recording media recording a program for executing the method may be further provided.

[0010] Other aspects, features and advantages other than those described above will become apparent from the following drawings, patent claims and detailed description of the invention.

[0011] According to the problem solving means of the present disclosure described above, by converting three-dimensional structural information into one-dimensional based text information, the processing speed can be improved in terms of analysis speed, and in terms of accuracy, an accuracy level similar to that of screening by three-dimensional structural analysis can be implemented.

[0012] In addition, the problem-solving means of the present disclosure can improve the derived effective substance and produce various derivatives by increasing the accuracy of the optimal binding structure and providing an efficient screening method.

[0013] Figure 1 is an exemplary diagram of an environment for discovering an effective substance according to one embodiment of the present invention.

[0014] Figure 2 is an exemplary diagram of an effective substance discovery platform according to one embodiment of the present invention.

[0015] FIG. 3 and FIG. 4 are diagrams for explaining a method for generating binding energy atomic characteristic information according to one embodiment of the present invention.

[0016] FIG. 5 is a diagram for explaining a method for generating atomic accessibility atomic characteristic information according to one embodiment of the present invention.

[0017] FIG. 6 and FIG. 7 are diagrams for explaining a method for determining an optimization limit value according to one embodiment of the present invention.

[0018] FIG. 8 is a drawing for explaining a method for determining selection compounds according to one embodiment of the present invention.

[0019] FIG. 9 is a drawing for explaining a method for discovering an effective substance according to one embodiment of the present invention.

[0020] Figure 10 is an exemplary diagram of a ligand derivative according to one embodiment of the present invention.

[0021] FIGS. 11 to 12 and FIGS. 13a to 13c are drawings for explaining a method for producing a ligand derivative according to one embodiment of the present invention.

[0022] FIGS. 14 and 15a to 15b are drawings for explaining a method of filtering a derivative according to one embodiment of the present invention.

[0023] FIGS. 16 to 18 are drawings for explaining a method for selecting a representative pose according to one embodiment of the present invention.

[0024] FIG. 19 is a drawing for explaining a method for determining an optimal pose according to one embodiment of the present invention.

[0025] Figure 20 is a drawing for explaining the characteristics of a bonding portion of an effective material according to one embodiment of the present invention.

[0026] FIG. 21 is a drawing for explaining a method of filtering an effective substance according to one embodiment of the present invention.

[0027] Figure 22 is a flowchart of a method for discovering effective substances through three-dimensional protein-ligand binding structure information according to one embodiment of the present invention.

[0028] Figure 23 is a block diagram of a device for discovering effective substances through three-dimensional protein-ligand binding structure information according to one embodiment of the present invention.

[0029] A first aspect of the present disclosure comprises: a step of determining selected compounds among compounds to be analyzed based on bond energy atomic characteristic information generated from a three-dimensional binding structure of a target protein and a ligand, and atom accessibility atomic characteristic information generated from a three-dimensional structure of the compounds to be analyzed; a step of discovering effective substances by screening the selected compounds based on a ratio of key interaction characteristic information generated through three-dimensional docking of the target protein and a positive control to interaction characteristic information generated through three-dimensional docking of the target protein and the selected compound; a step of generating conformers based on compound poses generated by docking the effective substances with a ligand derivative; And a step of inputting the conformers into a deep learning model to determine the optimal pose and binding affinity of the compound for the protein to be analyzed; wherein the binding energy atomic characteristic information and the compound atomic accessibility atomic characteristic information can provide a method for discovering effective substances through the three-dimensional binding structure and the three-dimensional binding structure information between the protein and the ligand, which is one-dimensional data in which the three-dimensional structure is compressed.

[0030] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments presented below, but can be implemented in various different forms, and it should be understood that it includes all transformations, equivalents, and substitutes included in the spirit and technical scope of the present invention. The embodiments presented below are provided to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the invention of the scope of the invention. In describing the present invention, if a detailed description of a related known technology is judged to obscure the gist of the present invention, the detailed description thereof will be omitted.

[0031] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0032] Some embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a given function. Furthermore, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented by algorithms that execute on one or more processors. Furthermore, the present disclosure may employ conventional techniques for electronic configuration, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.

[0033] Additionally, the connecting lines or connecting members between components depicted in the drawings are merely exemplary representations of functional connections and / or physical or circuit connections. In an actual device, connections between components may be represented by various functional connections, physical connections, or circuit connections that may be replaced or added.

[0034] The present disclosure will be described in detail with reference to the attached drawings below.

[0035] Figure 1 is an exemplary diagram of a system for discovering effective substances according to one embodiment of the present invention.

[0036] In one embodiment, the environment for discovering effective substances includes a system (1) for discovering effective substances, a device (hereinafter referred to as “device”) (10) for discovering effective substances, and a server (20). For example, the device (10) and the server (20) may be connected via wired or wireless communication to transmit and receive various data between them.

[0037] For convenience of explanation, FIG. 1 illustrates a system (1) including a device (10) and a server (20), but the present invention is not limited thereto. For example, the system (1) may include other external devices (not shown). In addition, the operations of the device (10) and the server (20) described below may be implemented by a single device (e.g., the device (10) or the server (20)) or by more devices.

[0038] The device (10) may be a computing device including a display device and a device for receiving user input (e.g., a keyboard, a mouse, etc.), and including a memory and a processor. In addition, the display device may be implemented as a touch screen and may perform a function for receiving user input. For example, the device (10) may be, but is not limited to, a notebook PC, a desktop PC, a laptop, a tablet computer, a smart phone, etc.

[0039] The server (20) may be a device that communicates with an external device (not shown) that includes the device (10). For example, the server (20) may be a device that stores various data, including known compounds such as ligands, conformers, proteins, and substituents, and characteristic information such as the structure and composition of the compounds. Alternatively, the server (20) may be a computing device that includes a memory and a processor and has its own computing capabilities. If the server (20) is a computing device, the server (20) may perform at least some of the operations of the device (10) or the effective substance discovery platform, which will be described later with reference to FIGS. 1 to 23. For example, the server (20) may be a cloud server, but is not limited thereto.

[0040] The device (10) can output information generated by the discovery of effective substances and provide it to the user (30). In addition, a report including the information can be output and provided to the user (30).

[0041] Meanwhile, for convenience of explanation, the device (10) is described as performing all operations throughout the specification, but this is not limited to this. For example, at least some of the operations performed by the device (10) may also be performed by the server (20).

[0042] Figure 2 is an exemplary diagram of an effective substance discovery platform according to one embodiment of the present invention.

[0043] Referring to Figure 2, when cloud users register a project and apply for use, the administrator (or management algorithm) reviews and approves the project and allocates resources for analysis. The allocated resources may include not only physical machine resources but also various parameters for executing automated workflows for analysis.

[0044] Specifically, when a user first creates an account, they can be granted guest user privileges and access only the billing page. Furthermore, this page allows users to view billing information and download documents for cloud service use. General users who have signed a formal user agreement can freely use the platform.

[0045] After reviewing the user's platform usage, the project creation page allows users to register a project by entering a simple title and description, then selecting the platform to use. The administrator verifies that the project is contracted and approves its use. Users then enter the parameters for the pipeline used for analysis online. The website for executing this project is designed to facilitate analysis for users unfamiliar with cloud environments.

[0046] In the case of websites that provide existing cloud environments, the analysis must be performed after the types and quantities of resources to be used and the workflows to be selected are written and selected according to the purpose of the analysis. However, the cloud environment of the present invention can be used by utilizing only simple parameters for pre-determined services.

[0047] Meanwhile, once the analysis begins, you can check the progress and detailed progress stages of the project on the project page, and when it is completed, you can go to the report page and view the analyzed data in various visual formats, as well as download a web application and documented analysis data in various formats.

[0048] The administrator approves the equipment required for the project based on the type of platform for the project registered by the user. It includes a resource request for executing the pipeline included in the project requested by the user, and a dynamic CPU, GPU, and storage resource management function is required to specify an appropriate quantity of resources and work storage for the request and to perform the workflow by utilizing optimal resources for each step. In the present invention, Argo's workflow management function and Kubernetes' application deployment function and container orchestration function are applied for this purpose.

[0049] When a user executes a project analysis after setting parameters, the platform sequentially executes according to the specified template and generates a report upon completion. The platform process, which sequentially executes according to the specified template, is described in detail in Figure 3 and below.

[0050] Users can download report results through the portal. After reviewing the analysis results, users can request a data backup from the administrator. The administrator then automatically migrates the data from the working storage to the backup storage using a pre-designed backup template. Report results are generated so that users can easily view and understand the analysis results required to discover new candidate substances. Using Docker container technology, an independent web service virtualization environment can be established for each project upon completion of the analysis. For the report web service, MongoDB can be introduced for fast data access and large-scale data processing. The website can also be configured as a single-page application, similar to a native app, enabling users to efficiently navigate and understand information.

[0051] Subsequently, data backups can be performed to safely store and deliver report results at the user's request, and storage resources can be recovered after the work is completed. This process can incorporate technologies for rapid report backup and policies to prevent loss and alteration of results.

[0052] Meanwhile, in the present invention, the administrator refers to an entity that manages the operation of the platform, and may be an operating personnel or an electronic system such as an operating program or an operating server.

[0053] Below, the detailed technical contents of the cloud platform and service method according to the present invention will be described in the order of the attached drawings.

[0054] FIG. 3 and FIG. 4 are diagrams for explaining a method for generating binding energy atomic characteristic information according to one embodiment of the present invention.

[0055] In one embodiment, the device can generate binding energy atomic characteristic information based on the three-dimensional binding structure of the target protein and the ligand. The ligand refers to a substance that specifically binds to the target protein.

[0056] At this time, the device can calculate the binding energy between each atom based on the core atom of the protein to be analyzed, the three-dimensional arrangement information of the atoms around the core atom of the protein, the core atom of the ligand, and the three-dimensional arrangement information of the atoms around the core atom of the ligand. As an example, the device can calculate the binding energy of the atoms constituting the ligand for each atom constituting the pocket of the protein to be analyzed using the three-dimensional structural information of the protein to be analyzed. In other words, the device can define the interaction between the protein to be analyzed and the ligand as one analysis unit based on the core atom pair of the ligand and the protein used to calculate the binding energy.

[0057] In one embodiment, the device can classify the type of core atoms of a ligand that constitute the interaction between the target protein and the ligand. For example, the device can classify the atoms constituting the ligand into N, O, and C. As an example, the device can replace Cl, S, P, F, and Br, which are classified as non-mainstream atoms, with C for a ligand composed of a combination of C, Cl, S, P, F, Br, N, and O. Accordingly, the ligand can be defined in one dimension (text) as in Mathematical Formula 1.

[0058]

[0059] However, depending on the method of substituting the atoms of the ligand, the ligand may be defined as [C, N, O, S, M], and is not limited to the examples described above.

[0060] Thereafter, the device can calculate the binding energy corresponding to the core atoms of the classified ligand. The device can calculate the binding energy between the above-mentioned atoms using, but is not limited to, AMBER, CHARMM, OPLS, MM / PBBSA, MM / GBSA, enva (Syntekabio), etc. For example, the device can calculate the binding energy of each of the core atoms of the ligand classified as any one of N, O, and C to the atoms constituting the pocket of the protein to be analyzed.

[0061] Afterwards, the device can map the core atoms of the ligand and the binding energies calculated for each core atom of the ligand and create a database. This is defined as binding energy atomic characteristic information.

[0062] At this time, the device can generate binding energy atomic characteristic information including the seed type of the atom in order to reflect the three-dimensional binding structure of the ligand with the target protein of analysis in the binding energy atomic characteristic information. For example, referring to FIG. 4, seed types 1 to 5 of core atoms are illustrated. At this time, the device can generate binding energy atomic characteristic information including the seed type of the core atom constituting the pocket of the target protein of analysis and the seed type of the core atom of the ligand bonded to the core atom of the target protein of analysis. At this time, the seed type can be expressed as text such as 1, 2, 3, 4, and 5. Meanwhile, the binding energy atomic characteristic information can further include information such as interatomic bond angle, charge, and type of bond (as an example, textualized one-dimensional information).

[0063] In one embodiment, the device can create a one-dimensional database of binding energy atomic characteristic information. As described above, a ligand can be defined one-dimensionally using core atoms, and the device can create a one-dimensional data structure of binding energy atomic characteristic information by mapping and database the binding energy corresponding to each core atom. That is, the binding energy atomic characteristic information can be a one-dimensional database of binding energies of atoms constituting a ligand for atoms constituting a pocket of a target protein. As an example, the binding energy characteristic information can be created in a format such as (-2-1-NOC). In this case, the number can correspond to the seed type of the atom.

[0064] By databaseizing binding energy atomic characteristic information into text-based one-dimensional data, it is possible to perform analysis of three-dimensional structures at high speed.

[0065] FIG. 5 is a diagram for explaining a method for generating atomic accessibility atomic characteristic information according to one embodiment of the present invention.

[0066] In one embodiment, the device can generate compound atomic accessibility atomic property information based on the three-dimensional structure of the compounds being analyzed. The compounds being analyzed are compounds being screened and may number from several thousand to several hundred.

[0067] In one embodiment, the device can extract three-dimensional structural information of a target compound from a complex bound to a target protein and calculate the atom accessibility of each atom constituting the extracted target compound to a solvent. The device can calculate the atom accessibility between the atoms described above using, but is not limited to, a molecular dynamics tool, enva (Syntekabio), machine learning, etc.

[0068] In the process of generating compound atomic accessibility atomic characteristic information by calculating atomic accessibility, classifying atoms constituting the target compound, defining the target compounds as one-dimensional data, and mapping the atoms constituting the target compound and the calculated atomic accessibility, any description overlapping with the process of generating binding energy atomic characteristic information will be omitted. However, the compound atomic accessibility atomic characteristic information can be generated to have the same data structure as the binding energy atomic characteristic information. Meanwhile, it should be noted that the compound atomic accessibility atomic characteristic information, unlike the binding energy atomic characteristic information, is information that reflects the atomic accessibility to a solvent calculated for each atom using only the three-dimensional structure of the compound. In other words, the compound atomic accessibility atomic characteristic information can be a one-dimensional database of the interaction tendencies of atoms constituting target compounds toward a solvent.

[0069] Similar to the binding energy atomic property information, in one embodiment, the device can create a one-dimensional database of compound atomic accessibility atomic property information. As described above, the compound to be analyzed can be defined one-dimensionally using core atoms, and the device can create atomic accessibility atomic property information in a one-dimensional data structure by mapping the atomic accessibility corresponding to each core atom and creating a database. In other words, the compound atomic accessibility atomic property information can be one-dimensional data in which the interaction tendency of the compound to be analyzed for a solvent is compressed.

[0070] In one embodiment, the device can generate ligand atom accessibility atomic property information based on the three-dimensional structure of the ligand. The ligand atom accessibility atomic property information for a ligand that specifically binds to a target protein can be generated in the same manner as the compound atom accessibility atomic property information for the target compound described above. That is, the device can generate ligand atom accessibility atomic property information by database-based on the interaction tendencies of the atoms constituting the ligand with respect to a solvent from the three-dimensional structure of the ligand.

[0071] Meanwhile, in one embodiment, the device can generate ligand characteristic information by databaseing the physical and chemical properties of the ligand from the three-dimensional binding structure of the target protein and the ligand. For example, the device can extract three-dimensional structural information of the ligand from the complex in which the target protein and the ligand are bound, and calculate the physical and chemical properties of the ligand from the extracted three-dimensional structural information of the ligand. In other words, the ligand characteristic information may be a database of the physical and chemical properties of the ligand calculated from the three-dimensional binding structure of the target protein and the ligand. As an example, the device can calculate the physical and chemical properties of the ligand using rdkit, but is not limited thereto, and the device can generate ligand characteristic information using various methods or tools.

[0072] Specifically, the ligand characteristic information may include at least one physical characteristic and / or chemical characteristic among the number of rings in the ligand molecule, the number of atoms of each constituent element of the molecule, molecular weight, charge, size of the molecule, partition coefficient of the molecule, pKa constant, number of hydrogen bond acceptor atoms, number of hydrogen bond donor atoms, polar surface area, and the ratio of sP3 carbon atoms.

[0073] Likewise, in one embodiment, the device can generate compound characteristic information by database the physical and chemical characteristics of the analyte compound from the three-dimensional bonding structure of the analyte protein and the analyte compound. At this time, the compound characteristic information may include at least one of the number of rings in the molecule of the analyte compound, the number of atoms of each constituent element of the molecule, the molecular weight, the charge, the size of the molecule, the partition coefficient of the molecule, the pKa constant, the number of hydrogen bond acceptor atoms, the number of hydrogen bond acceptor atoms, the polar surface area, and the ratio of sP3 carbon atoms. In the embodiment in which the device generates compound characteristic information, the description overlapping with the embodiment in which the ligand characteristic information is generated will be omitted.

[0074] In one embodiment, the device can determine at least some of the target compounds as selectable compounds based on an optimization constraint value.

[0075] FIG. 6 and FIG. 7 are diagrams for explaining a method for determining an optimization limit value according to one embodiment of the present invention.

[0076] In one embodiment, the device can determine a binding energy optimization limit value (610) and an atomic accessibility optimization limit value (620). The binding energy optimization limit value (610) is a reference value for extracting a binding energy characteristic profile from binding energy atomic characteristic information, and the atomic accessibility optimization limit value (620) can be a reference value for extracting an atomic accessibility characteristic profile from compound atomic accessibility atomic characteristic information. The device can ultimately determine at least some of the analysis target compounds as selected compounds through a similarity analysis between the extracted binding energy characteristic profile and the atomic accessibility characteristic profile using the binding energy optimization limit value (610) and the atomic accessibility optimization limit value (620). Hereinafter, an embodiment of determining the selected compounds will be described in detail.

[0077] First, in one embodiment, the device can search for an optimal combination of binding energy atomic characteristic information and ligand atom accessibility atomic characteristic information. As an example, the device can calculate the similarity between the binding energy atomic characteristic information and the ligand atom accessibility atomic characteristic information according to any binding energy value and any atom accessibility value within a predetermined range. At this time, the device can calculate the similarity according to various embodiments such as chi analysis, cosine similarity, tanimoto algorithm, Euclidean distance, etc., but is not limited thereto.

[0078] Referring to FIG. 6, an embodiment of determining a binding energy optimization limit value (610) and an atomic accessibility optimization limit value (620) using the tanimoto algorithm is illustrated.

[0079] In one embodiment, the device can transform the binding energy atomic property information and the ligand atom accessibility atomic property information into a set of binding energy IDs and a set of ligand atom accessibility IDs, which are binary vectors. For example, if the binding energy atomic property information or the ligand atom accessibility atomic property information is as shown in the following mathematical expression 2, the device can It can be converted into a binary vector such as .

[0080]

[0081] In one embodiment, the device can calculate a value obtained by dividing the number of elements belonging to the intersection of the binding energy ID set (BE_ID_set) and the ligand atom accessibility ID set (ligand_AA_ID_set) by the number of elements belonging to the union of the two ID sets. The calculation formula is as shown in Mathematical Formula 3 below, and the calculated value is defined as a similarity score.

[0082]

[0083] In one embodiment, the device can calculate the binding energy value and the atomic accessibility value that result in the highest similarity score. That is, the device can search for an optimal combination of the binding energy atomic property information and the ligand atomic accessibility atomic property information that maximizes the number of elements included in both the binding energy atomic property information and the ligand atomic accessibility atomic property information. As an example, as illustrated in FIG. 6, the device can search the binding energy atomic property information in the range of [-4.00, 0.00] at an interval of -0.25 to calculate the binding energy value that result in the highest similarity score. Furthermore, the device can search the ligand atomic accessibility atomic property information in the range of [0, 60] at an interval of 5 to calculate the atomic accessibility value that result in the highest similarity score. However, the device can determine the range and the search interval of the binding energy atomic property information and the ligand atomic accessibility atomic property information according to various embodiments including, but not limited to, machine learning such as grid search, random forest, and deep learning.

[0084] In one embodiment, the device may determine the binding energy value obtained as a search result as a binding energy optimization limit value (610), and may determine the obtained atomic accessibility value as an atomic accessibility optimization limit value (620). According to the embodiment illustrated in FIG. 6, the device may determine the binding energy optimization limit value (610) as -0.25, and the atomic accessibility optimization limit value (620) as 15.

[0085] Meanwhile, there may be cases where the number of binding energy atomic characteristic information and ligand atom accessibility atomic characteristic information according to the determined optimization constraints is insufficient. In this case, there is a risk that the similarity, which will be described later, will be calculated high despite the lack of information for defining the pocket environment of the target protein. Therefore, the device may exclude binding energy atomic characteristic information and / or ligand atom characteristic information (700) where the number of constituent atoms is below a specific value from the step of determining the optimization constraints. This process is illustrated in Figure 7.

[0086] In one embodiment, the device can generate a binding energy characteristic profile in which ligand characteristic information and binding energy atomic characteristic information are combined based on a binding energy optimization limit value (610). Specifically, the device can generate a binding energy characteristic profile by extracting only atoms that interact with each other less than or equal to the binding energy optimization limit value (610) among the atoms constituting the binding energy atomic characteristic information, and combining the extracted atoms with the ligand characteristic information corresponding to the extracted atoms.

[0087] Likewise, in one embodiment, the device can generate an atomic accessibility characteristic profile by combining compound characteristic information and compound atomic accessibility atomic characteristic information based on the atomic accessibility optimization limit value (620). Specifically, the device can generate an atomic accessibility characteristic profile by extracting only atoms that interact with each other greater than or equal to the atomic accessibility optimization limit value (620) among the atoms constituting the compound atomic accessibility atomic characteristic information, and combining the extracted atoms with the compound characteristic information corresponding to the extracted atoms.

[0088] FIG. 8 is a drawing for explaining a method for determining selection compounds according to one embodiment of the present invention.

[0089] In one embodiment, the device can determine selected compounds among the target compounds by analyzing the similarity between the binding energy characteristic profile and the atomic accessibility characteristic profile.

[0090] Figure 8 illustrates an embodiment in which the device determines selected compounds through chi analysis between binding energy characteristic profiles and atomic accessibility characteristic profiles. In one embodiment, the device can calculate a characteristic profile similarity score based on a p-value through chi analysis, as shown in Equation 4 below.

[0091]

[0092] Here, c is the degrees of freedom, is the observed value, is the expected value.

[0093] Similarity analysis like this can predict the likelihood of interaction between the target protein and the target compound. Specifically, a device with a lower characteristic profile similarity score exhibits greater differences in the atomic compositions that make up its atomic accessibility and binding energy characteristic profiles, indicating a lower likelihood of interaction with the target protein.

[0094] Accordingly, in one embodiment, the device can determine at least some of the target compounds as selected compounds by excluding those with low characteristic profile similarity scores based on a predetermined limit value. According to the process for determining selected compounds described in FIGS. 3 to 8, compounds can be quickly selected from thousands to hundreds of millions of target compounds using only one-dimensional compressed data while reducing computational load.

[0095] FIG. 9 is a drawing for explaining a method for discovering an effective substance according to one embodiment of the present invention.

[0096] In one embodiment, the device can generate interaction characteristic information (970) through three-dimensional docking (912) of the target protein (910) and the selection compound (920). At this time, the interaction characteristic information (970) may be a one-dimensional database of each of the interactions between the atoms constituting the target protein (910) and the atoms constituting the selection compound (920). Unlike the binding energy atomic characteristic information and the ligand / compound atomic accessibility atomic characteristic information, the interaction characteristic information (970) itself between the atoms constituting the target protein (910) and the atoms constituting the selection compound (920) is databased, so that an atom-atom pair can become a field of data.

[0097] As an example, interaction characteristic information (970) may be generated as 1_3_N4_C3_C4-G_N_CA_C. In this case, 1_3_N4_C3_C4 may correspond to the seed type of the ligand, N4_C3_C4 may correspond to the fragment type of the ligand (a substructure of a molecule exhibiting specific chemical characteristics), and G_N_CA_C may correspond to the fragment type of the protein to be analyzed. As another example, the interaction characteristic information (970) may further include interatomic bond angles.

[0098] In addition to the above description, any description (e.g., one-dimensional data structure) that overlaps with the example of generating binding energy atomic characteristic information and / or ligand / compound atomic accessibility atomic characteristic information in the method of generating interaction characteristic information (970) will be omitted.

[0099] In one embodiment, the device can generate key interaction characteristic information (960) through three-dimensional docking (913) of the target protein (910) and the positive control (930). At this time, the key interaction characteristic information (960) may be an interaction with high significance extracted from among the interactions between atoms constituting the target protein (910) and atoms constituting the positive control (930).

[0100] Specifically, the device can generate positive control atomic characteristic information for atoms constituting the positive control (930) through three-dimensional docking (913) of the target protein (910) and the positive control (930). At this time, the positive control (930) may refer to substances known to bind to the target protein (910). The positive control (930) may be obtained by extracting 127 ligands of PLK4 from the Chembl database and forming a conformer structure for each using rdkit, but is not limited thereto. That is, the device can generate atomic characteristic information for each of the poses (hereinafter referred to as “reference poses”) formed when the positive control (930) is three-dimensionally docked to the target protein (910). This is defined as key-atom characteristic information (940).

[0101] Thereafter, the device can calculate the frequency of occurrence (951) of the key-atom characteristic information (940) for the set of target proteins (910). That is, the device can calculate the number of atoms interacting with the target proteins (910) for the number of target proteins (910) as the frequency of occurrence (951). For example, the frequency of occurrence (951) can be calculated as in the following mathematical expression 5.

[0102]

[0103] is the number of data included in the key-atom characteristic information (940). is the number of positive controls (930).

[0104] In addition, the device can calculate the energy frequency (952) of the reference poses for the set of the target protein (910) for analysis by using the binding energy calculated from the binding structure of the target protein (910) and the reference poses as a weight. That is, the device can calculate the value obtained by multiplying the number of atoms interacting with the target protein (910) for the number of the target protein (910) for analysis by the binding energy of the corresponding interaction as the energy frequency (952). Accordingly, it is possible to prevent interactions that do not strongly bind to the target protein (910) for analysis from being included in the key-interaction characteristic information (960) and disturbing the data. For example, when BE is the binding energy, the energy frequency ( )(952) can be calculated as shown in mathematical formula 6 below.

[0105]

[0106] In one embodiment, the device can screen the selection compounds (920) based on the appearance frequency (951) and the energy frequency (952). As an example, the device can set an expected value for the presence of a corresponding interaction in the selection compound (920) based on the appearance frequency (951), and determine a value obtained by multiplying the expected value by an appropriate standard of the binding energy of each interaction as a reference value. At this time, the device can add an interaction having a lower energy frequency (952) than the reference value to the key-interaction characteristic information (960). More specifically, assuming that the number of positive controls (930) is 100 and each pose has 20 pieces of key-atom characteristic information (940), it can be seen that the expected value of the appearance frequency (951) of the key-atom characteristic information (940) present in a certain pose in the positive control group (930) is 20 / 100. Therefore, when the appropriate standard of binding energy having a certain importance in the interaction is set to -0.5, the device can determine -0.1, which is obtained by multiplying 0.2, which is the expected value of the previously calculated appearance frequency (951), as the reference value. Accordingly, the device can determine only the interactions having an energy frequency (952) smaller than -0.1 among the interactions between the atoms constituting the analysis target protein (910) and the atoms constituting the positive control group (930), as key-interaction characteristic information (960). However, the embodiment in which the device determines key-interaction characteristic information (960) is not limited thereto.

[0107] In one embodiment, the device can discover effective substances (980) by screening selected compounds (920) based on the ratio of key-interaction characteristic information (960) to interaction characteristic information (970). As an example, based on how much of the key-interaction characteristic information (960) is included in the generated interaction characteristic information (970) for any selected compound (920), if the selected compound (920) includes a ratio that is equal to or greater than a preset ratio, the selected compound (920) can be determined as an effective substance (980).

[0108] The method for discovering effective substances according to the aforementioned embodiment can analyze the possibility of a selected compound (920) interacting well with a target protein (910) by calculating possible types of interactions for docking structures for multiple positive controls (930) and extracting substances with a high proportion of similar interactions. Accordingly, screening can be performed even when a docking pose to be used as a reference pose does not exist.

[0109] In one embodiment, the device can generate conformers based on compound poses generated by docking effective molecules (980) with ligand derivatives.

[0110] Figure 10 is an exemplary diagram of a ligand derivative according to one embodiment of the present invention.

[0111] In one embodiment, the ligand derivative (1000) may be formed using a ligand structure (1010 to 1040) bound to a template protein as a template. In this case, the ligand derivative (1000) may be formed to include all compound binding sites within the template protein. The device can generate compound poses by three-dimensionally docking the ligand derivative (1000) with effective substances. That is, by using the ligand derivative (1000) as an input for three-dimensional docking, the accuracy of the optimal pose of the compound can be increased.

[0112] Below, the method by which the device generates a ligand derivative is described.

[0113] FIGS. 11 to 12 and FIGS. 13a to 13c are drawings for explaining a method for producing a ligand derivative according to one embodiment of the present invention.

[0114] In one embodiment, the method by which the device generates a ligand derivative may consist of the following three steps.

[0115] (A) A step of selecting a scaffold excluding the substitutable portion and the substitutable portion from the chemical structure of the effective compound.

[0116] (B) A step of setting a target space within the target protein centered on the region where the substitution target portion of the selected effective compound binds.

[0117] (C) A step of generating a derivative by selecting a substituent that replaces a substitutable portion of an acceptable effective compound within the target space of the set target protein.

[0118] FIG. 11 and FIG. 12 are drawings for explaining a method of selecting a scaffold to produce a ligand derivative according to one embodiment of the present invention.

[0119] Referring to FIG. 11, in one embodiment, the device can generate an interaction profile between an active substance and a target protein when the active substance is bound to the target protein. The interaction profile can be obtained using, but is not limited to, enva (Syntekabio), FEP+, MMPBSA, etc.

[0120] Referring to FIG. 12, in one embodiment, the device can cleave one of the bonds forming the active compound that binds to the target protein through the acquired interaction profile, thereby generating molecular fragments on either side of the cleavage site. This cleavage and generation of molecular fragments is performed across the entire single bond of the active compound, and may yield twice as many molecular fragments as there are single bonds.

[0121] In one embodiment, the device can select only molecular fragments containing atoms within a preset range from among the generated molecular fragments. This is to ensure that the device selects only molecular fragments that are valuable as replacement targets due to their minimal impact on the binding between the active substance and the target protein. In other words, if the number of atoms constituting the molecular fragment is too small, the possibility of deriving a new compound through replacement is low. Conversely, if the number of atoms constituting the molecular fragment is too large, the properties of the active substance after replacement may significantly change. In this case, the active substance with the molecular fragment cut off is called a scaffold.

[0122] Meanwhile, the device can select an atom on the scaffold that binds to the substitution target portion in the chemical structure of an effective compound that binds to a target protein as an anchor atom. Specifically, the device can calculate the interaction efficiency for the selected molecular fragments, as illustrated in Figure 12, and select a cleavage portion of the molecular fragment with a low influence on binding interaction as an anchor. At this time, the interaction efficiency can be calculated as the average value of the binding energy of each atom that constitutes the molecular fragment.

[0123] In one embodiment, the device can calculate the space within the pocket when the scaffold from which the molecular fragments have been removed binds within the pocket of the target protein.

[0124] As an example, referring to FIG. 13a, the device can generate a cylindrical region filter by extracting a cylindrical region of a preset size (e.g., length 10 Å, radius 10 Å) centered on the substitution target portion or anchor portion of the effective compound of the target protein for analysis. At this time, the size of the region filter can be changed depending on the structure of the scaffold and the resources of the computational system, but can be set so as to sufficiently cover the region of the scaffold that interacts with the target protein for analysis.

[0125] Meanwhile, the cylinder region filter is configured with equally spaced marker points (dots), and the marker points (dots) are distinguished by their interaction energy with target protein atoms. In Fig. 13, the marker points included in the cylinder region filter are distinguished by color according to their interaction energy with target protein atoms.

[0126] Meanwhile, the cylindrical region filter is configured with equally spaced marker points (dots), and the marker points can be distinguished by their interaction energy with atoms of the target protein. As an example, FIG. 13a illustrates an example in which marker points included in the cylindrical region filter are distinguished by color according to their interaction energy with atoms of the target protein.

[0127] Referring to FIG. 13B, in one embodiment, the device can be positioned so that the anchor portion of the scaffold approaches the anchor portion of the cylindrical area filter (see (A) and (B)). At this time, the scaffold bonding axis direction can maintain its original bonding direction.

[0128] Thereafter, the device can exclude from the cylindrical region filter the display points belonging to the region with large interaction energy among the display points, and leave only the display points belonging to the region with small interaction energy among the display points (see (C)).

[0129] Thereafter, the device can cluster the remaining marker points into regions by spatial unit (see (D)). At this time, the device can extract only at least some clustering regions among the clustered regions based on their proximity to the scaffold anchors and the sizes of the clustered regions. As an example, the top few (e.g., two) clustering regions with the largest sizes among the clustering regions connected to the scaffold anchors can be selected (see (E)).

[0130] In this way, the reason for deriving the target space within the protein to be analyzed is to select a substituent (replace atom group) to replace the target part according to the size and shape of the target space (see (F)).

[0131] Referring to FIG. 13C, in one embodiment, the step of generating a derivative as in step (C) described above by the device includes selecting a substituent that matches the size of the target space within the target protein to be analyzed generated in step (B). In this case, the substituent is a group of atoms that replaces the target portion of the molecule removed from the scaffold, and can be selected from a database storing various atomic configurations.

[0132] Thereafter, the device can generate a derivative by binding the selected substituent to the anchor of the scaffold. At this time, the device can generate a plurality of derivatives having different bonding structures for the same substituent by varying the bonding position within the substituent that binds to the scaffold of the effective compound. That is, as illustrated in Fig. 13c, derivatives having various bonding structures can be generated by varying the bonding position to the anchor atom of the scaffold within a single selected substituent. In addition, even for a substituent having the same bonding structure, derivatives having a plurality of different bonding forms can be generated by varying the bonding angle between the substituent and the scaffold of the effective compound.

[0133] In this case, the device can extract and generate a linker from the anchor portion of the target protein. The linker extracts only the portion proximal to the anchor, and can be configured to extract only up to a preset number of linking atoms from the anchor. In this way, the device can generate derivatives with various linking forms by diversifying the bonding configurations of the linker and substituent.

[0134] In one embodiment, the device can filter the binding forms of the anchor portions of the generated derivatives to exclude derivatives having binding forms with a low probability of existence.

[0135] FIGS. 14 and 15a to 15b are drawings for explaining a method of filtering a derivative according to one embodiment of the present invention.

[0136] Referring to FIG. 14, in one embodiment, the device can perform bond filtering, which filters the bonding forms of bonding groups and substituents based on their likelihood of existence in existing materials.

[0137] Referring to FIG. 15a, the device can use a method of comparing the bonding structures of existing substances in a database of existing compounds and proteins (e.g., ChEMBL) as a method of determining the possibility of the existence of various bonding forms of the bonding group and substituent in existing substances. Specifically, the device can select bonding structures having the same composition as the bonding group and substituent from the compound database and analyze their bonding forms. As a result of the analysis, the device can select only derivatives in which the bonding group and substituent are combined in bonding forms of existing substances.

[0138] Referring to FIG. 15b, in one embodiment, the device can select (filter) derivatives based on whether they collide and the amount of collision within the target space by considering multiple binding forms generated by rotating the substituent of the derivative bound to the target protein at a certain angle.

[0139] In one embodiment, the device can input conformers obtained as a result of docking of a ligand derivative and an effective substance generated by the aforementioned method into a deep learning model to determine the optimal pose of the compound, i.e., the effective substance, for the protein to be analyzed.

[0140] FIGS. 16 to 18 are drawings for explaining a method for selecting a representative pose according to one embodiment of the present invention.

[0141] In one embodiment, the device can input conformers (1620) into a deep learning model (1650) to obtain a convolution score (1660). The convolution score (1660) is a score reflecting the likelihood of binding to a target protein and may be a value output from a pre-trained deep learning model (1650).

[0142] In one embodiment, the deep learning model (1650) can be trained using binding information of proteins and compounds obtained from a public database (1610). Specifically, the device can use atomic density data (1630) related to binding sites within a protein-compound binding structure obtained from the public database (1610) to train the deep learning model (1650). Alternatively, the device can arrange binding sites of proteins and compounds obtained from the public database (1610) in a pre-divided three-dimensional space, directly calculate atomic density data (1630) in each partitioned space, and use this as training data to train the deep learning model (1650). At this time, the device can apply a transfer learning method that uses the initial weights at which model training begins when training is in progress as the weights of a pre-trained model (1640) obtained from the public database (1610).

[0143] In one embodiment, the device may select some of the conformers as top conformers based on their convolution scores. As an example, the top 100 conformers with high convolution scores may be selected. However, the number of top conformers is not limited thereto.

[0144] Referring to FIG. 17, in one embodiment, the device can generate two or more conformer clusters (1710, 1720) by performing clustering on the upper conformers. Specifically, the device can generate two conformer clusters (1710, 1720) by performing k-means clustering with k set to 2 on the upper conformers. This reflects the fact that compounds often have a symmetric structure due to their nature. However, the embodiment in which the device generates the conformer clusters (1710, 1720) is not limited thereto.

[0145] In one embodiment, the device can select a pose (1711, 1721) with the highest convolution score from each of the conformer clusters (1710, 1720). The selected pose is defined as a representative pose (1711, 1721). Since the representative poses (1711, 1721) are selected from each of the conformer clusters (1710, 1720), the number of representative poses (1711, 1721) can be determined as many as the number of conformer clusters (1710, 1720). That is, when k-means clustering with k as 2 is performed, two representative poses (1711, 1721) can be determined.

[0146] In one embodiment, the device can determine an optimal pose among the representative poses (1711, 1721).

[0147] Referring to FIG. 18, comparison results (1800, 1810) of the accuracy of poses derived according to a conventional technique (1801, 1811) and the accuracy of representative poses (1802, 1812) according to an embodiment of the present invention are shown. Referring to the comparison result (1800), it can be seen that the accuracy (1802) (62%) of the pose derived through docking using a ligand derivative according to an embodiment of the present invention is significantly improved over the accuracy (1801) (38%) of the pose derived according to the conventional technique. In addition, referring to the comparison result (1810), it can be seen that the accuracy (1812) (87%) of the representative pose according to an embodiment of the present invention including docking using a ligand derivative is also improved over the accuracy (1811) (68%) of the pose derived through another conventional technique that is more advanced than the conventional technique described above.

[0148] The accuracy of a pose can refer to the degree of closeness to the bonding structure of an actual substance. In other words, the accuracy of a pose can be an indicator of how similar a pose is to an actually existing bonding structure. In this case, the accuracy (1812) of a representative pose according to one embodiment of the present invention can refer to the accuracy of one of two representative poses derived according to one embodiment of the present invention (e.g., an optimal pose).

[0149] FIG. 19 is a drawing for explaining a method for determining an optimal pose according to one embodiment of the present invention.

[0150] In one embodiment, the device can generate MD conformers based on representative poses (1910, 1920) selected for each of the valid materials. For example, if two representative poses (1910, 1920) are selected for each valid material, the device can generate MD conformers for each representative pose (1910, 1920).

[0151] In one embodiment, the device can perform molecular dynamics simulations on MD conformers in a solvent environment and output snapshots. Although the process of selecting representative poses (1910, 1920) according to the aforementioned embodiment was performed in a solvent-free environment, molecular dynamics simulations in a solvent environment can effectively correct the results.

[0152] As an example, the device can generate poses of MD conformers through short MD ensemble simulation using amber MD simulation software. However, embodiments of generating MD conformers and poses of MD conformers are not limited thereto.

[0153] As an example, the device may perform a short-duration first simulation of MD conformers and use the results as input values ​​for a long-duration second simulation. Specifically, the device may perform the first simulation multiple times and use the best result among the results as input values ​​for the second simulation. For example, the first simulation may have a time scale of 500 ps, ​​and the second simulation may have a time scale of 10 ns. Meanwhile, the first simulation may correspond to the aforementioned short MD ensemble simulation.

[0154] In one embodiment, the device can obtain affinity scores of snapshots using a deep learning model. The deep learning model may be the same as, but not limited to, the deep learning model described above with reference to FIGS. 16 to 18 . The deep learning model can input snapshots and output an affinity score for each snapshot. The affinity score may be a score reflecting the binding affinity between the target protein and the active substance (a specific pose resulting from binding). Thereafter, the device can select one of the output snapshots based on the affinity score. Finally, one snapshot is selected for each representative pose (1910, 1920).

[0155] Referring to FIG. 19, in one embodiment, the device can calculate a sigmoid fit of the similarity between the binding affinities of MD conformers and the binding affinities of representative poses (1910, 1920) (e.g., selected snapshots). Furthermore, based on the sigmoid fit, one of the representative poses (1910, 1920) selected from each conformer cluster can be determined as the optimal pose.

[0156] Specifically, the device can calculate the RMSD (Root Mean Square Deviation) of the structures of MD conformers based on the optimal pose among the representative poses (1910, 1920). The RMSD of the structures of MD conformers can be calculated as a distance indicating how structurally similar each MD conformer is. Thereafter, the device can fit a sigmoid function to the RMSD to determine the representative pose (1910, 1920) that is closer to the sigmoid function fit as the optimal pose among the representative poses (1910, 1920). Accordingly, a conformer cluster that better fits the sigmoid function, which reflects thermodynamic characteristics, can be selected based on the representative pose having the most stable structure. Thereafter, the device can determine the binding affinity of the compound for the target protein from the optimal pose. That is, the binding affinity of the compound for the target protein in the optimal pose can be estimated from the fitting result of the sigmoid function. According to the above-described embodiment, there is an effect that can quickly and effectively generate a thermodynamically stable optimal pose.

[0157] Figure 20 is a drawing for explaining the characteristics of a bonding portion of an effective material according to one embodiment of the present invention.

[0158] In one embodiment, the device can determine the binding site of the target protein and the active substance as either the first binding site (2010) or the second binding site (2020) based on the solvent exposure degree. For example, the first binding site (2010) can be a binding site with a low solvent exposure degree (Deep), and the second binding site (2020) can be a binding site with a high solvent exposure degree (Shallow). As an example, the device can calculate the binding site of the target protein and the active substance and the solvent access value of the active substance using a tool such as enva, and divide the solvent access value for the binding site by the solvent access value for the active substance to calculate the solvent exposure degree. By classifying the binding sites according to the solvent exposure degree, the device can filter the active substances differently to discover the final active substance among the active substances.

[0159] In one embodiment, the device can calculate the total binding energy of the active substance (hereinafter referred to as “original BE”) and the total binding energy for the size of the active substance (hereinafter referred to as “normalized BE”). Original BE refers to the sum of binding energies between atoms constituting the active substance. Normalized BE refers to a value obtained by dividing the original BE by the size of the active substance. The size of the active substance may be determined by the size of the molecule, but may also be determined by the number of atoms with large sizes among the atoms constituting the active substance.

[0160] In one embodiment, the device may perform primary filtering based on the existing BE and secondary filtering based on the normalized BE in response to determining that the coupling unit is the first coupling unit (2010). Furthermore, in one embodiment, the device may perform primary filtering based on the normalized BE and secondary filtering based on the existing BE in response to determining that the coupling unit is the second coupling unit (2020).

[0161] This is because, when the solvent exposure of the binding site is low, most of the binding site exists within the protein being analyzed, and thus, the overall binding energy significantly influences the actual binding. Similarly, when the solvent exposure of the binding site is high, most of the binding site is exposed to the outside of the protein being analyzed, and therefore, the binding energy of specific atoms involved in the actual binding has a more significant influence on the actual binding than the overall binding energy.

[0162] FIG. 21 is a drawing for explaining a method of filtering an effective substance according to one embodiment of the present invention.

[0163] Referring to Fig. 21, a filtering process for the first joint is illustrated.

[0164] In one embodiment, the device can determine at least some of the valid substances as a result of the first filtering (2110) and second filtering (2120) as the final valid substances.

[0165] As described above, the device can perform primary filtering (2110) based on the existing BE for the first binding portion. Accordingly, the device can select an effective substance that is well bound to the pocket of the target protein.

[0166] Thereafter, the device can perform secondary filtering (2120) based on the normalized BE. Accordingly, the device can remove structures that are exposed outside the pocket of the target protein even though they have low binding energy.

[0167] Referring again to FIG. 18, a comparison result (1820) between the accuracy (1821) of an effective substance derived according to a conventional technique and the accuracy (1822) of a final effective substance according to an embodiment of the present invention is illustrated. Referring to the comparison result (1820), it can be seen that the accuracy (1822) (94%) of the final effective substance according to an embodiment of the present invention is improved compared to the accuracy (1821) (93%) of the effective substance derived according to a conventional technique. That is, according to the method for deriving an effective substance according to an embodiment of the present invention, a substance and its structure that are closer to the correct substance can be derived accurately and quickly through molecular dynamics simulation.

[0168] Figure 22 is a flowchart of a method for discovering effective substances through three-dimensional protein-ligand binding structure information according to one embodiment of the present invention.

[0169] Referring to FIG. 22, in step 2210, the device can determine selected compounds among the compounds to be analyzed based on bond energy atomic characteristic information generated from the three-dimensional bonding structure of the target protein and ligand, and compound atom accessibility atomic characteristic information generated from the three-dimensional structure of the compounds to be analyzed.

[0170] In one embodiment, the binding energy atomic property information and the compound atomic accessibility atomic property information may be the three-dimensional binding structure and one-dimensional data in which the three-dimensional structure is compressed.

[0171] In one embodiment, the binding energy atomic characteristic information may be a one-dimensional database of binding energies of atoms constituting the ligand to atoms constituting the pocket of the protein to be analyzed.

[0172] In one embodiment, the compound atomic accessibility atomic characteristic information may be a one-dimensional database of the interaction tendencies of the atoms constituting the compounds to be analyzed with respect to a solvent.

[0173] In one embodiment, the device can determine the selected compounds through a similarity analysis between a binding energy characteristic profile extracted from the binding energy atomic characteristic information based on a binding energy optimization cut-off value, and an atomic accessibility characteristic profile extracted from the atomic accessibility atomic characteristic information based on an atomic accessibility optimization cut-off value.

[0174] In one embodiment, the device can generate ligand characteristic information by database-ing physical and chemical properties of the ligand from the three-dimensional binding structure of the target protein and the ligand.

[0175] In one embodiment, the device can generate compound characteristic information by database-ing physical and chemical properties of the analyte compound from the three-dimensional binding structure of the analyte protein and the analyte compound.

[0176] In one embodiment, the device can generate ligand atom accessibility atomic characteristic information by database-based on the interaction tendencies of atoms constituting the ligand with respect to a solvent, which are derived from the three-dimensional structure of the ligand.

[0177] In one embodiment, the ligand characteristic information and the compound characteristic information may include at least one of information from among the number of rings in the molecule, the number of atoms of each constituent element of the molecule, molecular weight, charge, size of the molecule, distribution coefficient of the molecule, pKa constant, number of hydrogen bond acceptor atoms, number of hydrogen bond donor atoms, polar surface area, and ratio of sP3 carbon atoms.

[0178] In one embodiment, the device can determine the binding energy value and the atomic accessibility value calculated based on an optimal combination of the binding energy atomic characteristic information and the ligand atomic accessibility atomic characteristic information as the binding energy optimization limit value and the atomic accessibility optimization limit value.

[0179] In one embodiment, the device can generate the binding energy characteristic profile combining the ligand characteristic information and the binding energy atomic characteristic information based on binding energy optimization constraint values.

[0180] In one embodiment, the device can generate the atomic accessibility characteristic profile by combining the compound characteristic information and the atomic accessibility atomic characteristic information based on an atomic accessibility optimization constraint value.

[0181] In step 2220, the device can discover effective substances by screening the selected compounds based on the ratio of the key interaction characteristic information generated through 3D docking of the target protein and the positive control to the interaction characteristic information generated through 3D docking of the target protein and the selected compound.

[0182] In one embodiment, the interaction characteristic information may be a one-dimensional database of each of the interactions between atoms constituting the protein to be analyzed and atoms constituting the selected compound.

[0183] In one embodiment, the key-interaction characteristic information may be an interaction with high significance extracted from among the interactions between atoms constituting the protein to be analyzed and atoms constituting the positive control.

[0184] In one embodiment, the device can generate positive control atomic characteristic information of atoms constituting the positive control through three-dimensional docking of the target protein and the positive control.

[0185] In one embodiment, the device can calculate the frequency of occurrence of the reference poses included in the positive control atomic feature information for the set of target proteins.

[0186] In one embodiment, the device can calculate the energy frequency of the reference poses for the set of target proteins by using the binding energy calculated from the binding structure of the target protein and the reference poses as a weight.

[0187] In one embodiment, the device can discover the active substances by screening the selected compounds based on the occurrence frequency and the energy frequency.

[0188] In step 2230, the device can generate conformers based on compound poses generated by docking effective substances with ligand derivatives.

[0189] In one embodiment, the ligand derivative is formed to include all compound binding sites within the template protein using the ligand structure bound to the template protein as a mold, and the compound pose can be generated by three-dimensional docking with the active substances.

[0190] At step 2240, the device can input the conformers into a deep learning model to determine the optimal pose and binding affinity of the compound for the target protein.

[0191] In one embodiment, the device can input the conformers into the deep learning model to obtain a convolution score.

[0192] In one embodiment, the convolution score may be a score reflecting the likelihood that each of the conformers will bind to the protein to be analyzed.

[0193] In one embodiment, the device can select some of the conformers as top conformers based on the convolution score.

[0194] In one embodiment, the device can determine the optimal pose based on the upper conformers.

[0195] In one embodiment, the device can generate two or more conformer clusters by performing clustering on the above-mentioned upper conformers.

[0196] In one embodiment, the device can select the representative pose with the highest convolution score from each of the conformer clusters.

[0197] In one embodiment, the device can generate MD conformers based on the representative poses selected for each of the effective substances.

[0198] In one embodiment, the device can perform molecular dynamics simulations in a solvent environment for the MD conformers and obtain affinity scores of the output snapshots using the deep learning model.

[0199] In one embodiment, the device may select one of the output snapshots for each of the representative poses based on the affinity score.

[0200] In one embodiment, the device can determine one of the representative poses selected from each of the conformer clusters as the optimal pose based on a sigmoid fit of the similarity between the binding affinities of the MD conformers and the binding affinities of the representative poses.

[0201] In one embodiment, the device can determine the binding site of the target protein and the active substance as either the first binding site or the second binding site based on solvent exposure.

[0202] In one embodiment, the device can perform primary filtering based on the total binding energy of the effective material and secondary filtering based on the total binding energy for the size of the effective material in response to the bonding portion being determined to be the first bonding portion.

[0203] In one embodiment, the device can perform primary filtering based on the total binding energy for the size of the effective material and perform secondary filtering based on the total binding energy in response to the bonding portion being determined to be the second bonding portion.

[0204] In one embodiment, at least some of the effective substances may be determined as final effective substances as a result of the first filtering and the second filtering.

[0205] Figure 23 is a block diagram of a device for discovering effective substances through three-dimensional protein-ligand binding structure information according to one embodiment of the present invention.

[0206] Referring to FIG. 23, the device (2300) may include a communication unit (2310), a processor (2320), and a database (2330). Only components related to the embodiment are illustrated in the device (2300) of FIG. 23. Therefore, those skilled in the art will understand that other general components may be included in addition to the components illustrated in FIG. 23.

[0207] The communication unit (2310) may include one or more components that enable wired / wireless communication with an external server or device. For example, the communication unit (2310) may include at least one of a short-range communication unit (not shown), a mobile communication unit (not shown), and a broadcast reception unit (not shown). In one embodiment, the communication unit (2310) may receive data for generating a substitution value.

[0208] DB (2330) is hardware that stores various data processed within the device (2300), and can store a program for processing and controlling the processor (2320).

[0209] DB (2330) may include random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disk storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.

[0210] The processor (2320) controls the overall operation of the device (2300). For example, the processor (2320) can control the input unit (not shown), the display (not shown), the communication unit (2310), the DB (2330), etc., by executing programs stored in the DB (2330). The processor (2320) can control the operation of the device (2300) by executing programs stored in the DB (2330).

[0211] The processor (2320) can control at least some of the operations of the devices described above in FIGS. 1 to 22.

[0212] The processor (2320) may be implemented using at least one of application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.

[0213] In one embodiment, the device (2300) may be a server. The server may be implemented as a computer device or multiple computer devices that communicate over a network to provide commands, codes, files, content, services, etc.

[0214] Meanwhile, embodiments according to the present invention may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. At this time, the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories.

[0215] Meanwhile, the computer program may be specifically designed and constructed for the present invention, or may be one known and available to those skilled in the computer software field. Examples of computer programs may include not only machine language code, such as that generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.

[0216] According to one embodiment, the method according to various embodiments of the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0217] Unless the steps constituting the method according to the present invention are explicitly described in a specific order or are not described to the contrary, the steps may be performed in any appropriate order. The present invention is not necessarily limited to the order in which the steps are described. The use of all examples or exemplary terms (e.g., "for example," etc.) in the present invention is merely intended to illustrate the present invention in detail, and the scope of the present invention is not limited by the examples or exemplary terms unless otherwise defined by the claims. Furthermore, those skilled in the art will appreciate that various modifications, combinations, and variations can be configured according to design conditions and factors within the scope of the appended claims or their equivalents.

[0218] Therefore, the idea of ​​the present invention should not be limited to the embodiments described above, and all scopes equivalent to or equivalently modified from the following claims as well as the claims are considered to fall within the scope of the idea of ​​the present invention.

Claims

1. A step of determining selected compounds among the compounds to be analyzed based on bond energy atomic characteristic information generated from the three-dimensional bonding structure of the protein to be analyzed and the ligand, and atom accessibility atomic characteristic information generated from the three-dimensional structure of the compounds to be analyzed; A step of discovering effective substances by screening the selected compounds based on the ratio of key interaction characteristic information generated through 3D docking of the target protein and the positive control to interaction characteristic information generated through 3D docking of the target protein and the selected compound; A step of generating conformers based on compound poses generated by docking the above effective substances with ligand derivatives; and A step of inputting the above conformers into a deep learning model to determine the optimal pose and binding affinity of the compound for the protein to be analyzed; including, The above binding energy atomic characteristic information and the above compound atomic accessibility atomic characteristic information are, A method for discovering effective substances through the three-dimensional binding structure and the three-dimensional binding structure information between proteins and ligands, which is one-dimensional data in which the three-dimensional structure is compressed.

2. In paragraph 1, The above binding energy atomic characteristic information is, The binding energy of the atoms constituting the ligand to the atoms constituting the pocket of the protein to be analyzed is one-dimensionally databased. The above compound atomic accessibility atomic property information is, A method for discovering effective substances through three-dimensional protein-ligand binding structure information, wherein the tendency of the atoms constituting the above-mentioned target compounds to interact with a solvent is one-dimensionally databased.

3. In paragraph 2, The step of determining the above-mentioned screening compounds is: A method for discovering effective substances through three-dimensional protein-ligand binding structure information, comprising: a step of determining the selected compounds through a similarity analysis between a binding energy characteristic profile extracted from the binding energy atomic characteristic information based on a binding energy optimization cut-off value, and an atomic accessibility characteristic profile extracted from the atomic accessibility atomic characteristic information based on an atomic accessibility optimization cut-off value.

4. In paragraph 3, The step of determining the above-mentioned screening compounds is: A step of creating ligand atom accessibility atomic characteristic information by databaseing the interaction tendency of the atoms constituting the ligand with respect to the solvent, which is derived from the three-dimensional structure of the ligand; and A method for discovering effective substances through three-dimensional protein-ligand binding structure information, further comprising: determining binding energy values ​​and atomic accessibility values ​​calculated based on an optimal combination of the binding energy atomic characteristic information and the ligand atomic accessibility atomic characteristic information as the binding energy optimization limit value and the atomic accessibility optimization limit value.

5. In paragraph 4, The step of determining the above-mentioned screening compounds is: A step of creating ligand characteristic information by databaseing the physical and chemical characteristics of the ligand from the three-dimensional binding structure of the protein to be analyzed and the ligand; A step of creating compound characteristic information by databaseing the physical and chemical characteristics of the analysis target compound from the three-dimensional binding structure of the analysis target protein and the analysis target compound; A step of generating the binding energy characteristic profile in which the ligand characteristic information and the binding energy atomic characteristic information are combined based on the binding energy optimization limit value; and A method for discovering effective substances through three-dimensional protein-ligand binding structure information, comprising: a step of generating an atomic accessibility characteristic profile in which the compound characteristic information and the atomic accessibility atomic characteristic information are combined based on the atomic accessibility optimization limit value.

6. In paragraph 5, The above ligand characteristic information and the above compound characteristic information are, A method for discovering effective substances through three-dimensional protein-ligand bonding structure information, which includes at least one of the following information: number of rings in a molecule, number of atoms of each constituent element of the molecule, molecular weight, charge, size of the molecule, partition coefficient of the molecule, pKa constant, number of hydrogen bond acceptor atoms, number of hydrogen bond donor atoms, polar surface area, and ratio of sP3 carbon atoms.

7. In paragraph 1, The above interaction characteristic information is, Each of the interactions between the atoms constituting the above-mentioned target protein and the atoms constituting the above-mentioned selected compound is one-dimensionally databased. The above key-interaction characteristic information is, A method for discovering effective substances through three-dimensional protein-ligand binding structure information, wherein interactions with high significance are extracted from among the interactions between atoms constituting the above-mentioned target protein and atoms constituting the above-mentioned positive control.

8. In paragraph 1, The steps for discovering the above effective substances are: A step of generating positive control atomic characteristic information of atoms constituting the positive control through three-dimensional docking of the above analysis target protein and the positive control; A step of calculating the appearance frequency of the reference poses included in the positive control atomic characteristic information for the set of target proteins for analysis; A step of calculating the energy frequency of the reference poses for the set of the analysis target proteins by using the binding energy calculated from the binding structure of the analysis target proteins and the reference poses as a weight; and A method for discovering effective substances through three-dimensional protein-ligand binding structure information, comprising: a step of discovering effective substances by screening the selected compounds based on the above appearance frequency and the above energy frequency.

9. In paragraph 1, The above ligand derivatives are, The ligand structure bound to the template protein is used as a mold, and is formed to include all compound binding sites within the template protein. A method for discovering effective substances through three-dimensional protein-ligand binding structural information, wherein the compound pose is generated by three-dimensional docking with the effective substances.

10. In paragraph 1, The step of determining the above optimal pose is: A step of inputting the above conformers into the deep learning model to obtain a convolution score; A step of selecting some of the conformers as upper conformers based on the convolution score; and A step of determining the optimal pose based on the above upper conformers; including: The above convolution score is, A method for discovering effective substances through three-dimensional protein-ligand binding structure information, which is a score reflecting the possibility that each of the above conformers will bind to the protein to be analyzed.

11. In paragraph 10, The step of determining the above optimal pose is: A step of generating two or more conformer clusters by performing clustering on the above upper conformers; and A method for discovering effective substances using three-dimensional protein-ligand binding structure information, comprising: selecting a representative pose having the highest convolution score from each of the above conformer clusters.

12. In paragraph 11, A step of generating MD conformers based on the representative poses selected for each of the above effective substances; A step of performing molecular dynamics simulations in a solvent environment for the MD conformers and obtaining affinity scores of the output snapshots using the deep learning model; and A method for discovering effective substances through three-dimensional protein-ligand binding structure information, further comprising a step of selecting one of the output snapshots for each of the representative poses based on the affinity score.

13. In paragraph 12, The step of determining the above optimal pose is: A step of determining one of the representative poses selected from each of the conformer clusters as the optimal pose based on the sigmoid fit of the similarity between the binding affinities of the MD conformers and the binding affinities of the representative poses; and A method for discovering effective substances through three-dimensional protein-ligand binding structure information, comprising: a step of determining the binding affinity of the compound for the target protein based on the optimal pose; 14. In paragraph 11, A step of determining the binding site of the above-mentioned target protein and the above-mentioned effective substance as either the first binding site or the second binding site based on the solvent exposure level; In response to the above-mentioned bonding portion being determined as the first bonding portion, a step of performing first filtering based on the total binding energy of the effective material, and performing second filtering based on the total binding energy for the size of the effective material; In response to the above-mentioned bonding portion being determined as the second bonding portion, a step of performing first filtering based on the total binding energy for the size of the effective material, and performing second filtering based on the total binding energy of the effective material; and A method for discovering effective substances through three-dimensional protein-ligand binding structure information, comprising: a step of determining at least some of the effective substances as final effective substances as a result of the first filtering and the second filtering.

15. Memory in which at least one program is stored; and A processor that performs an operation by executing at least one program; The above processor, Based on the bond energy atomic characteristic information generated from the three-dimensional bonding structure of the target protein and ligand, and the compound atom accessibility atomic characteristic information generated from the three-dimensional structure of the target compounds, the selected compounds are determined among the target compounds. Based on the ratio of the key interaction characteristic information generated through 3D docking of the target protein and the positive control to the interaction characteristic information generated through 3D docking of the target protein and the selected compound, effective substances are discovered by screening the selected compounds, Conformers are generated based on the compound poses generated by docking the above effective substances with ligand derivatives, The above conformers are input into a deep learning model to determine the optimal pose and binding affinity of the compound for the target protein. A device for discovering effective substances through three-dimensional protein-ligand binding structure information, wherein the above binding energy atomic characteristic information and the above compound atomic accessibility atomic characteristic information are one-dimensional data.

16. A computer-readable recording medium recording a program for executing the method of Article 1 on a computer.

Citation Information

Patent Citations

  • System and method for simulating all atom-based polymer composite

    KR101400717B1

  • Method of using a water-based pharmacophore

    KR1020160128288A

  • Smart office reservation system using artificial intelligence and method thereof

    KR1020220122206A

  • A device that is generating information about new drugs candidate substance

    KR102566459B1