Targeted protein structure-based candidate substance mining system for predicting binding affinity and method of operation thereof
By providing a candidate substance excavation system based on protein structure on the cloud platform, combining artificial intelligence to predict the binding affinity of targeted proteins and ligands, the problem of difficult to predict binding affinity and data management in the prior art is solved, and rapid and economical candidate substance excavation and experimental concentration determination are achieved.
Patent Information
- Application Number
- CN202480003880.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-01-05
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to predict the binding affinity between the targeted protein and the ligand while discovering candidate substances based on protein structure, and there are problems with data sharing and consistent data management in the new drug candidate substance discovery system.
By providing a candidate substance discovery system based on protein structure on the cloud platform, the molecular docking and binding energy prediction of targeted proteins and ligands is realized, and artificial intelligence models are used to predict binding affinity, providing user-friendly interfaces and functions to simplify data management and sharing.
It realizes rapid prediction of binding affinity during the discovery of candidate substances, reduces the time and cost of determining experimental concentrations, and provides experimental concentration values through cloud platforms, which facilitates user access and collaboration.
Smart Images

Figure CN119948570A_ABST
Abstract
Description
Technical Field
[0001] The disclosure relates to a candidate substance discovery system based on protein structure and an operating method thereof, and more specifically, to a candidate substance discovery system based on targeted protein structure for predicting binding affinity and an operating method thereof. Background Art
[0002] Proteins have limited degrees of freedom and have a specific three-dimensional structure based on their constituent amino acid sequences. Their functions are determined by the specific three-dimensional structure of the protein. Therefore, once the target protein is determined, candidate substances that can bind to the active site of the specific protein and regulate the function of the protein can be searched by targeting its specific structure. As described above, substances designed based on structure can shorten the time to develop substances, and substances that only act on the targeted protein can be developed, thereby minimizing undesirable effects (i.e., side effects).
[0003] As mentioned above, the search for candidate substances based on targeted protein structure is usually mainly used for the discovery of new drug candidate substances, because it must be one of the representative fields that can reduce the time and cost required for the screening of initial lead substances or candidate substances and has no side effects, considering the characteristics of new drug development. The discovery of new drug candidate substances is equivalent to the initial stage of new drug development, and the In Silico screening method based on protein structure has attracted much attention. In Silico screening based on protein structure is a method of identifying potential new drug candidates in the compound database based on the three-dimensional structure of proteins to which new drug candidate substances can be attached. In Silico refers to computer programming in computer simulation experiments or virtual experiments, and In Silico screening based on protein structure can be used to simulate the interaction between proteins and compounds to predict the binding force of compounds. Compared with ligand-based screening, which relies on known ligands and is limited in finding completely new chemical structures because it searches for compounds similar to known ligands, In Silico screening based on protein structure can not only find compounds that accurately match the active site of the protein, but also can be used to find compounds in new chemical space.
[0004] For the discovery of candidate substances based on protein structure using the above principles, there is an increasing demand for integrating big data analysis technology of billions of compound libraries with artificial intelligence technology to use multiple analysis tools in parallel. In order to meet the above needs, various software are set up and run in research laboratories of research institutions such as pharmaceutical companies or universities. As an example, the European Molecular Biology Laboratory (EMBL) has developed a variety of tools required for bioinformatics for download or provides simple web applications.
[0005] As described above, even if a candidate substance based on protein structure is discovered, it is difficult to immediately start the development stage for commercialization. As an example, if a new drug candidate substance is discovered, a variety of cell experiments and animal experiments are used to verify and optimize the new drug candidate substance, and finally a clinical trial for human application is carried out. However, even if the new drug candidate substance is specified, it is difficult for users who only have biological knowledge and are not familiar with computing technology to directly apply the discovery results to clinical trials. This is because, usually, the experimental concentration value cannot be known based on the discovery results provided by the new drug candidate substance software. In this case, even after discovering the new drug candidate substance, additional experiments are required to find the experimental concentration value. Typically, this additional experiment includes performing multiple repeated experiments by serial dilution of the candidate drug, and analyzing it to determine the concentration of the therapeutic range that is judged to be effective, so it takes time and cost. Summary of the invention
[0006] Technical issues A technical problem to be solved is to provide a protein structure-based candidate substance discovery system that can predict the binding affinity between a target protein and a ligand while discovering candidate substances based on the protein structure, and provide it to users through a cloud platform.
[0007] Another technical problem to be solved is to provide a new drug candidate discovery system that can predict the binding affinity between the target protein and the ligand while discovering candidate substances, and provide it to users through a cloud platform.
[0008] Technical Solution According to an operating method of a protein structure-based candidate substance discovery system according to one embodiment, it is an operating method of a protein structure-based candidate substance discovery system that is implemented as a cloud platform that provides users with functions or services required for discovering protein structure-based candidate substances in the form of network services, which may include the following steps: displaying a first screen for receiving multiple ligands for docking a target protein; performing molecular docking on the multiple ligands in sequence and calculating estimated binding energies according to settings input by a user through the first screen; displaying a first list consisting of rows including a user interface control for user selection, the names of the multiple ligands, and the calculated binding energy values in a second screen; predicting the binding affinity corresponding to the protein ligand binding posture used in the binding energy calculation for the row selected by the user from the first list of the second screen using the user interface control; and displaying a second list consisting of rows including the names of the multiple ligands, the binding energy values, and the binding affinity values in a third screen.
[0009] In some embodiments, the binding energy value may be expressed as a value according to units representing energy, and the binding affinity value may be expressed as a value according to units of molar concentration.
[0010] In some embodiments, the value of the binding energy may be displayed in Kcal / mol, and the value of the binding affinity may be displayed in fM, pM, nM, μM, mM, or M.
[0011] In some embodiments, the step of displaying the second list in the third screen may include the step of arranging the rows constituting the second list in descending or ascending order based on the values of the binding affinity and displaying the rows in the third screen.
[0012] In some embodiments, the binding affinity can be predicted by an artificial intelligence model that is learned using experimentally measured values of at least one of a dissociation constant Kd, an inhibition constant Ki, and a half maximal inhibitory concentration IC50 as learning data.
[0013] In some embodiments, the artificial intelligence model may learn using the experimental measurements and the binding structure data of the target protein and the ligand as the learning data.
[0014] In some embodiments, the artificial intelligence model may include: a convolution layer portion, including a filter that encodes a pattern to be identified for predicting the binding affinity between the targeting protein and the ligand; and a dense layer portion that integrates features extracted by the convolution layer portion.
[0015] In some embodiments, the convolutional layer portion may include three convolutional layers having 64, 128, and 256 filters, respectively.
[0016] In some embodiments, the dense layer portion may include three dense layers having 1000, 500, and 200 neurons respectively.
[0017] In some embodiments, the artificial intelligence model may include a convolutional neural network (CNN) model or a residual network 3D (ResNet 3D) model.
[0018] According to one embodiment, a protein structure-based candidate substance discovery system is implemented using a cloud platform that provides users with functions or services required for discovering candidate substances in the form of network services, and includes: a project management module that generates projects to which tasks for performing protein structure-based candidate substance discovery can be added; a simulation management module that generates simulations required by users on the generated projects; a simulation setting module that sets a simulation process for the simulation received from the user using a simulation setting area that includes a task module selection area, the task module selection area including a first object that is dragged and dropped into a canvas area and changed to a first node for calculating an estimated binding energy based on the docking of a target protein with a ligand, and a second object that is dragged and dropped into the canvas area and changed to a second node for predicting a binding affinity corresponding to the protein-ligand binding pose used in the binding energy calculation calculated at the first node; and a simulation process management module that manages information about nodes that can be performed first or later in the simulation process setting.
[0019] In some embodiments, when a run button displayed on the first node is clicked, a first screen for receiving a plurality of ligands to be docked with the targeting protein may be displayed.
[0020] In some embodiments, the first node may perform molecular docking on the plurality of ligands in sequence according to the settings input by the user through the first screen, and calculate the estimated binding energy.
[0021] In some embodiments, after the calculation of the binding energy is completed, a first list consisting of rows including a user interface control for user selection, the names of the plurality of ligands, and the calculated values of the binding energy may be displayed in the second screen.
[0022] In some embodiments, the second node may predict the binding affinity corresponding to the protein-ligand binding pose used in the binding energy calculation for a row selected by the user from the first list of the second screen using the user interface control.
[0023] In some embodiments, after the prediction of the binding affinity is completed, a second list consisting of rows including the names of the plurality of ligands, the values of the binding energy, and the values of the binding affinity may be displayed in a third screen.
[0024] In some embodiments, displaying the second list in the third screen may include arranging the rows constituting the second list in descending or ascending order based on the binding affinity values and displaying them in the third screen.
[0025] In some embodiments, the value of the binding energy may be displayed in Kcal / mol, and the value of the binding affinity may be displayed in fM, pM, nM, μM, mM, or M.
[0026] In some embodiments, the binding affinity can be predicted by an artificial intelligence model that uses experimentally measured values of at least one of the dissociation constant Kd, the inhibition constant Ki, and the half-inhibitory concentration IC50 as learning data.
[0027] In some embodiments, the simulation process management module may manage information about nodes that can be executed first or later in the simulation process setting through metadata.
[0028] Effects of the Invention According to the embodiments, functions and user interfaces most suitable for discovering candidate substances based on protein structures can be provided, which improves the problem in the prior art that detailed tasks related to discovering candidate substances based on protein structures are provided as separate tools with low mutual compatibility, making it difficult to share data and manage data consistently. The simulation process can be easily managed by generating, changing and deleting nodes.
[0029] In this regard, the binding affinity between the target protein and the ligand is predicted and provided to the user, thereby helping the user to intuitively understand the result value, so that the user can easily calculate the molecular weight and volume of the culture medium of the derived candidate substance without additional additional experiments, and directly apply it to clinical trials such as cell experiments. Therefore, the time and cost consumed in determining the experimental concentration of the desired candidate substance can be significantly reduced. In particular, as described above, the experimental concentration is provided through the cloud platform, so it has the advantage that the user can easily access the experimental concentration value anywhere that can be connected to the Internet and facilitate collaboration with other users. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 FIG. 4 is a block diagram showing a protein structure-based candidate substance discovery system according to an embodiment.
[0031] Figures 2 to 6 FIG. 4 is a diagram showing an exemplary screen of a protein structure-based candidate substance discovery system according to an embodiment.
[0032] Figure 7 FIG. 1 is a diagram showing a simulation setting example of a protein structure-based candidate substance discovery system according to an embodiment.
[0033] Figures 8 to 10 FIG. 4 is a diagram illustrating the operation of a protein structure-based candidate substance discovery system according to an embodiment.
[0034] Fig.11 FIG. 1 is a diagram for explaining an operation method of a protein structure-based candidate substance discovery system according to an embodiment.
[0035] Fig.12 is a diagram showing an implementation example of an artificial intelligence model of a protein structure-based candidate substance discovery system according to an embodiment.
[0036] Fig.13 is a block diagram for illustrating a computing device according to an embodiment. DETAILED DESCRIPTION
[0037] The embodiments of the present invention are described in detail below with reference to the accompanying drawings so that a person with ordinary knowledge in the technical field to which the present invention belongs can easily implement it. However, the present invention can be implemented in various forms and is not limited to the embodiments described herein. In addition, in order to clearly illustrate the present invention, parts not related to the description are omitted in the accompanying drawings, and similar parts are given similar reference numerals throughout the specification.
[0038] Throughout the specification and claims, when it is mentioned that a part “includes” a certain constituent element, unless there is any special description to the contrary, it means that other constituent elements may also be included, rather than excluding other constituent elements.
[0039] In addition, terms such as "...part", "...device", "module" and the like recorded in the specification may represent a unit capable of processing at least one function or operation described in this specification, and may be implemented by hardware, software, or a combination of hardware and software.
[0040] In this specification, "candidate substance" or "candidate substance based on protein structure" refers to all substances derived using the candidate substance discovery system based on targeted protein structure provided in this specification, without being limited to its field. Preferably, it includes the field of pharmaceutical development, the field of food development, the field of livestock material discovery, etc. More preferably, it refers to candidate substances in the field of pharmaceutical development, but its interpretation is not limited to this.
[0041] Figure 1 FIG. 4 is a block diagram showing a protein structure-based candidate substance discovery system according to an embodiment.
[0042] According to an embodiment, the protein structure-based candidate substance discovery system 1 can be implemented as a cloud platform that provides users with functions or services required for discovering protein structure-based candidate substances in the form of network services. Specifically, the protein structure-based candidate substance discovery system 1 can provide users with a variety of functions or services, for example, when biologists with specific ideas for new drug development use the In Silico screening method, even without knowledge about other fields, errors (or mistakes) in protein structure files can be detected and removed, or the enzyme activity pocket (EAPDC: Enzymatically Active Pocket for Docking Calculation) used for docking calculation can be effectively detected from the protein structure and provided to the user, or the ranking of candidate substances can be provided to the user in real time based on docking binding energy during the execution of a docking simulation that requires a long time, or the docking binding energy can be predicted in two stages to improve reliability, or even the verification of the discovered candidate substances can be performed through contact with a verification company, etc.
[0043] The protein structure-based candidate substance mining system 1 can provide the same functions or services to users in various environments using a network interface. Specifically, for example, a certain user can receive services from the protein structure-based candidate substance mining system 1 using a mobile device such as a smart phone or a tablet computer running a mobile operating system, another user can receive services from the protein structure-based candidate substance mining system 1 using a notebook computer running a Windows operating system, and another user can receive services from the protein structure-based candidate substance mining system 1 using a desktop computer running a Linux operating system. That is, the protein structure-based candidate substance mining system 1 is provided in the form of a cloud platform implemented through a network service, so that users in different environments can use artificial intelligence neural networks to calculate In Silico candidate substances, and can use the same multiple functions or services required for subsequent experiments (e.g., preclinical experiments) for candidate substances, thereby improving compatibility and user convenience, and solving various problems that need to be improved in In Silico calculations that were previously performed through terminals on Linux.
[0044] Reference Figure 1 According to an embodiment, the protein structure-based candidate substance discovery system 1 may include a project management module 10 , a simulation management module 12 , a simulation setting module 14 and a simulation process management module 16 .
[0045] The project management module 10 can generate a project that can add a series of tasks for performing protein structure-based candidate substance discovery. In addition, the project management module 10 can display the generated project to the user by driving the display device of the computing device of the protein structure-based candidate substance discovery system 1.
[0046] In some embodiments, the project management module 10 can perform project name encryption. For projects with project name encryption enabled, the project management module 10 can display the unencrypted original project name to the user who generated the project and the user who has the permission to use the project, and display the encrypted project name to the user who does not have the permission to use the project. The project name may contain keywords that need to be kept safe as keywords related to protein structure-based candidate discovery. The project name is not directly exposed to users who are not involved in the same project on the cloud platform, thereby improving security in the protein structure-based candidate material discovery system 1 used by multiple users. Enabling or disabling project name encryption can be done not only when generating a project, but also by changing the option setting after the project is generated.
[0047] In some embodiments, the project management module 10 can support joining a project through an invitation code. For example, in the case where a user generates a project, the user can send the invitation code to another user, and the other user who receives the invitation code can participate in the corresponding project by entering the invitation code. That is, a user can participate in a project generated by others through an invitation code. Thus, users with knowledge in various fields that are different from each other can be supported to participate in a project on the cloud platform to perform candidate substance discovery based on protein structure. In addition, the project management module 10 can support the setting of permissions for project members. For example, the project management module 10 can grant administrator permissions to specific members of the members. Of course, the project management module 10 can also support users who have participated in a project to leave the project.
[0048] The simulation management module 12 can generate a simulation desired by the user on the project generated by the project management module 10. One simulation can include multiple task modules with detailed functions for performing candidate discovery based on protein structure, and one project can include multiple simulations. After generating the simulation, the user can upload protein structure data for executing multiple task modules.
[0049] When the user selects one or more simulations managed by the simulation management module 12, the simulation setting module 14 can display the simulation setting area 140 to the user by driving the display device of the computing device of the protein structure-based candidate substance discovery system 1. In the simulation setting area 140, the user can arrange and connect multiple task modules based on graphic calculations to set the simulation process to perform the simulation desired by the user. In particular, in the simulation setting area 140, the function of uploading protein structure data and the task modules for performing detailed tasks related to candidate substance discovery based on the uploaded protein structure data can be arranged so that the user can see them at a glance. By means of the above-mentioned structure, the problem that the detailed tasks related to the protein structure-based candidate substance discovery in the prior art are provided as separate tools with low mutual compatibility and are difficult to share data and manage data consistently can be improved. The simulation setting area 140 may include a protein structure data input area 142, a task module selection area 144, and a canvas area 146.
[0050] The protein structure data input area 142 may include more than one object that can be dragged and dropped into the canvas area 146. For example, the protein structure data input area 142 may include first to fourth objects. The first object may be dragged and dropped into the canvas area 146 and converted into a first node, and the first node may receive protein structure data in the form of a protein data bank (PDB) file from a user. The second object may be dragged and dropped into the canvas area 146 and converted into a second node, and the second node may receive protein structure data in the form of a PDB code from a user. The third object may be dragged and dropped into the canvas area 146 and converted into a third node, and the third node may receive protein structure data in the form of a protein sequence file from a user. The fourth object may be dragged and dropped into the canvas area 146 and converted into a fourth node, and the fourth node may receive protein structure data in the form of a protein sequence from a user.
[0051] The task module selection area 144 may also include one or more objects that can be dragged and dropped into the canvas area 146. For example, the task module selection area 144 may include the fifth object to the eighth object. The fifth object may be dragged and dropped into the canvas area 146 and converted into a fifth node, and the fifth node may perform the task of finding the best docking site in the targeted protein structure. In particular, the fifth node may use an artificial intelligence language model based on natural language processing (NLP: Natural Language Processing) to automatically find the active site (Active Sit) in the targeted protein and generate the best docking grid box (Docking Grid-Box). In addition, the fifth node may automatically correct various errors that may exist in the protein structure file (i.e., PDB file). That is, the fifth node can automatically perform the following tasks: detect and remove anisotropic B-factors in PDB files; detect alternative conformations in amino acid residue regions and modify them to non-alternative conformations; detect unusual amino acids in amino acid residue regions and modify them to 20 kinds of unspecific amino acids. The sixth object can be dragged and dropped into the canvas area 146 and converted into the sixth node, and the sixth node can predict the amino acid sequence as a protein tertiary structure. The seventh object can be dragged and dropped into the canvas area 146 and converted into the seventh node, and the seventh node can analyze the binding energy (e.g., in kcal / mol) of the target protein with the ligand, arrange them in the best order and provide them to the user. That is, the seventh node can perform a grid-based In Silico docking based on the input protein structure, and can use the Lamarckian Genetic Algorithm (LGA) to determine the chemical pose, and use the empirical scoring function to calculate the binding energy. The eighth object can be dragged and dropped into the canvas area 146 and converted to the eighth node, and the eighth node can convert the predicted binding energy between the target protein and the ligand into binding affinity and perform comparative analysis. For example, the eighth node can use an artificial intelligence model that learns protein-ligand structure and Kd / Ki / IC50 values to predict binding affinity in molar concentration units (e.g., in fM, pM, nM, μM, mM, or M). The concentration units predicted in this way can be displayed next to the binding affinity value so that a person of ordinary skill in the art can intuitively identify the desired value.
[0052] The first to eighth objects as described above can be dragged and dropped into the canvas area 146 and converted into the first to eighth nodes, respectively, and the user can freely arrange the first to eighth nodes in the canvas area 146 according to the desired operation order according to the purpose and environment of the simulation. In addition, the user can set the connection relationship between the first to eighth nodes arranged in the canvas area 146 by edges. The user can generate a candidate substance discovery simulation by arranging nodes and connecting the edges between the nodes. In addition, in a simulation process that increases in complexity as the number of nodes placed in the canvas area 146 and the number of edges between the nodes increases, the user can clearly and easily manage information about the protein structure data input by generating, changing and deleting nodes.
[0053] In some embodiments, the simulation setting area 140 may further include an external module providing area 148. The external module providing area 148 may include a ninth object, which is dragged and dropped into the canvas area 146 and may be converted into a ninth node having any function provided from outside the protein structure-based candidate substance discovery system 1. Thus, by setting the node in the canvas area 146, the functions provided by other systems connected to and operated by the protein structure-based candidate substance discovery system 1 can be easily included in the simulation.
[0054] As described above, the user can drag and place the desired object on the canvas area 146 after clicking it from the protein structure data input area 142, the task module selection area 144, and the external module providing area 148, so as to arrange and freely move the nodes, and in some embodiments, the nodes arranged in the canvas area 146 can display the node connection shape. During the state in which the user is in the state of clicking the node connection shape displayed on the right side of a certain node, the color or shape of the node connection shape displayed on the left side of the connectable node in other nodes can be changed. In other nodes, the user cannot establish a connection with the node whose color of the node connection shape displayed on the left side has not changed, but can only establish a connection with the node whose color of the node connection shape displayed on the left side has changed. Accordingly, the user can prevent the generation of an erroneous simulation process in advance by checking whether the color of the node connection shape changes without knowing whether a causal relationship can be established between each node.
[0055] The information used to determine whether the nodes are connectable can be managed by the simulation process management module 16. The simulation process management module 16 can manage information about nodes that can be preceded or followed in the simulation setting, and can use additional data structures such as metadata as needed. In addition, the simulation process management module 16 can also update the information when the information about the nodes that can be preceded or followed changes, such as reflecting it in the metadata.
[0056] When the user clicks the node connection shape displayed on the right side of a certain node and wants to connect to another node whose node connection shape displayed on the left side has changed color, the user can click the node connection shape of a certain node and then click the node connection shape of another node to connect, or can drag the node connection shape of a certain node to the node connection shape of another node in the clicked state to connect. If the connection is completed, a connection line can be displayed between the nodes, and if the X displayed on the connection line is clicked, the connection between the nodes can be released.
[0057] If the objects included in the task module selection area 144 are converted into nodes on the canvas area 146, a run button can be generated in the node. The user can click the run button to execute the task of the node. Before the task starts, the number of tokens required to execute the task can be displayed, and the tokens can be reduced after the user's confirmation to run the task. If the task is completed, the run button can be changed to a result button, and a download button can be added. The user can view the task execution results by clicking the result button, and can download the task execution results by clicking the download button.
[0058] As described above, in the discovery of candidate substances, the connection relationship of nodes having various functions can be set, thereby generating a simulation process suitable for developing the optimal compound. In addition, even if a function is added inside the protein structure-based candidate substance discovery system 1 or a new function is added from the outside, the connection relationship with the existing nodes can be easily set by generating a node corresponding to the added function.
[0059] Figures 2 to 6 FIG. 4 is a diagram showing an exemplary screen of a protein structure-based candidate substance discovery system according to an embodiment.
[0060] Reference Figure 2, the simulation setting area 140 may include a protein structure data input area 142, a task module selection area 144, a canvas area 146, and an external module providing area 148. As shown in the figure, the protein structure data input area 142 may include one or more objects related to the function of uploading protein structure data, the task module selection area 144 may include one or more objects related to the detailed tasks of candidate substance discovery performed based on the uploaded protein structure data, and the external module providing area 148 may include one or more objects related to any function provided from the outside. The objects as described above can be dragged and dropped by the user into the canvas area 146 and converted into nodes, the nodes can be connected by edges to form a graph, and the graph can directly represent the process of the simulation. Of course, the number or type of objects included in the task module selection area 144, the canvas area 146, and the external module providing area 148 shown in the figure are exemplified for the purpose of illustrating the implementation, and the scope of the present invention is not limited to the content of the diagram.
[0061] Reference Figure 3 , the protein structure data input area 142 may include a first object 1420, a second object 1421, a third object 1422, and a fourth object 1423. The first object 1420 represented as "PDB file upload" is related to the function of receiving protein structure data in the form of a PDB file, and the second object 1421 represented as "PDB code input" is related to the function of receiving protein structure data in the form of a PDB code. The third object 1422 represented as "protein sequence file (Fasta)" is related to the function of receiving protein structure data in the form of a protein sequence file, and the fourth object 1423 represented as "protein sequence (File)" is related to the function of receiving protein structure data in the form of a protein sequence.
[0062] In addition, the task module selection area 144 may include a fifth object 1440, a sixth object 1441, a seventh object 1442, and an eighth object 1443. The fifth object 1440 represented as "PocketFinder" is related to the function of automatically finding the best docking site, and the sixth object 1441 represented as "CaliciFold" is related to the function of predicting the amino acid sequence as the tertiary structure of the protein. The seventh object 1442 represented as "AI-Dock" is related to the function of analyzing the binding energy of the target protein and the ligand and automatically arranging them in the best order. The eighth object 1443 represented as "DeepCalici" is related to the function of converting the predicted binding energy between the target protein and the ligand into binding affinity and performing comparative analysis.
[0063] In some embodiments, the fifth object 1440 can detect and remove anisotropic B-factors from the PDB file in the protein structure data file input by the user, detect alternative conformations in the amino acid residue region and modify them to non-alternative conformations, and detect unusual amino acids in the amino acid residue region and modify them to 20 kinds of unspecific amino acids. For example, if the docking simulation is performed without removing the anisotropic B-factor in the PDB file, an error may occur that the PDB file format cannot be recognized or the PDB file cannot be read. If the docking simulation is performed when there are alternative conformations or unusual amino acids in the amino acid residue region, an error may occur due to unknown amino acids, which may reduce the accuracy when performing the In Silico screening method or increase the failure rate of candidate substance discovery. By automatically processing the error-causing factors as described above, not only can the inefficiency and inaccuracy that may occur due to manual modification of the PDB file by the user be prevented, and the collaboration with the structural biologist be omitted, but also the processing can be automatically performed internally in a manner that the user cannot recognize the process of preprocessing the PDB file, thereby providing an environment in which the user can focus only on the discovery of candidate substances. In addition, in some embodiments, the fifth object 1440 can perform modification of missing residues in the protein structure of the PDB file. Specifically, the missing residues are detected by checking the intervals between the residues in the protein structure of the PDB file. In the case of missing residues, a protein amino acid sequence suitable for completing the missing residues can be obtained by searching the sequence database, and the missing residues can be automatically completed in the protein amino acid sequence thus obtained. Thus, subsequent tasks can be performed based on error-free protein structure files in which errors that may occur in the simulation are eliminated.
[0064] The fifth object 1440 can detect the enzymatically active pocket for docking calculation (EAPDC) from the protein structure file and determine the docking calculation site. Specifically, the fifth object 1440 can use an artificial intelligence language model to predict the docking site (i.e., EAPDC) in the target protein structure. Specifically, the fifth object 1440 can calculate the depth value of the pocket based on the solvent accessible surface (SAS: Solvent Accessible Surface) of the target protein surface, generate a gradient class activation map (Gradient Class Activation Map) of amino acids that contribute to the process of predicting the activity of the target protein, and determine the site that has a great impact on the activity of the target protein as the docking calculation site by considering the depth value of the pocket and the value of the amino acid with high contribution in the gradient class activation map. Here, the gradient class activation map can be extracted from a graph convolutional network (GCN: Graph Convolutional Network) learned using the Enzyme Commission number (Enzyme Commission number) or the Gene Ontology number (Gene Ontology number) implemented by the natural language processing model embedding layer. Furthermore, the natural language processing model implemented by the embedding layer may be a Transformer natural language processing model.
[0065] In addition, the external module providing area 148 includes a ninth object 1480, and the ninth object 1480 represented as "CRO-Order" is related to a function of transmitting a verification request for a candidate substance desired by a user to a verification company server.
[0066] like Figure 3 As shown, for example, a user can drag and drop the first object 1420 represented as "PDB file upload" into the canvas area 146, which can be converted into a node N31. Node N31 may include a button that can perform a function of receiving protein structure data in the form of a PDB file. In addition, node N31 can also display relevant information such as the identification of the node, task execution status, etc. In addition, node N31 can also include a button that can delete itself. As shown in the example node N31, the second object 1421 to the ninth object 1480 can be dragged and dropped into the canvas area 146, and can be converted into a node that displays inherent buttons, information, etc.
[0067] like Figure 4As shown, the user can drag and drop the second object 1421 represented as "PDB code input" into the canvas area 146, thereby converting it into a node N41. The node N41 may include a button capable of executing a function of receiving protein structure data in the form of a PDB code. Next, the user can drag and drop the fifth object 1440 represented as "PocketFinder" into the canvas area 146, thereby converting it into a node N42. The node N42 may include a button capable of executing a function of automatically searching for the best docking site.
[0068] like Figure 5 As shown, the nodes arranged in the canvas area 146 may be displayed with node connection shapes. Specifically, a node connection shape CS1 may be displayed on the right side of the node N51, and a node connection shape CS2 may be displayed on the left side of the node N52. During the period when the user is in the state of clicking the node connection shape CS1 displayed on the right side of a certain node N51, the color or shape of the node connection shape CS1 displayed on the left side of the connectable node N52 in other nodes may change. In other nodes, the user cannot establish a connection to the node whose color of the node connection shape displayed on the left side has not changed, but can only establish a connection to the node whose color of the node connection shape displayed on the left side has changed. During the period when the user is in the state of clicking the node connection shape CS1 displayed on the right side of a certain node N51, the change of the color or shape of the node connection shape CS1 displayed on the left side of the connectable node N52 in other nodes can be determined based on the information about the nodes that can be preceded or followed provided from the simulation process management module 16.
[0069] like Figure 6 As shown, when the node connection shape CS1 displayed on the right side of a certain node N61 is clicked, if you want to connect to another node N62 whose color of the node connection shape CS2 displayed on the left side has changed, you can click the node connection shape CS1 of the certain node N61 and then click the node connection shape CS2 of another node N62 to connect, or you can drag the node connection shape CS1 of the certain node N61 to the node connection shape CS2 of another node N62 in the clicked state to connect. If the connection is completed, a connection line is displayed between the nodes, and if you click the X displayed on the connection line, the connection between the nodes can be released.
[0070] Figure 7 FIG. 1 is a diagram showing a simulation setting example of a protein structure-based candidate substance discovery system according to an embodiment.
[0071] Reference Figure 7, in an example of generating a certain simulation process, nodes N71 to N74 are shown. The node connection shape on the right side of node N71 that receives protein structure data in the form of PDB code can be connected to the node connection shape on the left side of node N72 that automatically searches for the best docking site by an edge, the node connection shape on the right side of node N72 can be connected to the node connection shape on the left side of node N73 that analyzes the binding energy between the target protein and the ligand and automatically arranges them by an edge, and the node connection shape on the right side of node N73 can be connected to the node connection shape on the left side of node N74 that converts the predicted binding energy between the target protein and the ligand into binding affinity to perform comparative analysis by an edge. As mentioned above, among other nodes, the user cannot establish a connection with the node whose color of the node connection shape displayed on the left side has not changed, but can only establish a connection with the node whose color of the node connection shape displayed on the left side has changed. Therefore, in the candidate substance discovery, the user does not need to consider whether the nodes can be preceded or followed, thereby increasing convenience.
[0072] In some embodiments, information about nodes that can precede or follow may be predetermined as follows.
[0073] [Table 1]
[0074] The nodes that can precede the node that automatically searches for the best docking site ("PocketFinder") may include: a node that receives protein structure data in the form of a PDB file ("PDB file upload"); a node that receives protein structure data in the form of a PDB code ("PDB code input"); a node that receives protein structure data in the form of a protein sequence file ("Protein Sequence File (Fasta)"); a node that receives protein structure data in the form of a protein sequence ("Protein Sequence (Text)"); and a node that predicts an amino acid sequence as a protein tertiary structure ("CaliciFold"), and the nodes that can be followed may include: a node that analyzes the binding energy of a targeted protein with a ligand and automatically arranges them ("AI-Dock"). The nodes that can precede the node that automatically searches for the best docking site ("CaliciFold") may include: a node that receives protein structure data in the form of a protein sequence file ("Protein Sequence File (Fasta)"); a node that receives protein structure data in the form of a protein sequence ("Protein Sequence (Text)"), and the nodes that can be followed may include: a node that automatically searches for the best docking site ("PocketFinder").
[0075] Nodes that can precede the node that analyzes the binding energy between the target protein and the ligand and automatically arranges them ("AI-Dock") may include: a node that automatically searches for the best docking site ("PocketFinder"), and nodes that can follow may include: a node that converts the predicted binding energy between the target protein and the ligand into binding affinity and performs comparative analysis ("DeepCalici").
[0076] Nodes that may precede a node that converts predicted binding energy between a target protein and a ligand into binding affinity and performs comparative analysis ("DeepCalici") may include: a node that analyzes the binding energy of a target protein and a ligand and automatically arranges them ("AI-Dock").
[0077] The information about the nodes that can be preceded or followed as described above can be managed by the simulation flow management module 16, and additional data structures such as metadata can be used as needed. In addition, the simulation process management module 16 can also update the information when the information about the nodes that can be preceded or followed changes, such as reflecting it in the metadata. As described above, in the candidate substance discovery, the connection relationship of the nodes with various functions can be set, thereby generating a simulation process suitable for developing the best compound.
[0078] According to the existing simulation method, in order to perform complex simulations on candidate substance excavation based on protein structure that needs to be tried many times in various ways, a lot of effort, time and cost are required, and it is difficult to set a satisfactory simulation. The simulation setting method based on graphical calculation described by the embodiment can intuitively and easily generate and manage complex simulation processes as examples by improving the existing method, and provide flexibility and convenience that are easy to change. In addition, highly complex and complex simulation processes can also be actually executed. In addition, functions and user interfaces optimized for candidate substance excavation are provided, which improves the problem that detailed tasks related to candidate substance excavation in the prior art are provided with separate tools with low mutual compatibility and are difficult to share data and manage data consistently, and the simulation process can be easily managed in a way of generating, changing and deleting nodes.
[0079] Figures 8 to 10 FIG. 4 is a diagram illustrating the operation of a protein structure-based candidate substance discovery system according to an embodiment.
[0080] Reference Figure 8 According to an embodiment, the protein structure-based candidate substance discovery system may display a first screen 30 for receiving a plurality of ligands to be docked with a target protein.
[0081] In this embodiment, the simulation setting module 14 of the candidate substance discovery system based on protein structure displays a simulation setting area 140 to the user. As mentioned above, the simulation setting area 140 may include a task module selection area 144 and a canvas area 146. The simulation setting module 14 may set a simulation process for a simulation based on the user dragging and dropping multiple task modules on the canvas area 146 and arranging and connecting multiple task modules based on graphical calculation. In particular, in this embodiment, the task module selection area 144 may include a first object and a second object. The first object may be dragged and dropped into the canvas area 146 and changed to a first node for calculating the estimated binding energy based on the docking of the target protein with the ligand, and the second object may be dragged and dropped into the canvas area 146 and changed to a second node for predicting the binding affinity corresponding to the protein ligand binding pose (Pose) used in the binding energy calculation calculated in the first node.
[0082] As described above, the first node and the second node may include a run button, respectively. The user may click the run button displayed on the first node to perform a task of calculating the estimated binding energy based on the docking of the target protein and the ligand, and may click the run button displayed on the second node to perform a task of predicting the binding affinity corresponding to the protein-ligand binding pose used in the binding energy calculation calculated in the first node.
[0083] When the run button displayed on the first node is clicked, the first screen 30 may be displayed. In addition, the user may set matters required for executing the task through the first screen 30. The first screen 30 may include a plurality of user interface elements. In this specification, the term "user interface elements" is used as a concept including all visual elements for user interaction such as buttons, labels, text boxes, images, sliders, drop-down menus, etc., and in distinction therefrom, "user interface controls" are used as a concept to represent elements such as buttons, check boxes, radio buttons, switches, etc. that receive user input and transmit commands to an application.
[0084] The plurality of user interface elements of the first screen 30 may include a first user interface element 301. The first user interface element 301 may be used to set the number of results to be displayed after the execution of the task of the first node is completed. For example, the user enters a value of "500" in the first user interface element 301, whereby after the execution of the first task is completed, the binding energy estimated based on the docking of the target protein with the ligand may be displayed up to 500.
[0085] The second user interface element 302 among the multiple user interface elements of the first screen 30 can be used to set the ligand to be docked with the targeted protein. For example, the second user interface element 302 can include "random" and "upload" as its value, and when the user selects "random", in the first task, the candidate substance discovery system based on protein structure can perform docking with the targeted protein based on the ligand library provided by itself. Different from this, when the user selects "upload", the user directly uploads the data about the ligand, and in the first task, the docking with the targeted protein can be performed based on the ligand data uploaded by the user.
[0086] The third user interface element 303 among the plurality of user interface elements of the first screen 30 can be used to set the type and quantity of ligands to be docked with the target protein. For example, the user enters "FDA", "all", "2115" in the third user interface element 303, so that the docking of all 2115 ligands of the FDA-approved drug library can be performed. Or, as another example, the user enters "MCULE", "in stock", "2000" in the third user interface element 300, so that the docking of 2000 in stock compounds in the Mcule library that are in stock and can be directly distributed can be performed.
[0087] The fourth user interface element 304 of the plurality of user interface elements of the first screen 30 can be used to set the number of GPUs to be used in the task of calculating the estimated binding energy based on the docking of the target protein and the ligand at the first node. For example, if the user inputs "1" in the fourth user interface element 304, one GPU thread can be used to execute the corresponding task, whereas if the user inputs "2", two GPU threads can be used to execute the corresponding task.
[0088] The fifth user interface element 305 of the plurality of user interface elements of the first screen 30 can display the number of tokens required for the task of running the operation performed by the first node according to the estimated binding energy of the docking of the target protein and the ligand. The candidate substance discovery system based on the protein structure is implemented to deduct the tokens that the user needs to pay for the specific task of the simulation for candidate discovery, or the tokens can be recharged when the user pays the fee through coupons or various payment methods, and the amount of tokens deducted can be determined by considering various factors about the detailed task, such as the type of the detailed task, the task amount of the detailed task, and the difficulty of the detailed task. Data on the deduction standard of the token amount can be stored in a storage medium or cloud accessible to the computing device in a form readable by the computing device. The user can use the token based on the deduction standard determined according to the detailed task to be performed. In particular, as described above, in the candidate substance discovery, the connection relationship of the nodes with various functions can be set to generate a simulation process suitable for developing the best compound, and the user can pay the tokens that are applicable to each node to use the corresponding node, and can also pay only the required tokens at the corresponding node used. For example, for the simulation task performed by the first node, such as for the case of the FDA-approved drug library, the fifth user interface element 305 may display that 2,115 tokens corresponding to the number of ligands are required. If, in the case where the value set in the fourth user interface element 304 is changed from "1" to "2", two GPU threads may be allocated to perform the simulation task for 2,115 ligands, thereby paying twice the fee and reducing the working time by half.
[0089] The sixth user interface element 306 of the plurality of user interface elements of the first screen 30 can submit the settings formed through the first screen 30 , perform molecular docking on a plurality of ligands in sequence according to the settings input by the user through the first screen 30 , and calculate the estimated binding energy.
[0090] Among them, the first user interface element 301 , the second user interface element 302 , the third user interface element 303 , the fourth user interface element 304 and the sixth user interface element 306 may be regarded as user interface controls differently from the fifth user interface element 305 .
[0091] Reference Fig. 9According to an embodiment of the protein structure-based candidate substance discovery system, a second screen 31 may be displayed to provide the user with the binding energy estimated based on the docking of the target protein and the ligand as a result. As described above, after the run button displayed on the first node is changed to a result button, the user can click the result button to confirm the results arranged in the order of binding energy through the second screen 31. In some embodiments, a generate download button may be added to the first node, and the user may also download the results arranged in the order of binding energy by clicking the download button.
[0092] The second screen 31 may include multiple user interface elements. The first user interface element 311 of the multiple user interface elements of the second screen 31 may be a user interface control (e.g., a check box) for user selection. In some embodiments, in order to facilitate the user to select multiple rows, a button may be added to display that all rows of the displayed page can be selected in batches ("Select All in Page"), or a button may be added to display that all results can be selected in batches ("Select All").
[0093] The second user interface element 312 among the plurality of user interface elements on the second screen 31 may display the names of a plurality of ligands on which molecular docking has been performed.
[0094] The third user interface element 313 among the multiple user interface elements of the second screen 31 can display the binding energy estimated by docking the targeted protein with the multiple ligands. The third user interface element 313 can display the value of the binding energy as a value according to the unit representing energy. For example, the third user interface element 313 can display the value of the binding energy in Kcal / mol.
[0095] The fourth user interface element 314 among the multiple user interface elements of the second screen 31 can display the structure of the resultant substance after molecular docking in the form of a character string. For example, the fourth user interface element 314 can display the structure of the resultant substance after molecular docking in the form of a character string according to the Simplified Molecular Input Line Entry System (SMILES) notation. In the SMILES notation, for example, atoms can be expressed by element symbols, and bonds can be expressed by specific characters including "=", "#", etc. In addition, intramolecular branches can be expressed using brackets, or ring structures can be expressed using numbers, or coordination structures can be expressed using specific symbols such as "@".
[0096] The second screen 31 displays the results as a first list consisting of a plurality of rows, and one row 315 may include a user interface control for user selection, the names of a plurality of ligands, and the calculated binding energy values.
[0097] A fifth user interface element 316 among the plurality of user interface elements of the second screen 31 may submit information about a row selected by the first user interface element 311 in the first list of the second screen 31 , and may allow calculation of a binding affinity corresponding to a protein-ligand binding pose used in binding energy calculation.
[0098] Different from the second user interface element 312 , the third user interface element 313 , and the fourth user interface element 314 , the first user interface element 311 and the fifth user interface element 316 may be regarded as user interface controls.
[0099] Reference Fig.10 According to an embodiment, the protein structure-based candidate substance discovery system can display a third screen 32 that provides the user with the predicted binding affinity corresponding to the binding energy corresponding to the row selected by the user as a result. The user can click the fifth user interface element 316 in the second screen 32 and confirm the results arranged in order of binding energy through the third screen 32. Alternatively, in the case of moving from the second screen 32 to the screen displaying the node, the result button displayed on the second node can be clicked, and the results arranged in order of binding affinity can be confirmed through the third screen 32. In some embodiments, a generate download button can be added on the second node, and the user can also download the results arranged in order of binding affinity by clicking the download button.
[0100] The third screen 32 may include a plurality of user interface elements. A first user interface element 321 among the plurality of user interface elements of the third screen 32 may display sequences arranged according to binding affinity.
[0101] The second user interface element 322 of the plurality of user interface elements of the third screen 32 may display the names of the plurality of ligands subjected to molecular docking. Specifically, among the names displayed by the second user interface element 312 of the second screen 31, only the name selected by the user through the first user interface element 311 may be displayed.
[0102] The third user interface element 323 among the plurality of user interface elements of the third screen 32 may display the binding energy estimated by docking the target protein with the plurality of ligands. Specifically, among the binding energies displayed as the third user interface element 313 of the second screen 31, only the binding energy selected by the user through the first user interface element 311 may be displayed.
[0103] A fourth user interface element 324 of the plurality of user interface elements of the third screen 32 may display a binding affinity corresponding to a protein-ligand binding pose used in the binding energy calculation. The fourth user interface element 324 may display the value of the binding affinity as a value based on a molar concentration unit. For example, the fourth user interface element 324 may display the value of the binding affinity in units of fM, pM, nM, μM, mM, or M.
[0104] The fifth user interface element 315 of the plurality of user interface elements of the third screen 32 may display the structure of the resultant substance after molecular docking in the form of a character string. Specifically, among the structures displayed as the fourth user interface element 314 of the second screen 31, only the structure selected by the user through the first user interface element 311 may be displayed.
[0105] The third screen 32 displays the results as a second list consisting of a plurality of rows, and one row 326 may include the names of a plurality of ligands, the values of binding energy, and the values of binding affinity.
[0106] In this embodiment, the rows constituting the second list can be arranged in descending or ascending order based on the value of binding affinity and displayed on the third screen. Of course, in the same third screen, the rows can also be arranged in descending or ascending order based on the value of binding energy.
[0107] In this way, the binding affinity between the target protein and the ligand is predicted, and the binding energy and binding affinity are provided to the user in a side-by-side form (i.e., as adjacent columns) on one screen, thereby helping the user's intuitive understanding, so that the user can easily calculate the molecular weight and volume of the culture medium of the candidate substance (e.g., a new drug candidate substance) without additional additional experiments, and directly apply it to clinical trials such as cell experiments. Therefore, the time and cost spent on determining the experimental concentration of the candidate substance can be significantly reduced.
[0108] Fig.11 FIG. 1 is a diagram for explaining an operation method of a protein structure-based candidate substance discovery system according to an embodiment.
[0109] According to one embodiment, the operating method of the candidate substance discovery system based on protein structure is implemented by using a cloud platform that provides users with the functions or services required for discovering candidate substances in the form of network services, which may include the following steps: displaying a first screen for receiving multiple ligands to be docked with a target protein; performing molecular docking on multiple ligands in sequence according to the settings input by the user through the first screen, and calculating the estimated binding energy; displaying a first list consisting of rows including user interface controls for user selection, names of multiple ligands, and calculated binding energy values in a second screen; predicting the binding affinity corresponding to the protein ligand binding posture used in the binding energy calculation for the row selected by the user from the first list of the second screen using the user interface controls; and displaying a second list consisting of rows including names of multiple ligands, binding energy values, and binding affinity values in a third screen. For more detailed content on the operating method of the candidate substance discovery system based on protein structure, please refer to the above reference. Figures 1 to 10 Therefore, repeated description will be omitted here.
[0110] In some embodiments, binding affinity can be predicted by an artificial intelligence model that uses experimental measurements of at least one of the dissociation constant Kd, the inhibition constant Ki, and the half maximal inhibitory concentration IC50 as learning data. The artificial intelligence model can learn experimental measurements and binding structure data of the target protein and the ligand as learning data. In some embodiments, the artificial intelligence model can include a convolutional neural network (CNN) model or a residual network 3D (ResNet 3D) model.
[0111] Reference Fig.11 , the learning of the artificial intelligence model can be performed according to steps S1101 to S1108, and the prediction of binding affinity based on the artificial intelligence model can be performed according to steps S1109 to S1111.
[0112] In step S1101, the binding structure of the protein and the ligand and the value measured by the experiment can be received as learning data. The binding structure of the protein and the ligand can be a three-dimensional structure or PDB data, and the value measured by the experiment can be at least one of Kd, Ki and IC50.
[0113] In step S1102, outliers present in the learning data may be removed. Outliers are values that are difficult to physically exist in terms of features and are likely to correspond to measurement errors or structural errors of proteins. For example, if outliers exist in the PDBBind dataset, a histogram-based outlier removal technique may be used to remove the outliers.
[0114] In step S1103, the missing residues in the protein structure and the ligand can be detected. In the case of missing residues, a protein amino acid sequence suitable for completing the missing residues is obtained by searching the sequence database, and the missing residues are automatically completed in the protein amino acid sequence thus obtained. The case of protein structure that cannot be modified can be deleted.
[0115] In step S1104, the data set may be separated so that similar structures are not mixed into the validation data set or the test data set. For example, when predicting the three-dimensional structure of a protein, the separation of the data set may be performed based on a TM score (Template Modeling Score) which is an indicator for evaluating information on the similarity of the predicted structure to the actual structure determined experimentally.
[0116] In step S1105, learning about the artificial intelligence model can be performed using the learning data provided through the previous steps.
[0117] In steps S1106 to S1108, in order to illustrate how each input feature of the artificial intelligence model contributes to the final prediction, feature engineering can be performed by calculating feature importance based on the Shapley Additive exPlanations (SHAP) value that numerically evaluates the contribution of each input variable to the individual prediction and adding or removing features through visual inspection using protein structure analysis tools.
[0118] In addition, in step S1109, the three-dimensional structural pose formed by the combination of the protein to be predicted and the candidate substance is input into the artificial intelligence model that has completed learning. In step S1110, the molar concentration-based binding affinity that can be dissolved in the drug by a user with only biological knowledge can be obtained as the prediction value of the artificial intelligence model. Fig.10 The binding affinity is displayed to the user in the associated illustrated screen.
[0119] Fig.12is a diagram showing an example of an artificial intelligence model of a protein structure-based candidate substance discovery system according to an embodiment.
[0120] Reference Fig.12 According to an embodiment, an artificial intelligence model for predicting binding affinity in a candidate substance discovery system based on protein structure may include a convolution layer part and a dense layer part. The convolution layer part may include a filter for encoding a pattern that needs to be recognized in order to predict the binding affinity between the target protein and the ligand. In addition, the dense layer part may merge features extracted by the convolution layer part.
[0121] In some embodiments, the artificial intelligence model can be a deep 4D convolutional neural network, specifically, a deep 4D convolutional neural network having a single output neuron for predicting the binding affinity between a target protein and a ligand.
[0122] The convolutional layer part can identify the patterns encoded by the filters of the convolutional layer and generate a feature map that emphasizes the spatial occurrences of each pattern within the data. Specifically, the convolutional layer part can include three convolutional layers with 64, 128 and 256 filters respectively, and the result of the final convolutional layer can be flattened and used as the input of the dense layer part.
[0123] The dense layer part may include three dense layers with 1000, 500 and 200 neurons respectively, and dropout with a drop probability of 0.5 may be used for all dense layers. In addition, an L2 normalization technique with a lambda value of 0.001 may be used.
[0124] In addition, the rectified linear unit (ReLU) can be used as the activation function in both the convolutional layer and the dense layer.
[0125] Fig.13 is a block diagram for illustrating a computing device according to an embodiment.
[0126] Reference Fig.13 The protein structure-based candidate substance discovery system according to the embodiment can be implemented using the computing device 50 .
[0127] The computing device 50 may include at least one of a processor 501, a memory 502, a storage device 503, a display device 504, a network interface device 505 for connecting to the network 40 to communicate with other objects, and an input / output interface device 506 for providing a user input interface or a user output interface. Of course, although Fig.13 Not shown in the figure, the computing device 50 may also include any electronic device required to implement the technical concepts recorded in this specification.
[0128] The processor 501 may be implemented in various types such as an application processor (AP), a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), etc., and may be any electronic device that executes programs or instructions stored in the memory 502 or the storage device 503. In particular, the processor 501 may be configured to implement Figures 1 to 12 Related to the aforementioned functions or methods, according to an embodiment of the present invention, artificial intelligence specific operations related to the protein structure-based candidate substance discovery system and its operating method can be processed on a GPU or NPU.
[0129] The memory 502 and the storage device 503 may include volatile or non-volatile storage media in various forms. For example, the memory 502 may include a read-only memory (ROM) or a random access memory (RAM), and the memory 502 may be located inside or outside the processor 501, and may be connected to the processor 501 in various known ways. In addition, examples of the storage device 503 may include a hard disk drive (HDD) or a solid state drive (SSD), etc. The scope of the present invention is not limited to the components listed above for illustration.
[0130] The protein structure-based candidate substance discovery system and the operation method thereof according to the embodiment can be implemented by a program or software running on the computing device 50 . The program or software described above can be stored in a computer-readable medium.
[0131] In addition, the protein structure-based candidate substance discovery system and the operation method thereof according to the embodiment may be implemented using the hardware of the computing device 50 , or may be implemented by separate hardware that can be electrically connected to the computing device 50 .
[0132] According to the embodiments described so far, functions and user interfaces optimized for candidate substance discovery are provided, which improves the problem in the prior art that detailed tasks regarding candidate substance discovery are provided as separate tools with low mutual compatibility, making it difficult to share data and manage data consistently. The simulation process can be easily managed by modifying and deleting nodes, and by generating, changing, and deleting nodes.
[0133] In this regard, the binding affinity between the target protein and the ligand is predicted and provided to the user, thereby helping the user to intuitively understand the result value, so that the user can easily calculate the molecular weight and volume of the culture medium of the candidate substance (e.g., a new drug candidate substance) without additional additional experiments, and directly apply it to clinical trials such as cell experiments. Therefore, the time and cost consumed in determining the experimental concentration of the desired candidate substance can be significantly reduced. In particular, as described above, the experimental concentration is provided through the cloud platform, so it has the advantage that the user can easily access the experimental concentration value anywhere that can be connected to the Internet and facilitate collaboration with other users.
[0134] The embodiments of the present invention are described in detail above, but the scope of rights of the present invention is not limited thereto. Various modifications and improvements of the present invention made by persons having general knowledge in the technical field to which the present invention belongs using the basic concepts of the present invention defined in the following claims also fall within the scope of rights of the present invention.
Claims
1. An operating method of a protein structure-based candidate substance discovery system, which is implemented by using a cloud platform that provides users with functions or services required for discovering candidate substances in the form of network services, and comprises the following steps: A first screen is displayed for receiving a plurality of ligands to be docked to a target protein; According to the settings input by the user through the first screen, molecular docking is performed on the plurality of ligands in sequence and estimated binding energies are calculated; Displaying on the second screen a first list consisting of rows including a user interface control (User Interface Control) for user selection, the names of the plurality of ligands, and the calculated binding energy values; For a row selected by a user from the first list of the second screen using the user interface control, predicting a binding affinity corresponding to the protein-ligand binding pose used in the binding energy calculation; as well as A second list consisting of rows including the names of the plurality of ligands, the values of the binding energies, and the values of the binding affinities is displayed in the third screen.
2. The method for operating the protein structure-based candidate substance discovery system according to claim 1, wherein: The binding energy value is displayed as a value according to the unit representing energy, The binding affinity values are shown as values based on molar concentration units.
3. The method for operating the protein structure-based candidate substance discovery system according to claim 2, wherein: The binding energy values are shown in Kcal / mol. The binding affinity values are shown in fM, pM, nM, μM, mM or M units.
4. The method for operating the protein structure-based candidate substance discovery system according to claim 3, wherein: The step of displaying the second list in the third screen comprises the following steps: The rows constituting the second list are arranged in descending or ascending order based on the binding affinity values and displayed on the third screen.
5. The method for operating the protein structure-based candidate substance discovery system according to claim 2, wherein: The binding affinity is predicted by an artificial intelligence model that is learned using experimentally measured values of at least one of a dissociation constant Kd, an inhibition constant Ki, and a half maximal inhibitory concentration IC50 as learning data.
6. The method for operating the protein structure-based candidate substance discovery system according to claim 5, wherein: The artificial intelligence model learns using the experimental measurement values and the binding structure data of the target protein and the ligand as the learning data.
7. The method for operating the protein structure-based candidate substance discovery system according to claim 6, wherein: The artificial intelligence model includes: a convolution layer portion, comprising a filter for encoding a pattern to be identified for predicting a binding affinity between the target protein and the ligand; and The dense layer part integrates the features extracted by the convolutional layer part.
8. The method for operating the protein structure-based candidate substance discovery system according to claim 7, wherein: The convolutional layer part includes three convolutional layers with 64, 128 and 256 filters respectively.
9. The method for operating the protein structure-based candidate substance discovery system according to claim 8, wherein: The dense layer part includes three dense layers having 1000, 500 and 200 neurons respectively.
10. The method for operating the protein structure-based candidate substance discovery system according to claim 5, wherein: The artificial intelligence model includes a convolutional neural network (CNN) model or a residual network 3D (ResNet 3D) model.
11. A protein structure-based candidate substance discovery system, which is implemented by using a cloud platform that provides users with functions or services required for discovering candidate substances in the form of network services, and includes: A project management module, which generates a project capable of adding a task for performing a discovery of candidate substances based on protein structures; A simulation management module generates simulations required by users on the generated projects; a simulation setting module for setting a simulation flow of the simulation received from a user using a simulation setting area including a task module selection area, the task module selection area including a first object dragged and dropped into a canvas area and changed to a first node for calculating an estimated binding energy based on docking of a target protein with a ligand, and a second object dragged and dropped into the canvas area and changed to a second node for predicting a binding affinity corresponding to a protein-ligand binding pose used in the binding energy calculation calculated at the first node; as well as The simulation process management module manages information about nodes that can be executed first or later in the simulation process setting.
12. The protein structure-based candidate substance discovery system according to claim 11, wherein: When the run button displayed on the first node is clicked, a first screen for receiving a plurality of ligands to be docked with the target protein is displayed.
13. The protein structure-based candidate substance discovery system according to claim 12, wherein: The first node sequentially performs molecular docking on the plurality of ligands according to settings input by a user through the first screen, and calculates estimated binding energies.
14. The protein structure-based candidate substance discovery system according to claim 13, wherein: After the calculation of the binding energy is completed, a first list consisting of rows including a user interface control for user selection, the names of the plurality of ligands, and the calculated values of the binding energy is displayed on the second screen.
15. The protein structure-based candidate substance discovery system according to claim 14, wherein: The second node predicts the binding affinity corresponding to the protein-ligand binding pose used in the binding energy calculation for a row selected by the user from the first list on the second screen using the user interface control.
16. The protein structure-based candidate substance discovery system according to claim 15, wherein: After the prediction of the binding affinity is completed, a second list consisting of rows including the names of the plurality of ligands, the values of the binding energy, and the values of the binding affinity is displayed on the third screen.
17. The protein structure-based candidate substance discovery system according to claim 16, wherein: The case where the second list is displayed on the third screen includes the case where the rows constituting the second list are arranged in descending or ascending order based on the values of the binding affinity and are displayed on the third screen.
18. The protein structure-based candidate substance discovery system according to claim 11, wherein: The binding energy values are shown in Kcal / mol. The binding affinity values are shown in fM, pM, nM, μM, mM or M units.
19. The protein structure-based candidate substance discovery system according to claim 11, wherein: The binding affinity is predicted by an artificial intelligence model that is learned using experimentally measured values of at least one of a dissociation constant Kd, an inhibition constant Ki, and a half-inhibitory concentration IC50 as learning data.
20. The protein structure-based candidate substance discovery system according to claim 11, wherein: The simulation process management module manages information about nodes that can be executed first or later in the simulation process setting through metadata.