Method and device for screening lead compound candidate group based on compound-protein binding prediction

The method addresses the limitations of static protein structures and high computational costs by using molecular dynamics simulations and AI modeling to predict compound-protein binding, enhancing the efficiency and accuracy of lead substance selection.

WO2026019050A1PCT designated stage Publication Date: 2026-01-22TINACLON CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/007114
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-15
Filing Date
2025-05-26
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing molecular docking techniques rely on static protein structures, limiting their ability to reflect dynamic changes in vivo, and molecular dynamics simulations require high computational costs and specialized personnel, making them unsuitable for rapid and economical prediction and screening of numerous candidate compounds.

Method used

A method involving molecular dynamics simulations to capture various protein structures, followed by artificial intelligence modeling to predict compound-protein binding affinities, using regression analysis and deep learning to select lead substance candidates based on true pose-based feature values.

Benefits of technology

Enables efficient selection of lead substances by predicting activity accurately and reducing resource wastage, improving prediction accuracy by utilizing high-confidence feature values derived from true poses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025007114_22012026_PF_FP_ABST
    Figure KR2025007114_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a technology for increasing screening efficiency of lead compounds in new drug development. More specifically, the present invention relates to a method for screening a lead compound candidate group, the method comprising the steps of: selecting a plurality of target protein structures on the basis of a molecular dynamics simulation of a target protein; training an artificial intelligence model on the basis of a learning library composed of derivatives of an active substance; and predicting activity of each compound included in an analysis target library composed of novel compounds to be analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for screening lead compound candidates based on compound-protein binding prediction

[0001] The present invention relates to a technology for increasing the efficiency of selection of a leading substance in the development of a new drug, and more specifically, to a technology for reducing cost and time and increasing the efficiency of selection of a leading substance by improving the prediction of binding between a 'compound and a target protein' in the early stages of new drug development.

[0002]

[0003] With the recent rapid advancements in computer hardware performance, the introduction of artificial intelligence technologies such as deep learning and machine learning is revolutionizing various fields. Active efforts are also being made to utilize these technologies in new drug development. In particular, various computing-based technologies are being developed to identify target proteins responsible for diseases based on biological mechanisms and design bioactive compounds, antibodies, RNAs, and other substances that bind to these proteins and inhibit or modulate their function. Among these, molecular docking and molecular dynamics simulations are widely used in compound-based new drug development to analyze the binding structure of target proteins for structure-based drug design.

[0004] Molecular docking is a useful technique for identifying compounds that can block specific regions of a target protein or modulate its function through binding. It is particularly crucial in virtual screening processes to identify novel compounds with potential pharmacological effects. Recently, its application has expanded significantly thanks to improvements in computer performance and its integration with artificial intelligence technology.

[0005] Meanwhile, molecular dynamics simulation is a technique that analyzes protein-compound interactions over time. It is used to calculate protein structural changes, binding stability, and free energy. This technology is used in diverse fields such as engineering, chemistry, biology, and physics, and plays a particularly important role in new drug development, enabling precise analysis of interactions at the molecular level.

[0006] However, molecular docking-based techniques often rely on static protein structures, limiting their ability to adequately reflect the dynamic changes in protein structures in vivo. Furthermore, while free energy calculations utilizing molecular dynamics simulations promise greater accuracy, they require repeated calculations for the same binding structure, resulting in significant computational costs. Furthermore, their use requires high-performance equipment and specialized personnel.

[0007] Moreover, most existing technologies focus on precisely predicting the binding affinity of individual compounds, making them unsuitable for economical and rapid prediction and screening of numerous candidate compounds.

[0008] Accordingly, there is a need to develop a technology that can efficiently predict binding affinities for a large number of compounds by utilizing various protein structures derived from molecular dynamics simulations and fusing them with artificial intelligence.

[0009]

[0010] The technical problem to be solved by the present invention is to provide a method for selecting and prioritizing a group of lead substance candidates based on activity prediction values ​​for new compounds to be analyzed.

[0011] Another problem that the present invention seeks to solve is to provide a regression-based artificial intelligence model that quantitatively predicts the activity of an effective substance based on true pose-based feature values.

[0012] The technical problems to be achieved in the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.

[0013]

[0014] In order to solve the above or other problems, according to one aspect of the present invention, a method for selecting a group of lead substance candidates is provided, comprising the steps of: selecting a plurality of target protein structures based on molecular dynamics simulations for the target protein; training an artificial intelligence model based on a training library composed of derivatives of effective substances; and predicting the activity of each compound included in an analysis target library composed of novel compounds to be analyzed.

[0015] The step of selecting the plurality of target protein structures may include the step of performing a molecular dynamics simulation on a complex of the target protein and the effective substance; the step of setting a stabilization interval based on the root mean square deviation (RMSD) of the backbone elements from the simulation results; and the step of selecting the plurality of sampled structures from the set stabilization interval.

[0016] The step of learning the artificial intelligence model may include: a step of molecular docking each derivative included in the learning library to the selected plurality of target protein structures; a step of selecting a pose that satisfies at least one of RMSD and CNN-based prediction value conditions among the molecularly docked derivatives as a true pose; and a step of learning the artificial intelligence model based on characteristic values ​​corresponding to the selected true pose.

[0017] The above AI model may include regression analysis, but other AI learning methods (such as utilizing false poses) may also be used.

[0018] The step of learning the artificial intelligence model may include a step of statistically selecting characteristic values ​​highly related to the experimental values ​​of the compound; and a step of learning based on the selected characteristic values.

[0019] The step of predicting the activity of the compound may include the step of selecting a pose that satisfies at least one of RMSD and CNN-based prediction value conditions among each compound included in the analysis target library as a true pose; the step of inputting the characteristic value of the selected true pose into the artificial intelligence model; and the step of predicting the activity value of the lead substance candidate compounds based on the output of the artificial intelligence model and determining their priority.

[0020]

[0021] The effects of the method and device for selecting a leading material according to the present invention are described as follows.

[0022] According to at least one of the embodiments of the present invention, there is an advantage in that a lead substance can be selected without wasting resources by predicting the activity of a novel compound in advance.

[0023] In addition, according to at least one of the embodiments of the present invention, there is an advantage in that prediction accuracy can be improved by utilizing only high-confidence feature values ​​derived based on true poses.

[0024] Further scope of the applicability of the present invention will become apparent from the detailed description below. However, since various modifications and variations within the spirit and scope of the present invention will become apparent to those skilled in the art, it should be understood that the detailed description and specific examples, such as preferred embodiments of the present invention, are given by way of example only.

[0025]

[0026] FIG. 1 is a diagram illustrating a control flowchart of a method for selecting a lead substance candidate group based on compound-protein binding prediction according to an embodiment of the present invention.

[0027] FIG. 2 is a drawing showing an example of a target protein (201) and an effective substance (202) prepared according to one embodiment of the present invention.

[0028] FIG. 3 is a drawing for explaining the concept of sampling multiple target proteins (201-1 to 201-3) in a molecular dynamics simulation process of a target protein (201) and an effective substance (202) complex according to one embodiment of the present invention.

[0029] Figure 4 shows the experimental values ​​(IC) of scaffold-based derivatives of the effective substance (202). 50 This is a diagram showing a configuration in which existing compounds are separated into a learning library (401) and a verification library (402) and utilized for learning an artificial intelligence model.

[0030] FIG. 5 is a diagram illustrating a process of molecular docking (520) of a separated learning library (401) and a verification library (402) to sampled target proteins (201-1 to 201-3), followed by deep learning pose prediction and collection of target protein and compound-derived feature values ​​according to the molecular docking (520) of the separated learning library (401) and the verification library (402) according to one embodiment of the present invention.

[0031] Fig. 6 is a diagram illustrating the configuration of a leading material selection device according to one embodiment.

[0032]

[0033] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably only for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. In addition, when describing the embodiments disclosed in this specification, if it is determined that a specific description of a related known technology may obscure the gist of the embodiments disclosed in this specification, a detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present invention.

[0034] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.

[0035] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0036] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0037] In this application, terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in the specification, but should be understood not to exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0038]

[0039] FIG. 1 is a diagram illustrating a control flowchart of a method for selecting a lead substance candidate group based on compound-protein binding prediction according to an embodiment of the present invention.

[0040] As used herein, the term "lead compound" refers to a compound whose activity has been improved to a level suitable for entering preclinical testing through a structural optimization process from an active substance (hit), and functions as a key substance in the stage prior to a clinical candidate.

[0041] Step S101 involves preparing structural data for target proteins and active substances and preprocessing them into a form usable for subsequent simulations and AI learning. During the new drug development process, target proteins associated with specific diseases are identified, active substances (compounds) with reported activity against these proteins are secured, and this information is then input into a computer-based system.

[0042] FIG. 2 is a drawing showing an example of a target protein (201) and an effective substance (202) prepared according to one embodiment of the present invention.

[0043] The target protein (201) is generally preprocessed into a form suitable for molecular dynamics simulation and molecular docking by obtaining an X-ray crystal structure, NMR structure, or AlphaFold-based predicted structure from a publicly available protein database (e.g., Protein Data Bank). The preprocessing process may include adding hydrogen atoms, adjusting ion states, and assigning binding parameters.

[0044] The active substance (202) is a representative compound known to exhibit a certain level of activity or higher against a target protein (201) and is a starting material based on a single scaffold structure. The active substance (202) includes a structure whose effectiveness has been experimentally confirmed and has the potential to be optimized as a leading substance. In addition, the active substance (202) functions as a basic structure that serves as the center of a learning compound library (learning library) to be constructed in a subsequent step, and structural analogs or derivatives are expanded based on this to construct the library.

[0045] The structure of the effective substance (202) input in this step is generally expressed in a 2D format (SMILES, InChI, etc.) or a 3D format (SDF, MOL2, etc.), and serves as the basis for feature extraction in subsequent docking and AI-based predictive modeling. In addition, a validation set of effective substance derivatives (validation library, validation set) for evaluating the performance of the artificial intelligence model, separate from the learning step, can also be defined in this step.

[0046]

[0047] Step S102 is a step of performing a molecular dynamics simulation (hereinafter referred to as MD simulation) on the input target protein (201) and sampling multiple target protein (201) structures that reflect a stable structural state among the structures of various time frames generated through the simulation.

[0048] FIG. 3 is a drawing for explaining the concept of sampling multiple target proteins (201-1 to 201-3) in a molecular dynamics simulation process of a target protein (201) and an effective substance (202) complex according to one embodiment of the present invention.

[0049] Typically, molecular docking-based binding affinity predictions are performed based on a fixed protein structure (the static state in the PDB). However, in vivo proteins are constantly in flux, with their structures constantly changing over time. Therefore, relying solely on fixed structures may not reflect actual binding affinity. To overcome this, the present invention proposes securing various structural states based on MD simulations to reflect the dynamic state of the target protein (201).

[0050] MD simulations are typically performed for a time period of tens to hundreds of nanoseconds. In the present invention, for example, simulations are performed for 200 ns, and then the RMSD (Root-Mean-Squared Deviation) value of the protein backbone is analyzed to identify a structurally stable region (310). In this stable region (310), multiple structures (three in the illustrated example, but in reality, more are selected) at which the target protein (201) is judged to be maintained without significant structural change are selected, and these structures are defined as protein structures to be used for subsequent docking and artificial intelligence learning.

[0051] Sampling criteria may include one or more of the following:

[0052] - Selection of stable interval based on RMSD or RMSF

[0053] - Selection of representative structures based on hydrogen bond maintenance patterns

[0054] - Accessibility of molecular docking technology (securing a large number of true poses of compounds)

[0055] - Clustering of compound-protein complexes or structural grouping using PCA (principal component analysis)

[0056] - Utilization of advanced sampling techniques such as replica exchange and umbrella sampling

[0057] The multiple structures (201-1 to 201-3) selected in this way are used for docking and binding predictions that reflect multiple states, and enable more realistic and accurate predictions than predictions based on fixed protein structures. Consequently, this step may be a step that provides a foundation for reflecting dynamic structural diversity in the docking and artificial intelligence modeling process. Recently, with the development of deep learning technologies such as AlphaFold3, it has become possible to sample protein-compound binding structures using deep learning technologies rather than molecular dynamics simulations, and the present invention may also include a configuration that utilizes such predicted-based binding structures as alternative inputs.

[0058]

[0059] S103 This step is a process of predicting the binding pose and binding characteristics that each derivative can form by docking chemically designed derivatives based on the scaffold structure of the effective substance (202) for various structures of the target protein (201), and securing the features required for subsequent artificial intelligence learning.

[0060] Figure 4 shows the experimental values ​​(IC) of scaffold-based derivatives of the effective substance (202). 50 This is a diagram showing a configuration in which compounds that exist (e.g., etc.) are separated into a learning library (401) and a verification library (402) and utilized for learning an artificial intelligence model.

[0061] FIG. 5 is a diagram illustrating a process of molecular docking (520) of a separated learning library (401) and a verification library (402) to sampled target proteins (201-1 to 201-3), followed by deep learning pose prediction and collection of target protein and compound-derived feature values ​​according to the molecular docking (520) of the separated learning library (401) and the verification library (402) according to one embodiment of the present invention.

[0062] The target protein (201) to be docked is a plurality of structural states (e.g., 201-1, 201-2, 201-3, etc.) sampled through molecular dynamics simulation in step S102, and the derivatives are included in the learning library (401) composed by variously modifying the substituents or functional groups while maintaining the scaffold of the representative effective substance (202) input in step S101. These derivatives are experimentally IC 50 These are compounds whose biological activities have been measured, and the correlation between docking results and experimental values ​​serves as the basis for learning the prediction model.

[0063] Each derivative is individually docked against multiple target protein (201) structures, and the following information is derived from this process:

[0064] - Binding position and orientation (binding pose) of the derivative

[0065] - Docking-based binding energy or interaction score

[0066] - Binding patterns such as hydrogen bonding and hydrophobic interactions with amino acid residues of the target protein (201)

[0067] - 2D properties (molecular weight, presence of hydrogen bonds, etc.) and 3D properties (radius of gyration, etc.) of the compound itself

[0068] Specifically, the present invention, rather than using a single pose, generates multiple poses (e.g., the top 10) per derivative to ensure structural diversity. In the subsequent step S104, only highly reliable combined poses (true poses) are selected based on defined criteria (such as root mean square deviation (RMSD) and CNN prediction values). By securing a large number of poses in this way, a high-quality dataset for AI learning can be constructed.

[0069]

[0070] Step S104 is a process of selecting only poses with a high probability of actual binding (true poses) among multiple binding poses generated through docking for the derivative compound in step S103, and extracting features required for artificial intelligence learning from the poses.

[0071] Multiple docking poses are generated for each derivative, and in order to use only poses that well reflect actual binding with the target protein (201) as input values ​​for the subsequent artificial intelligence model, the following filtering method is applied in this step. Each filtering method can be applied individually, or two filtering methods can be combined.

[0072] 1) RMSD-based structural similarity assessment

[0073] The positional similarity of the maximum common substructure (MCS) between the docked derivative pose and the binding reference structure (simulation-derived structure) derived from the target protein (201)-effective substance (202) complex structure obtained through molecular dynamics simulation (MD) is calculated as the root-mean-squared deviation (RMSD) value.

[0074] If this RMSD value is less than 2-3 Å, the pose can be considered a true pose candidate.

[0075] 2) Deep learning-based combined pose reliability assessment

[0076] This method predicts the combinatorial likelihood of each pose using a deep learning model based on a Convolutional Neural Network (CNN). Only poses with a CNN model output value (CNN score) of 0.6 to 0.7 or higher are considered to meet the reliability criterion.

[0077] Only poses that satisfy both of the above conditions or one of the conditions are defined as true poses, and feature extraction is performed only for these.

[0078] The characteristic values ​​include:

[0079] ·Docking-based variables: binding energy, docking score, etc.

[0080] ·Deep learning-based variables: CNN output values ​​(binding affinity prediction values, etc.)

[0081] ·2D molecular information: molecular weight, number of hydrogen-bondable elements, polar surface area, etc.

[0082] ·3D molecular information: molecular gyration radius, structural complexity, etc.

[0083] ·Interaction information: number of hydrogen bonds, number of hydrophobic interactions, number of bound amino acids, etc.

[0084] The feature values ​​extracted in this way are used as input variables for artificial intelligence regression model learning performed in step S105.

[0085] In other words, this step corresponds to the preprocessing step that determines the reliability and accuracy of the learning dataset.

[0086]

[0087] Step S105 is to extract various features based on the true pose in step S104, and then select the experimental value (IC) 50 This is the process of selecting only the feature values ​​with high correlation (etc.) and constructing a feature set to be used as input values ​​for subsequent artificial intelligence regression model learning.

[0088] The characteristic values ​​derived for each derivative of the effective substance (202) may include the following items:

[0089] ·Binding-related variables: binding energy obtained from molecular docking results, CNN-based binding prediction values, etc.

[0090] ·2D molecular structure information: molecular weight, number of hydrogen-bondable elements, polar surface area, etc.

[0091] ·3D molecular structure information: gyration radius, structural complexity, etc.

[0092] ·Interaction information with the target protein (201): number of hydrogen bonds, number of hydrophobic interactions, number of binding residues, etc.

[0093] Among these various characteristic values, the predicted value (e.g. IC 50 ) may also include variables that are less related to the model or may cause problems such as multicollinearity.

[0094] Accordingly, in the present invention, a feature selection procedure is performed to quantitatively analyze the relevance (correlation, explanatory power, etc.) that each feature value has with the predicted target value, and then select only statistically significant variables.

[0095] Specifically, each feature and IC are used using modules such as f_regression, r_regression, or SelectKBest provided by Sci-kit learn, a Python-based machine learning library. 50 Pearson correlation coefficient, analysis of variance (F-score), p-value, etc. are calculated, and only characteristic values ​​that satisfy a certain standard are selected.

[0096] The feature value set selected in this way is used as an optimized input value in the learning of artificial intelligence regression models after S106, and contributes to improving the learning efficiency of the model.

[0097]

[0098] Step S106 uses the characteristic values ​​selected in step S105 as input to obtain the experimental values ​​(IC) for the derivatives of each effective substance (202). 50 It is a process of learning a regression-based artificial intelligence model that can predict (etc.).

[0099] The input data to be learned may be (1) a derivative of a valid substance (202) having a true pose derived in step S104, and (2) data composed only of variables selected in step S105 among the characteristic values ​​extracted from the pose.

[0100] For each derivative satisfying these conditions, a regression-based artificial intelligence model is trained by mapping input values ​​(selected characteristic values) and output values ​​(experimental values), thereby building a model that can quantitatively predict the activity of new compounds in the future.

[0101] Model training typically involves one or more of the following regression algorithms:

[0102] · Linear regression

[0103] Ridge regression

[0104] · Lasso regression

[0105] · Support Vector Regression (SVR)

[0106] · Random Forest Regression

[0107] · Gradient Boosting Machine (GBM), etc.

[0108] The learned regression model produces predicted values ​​(activation values) from the characteristic values ​​of the true pose of each derivative, and the model performance is evaluated by the correlation (R) between the experimental values ​​and predicted values ​​for the separately secured verification derivatives (verification library) in step S107. 2 , Pearson's r, etc.) are evaluated. However, model learning is not necessarily limited to regression models, and other modeling methods such as deep learning can also be applied. In such cases, false pose information can also be included in the learning process.

[0109]

[0110] Step S107 is a process of analyzing the quantitative correlation between predicted values ​​and experimental values ​​by utilizing a verification library (402) consisting of separately separated verification compounds to evaluate the performance of the artificial intelligence regression model learned in step S106.

[0111] The verification library (402) is based on the scaffold of the effective substance (202) input in steps S101 to S102, but is composed of derivative compounds that were not used in model learning. All of these are based on experimental values ​​(IC 50 ) and serves as an independent test set to verify the generalization performance of artificial intelligence models.

[0112] The verification process consists of the following steps:

[0113] 1. Performing derivative docking for verification

[0114] As in the training step, docking was performed on multiple target protein structures (201-1 to 201-3) for each validation derivative.

[0115] Among the derived poses, only true poses that satisfy the RMSD and CNN prediction criteria are selected.

[0116] 2. True pose-based feature extraction

[0117] Composed of the same items as the characteristic value items selected in step S105

[0118] 3. Calculating model prediction values

[0119] The extracted feature values ​​are input into the regression model trained in step S106 to obtain the predicted value (IC 50 (Predicted value) Production

[0120] 4. Calculating performance evaluation indicators

[0121] Predicted values ​​and actual experimental values ​​(IC 50 ) compared, (1) coefficient of determination R 2 (2) Calculate various model performance indicators such as Pearson's correlation coefficient (Pearson's r) and (3) mean square error (MSE).

[0122] The purpose of this step is to quantitatively evaluate the reliability and generalization ability of the entire prediction system by determining how accurately the regression model can predict experimental values ​​even for untrained derivatives.

[0123]

[0124] Step S108 is a step to predict the activity of each compound included in the analysis target library consisting of new compounds whose activity has not been experimentally confirmed, by utilizing the artificial intelligence regression model that has been trained and verified through steps S106 and S107.

[0125] The library for analysis is composed of newly designed derivatives based on the scaffold structure of the active substance (202), and the compounds are composed for the purpose of evaluating their activity at a stage prior to actual synthesis.

[0126] Each compound included in the analysis target library is docked against multiple target protein (201) structures sampled in step S102, and among the multiple poses generated therefrom, only binding poses satisfying the RMSD and CNN prediction value criteria are selected as true poses. From the selected true poses, structural and interaction-based features identical to the feature values ​​selected in step S105 are extracted, and these features are provided as input values ​​for the artificial intelligence regression model trained in step S106.

[0127] The model provides quantitative predictions (IC) for each novel compound. 50 etc.) and this predicted value reflects the relative activity intensity. Based on the predicted activity value, a relatively low IC 50Compounds determined to possess value are considered to have high activity against the target protein (201) and can be selected as priority targets for subsequent synthesis or biological evaluation. This prediction-based candidate selection maximizes the efficiency of new drug discovery prior to actual experiments and minimizes unnecessary synthesis and evaluation resources.

[0128] In summary, this step is a key prediction step designed to simultaneously improve cost-effectiveness and time-efficiency in the new drug development process by pre-evaluating the activity of new compounds to be analyzed using an artificial intelligence-based prediction model and quickly selecting substances likely to exhibit high activity against the target protein (201).

[0129]

[0130] Meanwhile, the present invention can be equally applied to the purpose of virtual screening, which preemptively predicts the activity of a set of new compounds for which no experimental values ​​exist, as well as predicting the activity of derivatives of effective substances.

[0131] Virtual screening is a technology for selecting effective substances (hits) that can bind to a specific target protein (201) from a large compound library, and is generally used when it is necessary to derive candidate substances through structure-based prediction without experimental data.

[0132] The prediction system of the present invention is based on an artificial intelligence model learned as a derivative of an effective substance (202), but its application is not limited to the analysis target library, and can be expanded to an external compound library composed of new compounds of unknown effectiveness.

[0133] These compound libraries can be constructed from in-house designs, commercial large libraries (e.g., ZINC, Enamine), or generative AI-based compound candidates.

[0134]

[0135] Compounds without experimental data cannot be directly used for model training, but can be applied in the same way to the regression-based prediction model trained in step S106.

[0136] To this end, docking is performed on multiple target protein (201) structures for each new compound in the same manner as steps S103 to S104, a true pose is selected based on RMSD and CNN prediction values, and the feature values ​​derived based on the true pose are aligned in the same configuration as the feature set selected in step S105 and input into the model.

[0137] Each of the new compounds entered in this way has a quantitative activity prediction value (IC) without experimental data. 50 etc.) can be obtained, and compounds with excellent prediction values ​​can be selected as hit candidates before actual experiments, allowing rapid entry into the preclinical exploration phase.

[0138] This structure has the technical effect that the core structure of the present invention can be naturally applied to a virtual screening environment in that it provides a fast and efficient effective substance discovery route compared to conventional high-cost HTS (High Throughput Screening) experiments.

[0139] In summary, the present invention has the advantage of being able to apply the same technical configuration to a virtual screening pipeline for selecting initial substances at the hit level by performing quantitative predictions on novel compounds without experimental values ​​by utilizing learning results based on effective substance derivatives.

[0140]

[0141] Furthermore, the artificial intelligence-based activity prediction technology of the present invention can be equally applied to biopharmaceuticals of various modalities, such as peptides, antibody therapeutics, and other protein drugs, in addition to small molecule compounds.

[0142] The reason why this expansion is possible is that the core structure of the present invention constructs a prediction model based on features derived based on the binding pose between the target protein (201) and the candidate substance.

[0143] Since it is based on a generalized feature set based on the binding structure, such as binding energy, number of interactions, and structure-based variables, the binding molecule does not necessarily have to be a small molecule, and it can also be applied to peptides with certain conformational degrees of freedom or protein drugs with limited structures.

[0144] For example, although peptide-based materials have greater flexibility than conventional ligands, they can form binding poses based on molecular dynamics simulations and docking, and quantitative properties such as hydrogen bonding, hydrophobic interactions, binding surface area, and binding energy can be derived from those poses.

[0145] This is also technically compatible with securing a high-confidence pose through the true pose selection method (RMSD, CNN-based filtering) defined in the present invention.

[0146]

[0147] In addition, antibody therapeutics traditionally require structure-based design of antigen-antibody interactions, and docking and feature extraction of specific sequences (e.g., CDR) or fragments (Fab, scFv, etc.) of antibodies based on protein-protein binding (PPI) modeling are possible, so they can be directly applied to the artificial intelligence learning structure of the present invention.

[0148] These diverse modality candidates are a major trend in the development of new drugs, and the present invention, due to the flexibility and universality of its predictive logic, can be effectively utilized not only for small molecule compounds but also for the screening and optimization process of bio-derived polymer-based candidates.

[0149] In summary, the binding structure-based prediction technology of the present invention can be expanded to all modalities that can structurally bind to a target protein, such as peptides, antibodies, or other protein-based drugs, which can greatly improve the versatility and efficiency of the candidate selection and optimization process in the early stages of drug discovery.

[0150]

[0151] Fig. 6 is a diagram illustrating the configuration of a leading material selection device according to one embodiment.

[0152] Referring to FIG. 6, the leading material sorting device includes a processor (601) and a memory (602). The memory (602) stores one or more instructions executable by the processor (601). The processor (601) executes one or more instructions stored in the memory (602). By executing the instructions, the processor (601) can perform one or more operations described above with respect to FIGS. 1 to 5.

[0153]

[0154] Hereinafter, embodiments of the method and device for selecting a leading material according to the present invention have been described, but this is described as at least one embodiment, and the technical idea of ​​the present invention and its configuration and operation are not limited thereby, and the scope of the technical idea of ​​the present invention is not limited / restricted by the drawings or the description referring to the drawings. In addition, the concept and embodiment of the invention presented in the present invention may be used by a person having ordinary skill in the technical field to which the present invention pertains as a basis for modifying or designing another structure to perform the same purpose of the present invention, and an equivalent structure modified or changed by a person having ordinary skill in the technical field to which the present invention pertains is bound by the technical scope of the present invention described in the claims, and various changes, substitutions, and modifications are possible within the scope that does not depart from the spirit or scope of the invention described in the claims.

Claims

1. A step of selecting multiple target protein structures based on molecular dynamics simulation for the target protein; A step of learning an artificial intelligence model based on a learning library composed of derivatives of effective substances; and Comprising a step of predicting the activity of each compound included in the analysis target library composed of novel compounds to be analyzed, Method for selecting lead material candidates.

2. In the first paragraph, the step of selecting the plurality of target protein structures comprises: A step of predicting the combination of a target protein and an effective substance using molecular dynamics simulation or deep learning; and comprising the step of selecting a plurality of sampled structures; Method for selecting lead material candidates.

3. In the first paragraph, the step of learning the artificial intelligence model comprises: A step of molecular docking each derivative included in the above learning library to the selected plurality of target protein structures; A step of selecting a pose that satisfies at least one of the RMSD and CNN-based predicted value conditions among the above molecular docked derivatives as a true pose; and A step of learning an artificial intelligence model based on characteristic values ​​corresponding to the above-mentioned selected true pose and false pose, Method for selecting lead material candidates.

4. In paragraph 3, The above artificial intelligence model includes regression analysis, Method for selecting lead material candidates.

5. In the fourth paragraph, the step of learning the artificial intelligence model comprises: A step of statistically selecting characteristic values ​​highly related to the experimental values ​​of the compound; and Including a step of learning based on the above-mentioned selected characteristic values, Method for selecting lead material candidates.

6. In the first paragraph, the step of predicting the activity of the compound comprises: A step of selecting a pose that satisfies at least one of the RMSD and CNN-based prediction value conditions among each compound included in the above analysis target library as a true pose; A step of inputting the characteristic values ​​of the above-described selected true pose or false pose into the artificial intelligence model; and A step of determining priority as a leading material based on the output of the artificial intelligence model, Method for selecting lead material candidates.

7. In paragraph 1, The compounds included in the library to be analyzed are composed of compounds whose experimental values ​​are unknown. Applied to virtual screening that derives effective substances through the same prediction procedure, Method for selecting lead material candidates.

8. In the first paragraph, the analysis target library that is the prediction target is, Including a drug candidate of a modality comprising at least one of a peptide, an antibody, and an antibody therapeutic agent in addition to a small molecule compound; Method for selecting lead material candidates.

Citation Information

Patent Citations

  • Drug molecule processing method and apparatus based on artificial intelligence, and device, storage medium and computer program product

    EP4239640A1

  • Drug indication and response prediction systems and method using AI deep learning based on convergence of different category data

    KR101953762B1

  • Electric Crop Harvesting Moving Device for Driving Motor by Sensing Torque Generated by Movement of Foot and Method of Operation thereof

    KR1020230049517A

  • System and method for process simulation

    KR1020250143610A