Virtual compound screening
The Extra Trees model integrated with docking enhances ligand prediction efficiency and accuracy by training on diverse decoys, addressing the limitations of existing methods and enabling rapid identification of candidate ligands for protein targets, thus accelerating drug development.
Patent Information
- Application Number
- PCT/SG2025/050215
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-02
AI Technical Summary
Existing ligand prediction methods for protein targets from large chemical libraries are time-consuming, resource-intensive, and lack generalizability due to biased and insufficient training data, leading to inaccurate predictions and high costs.
A virtual screening method using an Extremely Randomized Trees (Extra Trees) machine learning model integrated with molecular docking, which predicts interacting compounds by training on diverse and structurally/chemically randomized decoys, enabling efficient identification of candidate ligands for protein targets.
The method significantly reduces the time and computational resources required for ligand prediction, allowing screening of ultra-large chemical libraries and identifying novel ligands for new targets, thereby accelerating drug development.
Smart Images

Figure SG2025050215_02102025_PF_FP_ABST
Abstract
Description
[0001] VIRTUAL COMPOUND SCREENING
[0002] Technical Field
[0003] The present invention relates, in general terms, to predicting interacting compounds (ligands) for protein targets from a chemical library. More particularly, the present invention relates to, but is not limited to, methods for predicting ligands for protein targets from large chemical libraries.
[0004] Background
[0005] Ligands bind to proteins often to induce or prevent the activity of the proteins. Identifying such ligands enables drug development and repurposing. However, doing so reliably by using experimental assays is time-consuming and expensive. Long-standing docking-based virtual screening approaches are faster but still scale poorly in terms of time and resources when creating predictions for ultra-large chemical libraries for multiple targets. Specifically, optimized approaches take at least 1 second to score each complex; screening multiple billion compounds in a practical time thus necessitates the use of parallelization in large computing clusters of thousands of processors, which can cost hundreds of thousands of US dollars to buy, and tens of thousands to rent. Knowledge-based models are faster yet but usually also use docking results as input. Therefore, there is a strong need for more efficient and accurate ligand prediction frameworks.
[0006] Machine Learning (ML) models have been proposed for this purpose. Unlike docking and preset knowledge-based approaches, the effectiveness of ML models is largely determined by the nature and quality of data used to algorithmically set ML model parameters during training. Unfortunately, popular datasets for this purpose have limitations, particularly as they relate to the small number and biases of the decoys included. Moreover, studies using decoys for developing and demonstrating ML models often fail to consider whether there is sufficiently diverse and dissimilar training and evaluation data. If that data is insufficient, reported performance can only substantiate model performance for the specific contexts covered by the training data rather than for arbitrary target-compound pairs - i.e., existing models are not generalised. Moreover, these studies typically use evaluation metrics misaligned with the practical objectives, where models are required to identify ligands with high precision in practical situations where they are vastly outnumbered by decoys. In general, while ML-based virtual screening can be both faster and more accurate than more traditional approaches, it is challenging to prove accuracy will transfer to future models for other tasks.
[0007] It would be desirable to provide a methodology for accelerating ligand prediction that overcomes or ameliorate at least one of the above-described problems, or at least to provide a useful alternative.
[0008] Summary
[0009] The methods disclosed herein streamline and accelerate the ligand (interacting compound) prediction process toward any given target. This involves developing a predictive ML model and integrating this ML model with conventional molecular docking-based virtual ligand screening. The ML models generated in accordance with present teachings can operate at the genome / superfamily scale.
[0010] The ML model and the integrated workflow will significantly reduce the need for 1) extensive molecular docking-based virtual ligand screening of ultra-large chemical libraries, and 2) time to develop novel ligands for specific targets.
[0011] Disclosed is a virtual screening method for predicting interacting compounds, from a chemical library, for proteins, the methodology comprising: providing an Extremely Randomized Trees (Extra Trees) machine learning model comprising an ensemble of decision trees, each decision tree comprising: at least one parent node, each parent node representing a decision in relation to a target-compound complex, based on a plurality of features corresponding to a protein target of the complex and a plurality of features corresponding to a compound of the complex; and and at least two leaf nodes, each leaf node classifying the targetcompound complex as either interacting or non-interacting, the Extra Trees model outputting a probability that target-compound complex is interacting, based on aggregated classifications from each said decision tree, receiving a further protein target; applying the Extra Trees model to the further protein target and compounds in the chemical library, to determine a probability, for each compound, that it will form an interacting complex with the further protein target; ranking the candidate compounds based on the respective probabilities, to identify candidate compounds having highest probability; and outputting the candidate compounds having highest probability, to a docking model.
[0012] Also disclosed is a virtual screening system for predicting interacting compounds, from a chemical library, for proteins, the system comprising: memory; a machine learning module comprising an Extremely Randomized Trees (Extra Trees) machine learning model comprising an ensemble of decision trees, each decision tree comprising: at least one parent node, each parent node representing a decision in relation to a target-compound complex, based on a plurality of features corresponding to a protein target of the complex and a plurality of features corresponding to a compound of the complex; and and at least two leaf nodes, each leaf node classifying the targetcompound complex as either interacting or non-interacting, the Extra Trees model outputting a probability that target-compound complex is interacting, based on aggregated classifications from each said decision tree; a ranking engine; and at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the system to: receiving a further protein target; applying the Extra Trees model of the machine learning module to the further protein target and compounds in the chemical library, to determine a probability, for each compound, that it will form an interacting complex with the further protein target; ranking the candidate compounds, using the ranking engine, based on the respective probabilities, to identify candidate compounds having highest probability; and outputting the candidate compounds having highest probability, to a docking model.
[0013] As used herein, the term "interacting compound" refers to a compound that interacts with (e.g., binds to) a protein target. Conversely, a "non-interacting compound" does not interact with the protein target. Similarly, an "interacting complex" refers to a combination of a protein and corresponding interacting compound.
[0014] As used herein, the terms "chemotypically diverse" and "structurally diverse" and similar, such as "chemotypically-randomised" and "structurally- randomised", relates to compound selection where chemotypical randomisation, also referred to as chemotypical diversity, involves preferencing selection of decoys from tranches regardless of the size of those tranches in the training data, and structural diversity or structural randomisation refers to selecting structures from diverse groups regardless of the number of compounds of a particular structure when compared with compounds of a different structure. This facilitates compound and structure selection from rarer tranches and groups, rather than disregarding those tranches and groups in favour of more common tranches and groups.
[0015] Brief description of the drawings Embodiments of the present invention will now be described, by way of nonlimiting example, with reference to the drawings in which :
[0016] Figure 1 illustrates a method for predicting interacting compounds for protein targets - i.e., compounds to form interacting complexes with protein targets.
[0017] Figure 2 shows protein classes in DUD-E and the BindingDB and ChEMBL subsets.
[0018] Figure 3 comprises images A and B, where image A shows the sequence identity between any two targets in DUD-E and BindingDB subset, and image B shows the maximum of the mean pairwise Tanimoto coefficient (Tc) values between the ligands of a target from one dataset and the ligands of any target from the other dataset.
[0019] Figure 4 provides the maximum of the mean pairwise Tc values between the ligands of a target from the ChEMBL subset and the ligands of any target from the BindingDB subset and DUD-E. One histogram of maximum mean pairwise Tc values is provided for each comparison, with that comparing the ChEMBL and BindingDB subsets' ligands displayed over that displaying the ChEMBL subset- DUD-E ligand comparison.
[0020] Figure 5 illustrates a random-style decoy selection procedure using drug-like molecules from the ZINC database.
[0021] Figure 6 comprises a sketch of the ET-Screen and IFP-GNN architectures.
[0022] Figure 7 is an overview of the four steps of a robust model evaluation pipeline.
[0023] Figure 8, comprises images A) and B), Logarithmic area under curve (LogAUC) and are under precision-recall curve (AUPR.) of five models under four training dataset, validation dataset, and decoy selection style configurations.
[0024] Figure 9 provides Per-target logAUC of ET-Screen for the ChEMBL subset.
[0025] Figure 10 shows the per-target difference in logAUC between BridgeDPI and IFP-GNN from ET-Screen for the ChEMBL subset.
[0026] Figure 11 provides the mean validation logAUC for each training dataset sizevalidation dataset size configuration pair under both decoy selection styles.
[0027] Figure 12 comprises images A and B. Image A shows the standard deviation of validation logAUCs calculated across the 50 trials of each training datasetvalidation dataset size configuration. Image B shows the mean percentage of targets shared between datasets across the 50 trials in each dataset sizes configuration.
[0028] Detailed description
[0029] Described below is a ML-based virtual screening method that efficiently and reliably identifies ligands from ultra-large chemical libraries for common protein targets at the superfamily / genome scale. Large and ultra-large chemical libraries refer to chemical libraries in the order of hundreds of thousands and tens of millions of protein targets, respectively. The ligand prediction framework described herein successfully identifies novel ligands for new targets that were not included in the training dataset.
[0030] The prediction framework is based on the Extremely Randomised Trees (Extra Trees - ET) model. Its outputs are used by a docking model to screen candidate compounds (hereinafter interchangeably referred to as "ligands") for binding energy, structural complementarity, binding pose, binding affinity and / or other measures, with a protein target. Thus, the framework is a compound screening process involving Extra Trees, for predicting compounds that interact with a protein target. The Extra Trees model is trained on a comprehensive set of features generated for each target-compound prediction pair. The predictions from the Extra Trees model thus accelerate docking screening. The Extra Trees model produces a likelihood, confidence or probability that a compound is an interacting compound with respect to a particular target protein - it predicts the candidate compounds (candidate interacting compounds) without regard to how they interact with the target. The docking model determines the nature of the interaction between the compound and protein target - e.g., which of those compounds interact or bind in the desired way or at the desired sites.
[0031] A method 100 for virtual screening of ligands to predict which interacting compounds (referred to interchangeably with "ligands" in the description), from a chemical library, could be used to bind to a particular protein target. The method 100 involves applying an Extra Trees machine learning framework to compounds from the chemical library, to identify candidate compounds - i.e. compounds that interact with the protein target and thus form complexes with the protein target that can be classified as "interacting".
[0032] The chemical library itself may be a single resource (e.g., ligands in ChEMBL), a portion of a single resource (e.g., ligands in a subset of ChEMBL) or an aggregate of multiple resources or portions of resources.
[0033] The best candidate compounds identified by the Extra Trees model are then processed using a docking model to determine the candidate compounds that interact with the protein target in the desired way.
[0034] The method 100 involves (step 102) providing an analysis model comprising the Extra Trees ML model, (step 104) receiving protein target information, (step 106) applying the model received in step 102 to the protein target information received at step 104 and compounds in the chemical library. Step 106 outputs a plurality of candidate compounds, each candidate compound being predicted to form an interacting target-compound complex with the target protein. Step 108 involves ranking the candidate compounds based on the probability that each compound will form an interacting complex with the protein target. The compounds with highest probability are than outputted to a docking model.
[0035] As used herein, "highest probability" means that the probability is either above a predetermined threshold, or that the probability is within the top N probabilities - e.g., compounds having a probability of forming an interacting complex with the protein target, that is within the top 5,000 probabilities.
[0036] Per step 110, the docking model can then process the candidate compounds having highest probability, with the protein target. The docking model ranks the candidate compounds based on a score corresponding to a predicted bond between the candidate compound and further protein target. The score may be based on binding energy or binding affinity, pose of the compound, and various other measures. Testing of complexes for a desired effect - i.e., that the compound has the desired effect on the target protein - can therefore start with the highest ranked candidate compound, and sequentially move on to the next lowest ranked compound. In an example implementation of step 110, standard molecular docking-based virtual screening can be performed on the top 5,000 chemicals (compounds) selected by ET-Screen (the model implementing steps 102 to 108) to suggest a final list of 10-20 compound or ligand candidates for experimental testing and the possible binding sites of these compounds.
[0037] The analysis model provided at step 102 includes an ET ML model. The ET ML model itself comprises an ensemble of decision trees. Each decision tree includes at least one parent node and two or more leaf nodes. Each parent node represents a decision in relation to a target-compound complex. The decisions are based on features corresponding to the protein target and features corresponding to the compound being analysed for interaction with the protein target. Each decision tree is traversed by making a decision in relation to the features at each parent node, those features being randomly selected. The features may be represented by embeddings that relate to whether the randomly selected features in the protein target and compound are indicative of likely interaction therebetween. The tree may produce a leaf node or nodes the embedding corresponding to one or both results of the decision is sufficiently strong (e.g., probability greater than a predetermined threshold) to infer interaction to be very likely or very unlikely.
[0038] Each decision tree includes at least two leaf nodes - where there are only two leaf nodes, one leaf node will classify the protein and compound as interacting, and the other leaf node as non-interacting.
[0039] The ensemble of decision trees in the Extra Trees model each output a classification of whether each compound forms an interacting complex with the target protein, or does not. The probability of the compound and target protein forming an interacting complex is determined from the aggregate outputs from all decision trees - e.g., the number of decision trees that classify the complex as interacting divided by the number of decision trees. In other embodiments, the aggregate may be weighted, such that classifications from particular decision trees are considered more influential in the aggregate, than the outputs of other trees. In this regard, a compound being considered to interact with a protein target, and the corresponding complex formed with that protein target being considered interacting, mean essentially the same thing.
[0040] The ET model is trained in a known manner on target-compound complexes, and classifications of whether those complexes are interacting or not, . Specifically, ET model training entails each of its decision trees independently learning its own set of decision rules. These rules are learned step by step, where each rule depends on the one before it, gradually narrowing down whether a protein and compound interact. At each stage, a tree makes a decision based on a specific feature, passing the complex down a branching path of further rules. These rules are set in a way that allows the tree to best separate interacting and non-interacting complexes, ensuring that each decision refines the classification and improves the tree's overall accuracy. The ET model, as an ensemble of these decision trees, makes predictions based on an aggregation of their classifications.
[0041] After training, the ET model is applied to a further, or "new", protein target per step 104. Applying the ET model to a protein target involves receiving information corresponding to that target - e.g., structural descriptors of the protein. Features are also received for each compound, from a chemical library of compounds. The features for each compound may include one or both of structural and physicochemical descriptors, such as structural and physicochemical descriptors created on the basis of target sequences and compound simplified molecular-input line-entry system (SMILES) entries.
[0042] Per step 106, the ET model is then applied to the features received at step 104. The ET model outputs a probability, for each compound, that it will form an interacting complex with the further protein target. This probability enables the compounds to be ranked at step 108.
[0043] The docking model used at step 110 on the ranked candidates generated at step 108, may be any docking model. For example, the docking model may be a genetic algorithm such as AutoDock or GOLD, a Monte Carlo algorithm (again, consider AutoDock), Tabu Search (PRO_LEADS), Swarm Optimization (SODOCK), or a fragmentation approach (FlexX, LUDI). The model may be selected based on the feature set of the target protein inputted at step 104. The docking models are simulators that computationally simulate molecular recognition using molecular docking.
[0044] In experiments, the Extra Trees-based model implementing steps 102 to 108, interchangeably referred to as ET-Screen, was found capable of screening a library of 10 million compounds for one target within 1 hour using only a 12- processor Intel Core i5-11600K CPU. Given this speed where each targetcompound complex is scored in less than 3.6 x 10-4seconds, a library of 10 billion compounds can be screened within 1 hour using 1000 similar CPUs. Thus, the present methodology enables ligand predictions for hundreds of targets within 1-2 years, dramatically accelerating drug development at the genome / superfamily scale.
[0045] The method 100 provides an integrated, predictive computational workflow, accelerating the ligand discovery and development process. This technology has been demonstrated on branched-chain ketoacid dehydrogenase kinase (BCKDK), a new target that was not included in the method development (see the other TD entitled "Discovery of small molecule inhibitors against branched- chain ketoacid dehydrogenase kinase (BCKDK) for cancer therapy").
[0046] Datasets
[0047] In example implementations, three datasets are used for training, validation and testing models including the Extra Trees model. These datasets are reflected in Table 1. Among them is the Directory of Useful Decoys, Enhanced (DUD-E) dataset containing 102 targets which each have an average of 224 ligands associated with them (a total of 22,886 ligands). The other two datasets, are original subsets of BindingDB and ChEMBL which have been constructed as described herein Figure 1.
[0048] Table 1: The datasets used to train and evaluate models.
[0049] The BindingDB chem ical database-derived subset contains 49 targets with an average of 92 ligands associated with each (for a total of 8316 ligands). The ligands in the BindingDB subset are selected to be drug-like - e.g., with a molecular weight between 200 and 500 Daltons, positive logP, and a binding affinity of 1000 nM or better (in terms of Ki, IC50, Kd, or EC50) with their target. In generating the dataset, ligands were selected based on a Tc of at most 0.70 with any other ligand selected for the same target. This cutoff prevents a few "archetypal" ligands having disproportionate influence on either training or validation with the dataset. This results in a topologically diverse set of ligands for each target, with any two ligands having a Tc of 0.31 on average (similar to the DUD-E's 0.35). Likewise, targets have low sequence identicalities from each other. After pairwise global alignment using a programme for performing global sequence alignment to find the optimal alignment of two sequences along their entire length - e.g., EMBOSS's Needle program that uses the Needleman-Wunsch algorithm - DUD-E targets have an identicality of 9.39% on average with each other, with a standard deviation of 6.98%. Similarly, the BindingDB subset's target sequences are on average 7.08% identical, with a standard deviation of 5.30%. While these two datasets are diverse within themselves, they are also distinct from each other. The sequence identicality between any two targets from DUD-E and the BindingDB subset after pairwise global alignment is 7.49% on average with a standard deviation of 5.17%.
[0050] Figure 2, image A, considers this further: the maximum sequence identicalities between targets in the two datasets is on average 18.30%, and is almost always less than 30%. There are only five targets within either of the two datasets that are more than 50% identical to a target in the other dataset. The chemical similarity between ligands of these two datasets is also compared. Figure 2, image B, illustrates this using the distribution of maximum mean pairwise ligand Tc values calculated between the two datasets. This histogram is generated in three steps. Firstly, the mean of ligand Tc values between ligands of a given target in one dataset is calculated regarding ligands of each target in the other dataset. For example, for one target in DUD-E, 49 such mean Tc values are calculated from the 49 targets in the BindingDB subset. Secondly, the maximum of these Tc values is taken as the representative for that target. Thirdly, the distribution is then created using a predetermined number (e.g., 151) such maximum mean pairwise Tc values, from each target in DUD-E and BindingDB subset. The median of the distribution is only 0.37, and no target's ligands have average chemical similarity higher than 0.50 Tc to the ligands of targets in the other dataset.
[0051] The ChEMBL chemical database-derived subset contains 22 targets with an average of 81 ligands associated with each for a total of 1691 ligands. It is designed to be more diverse within itself and from other two datasets than the DUD-E and the BindingDB subset. The sequences of each of these targets is at most 20% identical to any other target within ChEMBL and the other two datasets. Moreover, more than a third of these targets come from protein classes that are not included in DUD-E and the BindingDB subset, such as Ligases, ABC transporters, Amidases, and the Coronaviruses polyprotein lab family - Figure 2. In the manner of the ligand compilation for the BindingDB subset, the ChEMBL subset's ligands are selected to be drug-like, and have their affinities measured in binding assays with a predetermined maximal internal ChEMBL reliability score of nine. Moreover, ligands associated with each target are selected to have a chemical similarity of at most 0.50 (instead of 0.70, as with the BindingDB subset) Tc to any other ligand from the same target.
[0052] Figure 4 shows that the maximum of the mean pairwise Tc values between ligands of a target from the ChEMBL subset and those from the other two datasets is also low, being less than 0.4 on average. In fact, the medians for the BindingDB and DUD-E dataset are 0.32 and 0.35 respectively.
[0053] Given the dearth of experimentally-verified weakly- and non-binding ligands, algorithmically-selected decoys need to be used as stand-ins during model training and evaluation. For each dataset, two kinds of decoys are generated: one in the "DUD-E style" and the other in a random manner. The former refers to the method of decoy selection introduced and employed in the creation of DUD-E. Namely, decoys are selected from the ZINC database to be physico- chemically similar but topologically dissimilar to their ligands - i.e., the decoys are physico-chemically similar but topologically dissimilar to ligands for the same protein target. Selected decoys closely match their ligands in terms of molecular weight, logP, number of rotatable bonds, number of hydrogen bond acceptors and donors, and net charge. They also have as low a Extended Connectivity Fingerprint (ECFP - with, e.g., maximum diameter of 4) based Tc to the ligand as possible (usually, 0.20-0.40). This process is intended to create challenging decoys which are also unlikely to be false negatives. However, the minimization of Tc values between decoys and their ligands also may create a bias in the data models that use a featurization implying ligands' topology might achieve good performance on a dataset with DUD-E-style decoys by simply learning to distinguish these different topologies.
[0054] To understand the impact of this bias on model comparison and generalization, a random-style decoy selection method is presently used. In this method, the analysis model may be trained on target-compound complexes where a plurality of the non-interacting compounds (decoys) are selected by randomly sampling chemical libraries using an approach that filters for chemotypically-diverse and structurally-diverse compounds. As displayed in Figure 5, a predetermined number of drug-like molecules can be randomly assigned to each ligand of a given target as the corresponding decoys, to reflect the rarity of ligands in the drug discovery process. For example, 50 drug-like molecules may be selected from ZINC for this purpose. ZINC organizes its ligands into tranches defined by their logP and molecular weight. Some tranches are far larger than others. Taking simple random samples from ZINC would fail to represent compounds from rarer tranches. To encourage chemotypic diversity in terms of factors such as logP and molecular weight (and thus their related properties as well), the method can involve randomly selecting samples in a constrained fashion that gives each tranche equal representation in the decoy set. Thus, compounds can be selected from each tranche irrespective of the number of compounds in the tranche relative to all other tranches. Moreover, to also ensure topological diversity amongst decoys, decoys can be selected to have at most a predetermined Tc to any other decoy selected for the same target. For example, decoys may be selected to have an average Tc of between 0.1 and 0.5, or between 0.2 and 0.4, or of at most 0.30, to any other decoy for the same target. For training using these datasets, the method can involve under-sampling the decoys in a stratified (by target) random fashion to be of an equal number to the ligands. This parity between both groups guarantees that the latter will not be disregarded in model training. If training were instead conducted on the original datasets, models would likely ignore learning from examples regarding target-ligand pairs. This oversight could occur because the optimization of model parameters would predominantly focus on minimizing the total training loss, which would largely be influenced by targetdecoy pairs due to their numerical superiority in these datasets.
[0055] Evaluation metrics
[0056] Model performance is primarily quantified using the logAUC metric, that emphasizes the model's early enrichment ability. This metric addresses the inefficiencies of two other commonly-used ones, the first is the area under the recall-false positive rate curve (AUC). The AUC fails to penalize models which achieve high recall at the expense of poor precision in imbalanced datasets. Specifically, if there are a large number of negatives compared to positives in a dataset, a model with high recall can also have a low false positive rate even if its precision is poor. For virtual screening models which rank datasets with only a small minority of ligands, a poor precision also means a poor hit rate amongst candidate ligands. The second metric popular in the literature is the area under the precision-recall curve (AUPR). While AUPR does not disregard precision, it weights a model's performance equally across the entire database. This is suboptimal, because only a small subset constituting of the most highly ranked candidate ligands can practically be experimentally tested given time and cost constraints. To demonstrate the implications of these inefficiencies, AUC and AUPR will also be collected and compared with logAUC. Machine learning models
[0057] Various ML models are discussed herein, demonstrating the ET was found, after extensive problem solving and investigation, to be the optimal model. This is surprising given the structure of the targets and the known structures that can bind to those targets, since the focus on randomisation in the ET model disregards such structural relationships.
[0058] ET-Screen. ET-Screen as an Extra Trees based-model is an ensemble of decision trees that is similar to Random Forests. Extra Trees fits multiple unpruned decision trees on the complete training dataset and uses majority voting to predict whether the molecule is a ligand. The subset of features considered for each split point in every decision tree is sampled randomly. In the present embodiment, Extra Trees randomly selects the split point after the feature that minimizes entropy is chosen to split on, to further encourage diversity amongst trees. As described by Figure 6, image A, the one-dimensional target and ligand input representations are featurized into an approximately 700-dimensional vector. The interaction classification is then generated by taking the average of the binary predictions from each constituent decision tree. Thus, ET-Screen is trained on about 700 features describing the target and ligand, with two sets of features for the ligand: the first is derived from RDKit or another database, and can include physico-chemical features like one or more, and preferably all, of the number of Aliphatic rings, relative formula mass, Labute's Approximate Surface Area, heavy atom count, kappa, Bertz CT, etc. The second is from Mol2Vec, which is a pre-trained embedding for ligand SMILES that maps them to dense, chemically- meaningful vector representations based on their Morgan substructures. Target features, calculated using iFeature, describe the position of where the residues belonging to the polar, neutral and hydrophobic groups are located. Through validation on DUD-E and the BindingDB subset in Evaluation one of the robust evaluation pipeline (to be described), ET-Screen is set to use a predetermined number of trees, presently 500 trees, with at least two samples required to split an internal node, and with its splits made on the basis of improved entropy. The scikit-learn implementation of Extra Trees is used, and training takes about two minutes over 50,000 target-ligand pairs on a 12-processor Intel Core i5-11600K CPU. The number of trees, samples required to split internal nodes, and split criterion are set with reference model performance on its DUD-E training dataset.
[0059] XGBoost. The XGBoost algorithm of gradient-boosted decision trees, another ensembling approach that uses decision trees, is trained on the same features that ET-Screen is. XGBoost is set to use 100 trees with a learning rate of 0.01. The dmlc implementation of XGBoost employed takes about 15 minutes to train on 50,000 target-ligand pairs using a 12-processor Intel Core i5-11600K CPU.
[0060] IFP-GNN. The interaction fingerprint graph neural network (IFP-GNN) model described herein is inspired by RF-Score in its target-compound complex featurization. Specifically, docked compound poses as generated by Glide are used to create target-compound graphs representing complexes, where target pocket and compound atoms are represented by nodes, and edges are created between nodes whose corresponding atoms are at most 10 Angstroms apart. Each node contains information on the element of the atom it represents in the form of a one-hot vector. Namely, the elements Br, C, Cl, F, H, N, O, and S are each represented by one position in the nine-dimensional one-hot vector, with every other (less commonly-found) element being represented by the last position. The target-ligand interaction graph is then processed by a GNN using the convolutional operator - e.g., as defined by Kipf and Welling - that approximates spectral graph convolutions by, broadly speaking, adding a linear combination of a given node's first-order neighbourhood information. Repeatedly applying the operator has the effect of increasing the order of the neighbourhood from which information is aggregated. Thus, the interaction graph represents closely-contacting target and compound atoms and is processed without changing the graph's structure using a GNN. Specifically, the information in nodes (representing atoms) is updated using information from their neighbouring nodes multiple times in a learnable, non-linear fashion.
[0061] After these convolutions, the processed graph is flattened by taking the mean of each channel in the nodes' featurization. Thus, the processed graph is flattened by taking the mean of this information along each of its dimensions, concatenating the docking-predicted binding affinity to this mean vector, and passing it through an FNN to create the classification of the compound being a ligand. Glide's predicted binding affinity for the complex is concatenated to this vector, which is passed through an FNN of three layers to produce an interaction classification. Figure 6, image B, sketches this architecture. IFP-GNN is implemented on PyTorch 1.8.1, and training takes about 40 minutes over 50,000 examples with one Nvidia GeForce R.TX 3060 graphics card.
[0062] BridgeDPI. The Bridge Drug-Protein Interaction (BridgeDPI) model makes use of a graph neural network (GNN) along with convolutional and fully-connected neural networks (CNNs and FNNs). Target sequences are used to generate k- mer features processed using an FNN, and 1-grams created using Word2Vec which are processed using a TextCNN. Compounds are similarly represented in two ways: as Word2Vec 1-grams processed using another TextCNN, and by their ECFP4 fingerprints processed using an FNN. The resulting target and compound embeddings are then updated by a GNN using information from other similar targets and compounds. The final embeddings are then multiplied element-wise before being used by another FNN to create an interaction classification. BridgeDPI is implemented on PyTorch 1.6.0, and it takes about two hours to train on 50,000 target-ligand pairs with one Nvidia GeForce RTX 3060 graphics card.
[0063] Docking (Glide algorithm). Schrodinger's Glide is used in its Standard Precision mode to provide a molecular docking "baseline" to compare ML models to. As a docking program, Glide does not have a training step. However, Glide does still conduct preparatory procedures before generating scores for target-compound complexes. While scoring 50,000 complexes takes about 11 hours on a 12- processor Intel Core i5-11600K CPU, these procedures take about three hours themselves.
[0064] Evaluation Dioeline
[0065] As displayed in Figure 7, the evaluation pipeline features a battery of four evaluations, with the last being a case study where the top model as ascertained by the previous evaluations (ET-Screen) is applied to the identification of novel actives for targets.
[0066] Evaluation one: both-sided validation. The first evaluation in this pipeline involves validating all five models in consideration. These models are either trained on the DUD-E dataset of 102 targets and validated on the BindingDB subset of 49 targets or vice versa. These reserved datasets are used in pairs, with each dataset used for training and validation in alternation. This process is repeated for both decoy selection styles. To estimate the variability of model training caused by the randomness of model parameter initialization and the order in which training data is learned from, each training and validation is done five times under different random seeds. Moreover, the three models presented herein (ET-Screen, the XGBoost-based model using the same featurization, and IFP-GNN) have their hyperparameters tuned with reference to this Evaluation's results. Both DUD-E and random-style decoys are independently used in this and the second and third evaluations.
[0067] Evaluation two: testing on ChEMBL subset of the models trained on DUD-E and the BindingDB subset. Three of the four ML models (ET-Screen, BridgeDPI, IFP- GNN) are trained on both the DUD-E dataset and BindingDB subset, and tested on the dataset created based on ChEMBL with docking as the control. This ChEMBL subset is selected to be diverse and highly distinct from DUD-E and the BindingDB subset. Evaluation three: mixing-and-matching data. Evaluation three allows for a deeper comparison of model consistency and generalizability across different training and validation dataset sizes and decoy selection styles. While previous evaluations offer multiple points of reference for models' reliability and generalization performance, more can be done with the benchmarking datasets at hand for picking a single model and decoy selection style to identify novel ligands. As has already been alluded to in Evaluation one, the specific combination of targets in training relative to those in validation along with the decoy selection style used may have significant effects on model performance. Therefore, taking this further, many more such combinations can be considered by "mixing-and-matching" data from these datasets. The number of targets featured in these training and validation data combinations can also be varied to understand the role of dataset size on model performance. Briefly, this is done here in three steps. Firstly, a set number of targets and their compound data is randomly assigned to be the validation dataset. Secondly, the remaining targets and their data are randomly added in increments to the training dataset. Thirdly, after each addition, the model (ET-Screen) in consideration is retrained (i.e., iteratively trained) on the expanded training dataset, and the performance of this model on the fixed validation dataset is measured. Given that the random assignment of targets and their data to either training or validation can have a significant impact on model performance, this process is repeated multiple times under different random seeds. Algorithm 1 describes this process presently used. Here, the combination of DUD-E, the BindingDB subset, and the ChEMBL subset is used as the source dataset and the splitting-training-validation process is repeated with both decoy styles under 50 different seeds to better understand ET-Screen's characteristics.
[0068]
[0069] Algorithm 1 : Pseudocode of the procedure of Evaluation three, where combinations of data are made from DUD-E dataset, the BindingDB, and ChEMBL subsets. Under each random number generator initialization controlled by its seed, the validation targets are first selected. The training dataset is created in increments of 15 targets out of the rest of the data. At each increment, ET- Screen is retrained and then evaluated on the validation dataset, once for each decoy style. In total, Evaluation three involves 2300 trainingvalidations: 1150 each for each decoy style. From this, 50 logAUC values are generated from the performance of ET-Screen on each training dataset size, validation dataset size, and decoy selection style configuration.
[0070] Evaluation four: novel ligand identification. After the top model and decoy selection style configuration is selected based on Evaluations 1-3, they can be used to identify novel ligands for targets of interest. ET-Screen trained on the combination of DUD-E dataset, the BindingDB, and the ChEMBL subsets using random-style decoys is applied to this task. This model is then used to rank a drug-like subset from the ZINC database for a kinase target not used in its development as the final, practical test of its generalizability. The top-ranked compounds are then experimentally verified. Specifically, the trained model is used to screen a large drug-like subset of the ZINC database (~10 million compounds) against branched-chain ketoacid dehydrogenase kinase (BCKDK) that was not included in the above three datasets and in the training of ET- Screen. The top 5,000 compounds ranked by ET-Screen are subjected to standard molecular docking using Glide SP, to suggest potential protein binding sites and compound docking poses. 14 ligand candidates prioritized by docking are then experimentally tested. From this, two new ligands (i.e., pre-existing but novel with respect to the target) were confirmed.
[0071] Model performance demonstrated by the evaluation pipeline
[0072] Evaluation one: both-sided validation
[0073] Figure 8, comprising images A and B, shows the logarithmic area under curve (LogAUC) and area under precision-recall curve (AUPR) of five models under four training dataset, validation dataset, and decoy selection style configurations. Models (except for the docking program Glide) are trained on one dataset before being evaluated on the other under five different random number generator initialization seeds that influence model training.
[0074] With the exception of XGBoost, every ML model performs better than docking for each validation configuration. Comparing the two training and validation dataset settings, it can generally be seen that the ML models usually benefit from training on DUD-E and validating on the BindingDB subset by about 5-10 additional logAUC points, which may at least be partially due to the former's larger size (Evaluation three will more specifically describe the impact of dataset sizes on performance). Moreover, across these settings, ML models achieve better performance when using random-style decoys than DUD-E style decoys, often with improvements of around 10-20 logAUC points (Figure 7a).
[0075] The ordering of most ML models in terms of their performance is largely retained across configurations, with either ET-Screen and BridgeDPI usually comparably performing best and either docking or XGBoost being worst. A notable exception to this consistency is IFP-GNN : its performance when using DUD-E-style decoys is comparable to those of the top models but is markedly worse than those models when using random-style decoys. This result is surprising given the fact that, unlike other ML models, IFP-GNN does not use any ligand fingerprint-related information in its featurization and would not benefit from learning any topological bias in DUD-E-style decoys. Moreover, the fact that ET-Screen and BridgeDPI see far better performance with randomstyle decoys than DUD-E-style decoys also highlights that they do not "overrely" on learning this bias to achieve their performance.
[0076] Figure 8, image B, displays performance in terms of a more commonly used metric, AUPR. While the performance rankings of models are similar to those in terms of logAUC, AUPR's focus away from early enrichment to enrichment across the entire dataset predictions are made for exaggerates differences between configurations. For example, performance in terms of logAUC often decreases by approximately 20-40% when switching from random- to DUD-E-style decoys, whereas it does for So those same models by 60- 80% in terms of AUPR. And importantly, in absolute terms, the performance of most models using DUD-E- style decoys is near-null in terms of AUPR, while logAUC reveals their early enrichment performance is not. Ultimately, a model's ability to reliably identify ligands amongst the compounds it most confidently predicts as active is more relevant to its practical usefulness; relatively few ligand candidates proposed by models can feasibly be experimentally verified to be active. Considering performance in terms of AUC reveals this metric has the opposite problem — models' performances are reported as near-perfect and differences between models are flattened due to the metric ignoring prediction precision (data not shown).
[0077] Evaluation two: testing on ChEMBL subset
[0078] All models in consideration except for that based on the poorly-performing XGBoost are evaluated on the ChEMBL subset after being trained on both DUD- E and the BindingDB subset. The ChEMBL subset, designed to be significantly dissimilar to these datasets, presents a further challenging test of models' generalizability. Figure 9 displays ET-Screen's performance in terms of logAUC for each target in the dataset; results are similar to its best performance in Evaluation one: it achieves a mean logAUC of 45.20 ± 10.70 (mean ± standard deviation) across ChEMBL subset with DUD-E-style decoys, and 52.99 ± 9.23 with random-style decoys. In particular, it achieves a logAUC of 45.83 and 59.10 in enriching 7QBB (SAR.S- CoV-2 main protease) ligands with DUD-E-style and random-style decoys respectively. ET-Screen does better when it is trained and evaluated on datasets with random-style decoys than when DUD-E-style decoys are used. Here, the difference in performance is on average 7.79 logAUC points in favor of the former style. Nonetheless, for both decoy styles, performance is quite consistently good over all 22 targets.
[0079] Moreover, Figure 10, providing the per-target difference in logAUC between BridgeDPI and IFP-GNN from ET-Screen for the ChEMBL subset, shows that ET- Screen generally performs better than both BridgeDPI and IFP-GNN on the ChEMBL subset. IFP-GNN exhibits signs of being over-fit to Evaluation one with performance 10-20 logAUC points poorer than ET-Screen's even when using DUD-E-style decoys (recall it performed comparably to both ET-Screen and BridgeDPI in Evaluation One). Docking performance of Glide on the ChEMBL subset is similarly and consistently poorer than ET-Screen's by 20-30 points, in addition to being more inconsistent between targets (data not shown).
[0080] Overall, BridgeDPI does better than IFP-GNN and is still comparable to ET- Screen. It does more poorly than ET-Screen for 17 out of 22 targets when trained and tested using random-style decoys, and for only eight targets when applied using DUD-E-style decoys. However, in four of the latter 14 cases BridgeDPI does better, it does so by nearly zero logAUC points. Considering these results in tandem with those of Evaluation one, it can broadly be said that BridgeDPI is the best performer with the latter decoy style and ET-Screen with the former.
[0081] The question of which model should be used for Evaluation four, where novel ligands are sought for BCKDK, rests primarily on which decoy style prepares it better for the specific way this screening is conducted. Given the database ranked by the model is the complete set of drug-like ligands from ZINC without any pre-exclusion based on, say, a high Tc to known ligands, a model trained on random-style decoys is likely more appropriate. Any pre-exclusion is avoided primarily because doing so might exclude the few promising scaffolds available for the targets with only a small number of known ligands; moreover, the target BCKDK has too few ligands to create exclusions on the basis of. Evaluation three therefore considers and ultimately substantiates the case for using ET-Screen with random-style decoys.
[0082] Evaluation three: mixing-and-matching data
[0083] ET-Screen's consistency and generalization characteristics are further considered by training and validating it on mixed combinations of the three datasets (DUD-E, BindingDB subset, ChEMBI subset) 2300 times.
[0084] Figure 11 provides the mean validation logAUC for each training dataset sizevalidation dataset size configuration pair under both decoy selection styles. Error bars imply the standard deviation of validation logAUC for the particular pair calculated across the 50 trials. As could be expected, the model achieves more logAUC points when training and validating with random-style decoys than DUD-E-style decoys, and performance under both styles improves consistently with training dataset size. An important difference between ET-Screen's behaviour under both styles is that its rate of improvement is markedly higher with random-style decoys than with DUD-E-style decoys. In this Evaluation, the improvement is by about 4-8 logAUC points. Also, as mentioned before and predictably, ET-Screen shows an increase in performance as the amount (and therefore variety) of the training data increases, particularly as the dataset increases from 15 to 75 targets' worth of data in size. However, the mean logAUC achieved under every training dataset size is largely unaffected by the validation dataset size; the range of mean logAUC values achieved when the training dataset contains 15 targets is only 0.09 (with DUD-E-style decoys) and 0.76 (with random-style decoys) across the all three validation dataset size configurations, for example.
[0085] However, a significant finding of this Evaluation is that the rate of this increase is greater for validations with random-style decoys. Specifically, the logAUC gap between the two approaches increases in a mostly linear fashion from about 4 to at most 8 points as the training dataset becomes larger. While a gap itself is not evidence for one decoy style creating a better-generalizing model over another since both training and validation datasets use the same decoy style, the increase in this gap with size is evidence that random- style decoys allow ET-Screen to make better use of training datasets.
[0086] One aspect of performance that does not occur to be dependent on decoy style is logAUC variability across the 50 trials of each training dataset size-validation dataset size configuration : this variability decreases as the training and validation dataset sizes increase. Since the pool of targets to create datasets from is limited to 173, configurations where a greater number of targets are randomly selected to be reserved for validation have more similar validation datasets across their trials, resulting in more similar performance. Figure 12 image A shows the standard deviation of validation logAUCs calculated across the 50 trials of each training dataset-validation dataset size configuration. The values themselves and their trends are largely identical regardless of decoy selection style, and show a consistent decrease with the increase of training and validation dataset sizes. The decrease in variability due to this factor is particularly pronounced between configurations having validation datasets with 23 targets and those having 53; the configuration with 98 targets, on the other hand, does not show a significant decrease over the configuration with 53 for any given training dataset size.
[0087] Under each validation dataset size configuration, models are trained iteratively on training datasets comprised of the targets not selected for validation. As the number of targets included in training datasets increases, these datasets become more similar across trials, rendering the trained models more similar too. This also contributes to less variability in model performance across trials, as evidenced by the reduction in trial logAUC standard deviation within each of the three panels of Figure 12 image A.
[0088] The impact of these two homogenizing effects is similar under both decoy selection styles and culminates in a logAUC standard deviation of 1.27 and 0.99 points between trials under the DUD-E and random decoy selection styles when 75 and 98 targets are featured in the training and validation datasets respectively. Nonetheless, even in the most diverse case where training and validation datasets only have 15 and 23 targets, the trial logAUC standard deviation is only 3.52 and 2.90 points under both decoy selection styles respectively; this is testament to ET-Screen's inherent consistency.
[0089] Figure 12 image B shows the mean percentage of targets shared between datasets across the 50 trials in each dataset sizes configuration. This mean percentage remains constant for validation datasets within each of their three size configurations since they are kept fixed across training dataset sizes. Figure 12 image B displays the mean percentage of shared targets across trials for each training dataset size- validation dataset size configuration. This percentage is calculated in two steps for each configuration. First, the training datasets associated with each of the 50 trials of the configuration are compared pairwise. In each comparison, the percentage of targets shared between one pair from these 50 datasets is calculated. In the second step, the mean of these percentages is calculated across all comparisons. The same procedure is repeated for the 50 validation datasets associated with each of the three validation dataset size configurations. Given validation datasets are kept fixed as ET-Screen is trained iteratively on larger training datasets, their mean percentage stays constant as this training dataset size increases.
[0090] As expected, however, configurations with larger validation datasets have a higher mean percentage as the number of targets to draw from and to create them is fixed. For this same reason, the mean percentage for larger training datasets is greater than those for smaller training datasets. Nonetheless, in all training dataset size-validation dataset size configurations, at least one of the training or validation datasets is significantly diverse across trials. For example, when ET-Screen is trained on the largest training datasets in this experiment - those with 150 targets, having a mean percentage of shared targets of 86.73% -- the corresponding validation datasets with 23 targets share only 13.50% of targets on average. The utility of mixing-and-matching data to assess ET-Screen's consistency across diverse training and validation conditions is therefore largely maintained across all dataset sizes configurations.
[0091] In the above model evaluation pipeline, comprising three virtual evaluations designed to select the top- performing virtual screening model from those in consideration, the third evaluation is designed to elucidate the manner in which the model of interest learns. This is achieved by tracking the model's performance as the size of the training and validation datasets change, and thereby the evolution of the model's consistency and generalizability in context □^different decoy selection styles is found. This allows for comparisons between the learning effectiveness of different decoy styles which cannot be made with the more conventional and "static" first and second evaluations.
[0092] ET-Screen is directly applicable to drug discovery in commercial settings, and the evaluation pipeline (along with its related novel decoy selection style and original benchmarking datasets) can also be used in these settings to compare ET-Screen to new candidate models in consideration. More generally, ET- Screen can quickly be retrained on proprietary datasets of virtually any size, and this pipeline can also be very easily applied to any data and model context. ET- Screen can also be repurposed for use as one element of a wider prediction framework - say, to be a part of a virtual screening model ensemble.
[0093] It will be appreciated that many further modifications and permutations of various aspects of the described embodiments are possible. Accordingly, the described aspects are intended to embrace all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims.
[0094] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.
Claims
Claims1. A virtual screening method for predicting interacting compounds, from a chemical library, for proteins, the methodology comprising: providing an Extremely Randomized Trees (Extra Trees) machine learning model comprising an ensemble of decision trees, each decision tree comprising: at least one parent node, each parent node representing a decision in relation to a target-compound complex, based on a plurality of features corresponding to a protein target of the complex and a plurality of features corresponding to a compound of the complex; and and at least two leaf nodes, each leaf node classifying the target-compound complex as either interacting or non-interacting, the Extra Trees model outputting a probability that targetcompound complex is interacting, based on aggregated classifications from each said decision tree, receiving a further protein target; applying the Extra Trees model to the further protein target and compounds in the chemical library, to determine a probability, for each compound, that it will form an interacting complex with the further protein target; ranking the candidate compounds based on the respective probabilities, to identify candidate compounds having highest probability; and outputting the candidate compounds having highest probability, to a docking model.
2. The method of claim 1, further processing the candidate compounds having highest probability, and the further protein target, using a docking model, the docking model outputting a ranking of the candidate compounds based on a score corresponding to a predicted bond betweenthe candidate compound and further protein target.
3. The method of claim 1 or 2, further comprising training the Extra Trees machine learning model on target-compound complexes, wherein a plurality of non-interacting compounds are selected by randomly sampling chemical libraries using an approach that filters for chemotypically-diverse and structurally-diverse compounds.
4. The method of any one of claims 1 to 3, wherein training the Extra Trees machine learning model comprises using target-compound complexes, where the corresponding compound is selected to be physico-chemically similar, but topologically dissimilar, to ligands for the respective protein target.
5. The method of claim 3, wherein the compounds of the target-compound complexes are arranged in tranches, and wherein filtering for chemotypically-diverse compounds comprises selecting compounds from each tranche irrespective of a number of compounds in said tranche relative to all other tranches.
6. The method of claim 5, wherein selecting compounds from each tranche comprises randomly taking compounds stratified by tranche.
7. The method of claim 5, further comprising arranging the complexes in each said tranche based on one or both of logP and molecular weight.
8. The method of claim 1, wherein receiving the input data comprises receiving structural and physicochemical descriptors.
9. The method of claim 8, wherein the structural and physicochemical descriptors are created on the basis of target sequences and compound simplified molecular-input line-entry system (SMILES) entries.
10. The method of 3, wherein filtering for structurally-diverse compoundscomprises selecting, for each said target, non-interacting compounds that have an average Tanimoto coefficient of between 0.1 and 0.5 for the respective target.
11. The method of 8, wherein the average Tanimoto coefficient is between 0.2 and 0.4, and more preferably about 0.3.
12. The method of 2, wherein processing the candidate compounds having highest probability, using a docking model, comprises outputting each candidate compound to a simulator and using the simulator to computationally simulate molecular recognition using molecular docking.
13. A virtual screening system for predicting interacting compounds, from a chemical library, for proteins, the system comprising: memory; a machine learning module comprising an Extremely Randomized Trees (Extra Trees) machine learning model comprising an ensemble of decision trees, each decision tree comprising: at least one parent node, each parent node representing a decision in relation to a target-compound complex, based on a plurality of features corresponding to a protein target of the complex and a plurality of features corresponding to a compound of the complex; and and at least two leaf nodes, each leaf node classifying the targetcompound complex as either interacting or non-interacting, the Extra Trees model outputting a probability that target-compound complex is interacting, based on aggregated classifications from each said decision tree; a ranking engine; and at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the system to: receiving a further protein target; applying the Extra Trees model of the machine learning module to the further protein target and compounds in the chemical library, todetermine a probability, for each compound, that it will form an interacting complex with the further protein target; ranking the candidate compounds, using the ranking engine, based on the respective probabilities, to identify candidate compounds having highest probability; and outputting the candidate compounds having highest probability, to a docking model.
14. The system of claim 13, further comprising a docking module, the docking module comprising a docking model, wherein the docking model is configured to process the candidate compounds having highest probability, and the further protein target, and output a ranking of the candidate compounds based on a score corresponding to a predicted bond between the candidate compound and further protein target.
15. The system of claim 13 or 14, wherein the Extra Trees machine learning model is trained on target-compound complexes by randomly sampling chemical libraries to select a plurality of non-interacting compounds, using an approach that filters for chemotypically-diverse and structurally- diverse compounds.
16. The system of any one of claims 13 to 15, wherein the Extra Trees machine learning model is trained using target-compound complexes, where the corresponding compound is selected to be physico-chemically similar, but topologically dissimilar, to ligands for the respective protein target.
17. The system of claim 15, wherein the compounds of the target-compound complexes are arranged in tranches, and wherein the at least one processor is configured to filter for chemotypically-diverse compounds by selecting compounds from each tranche irrespective of a number of compounds in said tranche relative to all other tranches.
18. The system of claim 17, wherein the at least one processor is configured to select compounds from each tranche by randomly taking compounds stratified by tranche.
19. The system of claim 17, wherein the at least one processor is configured to arrange the complexes in each said tranche based on one or both of logP and molecular weight.
20. The system of claim 13, wherein the system receives the input data by receiving structural and physicochemical descriptors.
21. The system of claim 20, wherein the structural and physicochemical descriptors are created on the basis of target sequences and compound simplified molecular-input line-entry system (SMILES) entries.
22. The system of 15, wherein the at least one processor is configured to filter for structurally-diverse compounds by selecting, for each said target, non-interacting compounds that have an average Tanimoto coefficient of between 0.1 and 0.5 for the respective target.
23. The system of 22, wherein the average Tanimoto coefficient is between 0.2 and 0.4, and more preferably about 0.3.
24. The system of 14, wherein the docking model is configured to process the candidate compounds having highest probability by outputting each candidate compound to a simulator and using the simulator to computationally simulate molecular recognition using molecular docking.
Citation Information
Patent Citations
Drug-target relationship prediction method based on deep forest and PU learning
CN112652355A
Cited By
Double-strategy screening method and system for active ingredients and targets of medicinal and edible food
CN121641269A