Methods and systems for differential drug discovery

By employing a multi-target drug design strategy and utilizing receptor ensemble and computational methods to optimize small molecule compounds, the challenge of simultaneously binding multiple targets in vivo has been addressed, enabling more effective multi-pharmacological drug design and reducing unwanted interactions and the risk of drug failure.

CN111684532BActive Publication Date: 2025-11-04CYCLICA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201880075047.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-11-22
Filing Date
2018-11-22
Publication Date
2025-11-04
Estimated Expiration
2038-11-22

AI Technical Summary

Technical Problem

Designing small molecule drugs that can simultaneously target multiple targets presents challenges in vivo, particularly due to the difficulty in prediction caused by target diversity and conformational changes, and the potential for undesirable off-target interactions from conventional small molecule drugs.

Method used

Employing a multi-target drug design (MTDD) strategy, we construct receptor sets containing both targets and anti-targets. Using derivatization, molecular docking simulations, and scoring engines, we optimize small molecule compounds (SMCs) to simultaneously interact with multiple targets while avoiding anti-target interactions. We then redesign the SMCs using computational methods.

Benefits of technology

It increases the likelihood of discovering SMCs that interact with multiple targets simultaneously and reduces unwanted interactions, enhances the effectiveness of treating complex diseases, reduces the risk of drug failure due to single-target mutations, and reduces drug interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111684532B_ABST
    Figure CN111684532B_ABST
Patent Text Reader

Abstract

A method for differential drug discovery involves obtaining a set of receptors. The set of receptors specifies a plurality of targets and a plurality of anti-targets. The method also involves obtaining a small molecule compound (SMC) seed model and deriving a first plurality of candidate SMCs from the SMC seed model. For each of the candidate SMCs in the first plurality of candidate SMCs, a first desired interaction between the candidate SMC and each of the plurality of targets is simulated. Further, for each of the candidate SMCs in the first plurality of candidate SMCs, a first undesired interaction between the candidate SMC and each of the plurality of anti-targets is simulated. The method also involves obtaining, based on the first desired interactions and based on the first undesired interactions, a first SMC interaction score for each of the candidate SMCs in the first plurality of candidate SMCs; and determining, based on the first SMC interaction scores, whether at least a minimum fraction of drugs is reached.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 590,141, filed November 22, 2017, under 35 USC §119(e), which has at least one inventor with the same inventor and is entitled “Method and System for Differential Drug Discovery”. U.S. Provisional Application No. 62 / 590,141 is incorporated herein by reference. Background Technology

[0003] Many diseases involve complex biopathologies involving multiple pathways. Small molecule drugs are traditionally designed to target a single protein, but often exhibit multiple off-target protein interactions (i.e., multipharmacology). The average small molecule compound (SMC) is thought to bind to 30-300 different proteins in vivo. Typically, a drug molecule is only intended to bind to one of these, the stated target. The others are unintended off-target interactions, which can be beneficial or detrimental and can occur in homologous or non-homologous proteins. In the field of multi-target drug design, small molecule drugs that target multiple proteins can improve the treatment of complex pathologies. However, designing small molecule drugs capable of binding to multiple targets simultaneously is challenging due to various complications. Summary of the Invention

[0004] Typically, one or more embodiments relate to a method for differential drug discovery, the method comprising: obtaining a receptor set, wherein the receptor set specifies a plurality of targets and a plurality of anti-targets; obtaining a small molecule compound (SMC) seed model; deriving a first plurality of candidate SMCs from the SMC seed model; simulating a first desired interaction between each of the candidate SMCs and each of the plurality of targets for each of the first plurality of candidate SMCs; simulating a first undesired interaction between each of the candidate SMCs and each of the plurality of anti-targets for each of the first plurality of candidate SMCs; obtaining a first SMC interaction score for each of the first plurality of candidate SMCs based on the first desired interaction and based on the first undesired interaction; and determining, based on the first SMC interaction score, whether at least a minimum score for a drug has been reached.

[0005] Generally, one or more embodiments relate to a system for differential drug discovery, the system comprising: a derivatization engine configured to derive a plurality of candidate small molecule compounds (SMCs) from a SMC seed model; a molecular docking simulation engine configured to: simulate, for each of the plurality of candidate SMCs, a desired interaction between the candidate SMC and each of a plurality of targets specified in a receptor set; simulate, for each of the plurality of candidate SMCs, an undesired interaction between the candidate SMC and each of a plurality of anti-targets specified in the receptor set; a scoring engine configured to: obtain, based on the desired interactions and based on the undesired interactions, a SMC interaction score for each of the plurality of candidate SMCs; and determine, based on the first SMC interaction score, whether a minimum score for a drug is at least achieved.

[0006] Generally, one or more embodiments relate to a non-transitory computer readable medium comprising computer readable program code for differential drug discovery, the computer readable program code causing a computer system to: obtain a receptor set, wherein the receptor set specifies a plurality of targets and a plurality of anti-targets; obtain a small molecule compound (SMC) seed model; derive a first plurality of candidate SMCs from the SMC seed model; simulate, for each of the first plurality of candidate SMCs, a first desired interaction between the candidate SMC and each of the plurality of targets; simulate, for each of the first plurality of candidate SMCs, a first undesired interaction between the candidate SMC and each of the plurality of anti-targets; obtain, based on the first desired interactions and based on the first undesired interactions, a first SMC interaction score for each of the first plurality of candidate SMCs; and determine, based on the first SMC interaction score, whether a minimum score for a drug is at least achieved.

[0007] Other aspects of embodiments will be apparent on review of the description and the appended claims. BRIEF DESCRIPTION OF DRAWINGS

[0008] The present embodiments are illustrated by way of example and not intent to be limited by the figures of the drawing.

[0009] Figure 1 A block diagram of a system in accordance with one or more embodiments is shown.

[0010] Figure 2 A flow diagram in accordance with one or more embodiments is shown.

[0011] Figure 3AExamples of target lists are shown in accordance with one or more embodiments.

[0012] Figure 3B Compilation of protein pocket models is shown in accordance with one or more embodiments.

[0013] Figure 3C Examples of receptor sets are shown in accordance with one or more embodiments.

[0014] Figure 4 Examples of scored interactions are shown in accordance with one or more embodiments.

[0015] Figure 5A and Figure 5B Computing systems are shown in accordance with one or more embodiments. DETAILED DESCRIPTION

[0016] Specific embodiments disclosed herein will now be described in detail with reference to the attached figures. Like elements in the various figures can be denoted by the same reference numerals and / or names for consistency.

[0017] The following detailed description merely sets forth example embodiments of the presently disclosed embodiments and alternatives and is not intended to limit the presently disclosed embodiments or the application and uses of the presently disclosed embodiments. Moreover, no theory or the like is intended to be bound by any stated or implied theory or the like presented in the preceding background, summary, or detailed description.

[0018] In the following detailed description of some embodiments disclosed herein, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments disclosed herein. However, it will be apparent to one of ordinary skill in the art that the embodiments can be practiced without these specific details. In other instances, well-known features have not been described in detail to avoid obscuring the description.

[0019] Throughout the application, ordinal numbers (e.g., first, second, third, etc.) can be used as adjectives to refer to an element (i.e., any noun in the application). The use of ordinal numbers does not imply or create any particular order for the elements, nor limit any element to only a single element, unless expressly disclosed, such as through the use of the terms “before,” “after,” “single,” and other such terms. Rather, the use of ordinal numbers is to differentiate between elements. By way of example, a first element is distinct from a second element, and the first element can include more than one element and be after (or before) the second element in the order of the elements.

[0020] Multi- pharmacological drugs that interact with multiple targets (e.g., proteins) can be particularly valuable because many diseases are known to involve many proteins. Thus, drugs that can target multiple proteins associated with a disease can be more effective than drugs that are specific to only one protein.

[0021] However, finding small molecule compounds (SMCs) that interact with multiple targets can be challenging for various reasons. In particular, for example, the pockets (or more generally: interaction sites) of the target proteins can have different geometries and / or physicochemical profiles. Moreover, the structure of SMCs tends to be dynamic, e.g., there can be many conformations of the same SMC. Thus, predicting SMC-target compatibility can not be a simple task. In particular, while similar pockets are more likely to bind a common SMC, pocket similarity is not strictly required for compatibility. Instead, due to conformational changes, SMCs can be able to dock with very different pockets.

[0022] In one or more embodiments, a multi-target drug design (MTDD) strategy is used to identify small molecule compounds (SMCs) with desirable characteristics. The MTDD strategy discussed subsequently can identify drug treatments that target multiple targets simultaneously. Thus, the MTDD strategy can result in the identification of promiscuous drugs that affect disease networks rather than single targets.

[0023] More specifically, the methods and systems according to one or more embodiments leverage combinatorics to mitigate the risk of target incompatibility. That is, according to one or more embodiments, SMCs are not designed for any one particular target. Instead, a panel of receptors that can include several of the targets and anti-targets is employed. In this way, the MTDD strategy of one or more embodiments increases the likelihood of finding SMCs that can interact with at least one or some of the targets on the panel of receptors while avoiding interaction with the anti-targets. Thus, the MTDD strategy according to one or more embodiments can be used to develop a single drug with multiple affinities or as a combination of drugs for a combination therapy.

[0024] The MTDD strategy can be based on the redesign of SMCs using computational methods, where candidate drugs are designed from building blocks such as molecular fragments, atoms, etc.

[0025] Turning to Figure 1FIG. 1 shows a system (100) for differential drug discovery, in accordance with one or more embodiments. The system (100) includes a differential drug discovery engine (150). Inputs to the differential drug discovery engine (150) include a receptor set (102) and a small molecule compound (SMC) seed model (104). Outputs produced by the differential drug discovery engine (150) include one or more discovered SMCs (190). More specifically, the outputs include one or more models (formulas, descriptions) of the discovered SMCs. Each of these components will be described subsequently.

[0026] According to one or more embodiments, the receptor set (112) can specify targets (104) and anti-targets (108). A target (104) can be any cellular component in a cell of any living species to be modulated by an SMC. A target can be a macromolecule, such as a protein. Modulation of a target by a drug (SMC) can produce a beneficial outcome when treating a medical condition. In contrast, an anti-target (108) can be a macromolecule with which a drug under development should not interact. For example, modulation of an anti-target can not produce a known effect or an adverse or opposite effect of a drug action, such as toxicity. The receptor set (112) specifies targets (104) with which an SMC should interact and anti-targets (108) with which an SMC should not interact, and thus, based on these desired and undesired interactions, fitness goals are established for an SMC under development. The receptor set (112) can be established based on a certain goal, such as treatment or cure of a disease, or more generally, based on a goal to affect a living organism in a desired way. Based on this goal, the receptor set (102) can be constructed based on targets (104) and anti-targets (108). The construction of the receptor set is discussed in the flowchart of FIG. 2, and examples are provided in Figure 2 FIGS. 3, Figure 3A , Figure 3B and Figure 3C .

[0027] The targets can have multiple interaction sites that allow a drug to interact with the target. Each of these interaction sites can be represented in the receptor set (102) by one or more models (106) of the interaction site. Many models of interaction sites can be included in the receptor set (102), e.g., models of hundreds or even thousands of interaction sites. Those skilled in the art will appreciate that an interaction site can be any kind of structure or pocket of a protein that allows interaction with the protein (e.g., a binding site such as a pocket). Moreover, the interaction between the SMC and the protein is not limited to ligand-pocket binding interactions between the SMC and the protein. Rather, any kind of interaction between the SMC and the protein is within the scope of the present invention. Multiple models of an interaction site can be in the receptor set to accommodate multiple conformations. Thus, the receptor set (102) can be based on any number of targets (104), each of which can have any number of interaction sites that can be included as models (106) of interaction sites in the receptor set (102). Moreover, even a single interaction site can be modeled using multiple models to represent different conformational configurations. Some or all of the known models can be included.

[0028] According to one or more embodiments, the presence of multiple targets (104) as models (106) of interaction sites is assumed to increase the likelihood of finding an SMC that interacts with at least some of the interaction sites. For example, an SMC that is identified to interact with three of ten targets can be less challenging to identify than an SMC that is identified to interact with three of three targets. Thus, embodiments of the present invention benefit from combinatorics. Given the known difficulties associated with systematically predicting the likelihood of interaction based on known interactions between an SMC and a protein, a combinatorial approach can be particularly beneficial. In the case of ligand-pocket binding, the geometric features of the pocket (pocket volume, surface area, mouth size, etc.) are important. However, due to the many possible molecular conformations, a first ligand that binds well to a pocket does not necessarily mean that a second ligand with a very different geometry does not bind to the same pocket. Similarly, although a first ligand can bind well to a pocket, a second ligand that is only slightly different from the first ligand can not bind well to the same pocket. Given this potential poor predictability, the availability of many targets with potential for interaction with an SMC increases the likelihood of finding an SMC with acceptable performance characteristics.

[0029] Continuing the discussion of the receptor set (102), the anti-targets (108) can be specified in the same manner. However, the models (110) of interaction sites for the anti-targets are based on those proteins that were previously identified as not being targeted by the SMC under development.

[0030] In one or more embodiments, priority weights can be assigned to targets (104) and anti-targets (108). These weights can indicate the importance of interacting with the respective target (104) and the importance of avoiding interaction with the respective anti-target (108).

[0031] Further, in Figure 2 Step 200 provides a detailed description of how the panel of receptors is established, and in Figure 3A to Figure 3C Examples are provided in

[0032] According to one or more embodiments, the SMC seed model (112) is a candidate model for which the method of Figure 2 The method of Figure 2 As discussed in According to one or more embodiments, the SMC seed model (112) is a candidate model for which the method of

[0033] Continuing the discussion of the system (100), according to one or more embodiments, the differential drug discovery engine (150) accepts as input the panel of receptors (102) and the SMC seed model(s) (112) to ultimately provide as output the discovered SMCs (190). According to one or more embodiments, the method performed by the differential drug discovery engine (150) intends to obtain discovered SMCs (190) that interact with multiple targets (106) while avoiding interaction with anti-targets by exploiting the combinatorics produced by the large number of targets (104) in the panel of receptors (102). The differential drug discovery engine (150) can include a derivatization engine (152), a molecular docking simulation engine (154), and a scoring engine (156).

[0034] The derivatization engine (152) includes a set of machine-readable instructions configured to derive candidate SMCs from the seed model in the first iteration or from previously analyzed candidate SMCs in subsequent iterations. The derivation of candidate SMCs is described below in steps 204 and 206 of the method 200. Figure 2 The derivation of candidate SMCs is described below in steps 204 and 206 of the method 200.

[0035] The molecular docking simulation engine (154) includes a set of machine-readable instructions configured to simulate interactions between the candidate SMCs and the specified targets and anti-targets in the receptor set. The simulation is described below in step 208 of the method 200. Figure 2 The simulation of the simulated interactions is described below in step 210 of the method 200.

[0036] The scoring engine (156) includes a set of machine-readable instructions configured to score the simulated interactions of step 208 to obtain an individual score for each candidate SMC. The simulation of the simulated interactions is described below in step 210 of the method 200. Figure 2 The simulation of the simulated interactions is described below in step 210 of the method 200.

[0037] The derivatization engine (152), the molecular docking simulation engine (154), and the scoring engine (156) combine to iteratively produce candidate SMCs that can ultimately qualify as discovered SMCs (190) having the desired characteristics. A discussion of the iterative execution is provided below with reference to the method 200. Figure 2 The simulation of the simulated interactions is described below in step 210 of the method 200.

[0038] Figure 2 Flow diagrams in accordance with one or more embodiments are shown. Although the various steps in the flow diagrams are provided and described sequentially, one of ordinary skill in the art will appreciate that some or all of the steps can be executed in different orders, can be combined or omitted, and some or all of the steps can be executed in parallel. Furthermore, the steps can be executed actively or passively. For example, according to one or more embodiments, some steps can be executed using polling or be interrupt driven. By way of example, according to one or more embodiments, a determination step can not require processor processing instructions unless an interrupt is received to indicate that a condition exists. As another example, according to one or more embodiments, a determination step can be executed by performing a test, such as checking a data value to test whether the value is consistent with a condition being tested.

[0039] Figure 2A flowchart of the process illustrates a method for differential drug discovery according to one or more embodiments. The method for differential drug discovery is based on a fragment growing strategy (FGS) to computationally optimize SMCs using target interaction sites of specified 3D protein structures in a receptor set. The FGS can include at least the following steps: finding lead scaffolds that "fit" the pocket using docking, molecular dynamics (MD) simulations, or machine learning methods based on profiled SMCs and / or profiled interaction sites; modifying the SMCs, rescoring, and selecting the best new SMCs; and iterating to optimize results, as discussed subsequently. With each iteration, changes are made to the SMC(s) by redesigning some SMC fragments, rederiving some molecules, building onto or removing from molecules.

[0040] In one or more embodiments, the methods described subsequently optimize the SMC(s) for an entire receptor set of several targets and anti-targets. In an example where the receptor set includes 64 different targets, 36 of the 64 targets can include targets that can have positive value therapeutically, while the remaining 28 targets can include anti-targets that should be optimized against. Using the disclosed methods, a given SMC of the example can be evaluated for all 64 different targets of the receptor set to determine predicted interactions and create a polypharmacology score. The polypharmacology score can be computed such that it rewards predicted interactions with multiple targets and penalizes interactions with anti-targets.

[0041] Turning to the flowchart, in step 200, a receptor set is obtained. The receptor set can be obtained in its final format, as illustrated in the example of Figure 3C , or alternatively, the receptor set can be constructed. Constructing the receptor set can be performed as follows. First, a selection of proteins (or other targets) can be obtained, including targets and anti-targets, for example, as a list of proteins, as illustrated in the example of Figure 3A The selection of proteins provided can be established based on a desired therapeutic effect (e.g., when treating a disease). Next, the receptor set can be compiled by mapping 3D structures (e.g., a list of atoms that make up a protein with its 3D positions) for each of the proteins on the list. Each of these 3D positions can be an interaction site, such as a pocket. In addition, for each of the 3D positions in the mapping, a known conformation (as a result of a conformational change) can be obtained. Figure 3B An example (limited to a single protein) is provided. Subsequently, clustering can be performed to reduce the total number of 3D positions. Clustering can be performed using any similarity measure between 3D structures. A sample can be selected for each of the clusters, and the receptor set is obtained by compiling the samples of targets and anti-targets. An example of a receptor set is illustrated in Figure 3C .

[0042] In step 202, SMC seed model(s) are obtained. As previously described, the SMC seed models can be obtained as SMILES representations.

[0043] In step 204, candidate SMCs are derived. For the first execution cycle of the method, step 204 can be skipped, i.e., the steps following step 204 can operate directly on the SMC seed model(s). For subsequent execution cycles, derivatization is performed in accordance with one or more embodiments, as subsequently described. Figure 2

[0044] With each execution of step 204, the SMCs under consideration can be modified computationally by replacing functional groups with other chemical fragments. Thus, new SMCs are obtained from the parent SMCs, i.e., the SMCs or SMC seed models obtained from the previous execution cycle can be modified. More specifically, the SMCs can be modified by breaking the SMCs into fragments and by exchanging, adding, and / or removing fragments from the SMCs. These operations can be governed by a set of rules to ensure that the essential structural features are preserved. The set of rules can be based on, for example, the Reverse Synthesis Combinatorial Procedure (RECAP) or Breaking Interesting Chemical Substructures for Reverse Synthesis (BRICS). Those skilled in the art will recognize that the present application is not limited to the particular set of rules used to establish how the SMCs are modified computationally. Any method that can meaningfully modify the SMCs chemically can be used. Moreover, the modifications can be of any size, ranging from modification of a single atom to modification of larger chemical substructures. In one embodiment, fragmentation is performed exhaustively. For example, consider the molecule A-B-C. Exhaustive fragmentation can yield the fragments A-B, B-C, A, B, and C. The resulting fragments can then be modified by adding one or more other fragments obtained from a fragment library to obtain new candidate SMCs. Although derivatization of candidate SMCs can be performed randomly, certain restrictions can be imposed when modifying the SMCs. For example, a minimum similarity to the parent SMC can be required, a portion of the parent SMC can need to remain intact, etc.

[0045] Derivatization of step 204 can be performed on all of the SMCs obtained from the previous execution cycle, or on a subset of the SMCs. For example, only the top-ranked SMCs can be considered based on SMC interaction scores. The selection criteria can vary between the early (exploratory) and late (refinement) stages of optimization.

[0046] ​In step 206, unqualified candidate SMCs are removed from the candidate SMCs based on screening criteria. Screening criteria can include, but are not limited to, a requirement that the candidate SMCs have minimal similarity to known drugs, a requirement that the SMCs be synthesized using no more than a prescribed effort, a requirement that certain ADMET properties be present, and / or a requirement that other desirable computational properties, such as optimal lipophilicity or absence of labile chemical groups, etc. Selection criteria can vary between the early (exploratory) and late (refinement) stages of optimization.

[0047] Furthermore, in one or more embodiments, a clustering algorithm can be used to identify representative SMCs from a set of candidate SMCs having similar polypharmacological profiles to undergo a subsequent round of optimization.

[0048] In step 208, according to one or more embodiments, the interactions of the SMCs with the targets and anti-targets in the set of receptors are simulated. As previously described, the interactions can be docking of the SMCs with the pockets, or more generally, any kind of interaction of the SMCs with the interaction sites. The simulations can be performed for all combinations of SMCs and interaction sites of the targets and anti-targets. Each of these interactions can be scored to assess the degree of interaction. The simulations can be performed in various ways, as subsequently described. After completion of step 208, each of the SMCs can be evaluated for interaction with each of the interaction sites (targets and anti-targets) in the set of targets based on the scores obtained.

[0049] In one embodiment, a molecular docking method is used to simulate the interactions between the SMCs and the targets (and anti-targets). The molecular docking method can rely on Monte Carlo simulations to minimize the energy associated with the interactions between the SMCs and the interaction sites. An energy-based score can be obtained based on the pose of the SMC that results in the interaction.

[0050] In one embodiment, a molecular dynamics method is used to simulate the interactions between the SMCs and the targets (and anti-targets). A physics engine operating on the SMCs and the interaction sites can determine whether binding, or more generally, interaction, occurs. Once an interaction is detected, a priority-based score can be obtained for the configuration of the SMC and the interaction site.

[0051] In one embodiment, a machine learning method is used to model the interaction between SMCs and targets (and anti-targets) based on the characterized SMCs and the characterized targets / anti-targets. The machine learning method can use a prediction algorithm, such as a random forest, a convolutional neural network, or any other prediction algorithm capable of making quantitative predictions. The prediction can be binding affinity. The prediction algorithm can have been previously trained using historical data, where the interaction (or lack of interaction) between SMCs and interaction sites is known. Thus, the trained prediction algorithm can predict a binding affinity that indicates what degree of interaction the SMC under consideration will have with the interaction site under consideration. The predicted affinity can be as a score.

[0052] After completion of step 208, a score is available for each of the interactions between the SMCs under consideration and the interaction sites under consideration.

[0053] In step 210, SMC interaction scores for SMC-target / anti-target interactions are obtained by evaluating the scores obtained in step 208. Specifically, one SMC interaction score is obtained for each of the SMCs. The SMC interaction score can indicate the degree to which the SMC interacts with the targets in the receptor set while avoiding interaction with the anti-targets in the receptor set. Broadly, the SMC interaction score can reward predicted interactions with the targets in the receptor set (resulting in an increase in the SMC interaction score) while penalizing interactions with the anti-targets in the receptor set (resulting in a decrease in the SMC interaction score). The SMC interaction score can be calculated in various ways. For example, the weighted sum of the top three (or five) target interaction scores minus the top three (or five) anti-target interaction scores for the SMC can be used. Figure 4 An example using the top three interaction scores is shown in Table 2, described below. Further, in one or more embodiments, the SMC interaction score can be designed to reward receptor combinations from the same or different biological pathways, which are believed to provide synergistic therapeutic results. Different weights can also be applied to different interaction sites to emphasize / undermine the contribution of these interaction sites to the SMC interaction score. For example, a weight of 1.5 can be assigned to a primary interaction site, a weight of 1.0 can be assigned to a secondary interaction site, and a weight of 0.5 can be assigned to a tertiary interaction site to encourage exploration in more important interaction sites or targets. Based on the SMC interaction scores, the associated candidate SMCs can be ranked.

[0054] In step 212, the SMC interaction score is evaluated to determine whether one or more of the candidate SMCs are eligible as drugs. If the associated SMC interaction score meets or exceeds a minimum score, the SMC may be eligible as a drug. Other criteria that can be calculated for the SMC (e.g., molecular weight, solubility, and / or other relevant properties) may also be used to classify the SMC as a drug.

[0055] Step 214 is used to determine whether another iteration should be performed or whether the execution of the method should be terminated. This determination can be made based on whether at least one of the SMCs qualifies as a drug. The determination can also be based on convergence. Convergence can be evaluated based on the score obtained in step 208. Convergence can be detected once the score has reached a certain threshold, has stabilized (e.g., no significant improvement after two iterations), etc. Alternatively or additionally, cost may be a determining factor. Execution can be used... Figure 2 The method measures the cost based on the CPU time spent, and the simulation can be terminated after a certain amount of CPU time has been spent. If another iteration is to be performed, the execution of the method can proceed to step 204. Alternatively, the execution of the method can be terminated.

[0056] Figure 2 The steps can be performed on many candidate SMCs. For example, candidate SMCs of 10s, 100s, 1000s, or 10000s can be processed. These candidate SMCs may originate from a single or multiple SMC seed models.

[0057] Turn Figure 3A , Figure 3B and Figure 3C This provides examples of generating receptor groups according to one or more embodiments. Figure 3A The target list (300) is shown in the diagram. The target list (300) enumerates proteins. For example, when treating a disease, proteins can be selected based on the desired therapeutic effect. In this example, each protein is identified by a UniProt ID. Furthermore, a protein classification is associated with each protein. The classification indicates whether the protein is intended to be used as a target or as an anti-target.

[0058] Figure 3B The compilation of a protein pocket model (310) according to one or more embodiments is shown. Figure 3B In the example, only a protein pocket model of protein "Q6PL18" is shown. Although Figure 3B Not shown in the image, when executing Figure 2 Step 200 yields the target Figure 3A Protein pocket model of all proteins listed in the target list (300).

[0059] Figure 3C Examples of receptor sets (320) according to one or more embodiments are shown. For Figure 3A Each of the proteins listed in the target list (300) includes a representative protein pocket model of the receptor group (320), from which... Figure 3B The protein pocket model (310) was obtained during compilation.

[0060] Figure 4 An example of a rating interaction (400) according to one or more embodiments is shown. Figure 4 The results for five SMCs (C001-C005), 16 targets, and eight anti-targets are shown. Based on the SMC interaction scores obtained in step 210, the top three targets are labeled, and the interaction scores of the top three anti-targets are also labeled. Figure 4 As illustrated, according to one or more embodiments, each SMC may have a different set of advantageous targets and anti-targets.

[0061] The various embodiments have one or more of the following advantages. The embodiments of this disclosure utilize combination. Due to the use of a relatively large receptor set, which may also include multiple models of the same protein, the likelihood of identifying SMCs that successfully interact with at least some targets increases. A larger number of targets in the receptor set can also have the additional benefit of allowing the interaction fraction to be normalized, where high or low ligand fractions across all targets can be downregulated or upregulated accordingly to avoid selecting confounding ligands, i.e., those that are generally more sticky toward all targets. Furthermore, the use of anti-targets provides additional compounds for normalization and allows optimization for target interactions that may be problematic for a specific condition or SMC scaffold.

[0062] Pairing receptor sets with iterative optimization strategies allows for simultaneous exploration across both the chemical and receptor set target spaces. Each SMC can have its own distinct set of targets. In the next generation, derivatives of SMCs can improve upon the same target or identify new combinations of targets.

[0063] Examples may require only three-dimensional (3D) structures of the target and anti-target, and one or more substructures. The molecular flexibility of the ligand and receptor, due to the use of 3D structure-based molecular docking simulations, allows for the detection of compatible target pairs with different binding site geometries. Experimental target SMC binding data are not required. The methods and systems of one or more embodiments are computational. That is, the disclosed methods can be performed entirely by computer. However, in vitro experiments can be integrated without departing from the invention.

[0064] In therapeutic applications, the polypharmacological drugs obtained using the described methods can be used to treat diseases by modulating multiple targets. The polypharmacological drugs can be more effective than a single conventional drug, and can greatly reduce the risk of losing efficacy due to single target mutations. Another significant advantage of using a single polypharmacological drug instead of a mixture of individual drugs can be a reduced risk of drug-drug interactions.

[0065] Embodiments of the present disclosure can be implemented on a computing system. Any combination of mobile, desktop, server, router, switch, embedded device or other type of hardware can be used. For example, as shown in FIG. 5, a computing system (500) can include one or more computer processors (502), non-persistent storage (504) (e.g., volatile memory, such as random access memory (RAM), cache memory), persistent storage (506) (e.g., a hard disk, optical, floppy, or other disk drive, flash memory, etc.), a communication interface (512) (e.g., a Bluetooth, infrared, network, or other interface), and many others. Figure 5A

[0066] The computer processor(s) (502) can be integrated circuits for processing instructions. For example, the computer processor(s) can be one or more cores or micro-cores of a processor. The computing system (500) can also include one or more input devices (510), such as a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device.

[0067] The communication interface (512) can include an integrated circuit for connecting the computing system (500) to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, a mobile network, or any other type of network) and / or another device, such as another computing device.

[0068] In addition, the computing system (500) can include one or more output devices (508), such as a screen (e.g., a liquid crystal display (LCD), a plasma display, a touchscreen, a cathode ray tube (CRT) monitor, a projector, or other display device), a printer, an external storage device, or any other output device. One or more of the output devices can be the same as or different from the input device(s). The input and output device(s) can be locally or remotely connected to the computer processor(s) (502), the non-persistent storage (504), and the persistent storage (506). There are many different types of computing systems, and the aforementioned input and output device(s) can take other forms. ​

[0069] Software instructions for implementing embodiments of the disclosure in the form of computer readable program code can be stored in non-transitory computer-readable media, such as CD, DVD, storage devices, diskette, tape, flash memory, physical memory, or any other computer readable storage medium, either entirely or in part. In particular, the software instructions can correspond to computer readable program code that, when executed by the processor(s), is configured to perform one or more embodiments of the disclosure.

[0070] Figure 5A The computing system (500) in Figure 5B may be connected to or a part of a network of one or more other computing systems. For example, as shown in Figure 5A , the network (520) can include a plurality of nodes (e.g., node X (522), node Y (524)). Each node can correspond to a computing system, such as the computing system shown in Figure 5A or a combination of nodes can correspond to the computing system shown in By way of example, embodiments of the present disclosure can be implemented on a node of a distributed system that is connected to other nodes of the distributed system via a communications network. As another example, embodiments of the present disclosure can be implemented on a distributed computing system, where each part of the present disclosure can reside on a different node of the distributed computing system. Further, one or more elements of the aforementioned computing system (500) can be located at a remote location and connected to the other elements by a network.

[0071] Although not shown in Figure 5B , the nodes can correspond to blades in a server chassis that are connected to other nodes via a backplane. As another example, the nodes can correspond to servers in a data center. As another example, the nodes can correspond to computer processors or micro-cores of computer processors with shared memory and / or resources.

[0072] The nodes (e.g., node X (522), node Y (524)) in the network (520) can be configured to provide services to client devices (526). For example, the nodes can be part of a cloud computing system. The nodes can include functionality to receive requests (526) from and transmit responses (526) to client devices. The client devices (526) can be computing systems, such as the computing system shown in Figure 5A . Further, the client devices (526) can include and / or execute all or a portion of one or more embodiments of the present disclosure.

[0073] In Figure 5A and Figure 5BThe computing system or computing systems described in the preceding can include functionality to perform various operations disclosed herein. For example, the computing system(s) can perform communications between processes on the same or different systems. A variety of mechanisms can facilitate data exchange between processes on the same device using some form of active or passive communication. Examples representative of these inter-process communications include, but are not limited to, implementation of files, signals, sockets, message queues, pipes, semaphores, shared memory, message passing, and memory-mapped files. Additional details related to several of these non-limiting examples are provided below.

[0074] Based on the client-server network model, sockets can be used as interfaces or endpoints of communication channels to enable bidirectional data transfer between processes on the same device. First, following the client-server network model, a server process (e.g., a process that provides data) can create a first socket object. Next, the server process binds the first socket object, thereby associating the first socket object with a unique name and / or address. After creating and binding the first socket object, the server process then waits and listens for incoming connection requests from one or more client processes (e.g., processes that seek data). At this point, when a client process wishes to obtain data from the server process, the client process begins by creating a second socket object. The client process then proceeds to generate a connection request that includes at least the second socket object and the unique name and / or address associated with the first socket object. The client process then transmits the connection request to the server process. Depending on availability, the server process can accept the connection request, establish a communication channel with the client process, or the server process, busy processing other operations, can queue the connection request in a buffer until the server process is ready. The established connection notifies the client process that communication can begin. In response, the client process can generate a data request specifying the data that the client process wishes to obtain. The data request is then transmitted to the server process. Upon receiving the data request, the server process analyzes the request and collects the requested data. Finally, the server process then generates a reply that includes at least the requested data, and transmits the reply to the client process. More commonly, data can be transferred as datagrams or streams of characters (e.g., bytes).

[0075] Shared memory refers to the allocation of virtual memory space in order to demonstrate a mechanism by which data can be communicated and / or accessed by multiple processes. In implementing shared memory, an initialization process first creates a sharable segment in persistent or non-persistent storage. After creation, the initialization process then loads the sharable segment, and subsequently maps the sharable segment into an address space associated with the initialization process. After loading, the initialization process proceeds to identify and grant access to one or more authorized processes that can also write data to and read data from the sharable segment. Changes made to data in the sharable segment by one process can immediately affect other processes that are also linked to the sharable segment. Moreover, when one of the authorized processes accesses the sharable segment, the sharable segment is mapped to the address space of the authorized process. Typically, only one authorized process can load the sharable segment at any given time, in addition to the initialization process.

[0076] Other techniques can be used to share data between processes, such as the various data described in this application, without departing from the scope of the present disclosure. The processes can be part of the same or different applications, and can be executed on the same or different computing systems.

[0077] Instead of or in addition to sharing data between processes, a computing system executing one or more embodiments of the present disclosure can include functionality to receive data from a user. For example, in one or more embodiments, a user can submit data via a graphical user interface (GUI) on a user device. Data can be submitted via the graphical user interface by the user selecting one or more graphical user interface widgets or inserting text and other data into the graphical user interface widgets using a touchpad, keyboard, mouse, or any other input device. In response to selecting a particular item, information about the particular item can be obtained by a computer processor from persistent or non-persistent storage. The obtained data about the particular item can be displayed on the user device in response to the user's selection when the user selects the item.

[0078] As another example, a request for data regarding a particular item can be sent to a server that is operably connected to the user device over a network. For example, a user can select a uniform resource locator (URL) link within a web client of the user device, thereby initiating a hypertext transfer protocol (HTTP) or other protocol request to a network host associated with the URL. In response to the request, the server can extract data regarding a particular selected item and send the data to the device that initiated the request. Once the user device has received the data regarding the particular item, the content of the received data regarding the particular item can be displayed on the user device in response to a user selection. Following the above example, the data received from the server after selection of the URL link can provide a web page in hypertext markup language (HTML) that can be rendered by the web client and displayed on the user device.

[0079] Once the data is obtained, such as by using the techniques described above or from a storage device, the computing system, in performing one or more embodiments of the present disclosure, can extract one or more data items from the obtained data. For example, the computing system can extract one or more data items from the obtained data as follows. First, the organization pattern of the data (e.g., syntax, schema, layout) is determined, which can be based on one or more of: position (e.g., bit or column position, Nth token in a data stream, etc.); attribute (where an attribute is associated with one or more values); or hierarchical / tree structure (composed of layers of nodes at different levels of detail, such as in nested package headers or nested document sections). Then, in the context of the organization pattern, the raw, unprocessed stream of data symbols is parsed into a stream of tokens (or hierarchical structure) (where each token can have an associated token “type”). Figure 5A

[0080] Next, an extraction criterion is extracted for extracting one or more data items from the stream of tokens or structure, where the extraction criterion is processed according to the organization pattern to extract one or more tokens (or nodes from the hierarchical structure). For position-based data, the token(s) at the position(s) identified by the extraction criterion are extracted. For attribute / value-based data, the token(s) and / or node(s) associated with the attribute(s) that satisfy the extraction criterion are extracted. For hierarchical / layers data, the token(s) associated with the node(s) that match the extraction criterion are extracted. The extraction criterion can be as simple as an identifier string, or can be a query provided to a structured data repository (where the data repository can be organized according to a database schema or data format, such as XML).

[0081] The extracted data can be used for further processing by the computing system. For example, the computing system can perform one or more operations on the extracted data, such as: Figure 5A ​Computing systems of the present disclosure, when executing one or more embodiments of the present disclosure, can perform data comparisons. Data comparisons can be used to compare two or more data values (e.g., A, B). For example, one or more embodiments can determine whether A > B, A = B, A!= B, A < B, etc. Comparisons can be performed by submitting an operation code A, B specifying an operation related to the comparison to an arithmetic logic unit (ALU), i.e., a circuit that performs arithmetic and / or bitwise logical operations on two data values. The ALU outputs a numeric result of the operation and / or one or more status flags related to the numeric result. For example, a status flag can indicate that the numeric result is positive, negative, zero, etc. By selecting the appropriate operation code and then reading the numeric result and / or status flags, a comparison can be performed. For example, to determine whether A > B, B can be subtracted from A (i.e., A - B), and the status flags can be read to determine whether the result is positive (i.e., if A > B, then A - B > 0). In one or more embodiments, if A = B or if A > B, B can be considered a threshold and A is considered to satisfy the threshold, as determined using the ALU. In one or more embodiments of the present disclosure, A and B can be vectors, and comparing A to B requires comparing a first element of vector A to a first element of vector B, a second element of vector A to a second element of vector B, etc. In one or more embodiments, if A and B are strings, the binary values of the strings can be compared.

[0082] Figure 5A Computing systems in the present disclosure can implement and / or connect to data repositories. For example, one type of data repository is a database. A database is a collection of information configured to facilitate data retrieval, modification, reorganization, and deletion. A database management system (DBMS) is a software application that provides an interface for users to define, create, query, update, or manage a database.

[0083] A user or software application can submit a statement or query to a DBMS. The DBMS then interprets the statement. The statement can be a select statement requesting information, an update statement, a create statement, a delete statement, etc. Also, the statement can include parameters specifying data or data containers (databases, tables, records, columns, views, etc.), identifier(s), conditions (comparison operators), functions (e.g., join, full join, count, average, etc.), ordering (e.g., ascending, descending), or other. The DBMS can execute the statement. For example, the DBMS can access memory buffers, references, or index files for reading, writing, deleting, or any combination thereof in response to the statement. The DBMS can load data from persistent or non-persistent storage and perform computations to respond to the query. The DBMS can return result(s) to the user or software application.

[0084] Figure 5A The computing system of the present disclosure can include functionality to provide raw and / or processed data, such as results of comparisons and other processing. For example, providing data can be accomplished through a variety of presentation methods. In particular, data can be provided through a user interface provided by the computing device. The user interface can include a GUI that displays information on a display device (computer monitor or touch screen on a handheld computing device). The GUI can include a variety of GUI widgets that organize what data is shown and how the data is provided to the user. Further, the GUI can provide data directly to the user, for example, data provided as actual data values through text or visual representations of data presented by the computing device, such as through a visual data model.

[0085] For example, the GUI can first obtain a notification from a software application requesting that a particular data object be provided within the GUI. Next, the GUI can identify a data object type associated with the particular data object, for example, by obtaining data from a data property within the data object that identifies the data object type. The GUI can then determine any rules specified for displaying the data object type, for example, rules specified by a software framework for a data object class or rules specified according to any local parameters defined by the GUI for presenting the data object type. Finally, the GUI can obtain a data value from the particular data object and present a visual representation of the data value within the display device according to the specified rules for the data object type.

[0086] Data can also be provided through a variety of audio methods. In particular, data can be presented in an audio format and provided as sound through one or more speakers operably connected to the computing device.

[0087] Data can also be provided to the user through haptic methods. For example, haptic methods can include vibrations or other physical signals generated by the computing system. For example, a vibration generated using a handheld computing device having a predetermined duration and intensity of vibration can be used to provide data to the user to convey the data.

[0088] The above description of functionality represents only a few examples of the functionality that can be executed by the computing system of the present disclosure and / or a client device in the system of the present disclosure. Other functionality can be performed using one or more embodiments of the present disclosure. Figure 5A Figure 5B The above description of functionality represents only a few examples of the functionality that can be executed by the computing system of the present disclosure and / or a client device in the system of the present disclosure. Other functionality can be performed using one or more embodiments of the present disclosure.

[0089] Although the present disclosure has been described with respect to a limited number of embodiments, those skilled in the art having the benefit of this disclosure will appreciate that other embodiments can be devised which do not depart from the scope of the present disclosure as disclosed herein. Accordingly, the scope of the present disclosure should be limited only by the appended claims.

[0090] ​The embodiments and examples set forth herein are presented to best explain the present application and its particular application to one skilled in the art. However, those skilled in the art will recognize that the foregoing description and examples have been presented for the purpose of illustration and example only. The description set forth is not intended to be exhaustive or to limit the application to the precise forms disclosed.

[0091] While the present application has been described with respect to a limited number of embodiments, those skilled in the art will appreciate that other embodiments can be devised which, while not specifically described, fall within the scope of the present application as disclosed herein. Accordingly, the scope of the application should be limited only by the appended claims.

Claims

1. A computer-implemented method for differential drug discovery, the method comprising: obtaining a receptor set, wherein the receptor set specifies a plurality of targets and a plurality of anti-targets, wherein each of the targets and the anti-targets comprises a single interaction site or a plurality of interaction sites; wherein the receptor set comprises models of the interaction sites of the targets and the anti-targets, and wherein the receptor set further comprises a plurality of models describing a plurality of the interaction sites and / or a plurality of models describing different molecular conformations of one of the interaction sites; obtaining a small molecule compound (SMC) seed model; deriving a first plurality of candidate SMCs from the SMC seed model; for each of the candidate SMCs in the first plurality of candidate SMCs, simulating a first desired interaction between the candidate SMC and each of the plurality of targets; for each of the candidate SMCs in the first plurality of candidate SMCs, simulating a first undesired interaction between the candidate SMC and each of the plurality of anti-targets; based on the first desired interactions and based on the first undesired interactions, obtaining a first SMC interaction score for each of the candidate SMCs in the first plurality of candidate SMCs, wherein obtaining a first SMC interaction score comprises calculating the first SMC interaction score based on a score quantifying a degree of interaction between the candidate SMC and the interaction sites of the targets and the anti-targets listed in the receptor set, wherein the score quantifying a degree of interaction between the candidate SMC and a target's interaction site increases the first SMC interaction score, wherein the score quantifying a degree of interaction between the candidate SMC and an anti-target's interaction site decreases the first SMC interaction score; and determining whether a minimum score for a drug is at least achieved based on the first SMC interaction scores.

2. The method of claim 1, further comprising: deriving a second plurality of candidate SMCs from the first plurality of candidate SMCs; for each of the candidate SMCs in the second plurality of candidate SMCs, simulating a second desired interaction between the candidate SMC and each of the plurality of targets; for each of the candidate SMCs in the second plurality of candidate SMCs, simulating a second undesired interaction between the candidate SMC and each of the plurality of anti-targets; based on the second desired interactions and based on the second undesired interactions, obtaining a second SMC interaction score for each of the candidate SMCs in the second plurality of candidate SMCs; and determining whether a minimum score for a drug is at least achieved based on the second SMC interaction scores.

3. The method of claim 1, further comprising, prior to simulating the first desired interactions, updating the first plurality of candidate SMCs by removing a subset of candidate SMCs based on a screening criterion. ​ ​ 4. The method of claim 3, wherein the screening criteria comprises at least one selected from the group consisting of ADMET properties, synthesizability, and similarity to known drugs.

5. The method of claim 1, wherein deriving the first plurality of candidate SMCs from the SMC seed model comprises: segmenting the SMC seed model into fragments; and swapping at least one of the fragments.

6. The method of claim 5, wherein segmenting the SMC seed model into fragments is performed exhaustively.

7. The method of claim 1, wherein simulating the first desired interaction comprises: for each combination of a candidate SMC and a selected interaction site of an interaction site of the target listed from the receptor set, obtaining a score quantifying a degree of interaction between the candidate SMC and the interaction site.

8. A system for differential drug discovery, the system comprising: a derivation engine configured to derive a plurality of candidate small molecule compounds (SMCs) from a SMC seed model; a molecular docking simulation engine configured to: simulate, for each of a candidate SMC of the plurality of candidate SMCs, a desired interaction between the candidate SMC and each of a plurality of targets specified in a receptor set; simulate, for each of a candidate SMC of the plurality of candidate SMCs, an undesired interaction between the candidate SMC and each of a plurality of anti-targets specified in the receptor set; wherein each of the targets and the anti-targets comprises a single interaction site or a plurality of interaction sites; wherein the receptor set comprises models of interaction sites of the targets and the anti-targets, and wherein the receptor set further comprises a plurality of models describing a plurality of the models of interaction sites and / or a plurality of models describing different molecular conformations of one of the interaction sites; a scoring engine configured to: obtain, for each of a candidate SMC of the plurality of candidate SMCs, a first SMC interaction score based on the desired interactions and based on the undesired interactions, wherein the first SMC interaction score is computed by obtaining a score quantifying a degree of interaction between the candidate SMC and the plurality of targets and the plurality of anti-targets listed in the receptor set based on the score quantifying a degree of interaction between the candidate SMC and the interaction site of the plurality of targets and the plurality of anti-targets to obtain the first SMC interaction score, wherein the score quantifying a degree of interaction between the candidate SMC and a target interaction site increases the first SMC interaction score, wherein the score quantifying a degree of interaction between the candidate SMC and an anti-target interaction site decreases the first SMC interaction score; and determine, based on the first SMC interaction score, whether a minimum score of a drug is at least achieved.

9. A non-transitory computer readable medium comprising computer readable program code for differential drug discovery, the computer readable program code causing a computer system to: ​ obtaining a receptor set, wherein the receptor set specifies a plurality of targets and a plurality of anti-targets, wherein each of the targets and the anti-targets comprises a single interaction site or a plurality of interaction sites; wherein the receptor set comprises a model of the interaction sites of the targets and the anti-targets, and wherein the receptor set further comprises a plurality of models describing a plurality of the interaction sites and / or a plurality of models describing different molecular conformations of one of the interaction sites; obtaining a small molecule compound (SMC) seed model; deriving a first plurality of candidate SMCs from the SMC seed model; for each of the candidate SMCs in the first plurality of candidate SMCs, simulating a first desired interaction between the candidate SMC and each of the plurality of targets; for each of the candidate SMCs in the first plurality of candidate SMCs, simulating a first undesired interaction between the candidate SMC and each of the plurality of anti-targets; based on the first desired interactions and based on the first undesired interactions, obtaining a first SMC interaction score for each of the candidate SMCs in the first plurality of candidate SMCs, wherein the first SMC interaction score is computed by a score quantifying a degree of interaction between the candidate SMC and the interaction sites of the targets and anti-targets listed in the receptor set to obtain the first SMC interaction score, wherein a score quantifying a degree of interaction between the candidate SMC and the interaction sites of a target increases the first SMC interaction score, wherein a score quantifying a degree of interaction between the candidate SMC and the interaction sites of an anti-target decreases the first SMC interaction score; and and determining whether a minimum score of a drug is at least achieved based on the first SMC interaction score.

10. The non-transitory computer readable medium of claim 9, wherein the computer readable program code further causes the computer system to: deriving a second plurality of candidate SMCs from the first plurality of candidate SMCs; for each of the candidate SMCs in the second plurality of candidate SMCs, simulating a second desired interaction between the candidate SMC and each of the plurality of targets; for each of the candidate SMCs in the second plurality of candidate SMCs, simulating a second undesired interaction between the candidate SMC and each of the plurality of anti-targets; based on the second desired interactions and based on the second undesired interactions, obtaining a second SMC interaction score for each of the candidate SMCs in the second plurality of candidate SMCs; and determining whether a minimum score of a drug is at least achieved based on the second SMC interaction score.

11. The method of claim 1, the system of claim 8, or the non-transitory computer readable medium of claim 9, wherein the receptor set comprises a model of hundreds of interaction sites.

12. The method of claim 1, the system of claim 8, or the non-transitory computer- readable medium of claim 9, wherein the set of receptors comprises a model of thousands of interaction sites.

Citation Information

Patent Citations

  • Design of molecules

    US20120265514A1

  • Systems and methods for applying a convolutional network to spatial data

    US20160300127A1