Generative design of small molecules
An iterative generative design method for small molecules addresses the challenge of identifying drug candidates by optimizing binding affinity and molecular diversity, enhancing the efficiency of small molecule drug development.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-08
- Publication Date
- 2026-03-25
Smart Images

Figure 2026509837000001_ABST
Abstract
Description
[Technical Field]
[0001] Technical field The subject matter described herein generally relates to small molecule design, and more specifically, to iterative methods for generating sets of molecules that bind to specific targets. [Background technology]
[0002] background Small molecule drugs, typically with molecular weights of approximately 100 to 1,000 daltons, modulate biochemical processes to diagnose, treat, and prevent a wide range of diseases. Such molecules form the basis of modern pharmacology due to several compelling advantages, including remarkable flexibility in their formulation and pharmacokinetics. For example, small molecule drugs can penetrate cell membranes to reach intracellular targets. Furthermore, small molecule drugs are adaptable to a wide variety of therapeutic applications. For instance, they can be formulated as pills and capsules, intravenous or subcutaneous injections, inhalants, or suppositories. Consequently, small molecule drugs have been designed to suit a diverse range of therapeutic uses. The development of small molecule drugs typically involves achieving specific specificity in binding to particular targets, along with optimizing various pharmacokinetic properties, including release, absorption, distribution, metabolism, and excretion.
[0003] However, developing small molecule drugs that exhibit specific desirable properties, such as the ability to bind to a target site, is a challenging and resource-intensive task because, although the chemical space of at least all possible drug-like molecules is vast, it is very sparsely occupied by potential drug candidates that possess the necessary binding affinity to any given target. 33 ~10 60 Given that the vast majority of individual drug-like molecules do not possess medicinal properties, identifying the few molecules in that space that are potential drug candidates usually requires exploring as much of the chemical space as efficiently as possible.
[0004] Nevertheless, a brute-force approach that indiscriminately screens all drug-like molecules to identify candidate molecules capable of binding to a target site, even if performed in silico, is computationally too expensive to be a practical solution. [Overview of the project]
[0005] Summary of Disclosure Systems, methods, and products (including computer program products) are provided for the generative design of small molecules. In one embodiment, a system is provided comprising at least one processor and at least one memory. The at least one memory may contain program code that gives operation when executed by the at least one processor.
[0006] In another embodiment, a computer program product is provided that includes a non-temporary computer-readable medium for storing instructions. The instructions can cause an operation to be performed by at least one data processor.
[0007] In another embodiment, a system for the generative design of small molecules is provided, comprising at least one processor and at least one memory. The at least one memory may contain program code that gives behavior when executed by at least one processor.
[0008] In another aspect, a computer implementation method, (a) Obtaining an initial molecular population from a molecular structure database, (b) Modifying at least one molecule in the initial molecular population and selecting a predetermined number of molecules from the initial molecular population and at least one modified molecule that satisfy one or more goodness-of-fit scores to generate a first new molecular population, (c) Modifying at least two molecules in the first new molecular population, and generating a second new molecular population by selecting one or more other predetermined numbers of molecules from the first new molecular population and at least the second modified molecules that satisfy one or more goodness-of-fit scores, (d) In response to determining that the convergence criteria have not been met, repeat step (c) at least once using a second new molecular population as the first new molecular population, (e) In response to determining that the convergence criteria have been met, select a second subset of new molecular populations as candidates for synthesis and testing, A computer implementation method is provided, which includes the following.
[0009] In another aspect, a computer implementation method, (a) Obtaining multiple initial molecular populations from molecular structure databases, (b) Generating a plurality of first new molecular populations from each of a plurality of initial molecular populations, wherein each of the plurality of first new molecular populations contains one or more molecules obtained by modifying one or more molecules in the initial molecular population on which the first new population is based, and each of the plurality of first new molecular populations satisfies one or more goodness-of-fit scores, (c) Generating a plurality of second new molecular populations from each of a plurality of first new molecular populations, wherein each of the plurality of second new populations contains one or more molecules obtained by modifying one or more molecules in the first new molecular population on which the second new population is based, and the molecules in each of the plurality of second new molecular populations satisfy one or more goodness-of-fit scores, (d) In response to determining that the convergence criteria are not met, repeat step (c) at least once, such that each of the first new molecular populations is one of the second new molecular populations obtained from the precedents of step (c). (e) In response to determining that the convergence criteria have been met, select a subset of molecules containing at least one molecule from each of several second new molecular populations for synthesis and testing, A computer implementation method is provided, which includes the following.
[0010] In another aspect, a computer implementation method, (a) Obtaining an initial molecular population in which each molecule has a molecular structure constructed from a reaction database containing a reaction set, wherein each reaction from the reaction set is associated with two or more reagent sets, and each molecule in the initial molecular population is the product of a first reagent from the first reagent set, a second reagent from the second reagent set, and optionally a third reagent from the third reagent set, according to a reaction associated with both the first and second reagents, and an optional third reagent set. (b) The first new molecular population, By replacing the first reagent in which the first molecule is formed with another first reagent from the first reagent set associated with the reaction in which the first molecule is formed, and optionally replacing the second reagent in which the first molecule is formed with another second reagent from the second reagent set associated with the reaction in which the first molecule is formed, at least the first molecule in the initial molecular population is modified, thereby forming the first modified molecule, and Select a predetermined number of molecules from an initial molecular population and at least one modified molecule that satisfy one or more fitness scores for inclusion in a first new molecular population. To generate by, (c) A second new molecular population, at least, By replacing the first reagent on which the second molecule is formed with another first reagent from the first reagent set, and optionally replacing the second reagent on which the second molecule is formed with another second reagent from the second reagent set, at least the second molecule in the first new molecular population is modified, thereby forming the second modified molecule, and Select a predetermined number of other molecules from the first new molecular population and at least the second modified molecule that satisfy one or more fitness scores. To generate by, (d) In response to determining that the convergence criteria have not been met, repeat step (b) at least once using a second new molecular population as the new initial molecular population, (e) In response to determining that the convergence criterion has been met, selecting a subset of the second new molecular population as candidates for synthesis and testing; A computer-implemented method is provided that includes.
[0011] In another aspect, a computer-implemented method includes: (a) Obtaining an initial molecular population from a molecular structure database, where each molecule in the initial molecular population is a product of a first reaction between a first reagent selected from a first reagent set related to the first reaction, a second reagent selected from a second reagent set related to the first reaction, and optionally a third reagent from a third reagent set related to the first reaction; (b) Creating a first modified molecule by applying a modification to at least a first molecule in the initial molecular population, where the modification is selected from one or more of: replacing the first reagent from which the first molecule is formed with another first reagent selected from the first reagent set; replacing the second reagent from which the first molecule is formed with another second reagent selected from the second reagent set; and replacing the third reagent from which the first molecule is formed with another third reagent selected from the third reagent set; and generating a first new molecular population by selecting a predetermined number of molecules that meet one or more fitness scores from the initial molecular population and at least the first modified molecule; and selecting a predetermined number of molecules that meet one or more fitness scores from the initial molecular population and at least the first modified molecule; (c) Creating a second modified molecule by applying a modification to at least a second molecule in the second new molecular population, where the modification is selected from one or more of: replacing the first reagent from which the first molecule is formed with another first reagent selected from the first reagent set; replacing the second reagent from which the first molecule is formed with another second reagent selected from the second reagent set; Replacing the second reagent by which the first molecule is formed with another second reagent selected from a second set of reagents, and Replacing the third reagent by which the first molecule is formed with another third reagent selected from a third set of reagents, producing a second modified molecule selected from one or more of the above, and selecting another predetermined number of molecules, from the first new molecular population and at least the second modified molecule, that meet one or more fitness scores TIFF2026509837000002.tif4170 generating by (d) in response to determining that the convergence criterion is not met, repeating step (c) at least once using the second new molecular population as a new initial molecular population, and (e) in response to determining that the convergence criterion is met, selecting a subset of the second new molecular population as candidates for synthesis and testing, A computer-implemented method is provided that includes the above.
[0012] In another aspect, one or more fitness scores include one or more diversity scores.
[0013] In another aspect, a system is provided that includes at least one data processor and at least one memory storing instructions that, when executed by the at least one data processor, cause operations including any of the methods described herein.
[0014] In another aspect, a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations including any of the methods described herein is provided.
[0015] Implementations of this subject matter may include, but are not limited to, systems and methods that include one or more features described herein, and articles that include tangibly embodied machine-readable media capable of causing one or more machines (e.g., computing devices as described elsewhere herein) to perform the operations described herein. Similarly, computer systems that may include one or more processors and one or more memories coupled to one or more processors are also described. Memories that may include computer-readable storage media may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implementations of one or more implementations of this subject matter may be implemented by one or more data processors in a single computing system or in a group of computing systems. Such group of computing systems may be connected, but are not limited to, via one or more connections, including connections via networks (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), via direct connections between one or more of the group of computing systems, etc., and may exchange data and / or instructions or other commands, etc.
[0016] Details of one or more variations of the subject matter described herein are shown in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will become apparent from the description and drawings, as well as from the claims. Certain features of the subject matter currently disclosed are described for illustrative purposes in relation to the design of small molecules, but it should be readily understood that such features are not intended to be limiting. The claims following the disclosure define the scope of the subject matter protected. [Brief explanation of the drawing]
[0017] The accompanying drawings incorporated herein and constituting part of this specification illustrate specific aspects of the subject matter disclosed herein and, together with the description, help to illustrate some of the principles relating to the disclosed implementations. In the drawings,
[0018] [Figure 1] A system diagram is drawn showing an example of a small molecule design system according to several exemplary embodiments.
[0019] [Figure 2] This figure shows examples of chemically and computationally demonstrated two reagent sets for forming a product and the reaction between them, according to several exemplary embodiments.
[0020] [Figure 3A-3B] These are schematic diagrams of two variations of a method for modifying molecules constructed from reagents according to a reaction, based on several exemplary embodiments.
[0021] [Figure 4] This figure shows a single iteration of an example of a generative algorithm for small molecule design, with several exemplary embodiments.
[0022] [Figure 5] This figure shows multiple iterations of an example generative algorithm for small molecule design, in which one or more subsequent molecular populations are generated based on a single initial molecular population, according to several exemplary embodiments.
[0023] [Figure 6] This figure shows multiple iterations of an example generative algorithm for small molecule design, in which one or more subsequent molecular populations are generated based on multiple initial molecular populations, according to several exemplary embodiments.
[0024] [Figure 7] The graphs show the evolution of molecular properties across a continuous population of molecules, according to several exemplary embodiments.
[0025] [Figure 8A] This figure shows an example of the application of the generative small molecule design method described herein, in several exemplary embodiments, in which the generated molecule is evaluated for consistency with one or more known molecules.
[0026] [Figure 8B] This figure shows another example of the application of the generative small molecule design method described herein, in which the generated molecules are evaluated for novelty, according to several exemplary embodiments.
[0027] [Figure 9A] This flowchart shows an example of a process for designing small molecules, according to several exemplary embodiments.
[0028] [Figure 9B] This flowchart shows another embodiment of an exemplary process for designing small molecules, according to several exemplary embodiments.
[0029] [Figure 10] This diagram depicts a block diagram illustrating an example of a computing system in several exemplary embodiments.
[0030] In practical terms, similar reference numbers indicate similar structures, features, or elements. [Modes for carrying out the invention]
[0031] Detailed explanation of disclosure The subject of this specification is an iterative method for generating a set of molecules optimized to have properties that bind to a specific target (e.g., are active against it) and / or make them promising drug candidates. The method described herein addresses the problem of identifying a set of molecules that are diverse, readily available for synthesis (e.g., economical to manufacture), and are promising drug candidates, compared to known methods that generate structurally complex molecules that lack, for example, at least one of these properties, e.g., lack molecular diversity, and / or are chemically impractical to synthesize.
[0032] In some embodiments, the method described herein addresses this problem by iteratively applying a generation algorithm to two or more molecular populations to generate molecules of subsequent populations of the two or more populations, where generating molecules of subsequent populations includes, for example, selection based on a diversity score to maximize diversity between populations. In some embodiments, generating molecules of subsequent populations includes, for example, selection based on another diversity score to maximize diversity within populations. In some embodiments, the diversity score includes a goodness-of-fit score. The terms “diverse” and “dissimilar” are used interchangeably herein. In some embodiments, diverse molecules are structurally and / or chemically dissimilar.
[0033] In some embodiments, generating molecules in a subsequent population involves selection based on a fitness score that has properties likely to make the generated molecules promising drug candidates, for example, optimizing the binding of the generated molecules to a specific target and / or making the generated molecules promising drug candidates. In some embodiments, the fitness score is different from a diversity score used to maximize inter-population and / or intra-population diversity, for example. In some embodiments, generating molecules involves fabricating and / or modifying molecules. In some embodiments, the fitness score determines whether two or more molecules are (1) promising drug candidates, (2) optimally suited to binding to a specific target, and / or (3) likely to make them promising drug candidates. In some embodiments, the fitness score includes a diversity score.
[0034] In some embodiments, generating a subsequent population of molecules involves, for example, chemically reasonable modifications so that the molecules can be readily synthesized. In yet another embodiment, generating a subsequent population of molecules is based on, for example, a reaction and corresponding reagents so that the molecules are readily available for synthesis. In some embodiments, the molecules generated based on a reaction and corresponding reagents consist of two or more synthons (e.g., reagents), each synthon associated with a particular reaction.
[0035] The methods described herein enable a more efficient search for identifying (e.g., generating) better molecular diversity across different molecular populations, wherein the identified molecules have at least equal or greater likelihood of exhibiting desirable qualities determined by one or more fitness scores, and where the identified molecules are readily available for synthesis.
[0036] Systems, methods, and products (including computer program products) are provided for the generative design of small molecules. In some embodiments, the system comprises at least one processor and at least one memory. The at least one memory may contain program code that gives behavior when executed by the at least one processor.
[0037] In some embodiments, a computer program product includes a non-temporary computer-readable medium for storing instructions. Instructions may be executed by at least one data processor.
[0038] In some embodiments, a system for the generative design of small molecules comprises at least one processor and at least one memory. The at least one memory may contain program code that gives behavior when executed by the at least one processor.
[0039] In some embodiments, the computer implementation method is (b) Obtaining an initial molecular population from a molecular structure database, (b) Modifying at least one molecule in the initial molecular population and selecting a predetermined number of molecules from the initial molecular population and at least one modified molecule that satisfy one or more goodness-of-fit scores to generate a first new molecular population, (c) Modifying at least two molecules in the first new molecular population, and generating a second new molecular population by selecting one or more other predetermined numbers of molecules from the first new molecular population and at least the second modified molecules that satisfy one or more goodness-of-fit scores, (d) In response to determining that the convergence criteria have not been met, repeat step (c) at least once using a second new molecular population as the first new molecular population, (e) In response to determining that the convergence criteria have been met, select a second subset of new molecular populations as candidates for synthesis and testing, Includes.
[0040] In some embodiments, generating a first new molecular population of the Method includes (i) calculating one or more goodness-of-fit scores for each molecule in at least a first modified molecule and an initial molecular population, and (ii) forming a first new molecular population by selecting a predetermined number of molecules from the initial molecular population and at least a first modified molecule that have one or more goodness-of-fit scores satisfying one or more goodness-of-fit scores, wherein each instance of generating a second new molecular population includes (i) calculating one or more goodness-of-fit scores for each molecule in at least a second modified molecule and the first new molecular population, and (ii) selecting a predetermined number of molecules from the first new molecular population and at least a second modified molecule that have one or more goodness-of-fit scores satisfying one or more goodness-of-fit scores TIFF2026509837000003.tif4170 This includes forming a second new molecular population by selecting a predetermined number of molecules.
[0041] In some embodiments, modifying the first and second molecules of the Method involves applying one or more modifications selected from the group consisting of: substituting a non-hydrogen atom radical of the molecule with a part selected from a standard set of parts; substituting a hydrogen atom of the molecule with a part selected from a standard set of parts; or replacing a divalent fragment of the molecule with a divalent part selected from a standard set of parts (wherein the radical is a ring system containing 1 to 5 non-hydrogen atoms or 3 to 10 non-hydrogen atoms, the part is a ring system containing 1 to 5 non-hydrogen atoms or 3 to 10 non-hydrogen atoms, and the divalent fragment is a ring system containing 1 to 5 non-hydrogen atoms or 3 to 10 non-hydrogen atoms).
[0042] In some embodiments, one or more fitness scores include a diversity score. In some embodiments, the diversity score indicates the chemical similarity between molecular pairs, and the diversity score penalizes at least one molecule in a molecular pair that is structurally or chemically similar to each other. In some embodiments, the diversity score is calculated from a metric selected from the group consisting of the Tanimoto exponent, cosine coefficient, Dice distance, Euclidean distance, Sitiblock distance, Hamming exponent, and Tversky exponent. In some embodiments, one or more fitness scores include a calculated docking score for the molecules relative to the target binding site.
[0043] In some embodiments, the convergence criterion is selected from: a fixed number of iterations of step (c) of the method, an average goodness-of-fit score satisfying a threshold, calculated for a second new population of molecules obtained from the last example of step (c), and a specified number of molecules in the second new population of molecules obtained from the last example of step (c) having a goodness-of-fit score below the threshold.
[0044] In some embodiments, obtaining an initial molecular population of the method includes clustering corresponding sets of molecules in a molecular structure database into one or more molecular clusters based on a similarity metric, and selecting one or more molecules from each of the one or more molecular clusters that have a fitness score that satisfies one or more initial fitness scores.
[0045] In some embodiments, the similarity metric is the Tanimoto exponent, cosine coefficient, Dice distance, Euclidean distance, Cityblock distance, Hamming exponent, or Tversky exponent, and the similarity metric is based on a property selected from 2D similarity, 3D similarity, and a vector of physicochemical properties.
[0046] In some embodiments, clustering is performed by applying one or more clustering algorithms selected from the following: Butina, centroid, CLink, Gower, McQuitty, SLink, Unweighted Pair Group Method with Arithmetic Mean (UPMGA), Ward, and Jarvis-Patrick.
[0047] In some embodiments, one or more fitness scores for a molecule include one or more calculated values from the following molecular properties: solubility, permeability, selectivity score, efficiency score, toxicity, and physiologically based pharmacokinetic (PBPK) score.
[0048] In some embodiments, the initial molecular population may be formed by randomly selecting one or more molecules from a molecular structure database, selecting one or more molecules from the molecular structure database that have one or more physicochemical properties that meet threshold criteria, selecting one or more molecules that have a specific scaffold, or a combination thereof. It is obtained by a method selected from the following.
[0049] Some embodiments of this method further include (f) synthesizing at least one of the candidate molecules.
[0050] In some embodiments, a second new molecular population is another TIFF2026509837000004.tif4170 is generated by selecting a fixed proportion of molecules from a first new molecular population, as at least a portion of a predetermined number of molecules.
[0051] In some embodiments, one or more modifications to each molecule are identified by determining a set of applicable modifications to the molecule from a list of available modifications and randomly selecting one or more modifications from the set of applicable modifications.
[0052] In some embodiments, the computer implementation method is (a) Obtaining multiple initial molecular populations from molecular structure databases, (b) Generating a plurality of first new molecular populations from each of a plurality of initial molecular populations, wherein each of the plurality of first new molecular populations contains one or more molecules obtained by modifying one or more molecules in the initial molecular population on which the first new population is based, and each of the plurality of first new molecular populations satisfies one or more goodness-of-fit scores, (c) Generating a plurality of second new molecular populations from each of a plurality of first new molecular populations, wherein each of the plurality of second new populations contains one or more molecules obtained by modifying one or more molecules in the first new molecular population on which the second new population is based, and the molecules in each of the plurality of second new molecular populations satisfy one or more goodness-of-fit scores, (d) In response to determining that the convergence criteria are not met, repeat step (c) at least once, such that each of the first new molecular populations is one of the second new molecular populations obtained from the precedents of step (c). (e) In response to determining that the convergence criteria have been met, select a subset of molecules containing at least one molecule from each of several second new molecular populations for synthesis and testing, Includes.
[0053] In some embodiments, one or more goodness-of-fit scores include a diversity score. In some embodiments, the diversity score is based on the structural similarity between pairs of molecules, each containing one molecule selected from each of two molecular populations, and the diversity score penalizes at least one molecule in a pair of molecules that are structurally similar to each other. In some embodiments, the diversity score is calculated from a metric selected from the group consisting of Tanimoto exponents, cosine coefficients, Dice distance, Euclidean distance, Cityblock distance, Hamming exponents, and / or Tverski exponents. In some embodiments, the diversity score is calculated between any pair of molecules that can be formed from molecules in one population and molecules in another population. In some embodiments, for any pair of molecules whose diversity score does not meet a threshold, only one molecule from the pair meets one or more goodness-of-fit scores.
[0054] In some embodiments, generating multiple first new molecular populations of this method involves, for each first new molecular population, (i) calculating one or more goodness-of-fit scores for at least one modified molecule and each molecule in the corresponding initial molecular population, and (ii) having one or more goodness-of-fit scores from the corresponding initial molecular population and at least one modified molecule that satisfy one or more goodness-of-fit scores. TIFF2026509837000005.tif4170 The process includes performing the action of forming a first new molecular population by selecting a predetermined number of molecules, and each instance that generates a plurality of second new molecular populations performs the following for each second new molecular population: (i) calculate one or more goodness-of-fit scores for at least the second modified molecules and each molecule in the corresponding first new molecular population, and (ii) select from the corresponding first new molecular population and at least the second modified molecules one or more goodness-of-fit scores that satisfy one or more goodness-of-fit scores TIFF2026509837000006.tif4170 To form a second new molecular population by selecting a predetermined number of molecules.
[0055] In some embodiments, the computer implementation method is (a) Obtaining an initial molecular population in which each molecule has a molecular structure constructed from a reaction database containing a reaction set, wherein each reaction from the reaction set is associated with two or more reagent sets, and each molecule in the initial molecular population is the product of a first reagent from the first reagent set, a second reagent from the second reagent set, and optionally a third reagent from the third reagent set, according to a reaction associated with both the first and second reagents, and an optional third reagent set. (b) The first new molecular population, By replacing the first reagent in which the first molecule is formed with another first reagent from the first reagent set associated with the reaction in which the first molecule is formed, and optionally replacing the second reagent in which the first molecule is formed with another second reagent from the second reagent set associated with the reaction in which the first molecule is formed, at least the first molecule in the initial molecular population is modified, thereby forming the first modified molecule, and Select a predetermined number of molecules from an initial molecular population and at least one modified molecule that satisfy one or more fitness scores for inclusion in a first new molecular population. To generate by, (c) A second new molecular population, at least, By replacing the first reagent on which the second molecule is formed with another first reagent from the first reagent set, and optionally replacing the second reagent on which the second molecule is formed with another second reagent from the second reagent set, at least the second molecule in the first new molecular population is modified, thereby forming the second modified molecule, and Select a predetermined number of other molecules from the first new molecular population and at least the second modified molecule that satisfy one or more fitness scores. To generate by, (d) In response to determining that the convergence criteria have not been met, repeat step (b) at least once using a second new molecular population as the new initial molecular population, (e) In response to determining that the convergence criteria have been met, select a second subset of new molecular populations as candidates for synthesis and testing, Includes.
[0056] In some embodiments, the computer implementation method is (a) Obtaining an initial molecular population from a molecular structure database, wherein each molecule in the initial molecular population is a product of a first reaction between a first reagent selected from a first reagent set associated with a first reaction, a second reagent selected from a second reagent set associated with a first reaction, and optionally a third reagent from a third reagent set associated with a first reaction. (b) The first new molecular population, The first modified molecule is created by applying a modification to at least one molecule in the initial molecular population, wherein the modification is Replacing the first reagent on which the first molecule is formed with another first reagent selected from the first reagent set, Replacing the second reagent on which the first molecule is formed with another second reagent selected from the second reagent set, and Replacing the third reagent on which the first molecule is formed with another third reagent selected from the third reagent set, To create a first modified molecule selected from one or more of the following: Furthermore Select a predetermined number of molecules from the initial molecular population and at least one modified molecule that satisfy one or more fitness scores. To generate by, (c) A second new molecular group, The method involves creating a second modified molecule by applying a modification to at least a second molecule in a second new molecular population, wherein the modification is Replacing the first reagent on which the first molecule is formed with another first reagent selected from the first reagent set, Replacing the second reagent on which the first molecule is formed with another second reagent selected from the second reagent set, and Replacing the third reagent on which the first molecule is formed with another third reagent selected from the third reagent set, To create a second modified molecule selected from one or more of the following: Furthermore From the first new molecular population and at least the second modified molecule, one or more other molecules satisfying the fitness score TIFF2026509837000007.tif4170 Select a predetermined number of molecules. To generate by, (d) In response to determining that the convergence criteria have not been met, repeat step (c) at least once using a second new molecular population as the new initial molecular population, (e) In response to determining that the convergence criteria have been met, select a second subset of new molecular populations as candidates for synthesis and testing, Includes.
[0057] In some embodiments, generating a first new molecular population of the method includes modifying an additional first molecule in the initial molecular population by selecting a third reagent having a calculated similarity to a first reagent in which the additional first molecule is formed (wherein the third reagent is in an additional set of first reagents associated with the additional reaction, and the additional reaction is different from the reaction in which the additional first molecule is formed); constructing an additional first modified molecule by replacing the first reagent in which the additional first molecule is formed with the third reagent and replacing the second reagent in which the additional first molecule is formed with a fourth reagent selected from an additional set of second reagents associated with the additional reaction (wherein the fourth reagent has a calculated similarity to the second reagent); and selecting a predetermined number of molecules from the initial molecular population, the first modified molecule, and the additional first modified molecule that satisfy one or more goodness-of-fit scores.
[0058] In some embodiments, generating a second new molecular population of the method includes modifying an additional second molecule in the initial molecular population by selecting a third reagent having a calculated similarity to a first reagent in which the additional second molecule is formed (wherein the third reagent is in an additional set of first reagents associated with the additional reaction, and the additional reaction is different from the reaction in which the additional second molecule is formed); constructing an additional second modified molecule by replacing the first reagent in which the additional second molecule is formed with the third reagent and replacing the second reagent in which the additional second molecule is formed with a fourth reagent selected from a fourth additional set of second reagents associated with the additional reaction (wherein the fourth reagent has a calculated similarity to the second reagent); and selecting a predetermined number of molecules from the first new molecular population, the second modified molecule, and the additional second modified molecule that satisfy one or more goodness-of-fit scores.
[0059] In some embodiments, the system includes at least one data processor and at least one memory for storing instructions, and when an instruction is executed by the at least one data processor, it results in an operation including any of the methods described herein.
[0060] In some embodiments, a non-temporary computer-readable medium stores instructions that, when executed by at least one data processor, result in an operation including one of the methods described herein.
[0061] Figure 1 shows a system diagram illustrating an example of a small molecule design system 100 according to several exemplary embodiments. Referring to Figure 1, the small molecule design system 100 may include a design engine 110, a client device 120, and a data store 130. As shown in Figure 1, the design engine 110, the client device 120, and the data store 130 can be communicatively coupled via a network 140. The client device 120 may be a processor-based device including, for example, a workstation, desktop computer, laptop computer, smartphone, tablet computer, wearable device, etc. The data store 130 may be a database such as a molecular structure database, implemented as, for example, a relational database, graph database, in-memory database, non-SQL (NoSQL) database, file system, etc. The network 140 may be a wired network and / or wireless network including, for example, a local area network (LAN), virtual local area network (VLAN), wide area network (WAN), public land mobile network (PLMN), internet, etc.
[0062] Perhaps instead of an indiscriminate search of the vast chemical space of relevant molecules, the design engine is configured to apply a generative algorithm in which one or more new molecular populations are generated based on one or more initial molecular populations, and each new population is evaluated for its properties. In some embodiments, the generative algorithm is a genetic algorithm. In some embodiments, the design engine 110 may be configured to apply a generative algorithm 115 to generate one or more subsequent molecular populations based on one or more initial molecular populations through one or more iterations. In some cases, one or more initial molecular populations and one or more subsequent molecular populations may include small molecules having molecular weights of approximately 100 daltons to approximately 1,000 daltons. Furthermore, in some cases, at least one initial molecular population may include a random selection of molecules from a set of molecules (e.g., a set of all known molecules, such as those available from the Chemical Abstracts Service or a medicinal chemistry database, or a subset of a set of all known molecules pre-filtered for specific drug-like properties).
[0063] With each iteration of the generation algorithm, the design engine may modify one or more molecules from the initial molecular population and a pool of one or more modified molecules before selecting several molecules that exhibit specific desirable properties indicated by one or more fitness scores. In some embodiments, the modification of one or more molecules is a chemically rational modification. Examples of approaches in which molecules are modified in each generation are further described herein. An example of a fitness score is a docking score for a target of interest. In relation to molecular design and modeling, the docking score is a measure of the binding affinity between two molecules after they have undergone a docking process (e.g., shape complementarity analysis, docking simulation, etc.) to predict the preferred orientation of the two molecules when they bind to form a stable complex. Other examples of fitness scores include one or more of the following: diversity score, solubility score, permeability score, selectivity score, efficacy score, toxicity score, and physiological pharmacokinetic (PBPK) score. For example, evaluating the fit based on measures such as docking strength to the target, solubility, permeability, selectivity, efficacy, toxicity, and physiological pharmacokinetics (PBPK) increases the chances of obtaining a suitable molecule. Docking strength to the target can be calculated or estimated by one of several docking methods, algorithms, or programs. Thus, the molecules selected to form a new molecular population may include molecules from an initial molecular population determined to exhibit the desired properties, as well as modified molecules determined to exhibit the desired properties. This selection strategy for forming each new molecular population preferably maximizes the likelihood that the best-performing molecules from each molecular population, including those that may be unmodified, will withstand subsequent iterations of the generation algorithm. Subsequent molecular populations are generated by modifying the best-performing molecules from the previous molecular populations.
[0064] In this way, the design engine performs fit assessment and selection based on complete molecules, rather than molecules constructed by capping a framework or core structure with molecular fragments or simple groups such as methyl or phenyl, and the design engine can ensure the relevance of the fit score to the target of the problem.
[0065] In some embodiments, the design engine can perform multiple iterations of the generation algorithm. In some embodiments, the design engine generates one or more new molecular populations based on a single initial molecular population, while in other embodiments, multiple initial molecular populations can be optimized separately and in parallel with each other. The design engine continues to generate additional new molecular populations until one or more convergence conditions are met, at which point one or more molecules from the new molecular populations can be synthesized and tested, or identified as candidates for further consideration before synthesis and testing.
[0066] One or more convergence conditions may require that the goodness-of-fit scores of a certain predetermined number of molecules in one or more new molecular populations meet a threshold or convergence criterion. For example, a molecule's goodness-of-fit score may show a threshold-less improvement over the goodness-of-fit scores of other molecular populations (e.g., generated based on the same or different initial molecular populations). When goodness-of-fit scores for two or more properties are calculated, the goodness-of-fit scores for each property may be weighted according to an adjustable parameter that reflects the relative importance of those properties. In some embodiments, convergence is determined by the docking score of the molecular population, while other goodness-of-fit scores are used to exclude molecules with one or more less desirable properties.
[0067] In some embodiments, the design engine may impose one or more restrictions on the modifications made to each molecule when generating a new molecular population, further increasing the likelihood that the molecules in the new molecular population will exhibit the same desirable properties as those in the previous molecular population. For example, the modifications made to each molecule may be limited to a fixed number of modifications to ensure that the structure of one generation does not deviate significantly from the structure of the previous generation.
[0068] In some embodiments, the modifications made to each molecule are limited to one or more chemically valid modifications, i.e., modifications that would be acceptable to an organic chemist. An exemplary approach for identifying such sets of modifications is described in Polischchuk, et al., J. Cheminform., 12:28, 1-18, (2020), which is incorporated herein by reference.
[0069] A molecule can typically be defined as comprising one or more fragments, each of which is an atom or a group of atoms bonded to one another and located at defined bonding points between the fragment and the rest of the molecule. Various definitions of fragments are consistent with the methods herein. Thus, a chemically reasonable modification of a molecule involves replacing one fragment of a molecule thus defined with another fragment. For example, the concept of isosterism can be used to limit modifications that replace one fragment with another fragment of similar size. Thus, a pyridyl group can be a chemically reasonable modification of a phenyl group because the two are isosteric. Furthermore, properties of fragments such as polarity can be utilized. Thus, a chemically reasonable modification of a hydroxyl group could be an amino group or a thio group. In examples where a particular part of a molecule is designated as a core or scaffold, modifications made to the molecule may exclude changes to the core or scaffold itself and thereby be limited to changes to the fragments bonded to the scaffold. For example, molecules containing condensed heterocyclic systems such as indole can be modified so that only the substituents on the indole moiety are identified and modified as fragments, or so that the indole moiety itself is maintained in all molecules based on the original molecule.
[0070] In some embodiments, the types of fragments to be modified are selected from, and are not limited to, a list of functional groups well known to chemists, but can be illustrated by the following examples: Single-atom fragments contain halogen atoms; two-atom fragments bonded via one of two atoms include groups such as hydroxyl, cyano, and thio; three-atom fragments bonded via one of three atoms include amino, nitro, and carbenyl; and four-atom fragments bonded via a specific atom include carboxylic acids, sulfonates, and methyl. Within each category, a list of fragments can be created and made available as a standard set from which fragments can be arbitrarily selected or chosen according to several criteria or standards such as size or polarity. In instances where the list of known functional groups proves to be too restrictive, broader definitions can be used: for example, a fragment can be defined as identifying one or more bond points between the fragment and the rest of the molecule, with each of the one or more bond points containing an atom and one or more other atoms within a specific radius of that atom. The radius can be defined geometrically, such as the spatial penetration threshold distance if a 3D model of the fragment is available, or by several bonds that separate the bond points from other atoms in the fragment. Furthermore, the fragment can be defined as containing, for example, a ring or ring system such as cyclopropyl, phenyl, or indolyl.
[0071] In some examples, chemically reasonable modifications performed on a molecule may be selected from a group including substituting a non-hydrogen atom radical of the molecule with a part selected from a standard list of parts, substituting a hydrogen atom of the molecule with a part selected from a standard list of parts, and replacing a divalent fragment of the molecule with a divalent part selected from a standard list of parts. In this regard, a radical may be a ring system containing 1 to 5 non-hydrogen atoms or 3 to 10 non-hydrogen atoms. A part may contain 1 to 5 non-hydrogen atoms or is a ring system containing 3 to 10 non-hydrogen atoms. A divalent fragment may contain 1 to 5 non-hydrogen atoms or is a ring system containing 3 to 10 non-hydrogen atoms. The term radical is understood to mean a fragment that contains an unpaired electron or “dangling bond” at its next bonding site, and correspondingly, the site of the molecule to which it intends to bond contains a “dangling bond” at the point where the fragment intends to bond. The use of the term radical herein is for convenience, as the fragment itself does not represent an isolated molecule.
[0072] In some embodiments, the methods for constructing and modifying molecules in each generation are based on reactions and corresponding reagents rather than fragments. The reactions correspond to well-understood chemical transformations (e.g., an alcohol reacting with a carboxylic acid to form an ester) where each reagent allows for variation. For example, in an esterification reaction, the alcohol can be selected from alkyl alcohols (methanol, ethanol, propanol, etc.) and cycloalkyl alcohols (cyclopropanol, cyclobutanol, etc.). Two-component reactions are based on two sets of reagents (e.g., an alcohol and an acid) that can be combined in a specific manner. The reaction may also be a three-component reaction, in which case three sets of reagents are available for selecting each component. The methods herein are generalizable to four-component, and even five-component reactions, even though such schemes are rare in actual chemistry.
[0073] Figure 2 illustrates the formation of molecular products by combining two reagents, both chemically and computationally encoded. The upper panel of Figure 2 shows the amide formation reaction between benzoic acid and methylamine. In this scheme, reagent 1 (benzoic acid) is selected from a list of carboxylic acids, and reagent 2 (methylamine) is selected from a second list of amines. In the actual reaction, the hydrogen atom bonded to the amino group is "lost" as it combines with the hydroxyl group in the acid to form H2O. When constructing amide molecules computationally, the formation of water, a "byproduct," is not important. Computationally, such reagents are represented as "synthons" where the connection point between the reagent and its corresponding reagent for a given reaction is identified, rather than as fragments containing one or more atoms. In the lower panel of Figure 2, the amine and acid reagents are shown as computationally stored "synthons." For convenience, computationally, the H and OH groups from the respective amino and carboxylate moieties are deleted from the reagent molecules when forming their representative synthons. The reagents are stored as "synthons" in a terminal-cleavage form, and their bond sites are identified as being able to form bonds with their respective bond sites on other synthons. Thus, each individual synthon is an intramolecular component representing the starting reagent in the synthesis of that molecule.
[0074] Therefore, each molecule in the initial molecular population can be constructed as the product of a reaction between a first reagent selected from a first set of reagents related to the reaction and a second reagent selected from a second set of reagents related to the reaction, and optionally (in the case of a three-component reaction) as the product of a reaction between a first reagent selected from a third set of reagents related to the reaction. The population typically includes molecules constructed from several different reactions, each of which has its own set of reagents from which molecules can be constructed. For example, an amine and a carboxylic acid react to form an amide. In the case of an amide formation reaction, the molecule can then be formed by selecting from a first list of amines as the first reagent and from a second list of carboxylic acids as the second reagent. The population may include molecules formed by amide formation reactions in addition to molecules formed from esterification reactions.
[0075] Therefore, modifications made to each molecule when generating a new molecular population may include changing at least one of the first reagent, the second reagent, and, if applicable, the third reagent, for a given reaction. If the modification includes a change to the first reagent, the modification may be limited to replacing the first reagent with another first reagent selected from the first set of reagents associated with the reaction. If the modification further includes a change to the second reagent, the modification to the second reagent may be limited to replacing the second reagent with another second reagent selected from the second set of reagents associated with the reaction. When constructing an initial molecular population, the reactions may be randomly selected, for example, from a standard set of reactions. Furthermore, in some cases, each of the first and second reagents associated with a given reaction may also be randomly selected, for example, from the corresponding sets of first and second reagents.
[0076] A general description of the modified molecules formed by replacing one or more reagents is shown in Figure 3A. (Notation: sn(X)) i This means that a molecule is formed by reaction X. For reaction X that requires two or more reagents, Synton sn(X) is1(X)4 represents the nth reagent (n=1, 2, 3, etc.) and the i-th member of the nth reagent set. Thus, s1(X)4 means the fourth synthon from the first set of synthons for reaction X (first reagent). On the left side of Figure 3A are four representative molecules in the population. All molecules are formed from two components s1 and s2; the first two molecules are formed from reaction P, and the third and fourth are formed from reactions Q and R, respectively. In forming the next generation population, the third and fourth molecules remain unchanged (due to their overall calculated goodness of fit, as will be further explained herein). Both the first and second molecules are transformed, although the underlying reaction (P) on which both are formed remains unchanged. In the first molecule, only the second reagent is changed (to another reagent selected from the list of second reagents), and in the second molecule, only the first reagent is changed (to another reagent selected from the list of first reagents). Note that when proceeding from s1(P)3s2(P)4 to s1(P)3s2(P)2, only the second reagent changes, but reagent s2(P)2 must be selected from the second reagent set that allows reaction P to be carried out.
[0077] In further embodiments, a general description of modifying molecules by replacing both the reaction and one or more reagents is shown in Figure 3B, with the same notation as in Figure 3A. In Figure 3B, the population on the left is the start, the same as the start population in Figure 3A. In the embodiment of Figure 3B, the first, third, and fourth molecules proceed to the next generation population as in Figure 3A. For a second molecule formed from the first reagent s1(P)4 and the second reagent s2(P)2 according to the first reaction (P), the modification has a minimum threshold similarity to the first reagent s1(T)4, but a new first reagent s1(T) is used in the second reaction (T) which is different from the first reaction (P). sim This includes identifying the second reaction (T), which depends on the second set of the first reagent and the second set of the second reagent from (P). Thus, here, the new first reagent s1(T) sim=A modified molecule is constructed using a second reaction between the first reaction and a new second reagent (s2(T)2 in Figure 3B) selected from a second reagent set of second reagents related to reaction (T). In other embodiments not shown, the new second reagent may be selected to have at least minimal similarity to the second reagent used to form the initial molecule. Figure 3B then shows that the next generation of molecules includes at least some molecules that pass through the unmodified stage, at least some molecules that are modified according to the embodiment in Figure 3A, and at least some molecules that are modified according to the way in which the second molecule is modified. It will be further understood that the new population may further include other molecules constructed from reaction (T), as well as yet other molecules constructed from other reactions (U), (V), etc. The similarity between the reagents used in the embodiment in Figure 3B may follow the Tanimoto coefficient, based on atomic composition, or based on shape similarity, or based on any other suitable method described elsewhere in this specification.
[0078] In some embodiments, the application of the generation algorithm 115 includes multiple iterations in which a molecular population is generated. In a first such iteration, the design engine 110 modifies at least a first molecule from the initial molecular population. Furthermore, a single iteration during the application of the generation algorithm 115 includes the design engine 110 selecting a predetermined number of molecules from the initial molecular population and at least the first modified molecules to be included in a new molecular population based on one or more fitness scores. Examples of fitness scores include diversity scores, docking scores to targets, solubility scores, permeability scores, selectivity scores, efficacy scores, toxicity scores, and physiological pharmacokinetic (PBPK) scores. Thus, the molecules selected to form a new molecular population may exhibit certain desirable properties, as indicated by the molecular fitness scores. In doing so, the design engine 110 can maximize the chances that the best-performing molecules from each molecular population survive to subsequent iterations of the generation algorithm 115. The subsequent molecular population thus formed may include molecules generated by modifying the best-performing molecules from the previous molecular populations and therefore have a higher chance of exhibiting the same desirable properties.
[0079] The design engine 110 may cause the generation algorithm 115 to perform one or more iterations to generate one or more new molecular populations based on a single initial molecular population. Alternatively and / or additionally, the design engine 110 may cause the generation algorithm 115 to perform one or more iterations to generate one or more new molecular populations based on multiple different initial molecular populations. In some cases, the design engine 110 may cause the generation algorithm to continue generating additional new molecular populations until one or more conditions are met. Another exemplary condition may include the design engine 110 having generated a predetermined number of molecules whose goodness-of-fit scores satisfy one or more thresholds. Other examples of conditions include the goodness-of-fit scores of new molecular populations that satisfy a threshold or convergence criterion, where the goodness-of-fit scores of the new molecular populations show improvement over the goodness-of-fit scores of other molecular populations (e.g., generated based on the same or different initial molecular populations). Such improvement may be quantified as falling below a predetermined value (threshold) of a certain amount.
[0080] When one or more conditions are met, at least one molecule from one or more new molecular populations may be synthesized and tested, or identified as a candidate for synthesis and testing. In some cases, the molecules synthesized and tested, or identified as candidates for synthesis and testing, may be determined as a result of the design engine 110 applying one or more filters or threshold criteria. Furthermore, in some examples, the design engine 110 may generate a user interface 125 that provides a visual representation of the molecules identified as candidates for synthesis and testing for display on a client device 120.
[0081] To further illustrate, Figure 4 shows a single iteration 200 of an example of a generative algorithm 115 for small molecule design. As shown in Figure 4, a single iteration 200 of the generative algorithm 115 may include starting from an initial molecular population 210 (shown as a circle numbered "1" in Figure 4). In some cases, the initial molecular population 210 may be obtained from the data store 130 by random selection, selection of one or more molecules exhibiting one or more physicochemical properties that meet a threshold criterion, selection of one or more molecules having a particular scaffold, etc. Typically, the initial molecular population consists of N molecules selected by the user. N may be a number such as 500, 1,000, 2,000, 5,000, or more, depending on the available computing resources. The behavior and operation of the algorithms herein are independent of the size of the initial population.
[0082] The generation algorithm includes modifying at least one first modified molecule from the initial population 210 to generate at least one first modified molecule (shown as a circle numbered "2" in Figure 4). As described elsewhere in this specification, a given molecule may be susceptible to two or more modifications, and as a result, two or more different modified molecules may be produced from the step of applying a single modification to that molecule.
[0083] In some examples, a portion of the first molecule from the initial molecular population 210 may be designated as the core of the first molecule. Such a core is typically a scaffold or framework to which various fragments are joined together. For example, the core of the first molecule may be associated with a specific desirable property, in which case modifications made to the first molecule should avoid changing the core of the first molecule. Thus, when modifying the first molecule, the design engine 110 can avoid making changes to the core of the first molecule. By avoiding changes to the core, it is possible to increase the likelihood that the desired property of the first molecule will be preserved when the first molecule is modified.
[0084] In some exemplary embodiments, each molecule from the initial molecular population 210 may be the product of a reaction between a first reagent selected from a first set of reagents associated with a first reaction and a second reagent selected from a second set of reagents associated with a first reaction. In some cases, the first and second reagents may be represented as synths or individual components selected from corresponding sets of synths associated with the reaction. Thus, in order to modify the first molecule from the initial molecular population 210, in one embodiment, the design engine 110 may modify at least one of the first and second reagents that form the first molecule. For example, if the modification involves modifying the first reagent, the modification may be limited to replacing the first reagent in the first molecule with another first reagent selected from the first set of reagents associated with the reaction. If the modification involves further modifying the second reagent, the modification may be limited to replacing the second reagent in the first molecule with another second reagent selected from the second set of reagents associated with the reaction.
[0085] It should be understood that any one molecule in the initial molecular population can undergo two or more modifications, and therefore can result in two or more different modified molecules. For example, a given molecule can be modified in two or more different ways, corresponding to two or more chemically valid modifications (such as changes to two or more different fragments). In another example, since either of the two reagents can be substituted independently, and each reagent can be separately replaced by two or more reagents from the reagent set, a molecule constructed from two reagents can be modified two or more times according to a particular reaction (thus producing two or more modified molecules).
[0086] It should also be understood that multiple molecules in the initial molecular population 210 typically undergo one or more modifications of themselves.
[0087] Next, the initial molecular population 210 and at least the first modified molecules undergo fitness evaluation and selection to form a first new molecular population 215. For example, in some cases, the design engine 110 may select a predetermined number of molecules from the initial molecular population 210 and at least the first modified molecules based on at least a first set of fitness scores associated with each molecule in the initial molecular population 210 and at least the first modified molecule. For example, a predetermined number of molecules may be selected to be included in the first new molecular population 215, each molecule having a fitness score that satisfies one or more thresholds. Alternatively, to generate the first new molecular population 215, the design engine 110 may select a predetermined number of molecules having the highest fitness scores. In some cases, a predetermined number of molecules may be selected by comparing the fitness scores of one or more randomly selected molecular pairs from at least the initial molecular population 210 and at least the first modified molecules, and including the molecule with the higher fitness score from each molecular pair in the first new molecular population 215. In some cases, a predetermined number of molecules selected to be included in the first new molecular population 215 may include a fixed proportion of molecules from the first initial molecular population 210.
[0088] The first new molecular population can be fixed to have the same number of molecules N as the initial molecular population. This ensures that the overall molecular population under consideration does not decrease, but does not expand to the point of reducing the overall efficiency of the algorithm.
[0089] Since the initial molecular population 210 and at least the first modified molecules undergo fitness evaluation as part of the selection process, the resulting first new molecular population 215 is preferably comprised of molecules exhibiting specific desirable properties. For example, the fitness scores of molecules in the first new population may include one or more of the following: diversity score, docking score to target, solubility score, permeability score, selectivity score, efficacy score, toxicity score, and physiological pharmacokinetic (PBPK) score. Thus, molecules selected for inclusion in the first new molecular population 215 may exhibit specific desirable properties with respect to diversity, target binding affinity, solubility, permeability, selectivity, efficacy, toxicity, and physiological pharmacokinetic (PBPK).
[0090] If certain molecules from the initial molecular population 210 already exhibit desirable properties, those molecules may survive the selection process and be included in the first new molecular population 215 to undergo one or more additional iterations of the generation algorithm 115. For example, an example of the first new molecular population 215 shown in Figure 4 includes molecules from the initial molecular population 210 (shown as circles marked with the number "1") as well as molecules generated by modifying molecules from the initial molecular population 210 (shown as circles marked with the number "2" in Figure 4).
[0091] In some embodiments, the design engine 110 may impose one or more diversity criteria (e.g., diversity scores) to ensure that the composition of the first new molecular population 215 is not dominated by multiple molecules having similar structures to one another, thereby optimizing the exploration of the chemical space. For example, the exploration of the chemical space may be increased by maximizing the intrapopulation diversity within the first new molecular population 215. As used herein, the term “intrapopulation diversity” may refer to the dissimilarity (e.g., structural and / or chemical dissimilarity) between molecules contained in the molecular population. Thus, to maximize intrapopulation diversity, the design engine 110 may cluster the initial molecular population 210 and at least the first modified molecules into one or more similar molecular clusters (e.g., structurally and / or chemically similar molecules). In some cases, the design engine 110 can apply a recognized clustering algorithm (e.g., Butina, centroid, Clink, Gower, McQuitty, Slink, unweighted coupling method (UPMGA), Ward, Jarvis-Patrick, etc.) to cluster molecules based on their diversity score (e.g., Tanimoto exponent, Dice exponent, cosine coefficient, Sogel distance, Euclidean distance, Cityblock distance, Hamming exponent, Tversky exponent, etc.). Thus, a first molecule in a first cluster may be similar to a second molecule in the first cluster, but dissimilar to a third molecule from the second cluster. Therefore, having generated one or more similar molecular clusters, the design engine 110 may select one or more molecules from each similar molecular cluster that have the best goodness-of-fit score (e.g., as calculated elsewhere in this specification) to include in a first new molecular population 215. For example, one or more molecules selected to be included in the first novel molecular population 215 may be associated with one or more of the best docking scores, solubility scores, permeability scores, selectivity scores, efficacy scores, toxicity scores, and physiological pharmacokinetic (PBPK) scores for a target. Thus, the resulting first novel molecular population 215 may contain diverse molecules that exhibit specific desirable properties relative to one another.
[0092] Referring again to Figure 4, in some exemplary embodiments, the design engine 110 may increase the likelihood that molecules in the first new molecular population 215 will exhibit properties comparable to those of molecules in the initial molecular population 210 by imposing one or more restrictions on the types of modifications made to any molecule in the initial molecular population 210. For example, the design engine 110 may restrict the modifications made to molecules in the initial molecular population 210 to those selected from a set of chemically valid modifications, as further described elsewhere in this specification.
[0093] Furthermore, in order to make the number of modified molecules a practical number, the design engine 110 may limit the number of modifications performed on molecules in the initial molecular population 210 to a threshold number of modifications, such as 1, 2, or 3.
[0094] In a preferred embodiment, the design engine 110 performs multiple iterations of the generation algorithm 115 to generate multiple successive new molecular populations, starting from a single initial molecular population. Thus, the new molecular populations generated by performing one iteration of the generation algorithm 115 are used as the starting population of molecules for subsequent iterations of the generation algorithm 115. That is, each successive iteration of the generation algorithm 115 is performed using the molecular populations corresponding to the new molecular populations generated by the previous iteration of the generation algorithm 115.
[0095] To further illustrate, Figure 5 shows an example of a generative algorithm 115 for small molecule design where multiple iterations are performed. Referring to Figure 5, where iteration 200 is reproduced from Figure 4, the design engine 110 performs two or more additional iterations of the generative algorithm 115 following the execution of iteration 200. Iteration 200 of the generative algorithm 115 is performed based on the initial molecular population 210, while subsequent iterations 300 of the generative algorithm 115 are performed based on a first new molecular population 215. When iteration 300 is first performed, the first new molecular population 215 is generated by iteration 200. In subsequent iterations 227, the second new population becomes the first new population to which modifications and goodness-of-fit evaluations are applied.
[0096] Referring again to Figure 5, iteration 300 of the generation algorithm 115 includes a design engine 110 that modifies at least one molecule from the first new molecular population 215 to generate one or more second modified molecules (shown as circles numbered "3" in Figure 5). (As shown in Figure 4, the first new molecular population 215 may include molecules from the initial molecular population 210, as well as molecules formed by modifying molecules from the first initial molecular population 210; therefore, modifications may be applied to any such molecule, whether it originates from the initial population or arises as a modified molecule, and both may be selected for inclusion in the first new population.)
[0097] To generate a second new molecular population 225, the design engine 110 selects a number of molecules from the first new molecular population 215 and at least one second modified molecule based on a second set of fitness scores associated with each molecule in the first new molecular population 215 and at least one second modified molecule. For example, if a molecule has a fitness score that satisfies one or more thresholds, the molecule may be selected to be included in the second new molecular population 225. Alternatively, the design engine 110 may select several molecules with the highest fitness scores to be included in the second new molecular population 225. For example, in some cases, the design engine 110 may select a number of molecules by at least comparing the fitness scores of one or more randomly selected molecular pairs from the first new molecular population 215 and at least one second modified molecule, and including the molecule with the higher fitness score from each molecular pair in the second new molecular population 225. In some cases, a predetermined number of molecules selected to be included in the second new molecular population 225 may include a fixed proportion of molecules derived from the first new molecular population 215.
[0098] As shown in Figure 5, after reaching certain convergence criteria, the design engine 110 further selects one or more molecules from a second new population 225, for example by applying one or more filters and threshold criteria, to form a subset 230 of the first molecules. In some cases, the first subset 230 of molecules is identified as a candidate for synthesis and testing.
[0099] In some exemplary embodiments, to further maximize the exploration of a vast chemical space, the design engine may also impose one or more diversity criteria (e.g., diversity scores) when generating one or more new molecular populations derived from the same or different initial molecular populations. For example, to ensure diversity between successive new molecular populations generated based on the same initial molecular population, a fitness score that determines whether a molecule is selected to be included in a subset of molecules may include a diversity score (e.g., Tanimoto exponent, Dice exponent, cosine coefficient, Sogel distance, Euclidean distance, Cityblock distance, Hamming exponent, Jackard exponent, or Tversky exponent) that indicates the structural and / or chemical similarity between the molecule and each molecule in the initial molecular population. Accordingly, in some embodiments, the design engine 110 may impose one or more diversity criteria (e.g., diversity scores) when generating the second new molecular population 225 in order to maximize the intrapopulation diversity within the second new molecular population 225, as well as the interpopulation diversity between the second new molecular population 225 and one or more previous molecular populations such as the initial molecular population 210 and the first new molecular population 215. To maximize the intrapopulation diversity within the second new molecular population 225, the design engine 110 may cluster the first new molecular population 215 and at least the second modified molecules into one or more similar molecular clusters (e.g., structurally and / or chemically similar molecules), for example, by applying a clustering algorithm further described herein. The design engine 110 can then select one or more molecules from each similar molecular cluster that have the best goodness-of-fit score for inclusion in the second new molecular population 225.
[0100] As used herein, the term “inter-population diversity” may refer to the dissimilarity (e.g., structural and / or chemical dissimilarity) between molecules contained in different molecular populations. Therefore, to maximize inter-population diversity, the design engine 110 may select a second new molecular population 225 based on the diversity score of each molecule contained in the first new molecular population 215 and at least the second modified molecule. In some examples, the molecular diversity score may be a metric such as the Tanimoto exponent, Dice exponent, cosine coefficient, Sogel distance, Euclidean distance, Cityblock distance, Hamming exponent, or Tversky exponent, indicating the similarity (e.g., structural and / or chemical similarity) between the molecules in the first new molecular population 215 and each molecule and at least the second modified molecule. By imposing this diversity criterion (e.g., diversity score), the design engine 110 can ensure that the most different (e.g., structurally and / or chemically dissimilar) molecules are selected to be included in the second new molecular population 215. This is advantageous because, as will be described in more detail below, the design engine 110 can also impose a diversity criterion (e.g., a diversity score) to maximize inter-population diversity between a second new molecular population 225 and molecular populations generated by other iterations of the generation algorithm 115 performed using one or more different initial molecular populations.
[0101] In another embodiment, the design engine 110 performs multiple iterations of the generation algorithm 115 to generate multiple subsequent molecular populations from multiple distinct initial molecular populations. Figure 6 shows four iterations of the generation algorithm 115, two of which are performed on two distinct initial molecular populations of molecules. Subsequent iterations of the generation algorithm 115, applied to a second new population 225 and a fourth new population 425, generating subsequent new populations for each of populations 225 and 425, are indicated by curved arrows pointing from left to right. In some embodiments not shown herein, the design engine 110 performs multiple iterations of the generation algorithm 115 on three or more diverse initial molecular populations to generate three or more new molecular populations having specific properties. For example, Figure 6 shows that in addition to performing at least one iteration of the generation algorithm 115 based on a first initial molecular population 210 (e.g., iteration 200 and subsequent iterations 300), the design engine 110 may perform one or more iterations of the generation algorithm 115 based on a second initial molecular population 410 different from the first initial molecular population 210. The second initial population 410 may be obtained, for example, from a data store 130. In some embodiments, one or more molecules in the second initial population 410 are modified as described herein, and one or more modified or unmodified molecules are selected based on a diversity score relative to the modified or unmodified molecules of the first initial population 210 (indicated by a vertical arrow between these two populations). In some embodiments, modified molecules in the second initial population 410 are selected only if their diversity score exceeds a certain threshold. In iteration 400 of the generation algorithm 115, for example, the design engine 110 may modify one or more molecules from the second initial molecular population 410 and one or more modified molecules before selecting a predetermined number of molecules based on a third set of fitness scores associated with each molecule in the second initial molecular population 410 and one or more modified molecules. In doing so, the design engine 110 generates a third new molecular population 415.In the example shown in Figure 6, the design engine 110 continues to perform another iteration 450 of the generation algorithm 115 based on the third new molecular population 415. For example, an iteration 450 of the generation algorithm 115 may include the design engine 110 modifying one or more molecules in the third new molecular population 415 before selecting another predetermined number of molecules from the third new molecular population 415 and one or more modified molecules that satisfy one or more goodness-of-fit criteria (e.g., goodness-of-fit scores) for inclusion in a fourth new molecular population 425.
[0102] Furthermore, in order to ensure diversity between subsequent molecular populations generated from different initial molecular populations, the goodness-of-fit score used to determine whether a molecule is selected to be included in a subset of molecules may include another diversity score (e.g., Tanimoto exponent, Dice exponent, cosine coefficient, Sogel distance, Euclidean distance, Cityblock distance, Hamming exponent, Jackard exponent, Tversky exponent, etc.) that indicates the chemical and / or structural similarity between the molecule and each molecule in different initial molecular populations and / or molecular populations generated by iterations of the generation algorithm 115.
[0103] Therefore, in some exemplary embodiments, the design engine 110 may impose one or more diversity criteria (e.g., diversity scores) to maximize inter-population diversity between the third new molecular population 415 and the molecular populations generated by other iterations of the generation algorithm 115 performed using one or more different initial molecular populations. In the example shown in Figure 6, the third new molecular population 415 may be selected from the second initial molecular population 410 and one or more molecules generated by modifying molecules from the second initial molecular population 410 based on diversity scores associated with each molecule. In some embodiments, modified or unmodified (i.e., initial) molecules in the third new population 415 are selected only if their diversity scores exceed a certain threshold. For example, modifications that are completely different from those applied in iteration 200 to generate modified molecules in the initial population 210 are more likely to result in selected molecules for the third new population 415 because the molecular diversity scores for the modified molecules in the initial population 210 exceed a certain threshold. This diversity score may be a metric indicating the similarity (e.g., structural and / or chemical similarity) between two or more molecules, such as the Tanimoto exponent, Dice exponent, cosine coefficient, Sogel distance, Euclidean distance, Sitiblock distance, Hamming exponent, or Tversky exponent. Similarly, the design engine 110 may select a fourth new molecular population 425 based on a second diversity score relating to molecules in a third new molecular population 415 and molecules generated by modifying one or more molecules in the third new molecular population 415. This second diversity score may also be a metric indicating the similarity (e.g., structural and / or chemical similarity) between two or more molecules, such as the Tanimoto exponent, Dice exponent, cosine coefficient, Sogel distance, Euclidean distance, Sitiblock distance, Hamming exponent, or Tversky exponent.
[0104] As described above, the design engine 110 can perform one or more iterations of the generation algorithm 115 to generate one or more subsequent molecular populations, such as a first new molecular population 215, a second new molecular population 225, a third new molecular population 415, and a fourth new molecular population 425. By generating each subsequent molecular population to contain molecules with satisfactory goodness-of-fit scores, it is ensured that the same desirable properties are preserved and propagated through the molecules of each subsequent generation. On the other hand, by penalizing molecules that are too similar (e.g., having a diversity score below a certain threshold), and thereby maximizing diversity within and / or between populations, it is ensured that the molecules included in each subsequent molecular population are novel and diverse (e.g., structurally and / or chemically dissimilar) to each other and to molecules in other molecular populations. As shown in Figure 6, in some cases, the design engine 110 can select molecules of a second subset 235 from the fourth new molecular population 425 as candidates for synthesis and testing.
[0105] Maximizing the exploration of chemical space (or a particular part thereof) can also be achieved, for example, by imposing one or more diversity criteria (e.g., diversity scores) to maximize the intra- and inter-population diversity of molecules generated by the generation algorithm 115 over subsequent generations, thereby further ensuring that the molecules generated by the generation algorithm 115 are novel compared to previously known molecules. [Examples]
[0106] To further illustrate, Figure 7 shows a graph illustrating the evolution of each of the three molecular properties across approximately 20 molecular populations evolved by performing multiple iterations of the optimized generative algorithm 115 for the same target. In this optimization scheme, the scoring function is a weighted combination of docking scores and diversity scores. According to the diversity score, the similarity of molecular pairs is scored as a penalty.
[0107] The three panels in Figure 7 each show the progression of a particular molecular property as subsequent generations of a population of 20 molecules are created and tested for goodness of fit by calculations of various properties. In a given panel, each row represents a particular run of the algorithm as subsequent generations are created starting from the initial population, and the particular value in each generation represents the average of that property for molecules in the population that have the best value of that property. As shown in Graph 710 of Figure 7, each subsequent generation of molecules produced by the application of the generation algorithm 115 exhibits a continuously improving target binding affinity, as indicated by the decrease in the average docking score of the molecules. (The docking score is the log of the binding energy calculated at any energy scale.) 10 This is expressed as follows: The more negative the bond energy, the better the molecule binds to the target. However, although docking improves overall, each subsequent generation of the molecule shows significant structural variation, as indicated by the average molecular weight of molecules in a given population in panel 720 and the diversity score of the population in graph 730.
[0108] The main advantage of the algorithm is that the populations are inherently diverse. Most populations have a diversity score greater than 0.5 when compared to other populations, meaning that the molecules are structurally diverse from other molecules in other populations. In panel 730, each line represents the average distance of the molecules in the population from all other populations.
[0109] A diversity score of 0.2–0.3 indicates a significant lack of diversity. Most populations have a diversity in the range of 0.7–0.8, representing a very high level of diversity. Variation fluctuations occur during algorithm execution because generative algorithms are used that do not guarantee smooth convergence. Furthermore, the convergence criterion is primarily governed by the characteristics of individual molecules (such as docking scores). Therefore, there is also competition between maintaining low (good) docking scores and maintaining high diversity. This competition can lead to considerable variation in diversity, as a significant improvement in one category may lead to disruption in the performance of the other category.
[0110] Examples - Positive Controls The performance of the generation algorithm 115 can be further evaluated against various benchmarks. Figure 8A shows a certain benchmark for evaluating whether a molecule (e.g., the first molecule 805) generated by the generation algorithm 115 (starting with a random selection of a starting molecule) is structurally similar to a known molecule (e.g., molecule 800) that has already been established to have a desirable property (e.g., an inhibitor that can bind to the same target as the first molecule and block its activity). Molecules 800 and 805 shown in Figure 8A are very similar to each other.
[0111] Examples Another exemplary benchmark quantifies the novelty of molecules generated by generation algorithm 115. This is done by calculating a similarity metric between 20 successor generations of the molecule and a specific known nearest inhibitor of the target in question from ChEMBL (in this case, ROCK1 kinase). Such similarity metrics may be, for example, the Tanimoto exponent, the Dice exponent, the cosine coefficient, the Sogel distance, the Euclidean distance, the Cityblock distance, the Hamming exponent, the Tversky exponent, etc.
[0112] Figure 8B shows the distribution of the nearest neighbor similarity between the molecules in each of the 20 groups and the known molecules. Each bar in the histogram is the number of groups for which the average similarity (w.r.t. a particular group, the "known" inhibitors) is within a particular range of similarity (0.200 - 0.299; 0.300 - 0.399, etc.). The solid line superimposed on the figure is merely for guiding the eye. Most groups have an average overall similarity that falls within the range of Tanimoto of 0.2 - 0.4, suggesting low similarity with existing inhibitors (i.e., a higher likelihood of novelty). However, there are 1 - 2 groups in the range of 0.5 - 0.7, suggesting higher similarity with known inhibitors. The summary from this plot is that the algorithms herein can generate both molecules that are structurally different from known inhibitors as well as molecules that are very similar to known inhibitors. That is, the molecules generated by the generation algorithm 115 have a high likelihood of being novel and exhibiting the same desirable properties as previously known molecules (such as inhibitors for the ROCK1 kinase identified in ChemBL).
[0113] Example Another important benchmark is to evaluate whether the application of the generation algorithm 115 appropriately explores the chemical space of drug - like molecules. Since the algorithms herein do not comprehensively sample all possible compounds within a given space, the inventors want to know how well they perform compared to an exhaustive enumeration and docking of all possible compounds. Thus, the goal of this example is to compare the results from the algorithms herein with the results obtained from docking all possible compounds into a particular space. Such a comparison is not possible with a database of approximately 10 10 available molecules, so a subspace where docking of all molecules is possible is selected.
[0114] Therefore, in this example, a much larger but small subset of the available chemical space was evaluated. In fact, this subset was approximately 10 10Approximately 10 molecules were randomly selected from a database of available molecules (other chemical databases would suffice, but in this case, those found in the Really Accessible (REAL) chemical space (see the internet site enamine.net / compound-collections / real-compounds / real-space-navigator)). 6 It contained 10 molecules. Each of these molecules in a randomly selected subset was docked to a target. The results of these dockings are called "ground truth" in terms of the molecular scoring function (in this case, the docking score). 6 We identified five top-docking molecules from a list ranked by molecular type.
[0115] For comparison, the starting point for applying the generation algorithms described herein is 10 6 These are random compounds selected from a subspace of molecules. The reactions used to form these molecules are used as a starting point in the application of the algorithm, and as a result, further molecules are formed by changing the reagents that can react according to those reactions, and / or by changing the reactions (and the corresponding reagents) themselves. After iteration until convergence, the inventors show that the algorithm generates and recovers four molecules from the top five molecules obtained from the random subset of molecules.
[0116] Figure 9A shows a flowchart illustrating an example of process 700 for designing small molecules. Referring to Figures 1 and 9A, process 700 may be carried out by the design engine 110 to generate one or more molecules, such as small molecules having molecular weights of approximately 100 daltons to approximately 1,000 daltons, for synthesis or as candidates for synthesis.
[0117] In step 702, the design engine 110 performs a first iteration of the generation algorithm 115 to generate a first new molecular population based on the initial molecular population. For example, as shown in Figure 4, the design engine 110 may perform a first iteration 200 of the generation algorithm 115 to generate a first new molecular population 215 based on the first initial molecular population 210. As shown in Figure 4, iteration 200 of the generation algorithm 115 may include the design engine 110 modifying one or more molecules from the first initial molecular population 210. Furthermore, as shown in Figure 4, iteration 200 of the generation algorithm 115 may include the design engine 110 selecting a predetermined number N molecules from the first initial molecular population 210 and one or more modified molecules that satisfy one or more goodness-of-fit criteria (e.g., goodness-of-fit scores) to be included in the first new molecular population 215.
[0118] In step 704, the design engine 110 may perform a second iteration 300 of the generation algorithm 115 to generate a second new molecular population 225 based on the first new molecular population 215. As shown in Figure 5, the design engine 110 may modify one or more molecules from the first new molecular population 215 and one or more modified molecules to include in the second new molecular population 225 if they satisfy one or more goodness-of-fit criteria (e.g., goodness-of-fit score). TIFF2026509837000008.tif4170 By selecting a predetermined number of molecules, a second new molecular population 225 can be generated.
[0119] In 706, the design engine 110 may determine whether one or more conditions are met. In some exemplary embodiments, the design engine 110 may continue to perform additional iterations of the generation algorithm 115 until one or more conditions are met. One example of a condition may include that the design engine 110 has generated a threshold number of subsequent new molecular populations. Another exemplary condition may include that the design engine 110 has identified a threshold number of molecules that satisfy one or more goodness-of-fit criteria (e.g., goodness-of-fit scores). Other examples of conditions may include the goodness-of-fit scores of a first new molecular population 215 and / or a second new molecular population 225 that satisfy a threshold or one or more convergence criteria, where the goodness-of-fit scores show an improvement of less than a threshold compared to the goodness-of-fit scores of other molecular populations (e.g., generated based on the same initial molecular population 210 or different initial molecular populations).
[0120] If, in 708, the design engine 110 determines that one or more conditions are met, it may identify at least one molecule from the second new molecular population as a candidate for synthesis. For example, if the design engine 110 determines that one or more conditions are met, it may, for example, select a subset 230 of molecules from the second new molecular population 225 as candidates for synthesis and testing. In some cases, the design engine 110 may select the subset 230 of molecules by applying one or more filters or threshold criteria.
[0121] If, in 710, the design engine 110 determines that one or more conditions are not met, it may perform a third iteration of the generation algorithm 115 to generate a third new molecular population based on the second new molecular population or the second initial molecular population. In some exemplary embodiments, if the design engine 110 determines that one or more conditions are not met, it may perform one or more additional iterations of the generation algorithm 115 to generate one or more additional new molecular populations.
[0122] In 712, the design engine 110 identifies one or more molecules from a new population of molecules that satisfy one or more convergence conditions as candidates for synthesis and / or testing. For example, molecules to be synthesized or identified as candidates for synthesis and / or testing may be determined by the design engine 110 by applying one or more filters or threshold criteria. In a preferred embodiment, one or more molecules identified as candidates for synthesis and / or testing may undergo synthesis.
[0123] In other preferred embodiments, one or more molecules identified as candidates for synthesis and / or testing may be tested. These identified molecules may be tested as a further filtering step prior to more labor-intensive tests, such as preclinical trials, by screening for specific measured physicochemical properties as well as additional calculated and predicted properties.
[0124] Figure 9B shows the steps of process 750 corresponding to a single iteration of the generation algorithm 115 performed by the design engine 110 in operations 702, 704, and 710 of process 700 described with respect to Figure 9A, for example.
[0125] In 752, the design engine 110 may modify one or more molecules from a given molecular population. For example, in the example of iteration 200 shown in Figure 4, the design engine 110 may generate a first new molecular population 215 by modifying one or more molecules from a first initial molecular population 210. To increase the likelihood that the molecules in the subsequent molecular population will exhibit the same desirable properties as those in the previous molecular population, the design engine 110 may limit the modifications made to each molecule to, for example, one or more chemically valid mutations. In another example, if the molecule is the product of a first reaction between a first reagent selected from a first set of reagents associated with a first reaction and a second reagent selected from a second set of reagents associated with a first reaction, the design engine 110 may modify the molecule by changing one or more of the first reaction, the first reagent, and the second reagent.
[0126] In 754, the design engine 110 may select a predetermined number of molecules from a molecular population and one or more modified molecules that satisfy one or more goodness-of-fit criteria (e.g., goodness-of-fit scores) for inclusion in a subsequent molecular population. As further described elsewhere in this specification, a predetermined number of molecules may be selected for inclusion in a subsequent molecular population based on each molecule satisfying one or more goodness-of-fit criteria (e.g., goodness-of-fit scores) relating to, for example, diversity, docking to targets, solubility, permeability, selectivity, efficacy, toxicity, physiological pharmacokinetic properties (PBPK), etc.
[0127] Implementation form of calculation Figure 10 is a block diagram illustrating an example of a computing system 800 in the implementation form of this subject. Referring to Figures 1 and 10, the computing system 800 may incorporate the design engine 110 and / or any components thereof.
[0128] As shown in Figure 10, the computing system 800 may include one or more processors 810, memory 820 which typically includes both high-speed random-access memory and non-volatile memory (such as one or more magnetic disk drives), one or more storage devices 830, and one or more input / output devices 840. The processors 810, memory 820, storage devices 830, and input / output devices 840 may be interconnected via a system bus 850.
[0129] Each of the one or more processors 810 can process instructions to be executed within the computing system 800. Such instructions to be executed may, for example, implement one or more components of the design engine 110. In some implementations of this subject, the processor 810 may be a single-threaded processor. Alternatively, the processor 810 may be a multi-threaded processor. The processor 810 can process instructions stored in memory 820 and / or one or more storage devices 830 to display outputs such as graphic information of a user interface provided via one or more input / output devices 840.
[0130] Memory 820 is a computer-readable medium, volatile or non-volatile, for storing information within the computing system 800. Memory 820 can store data structures representing molecular structures, molecular properties, and chemical reactions, for example.
[0131] One or more storage devices 830 can provide persistent storage for the computing system 800. The storage devices 830 can be removable devices such as hard disk devices, optical disk devices, or floppy disk devices, or tape devices, or other suitable persistent storage means.
[0132] One or more input / output devices 840 provide input / output functionality to the computing system 800. In some implementations of this subject, one or more input / output devices 840 include one or more of the following: a keyboard, a pointing device, and other methods of providing user input. In various implementations, one or more input / output devices 840 include a display device for displaying a graphical user interface.
[0133] According to some implementations of this subject, one or more input / output devices 840 can provide input / output operations for network devices. For example, one or more input / output devices 840 may include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., local area networks (LANs), wide area networks (WANs), the Internet). Thus, the input / output devices can communicate with the network 140 and one or more client devices 120.
[0134] The results of calculations (such as population diversity) produced by the techniques described herein, as well as the generated molecular structures themselves, can be displayed in tangible form on one or more computer displays such as monitors, laptop displays, or tablet, notebook, netbook, or mobile phone screens. Numerical results may be further printed on paper, stored on computer-readable media, stored as electronic files in a format for transfer or sharing between computers, or projected onto screens in the audience during presentations, etc.
[0135] In some implementations of this subject, the computing system 800 may be used to run various interactive computer software applications (e.g., Microsoft Excel®, and / or other types of software) that can be used for organizing, analyzing, and / or storing data in various (e.g., tabular) formats. Alternatively, the computing system 800 may be used to run any type of software application. These applications may be used to perform various functions, planning functions, etc. (e.g., generating, managing, and editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications may include various add-in functions or may be standalone computing products and / or functions. When activated within an application, functionality may be used to generate a user interface provided via the input / output device 840. The user interface may be generated by the computing system 800 and presented to the user (e.g., on a computer screen monitor).
[0136] As used herein, the term “machine-readable medium” means any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as magnetic disks, optical disks, memory, and programmable logic devices (PLDs), and includes machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signals” means any signals used to provide machine instructions and / or data to a programmable processor. Machine-readable medium can store such machine instructions non-temporarily, for example, non-temporarily, solid-state memory, magnetic hard drives, or any equivalent storage medium. Alternatively or additionally, machine-readable medium can store such machine instructions temporarily, for example, a processor cache or other random-access memory associated with one or more physical processor cores.
[0137] To provide user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having, for example, a display device such as a cathode ray tube (CRT), liquid crystal display (LCD), or light-emitting diode (LED) monitor for displaying information to the user, and a keyboard, and a pointing device such as a mouse or trackball by which the user can provide input to the computer. User interaction may also be provided using other types of devices. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic input, voice input, or tactile input. Other possible input devices include touchscreens, with or without styluses, or single or multipoint resistive or capacitive trackpads, speech recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, speech recognition devices, gesture recognition technology, human fingerprint readers, or other touch-sensing devices such as other inputs based on the user's eye movements.
[0138] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuits, integrated circuits, specially designed application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementations in one or more computer programs executable and / or interpretable on a programmable system which includes at least one programmable processor, which may be for special-purpose or general-purpose use, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device. The programmable system or computing system may include a client and a server. The client and server are generally located geographically separated from each other and typically interact through a communication network. The client-server relationship arises from computer programs running on each computer and having a client-server relationship with each other.
[0139] The methods described herein are preferably implemented as computer programs executed by one or more computer systems, and such implementation is within the scope of the skills of those skilled in the art. In particular, computer functions for manipulating digitally stored molecular structures can be developed and implemented by those skilled in the art. These computer programs, also called programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and can be implemented in many different programming languages, including, in some cases, mixed implementations (i.e., relying on separate parts written in two or more computing languages appropriately configured to communicate with each other). For example, programs, as well as any necessary scripting functions, can be programmed in several compiled or uncompiled languages, including, but not limited to, high-level procedural and / or object-oriented programming languages, and / or in assembly / machine languages. Such languages can be selected from: C, C++, Java, JavaScript, Visual Basic, Tcl / Tk, Python, Perl, .Net languages such as C#, and other equivalent languages. The capabilities of the art are not limited to, and do not depend on, the underlying programming language used to implement or control access to the basic functions. Alternatively, the functionality could be implemented from higher-level features, such as toolkits, that rely on previously developed features for manipulating bitstrings and fingerprints.
[0140] To the extent that a given implementation relies on other already implemented software components, such as functions for calculating molecular similarity functions, reading molecular databases, and calculating the fit of molecular structures of protein active sites, it can be assumed that these functions are accessible to those skilled in the art.
[0141] When generating random or pseudorandom numbers, the process is preferably not based solely on a formula or process. Preferably, the selection of random numbers, or the seed for the random number generation method, is derived from real-world fluctuations such as electrical potential within the computing device being used.
[0142] The method of this technology can also retrieve functions contained in one or more dynamically linked libraries stored in either memory 820 or disk 830, although this is not shown in Figure 10.
[0143] When the operation mode of this technology is reduced to one or more software modules, functions, or subroutines, it may operate in batch mode, either by processing in batches on a database storing molecular structures or by interacting with a user who inputs specific instructions for a single molecular structure.
[0144] Various implementations of the techniques described herein may be intended to be carried out on computing devices of varying complexity, including but not limited to workstations, PCs, laptops, notebooks, tablets, netbooks, and other mobile computing devices, including mobile phones and personal digital assistants. The computing devices may have appropriately configured processors, including but not limited to graphics processors and numerical coprocessors, for running software that performs the methods described herein. In addition, certain computing functions are typically distributed across two or more computers, for example, one computer accepts input and instructions, and a second or additional computer receives the instructions via a network connection, performs processing remotely, and optionally sends the results or output back to the first computer.
[0145] Finally, it should be understood that the executable instructions for causing a appropriately programmed computer to execute the generative design methods for molecular structures described herein can be stored and delivered in any suitable computer-readable format. This includes, but is not limited to, portable computer-readable drives such as high-capacity “hard drives” or “pen drives,” for example, connections to a computer’s USB port, and internal drives to a computer, as well as CD-ROMs or optical discs. While the executable instructions may be stored on portable computer-readable media and delivered to the purchaser or user in such a tangible form, it should be further understood that the executable instructions can also be downloaded to the user’s computer from a remote location via an internet connection, which may itself partially rely on wireless technologies such as WiFi. Such aspects of the technology do not mean that the executable instructions take the form of signals or other non-tangible embodiments. The executable instructions may also be executed as part of a “virtual machine” implementation.
[0146] The subject matter described herein may be embodied in systems, apparatus, methods, and / or articles, depending on the desired configuration. The implementations described above do not necessarily represent all implementations that correspond to the subject matter described herein. Rather, they are merely some examples that correspond to aspects related to the described subject matter. While some variations have been detailed above, other modifications and additions are possible. In particular, further features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features, and / or combinations and subcombinations of some of the further features disclosed above. In addition, the logical flows depicted in the accompanying figures and / or described herein do not necessarily require a specific order or sequence shown to achieve the desired result. Other implementations may be within the scope of the following claims.
Claims
1. A computer implementation method, (a) Obtaining an initial molecular population from a molecular structure database, (b) Modifying at least a first molecule in the initial molecule population, and generating a first new molecule population by selecting a predetermined number of molecules from the initial molecule population and at least the first modified molecule that satisfy one or more goodness-of-fit scores, (c) Modifying at least a second molecule in the first new molecular population, and generating a second new molecular population by selecting a predetermined number of other molecules from the first new molecular population and at least a second modified molecule that satisfy the one or more fitness scores, (d) In response to determining that the convergence criteria are not met, the process (c) is repeated at least once using the second new molecular population as the first new molecular population, (e) In response to determining that the convergence criteria have been met, a subset of the second new molecular population is selected as a candidate for synthesis and testing, Computer implementation methods, including those mentioned above.
2. The generation of the first new molecular population is (i) Calculate one or more goodness-of-fit scores for at least the first modified molecule and each molecule in the initial molecule population, (ii) Forming the first new molecular population by selecting a predetermined number of molecules from the initial molecular population and at least the first modified molecules, each having one or more fitness scores that satisfy the one or more fitness scores, Each example that generates the second new molecular population is, (i) Calculate one or more goodness-of-fit scores for at least the second modified molecule and each molecule in the first new molecular population, (ii) From the first new molecular population and at least the second modified molecule, one or more molecules having a fitness score that satisfies the one or more fitness scores. By selecting a predetermined number of molecules, the second new molecular group is formed, The method according to claim 1, including the method described in claim 1.
3. Modifying each of the first molecule and the second molecule is The modification of the molecule comprises applying one or more modifications to the molecule's structure selected from the group consisting of: substituting a non-hydrogen atom radical of the molecule with a part selected from a standard set of parts; substituting a hydrogen atom of the molecule with a part selected from the standard set of parts; and replacing a divalent fragment of the molecule with a divalent part selected from the standard set of parts. The radical is a ring system containing 1 to 5 non-hydrogen atoms, or 3 to 10 non-hydrogen atoms. The portion is a ring system containing 1 to 5 non-hydrogen atoms, or 3 to 10 non-hydrogen atoms. The divalent fragment is a ring system containing 1 to 5 non-hydrogen atoms, or 3 to 10 non-hydrogen atoms. The method according to claim 2.
4. The method according to claim 2, wherein one or more goodness-of-fit scores include a diversity score.
5. The method according to claim 4, wherein the diversity score indicates chemical similarity between molecular pairs, and the diversity score penalizes at least one molecule in molecular pairs that are structurally or chemically similar to each other.
6. The method according to claim 4, wherein the diversity score is calculated from a metric selected from the group consisting of the Tanimoto exponent, cosine coefficient, Dice distance, Euclidean distance, Cityblock distance, Hamming exponent, and Tversky exponent.
7. The method according to claim 2, wherein the one or more fitness scores include a calculated docking score of the molecule for a target binding site.
8. The aforementioned convergence criteria are, Repeating a fixed number of steps (c), An average goodness-of-fit score that satisfies a threshold, calculated for the second new molecular population obtained from the last example of process (c), and The number of specified molecules in the second new molecular population obtained from the last example of step (c) has a goodness-of-fit score that is below a threshold. The method according to claim 1, selected from the following.
9. Obtaining the aforementioned initial molecular population is possible. Clustering corresponding sets of molecules in the molecular structure database into one or more molecular clusters based on a similarity metric, From each of the one or more molecular clusters, select one or more molecules having a fitness score that satisfies one or more initial fitness scores, The method according to claim 1, including the method described in claim 1.
10. The method according to claim 9, wherein the similarity metric is the Tanimoto exponent, cosine coefficient, Dice distance, Euclidean distance, Cityblock distance, Hamming exponent, or Tversky exponent, and the similarity metric is based on a characteristic selected from 2D similarity, 3D similarity, and a vector of physicochemical properties.
11. The method according to claim 9, wherein clustering is performed by applying one or more clustering algorithms selected from Butina, Centroid, CLInk, Gower, McQuitty, SLink, Unweighted Pair Group Method with Arithmetic Mean (UPMGA)), Ward, and Jarvis-Patrick.
12. The method according to claim 2, wherein the one or more fitness scores for a molecule include one or more calculated values from the following molecular properties: solubility, permeability, selectivity score, efficiency score, toxicity, and physiological pharmacokinetic (PBPK) score.
13. The aforementioned initial molecular population Randomly selecting one or more molecules from the aforementioned molecular structure database, Selecting one or more molecules from the molecular structure database that have one or more physicochemical properties that meet the threshold criteria, and Selecting one or more molecules that have a specific scaffold, Or a combination of those, The method according to claim 1, obtained by a method selected from the following.
14. (f) The method according to claim 1, further comprising synthesizing at least one candidate molecule.
15. The second new molecular group is the other The method according to claim 2, wherein a fixed proportion of molecules derived from the first new molecular population is selected as at least a portion of a predetermined number of molecules.
16. The one or more modifications to each molecule Determining a set of modifications applicable to the molecule from a list of available modifications, and Randomly selecting one or more modifiers from the set of applicable modifiers, The method according to claim 3, as specified by [the specified method].
17. A computer implementation method, (a) Obtaining multiple initial molecular populations from molecular structure databases, (b) Generating a plurality of first new molecular populations from each of the plurality of initial molecular populations, wherein each of the plurality of first new molecular populations includes one or more molecules obtained by modifying one or more molecules in the initial molecular population on which the first new population is based, and the molecules in each of the plurality of first new molecular populations satisfy one or more fitness scores, (c) Generating a plurality of second new molecular populations from each of the plurality of first new molecular populations, wherein each of the plurality of second new populations includes one or more molecules obtained by modifying one or more molecules in the first new molecular population on which the second new population is based, and the molecules in each of the plurality of second new molecular populations satisfy the one or more goodness-of-fit scores, thereby generating a plurality of second new molecular populations. (d) In response to the determination that the convergence criteria are not met, the process (c) is repeated at least once, such that each of the first new molecular populations is one of the second new molecular populations obtained from the prior examples of the process (c). (e) In response to determining that the convergence criteria have been met, a subset of molecules containing at least one molecule from each of the plurality of new second molecular populations is selected for synthesis and testing, Computer implementation methods, including those mentioned above.
18. The method according to claim 17, wherein one or more goodness-of-fit scores include a diversity score.
19. The method according to claim 18, wherein the diversity score is based on the structural similarity between molecular pairs, each containing one molecule selected from each of two molecular populations, and the diversity score penalizes at least one molecule in molecular pairs that are structurally similar to each other.
20. The method according to claim 18, wherein the diversity score is calculated from a metric selected from the group consisting of the Tanimoto exponent, cosine coefficient, Dice distance, Euclidean distance, Cityblock distance, Hamming exponent, and / or Tversky exponent.
21. The method according to claim 18, wherein the diversity score is calculated among any molecular pairs that can be formed from molecules in one population and molecules in another population.
22. The method according to claim 18, wherein, for any pair of molecules whose diversity score does not meet the threshold, only one molecule from the pair satisfies the one or more goodness-of-fit scores.
23. The generation of the aforementioned multiple new first molecular populations is performed as follows for each new first molecular population: (i) Calculate one or more goodness-of-fit scores for at least the first modified molecule and each molecule in the corresponding initial molecule population, (ii) From the corresponding initial molecule population and at least the first modified molecule, one or more molecules having a fitness score that satisfies the one or more fitness scores By selecting a predetermined number of molecules, the first new molecular group is formed, This includes implementing, and Each example of generating the aforementioned multiple second new molecular populations is as follows for each second new molecular population: (i) Calculate one or more goodness-of-fit scores for each molecule in at least the second modified molecule and the corresponding first new molecular population, (ii) From the corresponding first new molecular population and at least the second modified molecule, another having one or more fitness scores that satisfy the one or more fitness scores. By selecting a predetermined number of molecules, the second new molecular group is formed, The method according to claim 17, which includes carrying out the following.
24. A computer implementation method, (a) Obtaining an initial molecular population in which each molecule has a molecular structure constructed from a reaction database containing a reaction set, wherein each reaction from the reaction set is associated with two or more reagent sets, and each molecule in the initial molecular population is the product of the first reagent from the first reagent set, the second reagent from the second reagent set, and optionally the third reagent from the third reagent set, according to a reaction associated with both the first and second reagents, and an optional third reagent set. (b) The first new molecular population, Modify at least the first molecule in the initial molecule population by replacing the first reagent on which the first molecule is formed with another first reagent from the first reagent set associated with the reaction on which the first molecule is formed, and optionally replacing the second reagent on which the first molecule is formed with another second reagent from the second reagent set associated with the reaction on which the first molecule is formed, thereby forming the first modified molecule, and Select a predetermined number of molecules from the initial molecular population and at least the first modified molecules that satisfy one or more fitness scores for inclusion in the first new molecular population. To generate by, (c) A second new molecular population, at least, Modify at least the second molecule in the first new molecular population by replacing the first reagent on which the second molecule is formed with another first reagent from the first reagent set, and optionally replacing the second reagent on which the second molecule is formed with another second reagent from the second reagent set, thereby forming the second modified molecule, and Selecting a predetermined number of other molecules from the first new molecular population and at least the second modified molecules that satisfy the one or more fitness scores. To generate by, (d) In response to determining that the convergence criteria are not met, the process (b) is repeated at least once using the second new molecular population as a new initial molecular population, (e) In response to determining that the convergence criteria have been met, a subset of the second new molecular population is selected as a candidate for synthesis and testing, Computer implementation methods, including those mentioned above.
25. A computer implementation method, (a) Obtaining an initial molecular population from a molecular structure database, wherein each molecule in the initial molecular population is a product of a first reaction between a first reagent selected from a first reagent set related to the first reaction, a second reagent selected from a second reagent set related to the first reaction, and optionally a third reagent from a third reagent set related to the first reaction. (b) The first new molecular population, The first modified molecule is produced by applying a modification to at least one molecule in the initial molecular population, wherein the modification is Replacing the first reagent on which the first molecule is formed with another first reagent selected from the first reagent set, Replacing the second reagent on which the first molecule is formed with another second reagent selected from the second reagent set, and Replacing the third reagent on which the first molecule is formed with another third reagent selected from the third reagent set, To create a first modified molecule selected from one or more of the following: Furthermore Select a predetermined number of molecules from the initial molecular population and at least the first modified molecule that satisfy one or more fitness scores. To generate by, (c) A second new molecular population, The method involves creating a second modified molecule by applying a modification to at least a second molecule in the second new molecular population, wherein the modification is Replacing the first reagent on which the first molecule is formed with another first reagent selected from the first reagent set, Replacing the second reagent on which the first molecule is formed with another second reagent selected from the second reagent set, and Replacing the third reagent on which the first molecule is formed with another third reagent selected from the third reagent set, To create a second modified molecule selected from one or more of the following: Furthermore From the first new molecular population and at least the second modified molecule, another molecule satisfying the one or more fitness scores. Select a predetermined number of molecules. To generate by, (d) In response to determining that the convergence criteria are not met, the process (c) is repeated at least once using the second new molecular population as a new initial molecular population, (e) In response to determining that the convergence criteria have been met, a subset of the second new molecular population is selected as a candidate for synthesis and testing, Computer implementation methods, including those mentioned above.
26. The first new molecular population is generated The additional first molecule in the aforementioned initial molecular population, Selecting a third reagent having a calculated similarity to the first reagent on which the additional first molecule is formed, wherein the third reagent is in an additional set of first reagents related to an additional reaction, and the additional reaction is different from the reaction on which the additional first molecule is formed. Constructing an additional first modified molecule by replacing the first reagent on which the additional first molecule is formed with the third reagent, and replacing the second reagent on which the additional first molecule is formed with a fourth reagent selected from an additional second reagent set related to the additional reaction, wherein the fourth reagent constructs an additional first modified molecule having a calculated similarity to the second reagent. Furthermore Select a predetermined number of molecules from the initial molecular population, the first modified molecule, and the additional first modified molecule that satisfy one or more fitness scores. Modify by The method according to claim 25, further comprising:
27. To generate a second new molecular population, An additional second molecule in the aforementioned initial molecular population, Selecting a third reagent having a calculated similarity to the first reagent in which the additional second molecule is formed, wherein the third reagent is in an additional set of first reagents related to an additional reaction, and the additional reaction is different from the reaction in which the additional second molecule is formed. Constructing an additional second modified molecule by replacing the first reagent on which the additional second molecule is formed with the third reagent, and replacing the second reagent on which the additional second molecule is formed with a fourth reagent selected from a fourth additional second reagent set related to the additional reaction, wherein the fourth reagent constructs an additional second modified molecule having a calculated similarity to the second reagent. Furthermore Select a predetermined number of molecules from the first new molecular population, the second modified molecule, and the additional second modified molecule that satisfy one or more fitness scores. Modify by The method according to claim 26, further comprising:
28. It is a system, At least one data processor, A memory for storing instructions that, when executed by the at least one data processor, result in an operation including the method according to any one of claims 1 to 27, A system equipped with these features.
29. A non-temporary computer-readable medium for storing instructions that, when executed by at least one data processor, result in an operation including the method according to any one of claims 1 to 27.