Pharmacophore-based molecule design by learning on voxel grids
The molecule design computation model enhances pharmacophore-shape based molecule design by generating linear representations from three-dimensional pharmacophore profiles, enabling efficient two-dimensional similarity searches to discover diverse binders with optimal properties within computational limits.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-02
AI Technical Summary
Conventional search-based approaches for molecule design face limitations in exploring the vast molecular space, leading to overlooked molecules with optimal properties due to computational constraints and inefficient search strategies, particularly in pharmacophore-shape based molecule design, which is time and resource intensive.
A molecule design computation model operates on a three-dimensional pharmacophore-shape profile to generate linear representations of molecular variants, translating them into two-dimensional molecular fingerprints for efficient similarity searches in molecular libraries, reducing computational overhead and expanding the search scope without exceeding budget constraints.
This approach enables a more comprehensive search of molecular libraries for molecules with similar pharmacophore-shape profiles, increasing the likelihood of discovering diverse binders with minimal liabilities and optimizing them into successful therapeutics, while maintaining computational efficiency.
Smart Images

Figure US2025048189_02042026_PF_FP_ABST
Abstract
Description
Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 PHARMACOPHORE-BASED MOLECULE DESIGN BY LEARNING ON VOXEL GRIDS CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 700,306, entitled “PHARMACOPHORE-BASED MOLECULE DESIGN BY LEARNING ON VOXEL GRIDS” and filed on September 27, 2024, and to U.S. Provisional Application No. 63 / 724,118, entitled “PHARMACOPHORE-BASED MOLECULE DESIGN BY LEARNING ON VOXEL GRIDS” and filed on November 22, 2024, the disclosures of which are incorporated herein by reference in their entireties. TECHNICAL FIELD
[0002] The subject matter described herein relates generally to generative artificial intelligence and more specifically to pharmacophore-based generative molecule design. INTRODUCTION
[0003] A molecule is a group of two more atoms held together by chemical bonds. Molecules form the smallest identifiable unit into which a pure substance can be divided while still retaining the composition and chemical properties of that substance. One example of a molecule is a small molecule, which is a low-weight compound having a molecular weight between approximately 100 Daltons and 1000 Daltons. Small molecule therapeutics, which modulate biochemical processes to diagnose, treat, and prevent a gamut of illnesses, have been a cornerstone in modern pharmacology due to a number of compelling advantages. For example, small molecule drugs are capable of penetrating cell membranes to reach intracellular targets. Moreover, small molecule drugs are adaptable to a wide variety of therapeutic applications. For instance, a small molecule drug may be formulated as pills and capsules, intravenous or subcutaneous injectables, inhalational medicines, or suppositories. The development of the smallAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 molecule drug may further extend to tailoring various pharmacokinetic properties including liberation, absorption, distribution, metabolism, potency, efficacy, phenotypic effects, and excretion.
[0004] By contrast, large molecules (also known as biopharmaceuticals, biologicals, or biologics) can range between approximately 3000 Daltons and 150,000 Daltons in molecular weight. Large molecule drugs are often derivatives of natural human proteins, which modulate many essential cellular functions such as enzymatic reactions, transport of molecules, regulation and execution of a number of biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. It is common for a single large molecule to have more than 1,300 amino acid residues, which are linked by peptide bonds to form one or more polypeptide. Due to their size and complexity, large molecule drugs are recombinantly produced by engineered cells instead of being chemically synthesized like the majority of small molecule drugs. Moreover, large molecule therapeutics are usually delivered through injection or infusion due to the ineffectiveness of oral administration. The development of a large molecule drug may entail designing one or more sequences of amino acid residues capable of binding to a target (e.g., a protein, a nucleic acid, and / or the like) with sufficient specificity and absent undesirable traits such as immunogenicity, self-association, instability, and / or the like. SUMMARY
[0005] Systems, methods, and articles of manufacture, including computer program products, are provided for pharmacophore based generative molecule design. For example, in some cases, a molecule design computation model may be trained to operate on the three- dimensional representation of an input molecule, such as a voxelized pharmacophore-shapeAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 profile of the input molecule, and output linear representations of molecular variants that can exhibit the same pharmacophore-shape profile as the input molecule. One or more two- dimensional analogs of the input molecule may be identified by at least determining a two- dimensional similarity between the two-dimensional representation of each molecular variant and the two-dimensional representations of molecules in one or more molecular libraries. In some cases, those two-dimensional analogs may undergo further screening to identify those that exhibits a sufficient degree of three-dimensional similarity to the input molecule.
[0006] In one aspect, there is provided a system for pharmacophore based generative molecule design. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one data processor. The operations may include: receiving a three-dimensional representation of an input molecule, wherein the three-dimensional representation of the input molecule specifies a location of one or more pharmacophores present in the input molecule; generating a first molecular variant of the input molecule by at least applying a molecule design computation model that has been trained to operate on the three-dimensional representation of the input molecule and output a linear representation of the first molecular variant; translating the linear representation of the first molecular variant into a two-dimensional representation of the first molecular variant; identifying a first molecule from a molecular library as a first two- dimensional analog of the input molecule based at least on a two-dimensional representation of the first molecule exhibiting a threshold degree of two-dimensional similarity to the two- dimensional representation of the first molecular variant of the input molecule; and generating, as an output molecule, the first two dimensional analog of the input molecule based at least on a three-dimensional representation of the first two-dimensional analog exhibiting a thresholdAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 degree of three-dimensional similarity to the three-dimensional representation of the input molecule.
[0007] In another aspect, there is provided a computer-implemented method for pharmacophore based generative molecule design. The method may include: receiving a three- dimensional representation of an input molecule, wherein the three-dimensional representation of the input molecule specifies a location of one or more pharmacophores present in the input molecule; generating a first molecular variant of the input molecule by at least applying a molecule design computation model that has been trained to operate on the three-dimensional representation of the input molecule and output a linear representation of the first molecular variant; translating the linear representation of the first molecular variant into a two-dimensional representation of the first molecular variant; identifying a first molecule from a molecular library as a first two-dimensional analog of the input molecule based at least on a two-dimensional representation of the first molecule exhibiting a threshold degree of two-dimensional similarity to the two-dimensional representation of the first molecular variant of the input molecule; and generating, as an output molecule, the first two dimensional analog of the input molecule based at least on a three-dimensional representation of the first two-dimensional analog exhibiting a threshold degree of three-dimensional similarity to the three-dimensional representation of the input molecule.
[0008] In another aspect, there is provided a computer program product for pharmacophore based generative molecule design. The computer program product may include a non-transitory computer readable medium storing instructions that result in operations when executed by at least one data processor. The operations may include: receiving a three- dimensional representation of an input molecule, wherein the three-dimensional representation ofAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 the input molecule specifies a location of one or more pharmacophores present in the input molecule; generating a first molecular variant of the input molecule by at least applying a molecule design computation model that has been trained to operate on the three-dimensional representation of the input molecule and output a linear representation of the first molecular variant; translating the linear representation of the first molecular variant into a two-dimensional representation of the first molecular variant; identifying a first molecule from a molecular library as a first two-dimensional analog of the input molecule based at least on a two-dimensional representation of the first molecule exhibiting a threshold degree of two-dimensional similarity to the two-dimensional representation of the first molecular variant of the input molecule; and generating, as an output molecule, the first two dimensional analog of the input molecule based at least on a three-dimensional representation of the first two-dimensional analog exhibiting a threshold degree of three-dimensional similarity to the three-dimensional representation of the input molecule.
[0009] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.
[0010] In some variations, a two-dimensional similarity metric quantifying a similarity between the two-dimensional representation of the first molecule and the two- dimensional representation of the first molecular variant of the input molecule is determined. That the two-dimensional representation of the first molecule exhibits the threshold degree of two-dimensional similarity to the two-dimensional representation of the first molecular variant of the input molecule is determined based at least on the two-dimensional similarity metric.
[0011] In some variations, the two-dimensional representation of the first molecule is determined to exhibit the threshold degree of two-dimensional similarity to the two-dimensionalAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 representation of the first molecular variant of the input molecule based at least on the two- dimensional similarity metric satisfying one or more thresholds.
[0012] In some variations, the two-dimensional similarity metric includes a Tanimoto index, a Jaccard coefficient, a cosine coefficient, a Dice metric, a Euclidean metric, a city-block metric, a Hamming index, and / or a Tversky index.
[0013] In some variations, the two-dimensional representation of the first two- dimensional analog of the input molecule is translated into the three-dimensional representation of the first two-dimensional analog. A three-dimensional similarity metric quantifying a similarity between the three-dimensional representation of the first two-dimensional analog of the input molecule and the three-dimensional representation of the input molecule is determined. That the three-dimensional representation of the first two-dimensional analog exhibits the threshold degree of similarity to the three-dimensional representation of the input molecule is determined based at least on the three-dimensional similarity metric.
[0014] In some variations, the determining the three-dimensional similarity metric includes aligning the three-dimensional representation of the first two-dimensional analog of the input molecule and the three-dimensional representation of the input molecule, determining an overlap between the three-dimensional representation of the first two-dimensional analog aligned with the three-dimensional representation of the input molecule, and determining, based at least on the overlap, the three-dimensional similarity metric quantifying the similarity between the three-dimensional representation of the first two-dimensional analog of the input molecule and the three-dimensional representation of the input molecule.
[0015] In some variations, the three-dimensional similarity metric includes a Tanimoto Combo score.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0016] In some variations, the three-dimensional similarity metric quantifies a similarity in an arrangement of atoms in three-dimensional space and a spatial overlap in one or more functional groups.
[0017] In some variations, a second molecule is identified, within a same molecular library or a different molecular library, as a second two-dimensional analog of the input molecule based at least on a two-dimensional representation of the second molecule exhibiting the threshold degree of two-dimensional similarity to the two-dimensional representation of the first molecular variant of the input molecule. The second two dimensional analog of the input molecule is generated as another output molecule based at least on a three-dimensional representation of the second two-dimensional analog exhibiting the threshold degree of three- dimensional similarity to the three-dimensional representation of the input molecule.
[0018] In some variations, the molecule design computation model is applied to generate, based at least on the three-dimensional representation of the input molecule, a plurality of molecular variants including the first molecular variant and a second molecular variant.
[0019] In some variations, a linear representation of the second molecular variant output by the molecule design computation model is translated into a two-dimensional representation of the second molecular variant. A second molecule is identified, within a same molecular library or a different molecular library, as a second two-dimensional analog of the input molecule based at least on a two-dimensional representation of the second molecule exhibiting the threshold degree of two-dimensional similarity to the two-dimensional representation of the second molecular variant of the input molecule. The second two dimensional analog of the input molecule is generated as another output molecule based at least on a three-dimensional representation of the second two-dimensional analog exhibiting theAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 threshold degree of three-dimensional similarity to the three-dimensional representation of the input molecule.
[0020] In some variations, the linear representation of the first molecular variant includes a simplified molecular input line entry system (SMILES) string or a self-referencing embedded string (SELFIES).
[0021] In some variations, the three-dimensional representation of the input molecule specifies the location of one or more pharmacophores as one or more densities across a three- dimensional voxel grid, and wherein each density is centered around a location of a corresponding pharmacophore.
[0022] In some variations, the three-dimensional representation of the input molecule further specifies a location of one or more atoms present in the input molecule.
[0023] In some variations, the three-dimensional representation of the input molecule specifies the location of the one or more atoms as one or more densities across a three- dimensional voxel grid, and wherein each density is centered around a location of a corresponding atom.
[0024] In some variations, the three-dimensional representation of the input molecule includes a plurality of channels, and wherein each channel of the plurality of channels corresponds to a different type of pharmacophore.
[0025] In some variations, the plurality of channels further includes a channel specifying a location of each atom in the input molecule.
[0026] In some variations, each channel of the plurality of channels corresponds to a different one of a hydrogen bond donor, a hydrogen bond acceptor, an aromatic ring, a cation, an anion, and a hydrophobic moiety.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0027] In some variations, each of the two-dimensional representation of the first molecule and the two-dimensional representation of the first molecular variant of the input molecule comprise a molecular fingerprint.
[0028] In some variations, the molecular fingerprint includes a binary code. Each bit in the binary code is set to a value indicating a presence or an absence of a chemical substructure.
[0029] In some variations, the molecular fingerprint includes a circular fingerprint, a path-based fingerprint, or a RDKit fingerprint.
[0030] In some variations, the molecular fingerprint includes a Morgan fingerprint or a RDKit fingerprint.
[0031] In some variations, the molecular design computation model includes an encoder that has been trained to compress, into a vector representation, the three-dimensional representation of the input molecule.
[0032] In some variations, the molecular design computation model further includes a decoder that has been trained to decode the vector representation of the input molecule to generate the linear representation of the first molecular variant.
[0033] In some variations, the molecule design computation model is trained based at least on a training dataset.
[0034] In some variations, the training dataset includes a plurality of training molecules. The training dataset includes, for each training molecule, a three-dimensional representation of the training molecule and a ground truth linear representation of the training molecule.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0035] In some variations, the training dataset is augmented by at least translating and / or rotating the three-dimensional representation of one or more training molecules.
[0036] In some variations, the training of the molecule design computation model includes applying the molecule design computation model to generate, for each training molecule in the training dataset, a linear representation of the training molecule; and adjusting one or more parameters of the molecule design computation model to reduce a difference between the ground truth linear representation of the training molecule and the linear representation of the training molecule generated by the molecule design computation model.
[0037] In some variations, the output molecule is synthesized to generate one or more synthesized molecules.
[0038] In some variations, the one or more synthesized molecules are analyzed to determine one or more properties of the output molecule.
[0039] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non- transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multipleAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.
[0040] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to the pharmacophore-based generative design of drug molecules, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter. DESCRIPTION OF DRAWINGS
[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,
[0042] FIG.1 depicts a system diagram illustrating an example of a molecule design system, in accordance with some example embodiments;
[0043] FIG.2 depicts a flowchart illustrating an example of a process for pharmacophore-based generative molecule design, in accordance with some example embodiments;Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0044] FIG.3A depicts a molecule graph of an example of a molecule, in accordance with some example embodiments;
[0045] FIG.3B depicts a three-dimensional shape profile of an example of a molecule, in accordance with some example embodiments;
[0046] FIG.3C depicts a three-dimensional pharmacophore profile of an example of a molecule, in accordance with some example embodiments;
[0047] FIG.4 depicts a flowchart illustrating an example of a process for generating a molecule, in accordance with some example embodiments;
[0048] FIG.5A depicts a block diagram illustrating the architecture of an example of an encoder for compressing a voxelized pharmacophore shape profile into a vector representation, in accordance with some example embodiments;
[0049] FIG.5B depicts a block diagram illustrating the architecture of an example of a decoder for mapping a vector representation of a voxelized pharmacophore shape profile to a tokenized a simplified molecular input line entry system (SMILES) string, in accordance with some example embodiments;
[0050] FIG.6 depicts a graph illustrating the correlation between three-dimensional similarity and two-dimensional similarity in molecules, in accordance with some example embodiments; and
[0051] FIG.7 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.
[0052] When practical, similar reference numbers denote similar structures, features, or elements.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 DETAILED DESCRIPTION
[0053] A molecule may be designed to exhibit multiple desirable properties including, in the case of therapeutics, drug-like properties such as binding affinity, specificity, biological activity, developability, and / or the like. Many conventional computational techniques for generating molecules with drug-like properties resort to searching the molecular space (or chemical space) occupied by every possible chemical compound (e.g., every possible combination of atoms of two or more chemical elements). For example, some search-based approaches may include scoring and ranking different molecules in the molecular space based on one or more drug-like properties, such as affinity, specificity, biological activity, anddevelopability. However, the aforementioned molecular space, which is estimated to contain10^^ possible chemical compounds, is prohibitively large and scales exponentially withmolecule size (e.g., the number of constituent atoms). Even with state-of-the-art computational resources, conventional search-based approaches are capable of exploring only a small fraction of the molecular space, such as small regions of the molecular space selected based on prior domain knowledge. This limitation in search scope means that conventional search-based approaches are likely to overlook molecules with more optimal properties. Moreover, conventional search-based approaches do not explore the molecular space in a principled manner, which prevents the generative process from being conditioned upon specific properties.
[0054] Ligand-based drug discovery (LBDD) is one variation of molecule design that aims to strategically reduce the search space by leveraging known molecules (or ligands). For example, instead of an indiscriminate search of the molecular space (or chemical space), ligand- based drug discovery (LBDD) focuses its efforts to exploring close variants of molecules (or ligands) known to interact with a certain biological target (e.g., binding to a site, such as aAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 receptor, on a target protein) are used to identify new, potentially better candidate compounds. While ligand-based drug discovery (LBDD) may be especially useful when the structure of the biological target is unknown, it may also complement structure-based drug design where at least some target information is available by facilitating tasks such as virtual screening and lead diversification. As described in more details below, pharmacophores are a critical element of ligand-based drug discovery (LBDD).
[0055] As used herein, a pharmacophore is a feature of a molecule, located in three- dimensional space, that may be involved in the binding interaction between the molecule and a biological target (e.g., a site, such as a receptor, on a target protein). Pharmacophores can include electronic features, steric features, and / or the like. Examples of pharmacophores include hydrogen bond donors, hydrogen bond acceptors, aromatic rings, cations, anions, hydrophobic moieties, and / or the like. Drug discovery campaigns, including the aforementioned ligand-based drug discovery (LBDD), rely on pharmacophore-shape based molecule design for finding a diverse pool of molecules with a high likelihood of binding to a biological target. For example, additional molecules capable of binding to the biological target may be found based on the pharmacophores of an input molecule (e.g., a lead compound, a query molecule, and / or the like) known to bind to the biological target. Having a diverse pool of binders increases the likelihood of that at least one of those molecules exhibit minimal labilities (e.g., chemical or physical instabilities, poor pharmacokinetics, immunogenicity, and / or the like) and has the potential to be optimized into a successful therapeutic.
[0056] Conventional pharmacophore-shape based molecule design includes a brute force search of a library of molecules (e.g., a library of synthetically accessible molecules) to identify those whose pharmacophore-shape profile, which describes the distribution ofAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 pharmacophore features in three-dimensional space, is sufficiently similar to the pharmacophore- shape profile of an input molecule known to bind to a biological target. For example, in some cases, the pharmacophore-shape profile of the input molecule may be aligned with the pharmacophore-shape profile of each molecule in the library. Moreover, in some cases, the overlap in three-dimensional space between the pharmacophore-shape profile of the input molecule and that of each molecule in the library may be computed. Aligning two pharmacophore-shape profiles and computing the overlap therebetween is time and resource intensive, especially because molecule libraries can contain millions and sometimes billions of molecules. Repeating the same process for each new input molecule further strains available time and computational resources. Running brute force searches against some larger molecular libraries may in fact be intractable. Thus, in order to operate within the available compute budget, conventional approaches pharmacophore-shape based molecule design impose certain limits, such as on the number of input molecules that are screened and the number of libraries that are searched. These constraints on the scope of pharmacophore-shape based molecule design reduce the likelihood of discovering molecules with better properties.
[0057] Various embodiments of the present disclosure expands the scope of pharmacophore-shape based molecule design without imposing significant computational overhead. In some example embodiments, a molecule design engine may implement a workflow that approximates a search for similar pharmacophore-shape profiles in three-dimensional space. For example, in some cases, the molecule design engine may apply a molecule design computation model to generate, based on an input molecule, one or more molecular variants of the input molecule. In some cases, the molecule design computation model may be trained to ingest a three-dimensional pharmacophore-shape profile of the input molecule and output theAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 linear (or one-dimensional) representation of each molecular variant. Moreover, in some cases, the molecule design engine may determine, based on the linear representation of each molecular variant, a corresponding two-dimensional representation. As described in more details below, instead of searching one or more molecular libraries based on the three-dimensional pharmacophore-shape profile of each molecular variant, which requires expensive and time- consuming computations in three-dimensional space, the one or more molecular libraries may be searched based on the two-dimensional representation of each molecular variant.
[0058] In some example embodiments, the molecule design computation model may operate on a voxelized representation of the pharmacophore-shape profile of the input molecule in which the input molecule is represented, in three-dimensional space, as continuous (e.g., Gaussian-like) densities centered around the location of each pharmacophore present in the input molecule. In some cases, the molecule design computation model may be trained to generate molecular variants that exhibit a similar pharmacophore-shape profile as the input molecule by at least predicting structures corresponding to the pharmacophore-shape profile of the input molecule. For example, in some cases, the molecule design computation model may be trained to output one or more linear representations (e.g., simplified molecular-input line-entry system (SMILES) strings, self-referencing embedded strings (SELFIES), and / or the like), each of which specifying a molecular structure with a pharmacophore-shape profile that is similar to the pharmacophore-shape profile of the input molecule. In some cases, each linear representation may be translated into a two-dimensional molecular graph by computing, for example, a corresponding molecular fingerprint (e.g., a circular fingerprint such as a Morgan fingerprint, a path-based fingerprint, a RDKit fingerprint, and / or the like). Instead of a three-dimensional pharmacophore profile of each molecular variant of the input molecule, one or more molecularAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 libraires may be searched using the molecular fingerprint to reduce the computational burden of the search. For instance, in some cases, one or more two-dimensional analogs of the input molecule may be identified based at least on a two-dimensional similarity metric (e.g., Tanimoto index) quantifying a similarity between the molecular fingerprint of each molecular variant and the molecular fingerprints of the molecules in one or more molecular libraries. Moreover, in some cases, the molecule design engine may perform additional screening of the two- dimensional analogs that includes, for example, a comparison between the three-dimensional pharmacophore-shape profile of the two-dimensional analogs and that of the input molecule. In some cases, the molecule design engine may generate, as an output molecule, a two-dimensional analog whose three-dimensional pharmacophore-shape profile exhibits a threshold degree of similarity (e.g., as quantified by a three-dimensional similarity metric such as Tanimoto Combo score) relative to the three-dimensional pharmacophore-shape profile of the input molecule. In some cases, the output molecule (e.g., the two-dimensional analog of the input molecule) may be undergo synthesis, purification, and / or analysis.
[0059] FIG.1 depicts a system diagram illustrating an example of a molecule design system 100, in accordance with some example embodiments. As shown in FIG.1, the molecule design system 100 may include a molecule design engine 110, a data store 120, one or more laboratory equipment 130, and a client device 140 with a user interface 145. In the example of the molecule design system 100 shown in FIG.1, the molecule design engine 110, the data store 120, the one or more laboratory equipment 130, and the client device 140 may be communicatively coupled via a network 150. It should be appreciated that the data store 120 may be a database including, for example, a relational database, a NoSQL database, a columnar database, an objected-oriented database, a key-value database, a hierarchical database, aAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 document database, a graph database, and / or the like. The one or more laboratory equipment 130 may include any wet lab and dry lab equipment capable of synthesis, purification, and / or analysis. Examples of the one or more laboratory equipment 130 may include synthesizers, including standard equipment (e.g., fume hoods, glassware, heating and cooling devices, stirrers), automated and specialized synthesis platforms (e.g., automated synthesizers, parallel synthesis workstations, and high-throughput experimentation (HTE) for efficient reaction optimization, and specialized reactors (e.g., microwave, flow, photochemistry). The one or more laboratory equipment 130 may also include tools to support various purification techniques such as chromatography (e.g., silica gel chromatography, preparative high performance liquid chromatography (Prep HPLC)), centrifugal, crystallization / recrystallization, and / or the like. In some cases, the one or more laboratory equipment 130 may also include analytical tools such as microscopes, spectroscopes, spectrometers (e.g., nuclear magnetic resonance (NMR) spectrometers, mass spectrometers), balances, pH meters, elemental analyzers, and / or the like. The client device 140 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable apparatus, and / or the like. The network 150 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like.
[0060] In some example embodiments, the molecule design engine 110 may include a molecule design computation model 115 and a query controller 117. In some cases, the molecule design computation model 115 may be trained to generate, based at least on an input molecule 101, one or more molecular variants 103 of the input molecule 101 including, for example, a first molecular variant 103a, a second molecular variant 103b, and / or the like. ForAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 example, in some cases, the molecule design computation model 115 may be trained to operate on a three-dimensional representation of the input molecule 101, such as a voxelized three- dimensional pharmacophore-shape profile of the input molecule 101, and output, for each molecular variant 103 of the input molecule 101, a corresponding linear representation (e.g., a simplified molecular-input line-entry system (SMILES) string, a self-reference embedded string (SELFIES), and / or the like). In some cases, the linear representation (e.g., SMILES string, SELFIES, and / or the like) of each molecular variant 103 may be converted into a two- dimensional representation (e.g., a molecular graph) by computing a corresponding molecular fingerprint (e.g., a circular fingerprint such as a Morgan fingerprint, a path-based fingerprint, a RDKit fingerprint, and / or the like). Moreover, in some cases, the query controller 117 may identify, based at least on the molecular fingerprint of each molecular variant 103, one or more two-dimensional analogs of the input molecule 101 in the one or more molecular libraries 125 stored at the data store 120. In some cases, the query controller 117 may generate at least one output molecule 105 corresponding to a two-dimensional analog of the input molecule 101 whose three-dimensional pharmacophore-shape profile exhibits a threshold degree of similarity (e.g., as quantified by a three-dimensional similarity metric such as Tanimoto Combo score) relative to the three-dimensional pharmacophore-shape profile of the input molecule 101.
[0061] FIG.2 depicts a flowchart illustrating an example of a process 200 for pharmacophore-based generative molecule design, in accordance with some example embodiments. Referring to FIGS.1-2, in some example embodiments, the process 200 may be performed by the molecule design engine 110 to search the one or more molecular libraries 125 for two-dimensional analogs of the input molecule 101. In some cases, the input molecule 101 may be known to bind to a biological target, meaning that the two-dimensional analogs of theAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 input molecule 101 may exhibit a high likelihood of also binding to the same biological target. In some cases, the process 200 leverages more compact structural representation of molecules to search the one or more molecular libraries 125 to reduce the concomitant resource overhead such that a comprehensive search of the one or more molecular libraries 125 may be conducted within the available time and compute budget. For example, in some cases, the molecule design engine 110 may apply the molecule design computation model 115 to generate, based at least on a three- dimensional pharmacophore-shape profile of the input molecule 101, one or more linear representations (e.g., simplified molecular input line entry system (SMILES) strings, self- referencing embedded stings (SELFIES), and / or the like) representative of one or more molecular variant 103 of the input molecule 101. In some cases, each linear representation may be converted to a molecular fingerprint that captures a two-dimensional representation (e.g., a molecular graph) of the corresponding molecular variant 103 such that the query engine 117 may search, based at least on the molecular fingerprint of each molecular variant 103, the one or more molecular libraries 125 for two-dimensional analogs of the input molecule 101.
[0062] By using the molecular fingerprints (or another two-dimensional representation) of the molecular variants 103 to search for two-dimensional analogs of the input molecule 101 in the molecular libraries 125, the process 200 imposes exponentially less computational overhead than conventional techniques that rely solely on comparisons of three- dimensional pharmacophore-shape profiles. In particular, to the extent the process 200 includes searching entire molecular libraries for the two-dimensional analogs of the input molecule 101, this search is conducted in two -dimensional space using more compact two-dimensional representations (e.g., molecular fingerprints) . With the process 200, more computationally expensive and protracted comparisons in three-dimensional space using three-dimensionalAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 pharmacophore-shape profiles are limited to a subset of molecules that have been identified as two-dimensional analogs. Doing so may increase the scope of the search (e.g., entire molecular libraries can be searched for two-dimensional analogs) without exceeding the compute budget while still ensuring that the pharmacophore-shape profile of the one or more output molecules 105 ultimately generated by the molecule design engine 110 are sufficiently similar to that of the input molecule 101.
[0063] At 202, a three-dimensional representation of an input molecule specifying a location of one or more pharmacophores present in the input molecule is received. In some example embodiments, the pharmacophore-shape profile of an input molecule may specify the location, in three-dimensional space, of one or more pharmacophores present in the input molecule. As noted, a pharmacophore may be a feature (e.g., an electronic feature, a steric feature, and / or the like) in the input molecule that may be involved in a binding interaction between the input molecule and a biological target (e.g., a site, such as a receptor, on a target protein). Examples of pharmacophores include hydrogen bond donors, hydrogen bond acceptors, aromatic rings, cations, anions, hydrophobic moieties, and / or the like. The spatial arrangement of pharmacophores may determine the biological activities of the input molecule including, for example, its ability to bind with and modulate the biological target. Accordingly, where the input molecule is known to bind to a biological target, additional molecules capable of binding to the same biological target may be generated by identifying molecules with similar pharmacophore-shape profile as the input molecule. Doing so may generate a diverse pool of binders, thus increasing the likelihood that at least one of those molecules exhibit minimal labilities (e.g., chemical or physical instabilities, poor pharmacokinetics, immunogenicity, and / or the like) and can be optimized into a successful therapeutic.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0064] At 204, a molecular variant of the input molecule is generated by at least applying a molecule design computation model that has been trained to operate on the three- dimensional representation of the input molecule and output a linear representation of the molecular variant. In some example embodiments, the molecule design computation model may be trained to generate molecular variants of the input molecule by at least operating on the pharmacophore-shape profile of the input molecule to generate molecular variants with a similar pharmacophore-shape profile. It should be appreciated that the pharmacophore-shape profile of the input molecule specifies the location, in three-dimensional space, of the pharmacophores present in the input molecule. For example, in some cases, the pharmacophore-shape profile of the input molecule may be a voxelized representation in which the input molecule is represented as densities (e.g., Gaussian-like densities) across a three-dimensional grid of voxels (or three- dimensional pixels), with each density being centered around the location of a corresponding pharmacophore present in the input molecule. In some cases, the voxelized representation of the input molecule may include a separate channel for each type of pharmacophore (e.g., hydrogen bond donors, hydrogen bond acceptors, aromatic rings, cations, anions, hydrophobic moieties, and / or the like). For instance, the voxelized representation of the input molecule may include a first channel depicting the spatial arrangement of a first pharmacophore type (e.g., hydrogen bond donors) as densities across a three-dimensional grid of voxels, with each density being centered around the location a first pharmacophore type (e.g., hydrogen bond donors) present in the input molecule. Furthermore, the voxelized representation of the input molecule may include a second channel depicting the spatial arrangement of a second type of pharmacophores (e.g., hydrogen bond acceptors) as densities across the three-dimensional grid of voxels, with eachAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 density being centered around the location of a second pharmacophore type (e.g., hydrogen bond acceptor) present in the input molecule.
[0065] In some example embodiments, the molecule design computation model may generate molecular variants of the input molecule that have a similar pharmacophore-shape profile as the input molecule but different molecular structures. For example, in some cases, the molecule design computation model may be trained to generate molecular structures that exhibit the pharmacophore-shape profile of the input molecule. It should be appreciated that the pharmacophore-shape profile of the input molecule specifies the types and spatial arrangement of pharmacophores present in the input molecule but not the molecular structure of the input molecule, which would be specified in terms of the types of atoms, bonds, rings, aromaticity, branching, stereochemistry, and isotopes present in the input molecule. Accordingly, in some cases, the molecule design computation model may be trained to determine, based at least on the three-dimensional pharmacophore-shape profile of the input molecule, the linear representation a corresponding molecular structure. In some cases, the linear representation may be a simplified molecular input line entry system (SMILES) string (or a self-referencing embedded string (SELFIES)) that specifies the types of atoms, bonds, rings, aromaticity, branching, stereochemistry, and isotopes that is present in a molecular structure with a similar pharmacophore-shape profile as the input molecule. Moreover, it is possible for multiple molecular structures to exhibit a similar pharmacophore-shape profile as the input molecule, meaning that the molecule design computation model may be applied to generate, based on the pharmacophore-shape profile of the input molecule, multiple molecular variants of the input molecule. That is, the molecule design computation model may be applied to the same pharmacophore-shape profile of the input molecule multiple times and generate a differentAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 molecular variant, as specified by a corresponding linear representations (e.g., SMILES string, SELFIES, and / or the like), each time. As described in more details below, the molecular variants of the input molecule may undergo screening, including by being matched to the molecules in one or more molecular libraries.
[0066] At 206, the linear representation of the molecular variant is translated into a two-dimensional representation of the molecular variant. In some example embodiments, the linear representation of each molecular variant of the input molecule may be converted into a molecular fingerprint. For example, in some cases, the linear representation (e.g., simplified molecular input line entry system (SMILES) string, self-referencing embedded string (SELFIES), and / or the like) of a molecular variant of the input molecule may be converted to a molecular fingerprint (e.g., a circular fingerprint such as a Morgan fingerprint, a path-based fingerprint, a RDKit fingerprint, and / or the like) that specifies the molecular structure of the molecular variant as a two-dimensional molecular graph. For instance, in some cases, the molecular fingerprint of a molecule variant of the input molecule may be a binary code in which each bit corresponds to the presence (or absence) of a chemical substructure or pattern within the molecule. As described in more details below, the molecular fingerprint of the molecular variant, which specifies the molecular structure of the molecular variant as a two-dimensional molecular graph, may be used to identify two-dimensional analogs of the input molecule with significantly less computational overhead than the three-dimensional comparisons required to determine the overlap between two three-dimensional pharmacophore-shape profiles. That is, in some cases, the molecular fingerprint of each molecular variant of the input molecule may be compared to the molecular fingerprint of every molecule in one or more molecular librariesAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 without exceeding the available compute budget whereas such a brute force search would have been an intractable task with conventional three-dimensional comparisons.
[0067] At 208, a molecule from a molecular library is identified as a two-dimensional analog of the input molecule based at least on a two-dimensional representation of the molecule exhibiting a threshold degree of similarity to the two-dimensional representation of the molecular variant of the input molecule. In some example embodiments, the set of two-dimensional analogs of the input molecules may include those molecules in the one or more molecular libraries whose molecular fingerprints (e.g., circular fingerprints such as a Morgan fingerprints, a path-based fingerprint, a RDKit fingerprint, and / or the like) are sufficiently similar to the molecular fingerprint of a molecular variant of the input molecule generated by the molecule design computation model. For example, in some cases, a two-dimensional similarity metric may be computed to quantify the similarity between the molecular fingerprint of each molecule in the one or more molecular libraries and the molecular fingerprint of a molecular variant of the input molecule. Examples of the two-dimensional similarity metric may include Tanimoto index, Jaccard coefficient, cosine coefficient, Dice metric, Euclidean metric, city-block metric, Hamming index, Tversky index, and / or the like. In some cases, the set of two-dimensional analogs of the input molecule may include those molecules from the one or more molecular libraries whose similarity metric satisfies one or more criteria. For instance, in some cases, the set of two-dimensional analogs of the input molecule may include molecules from the one or more molecular libraries whose similarity metric satisfies one or more thresholds. Alternatively and / or additionally, in some cases, the set of two-dimensional analogs of the input molecule may include a threshold quantity of molecules from the one or more molecular libraries with the highest similarity metric.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0068] At 210, the two-dimensional analog of the input molecule is generated as an output molecule based at least on a three-dimensional representation of the two-dimensional analog exhibiting a threshold degree of similarity to the three-dimensional representation of the input molecule. In some example embodiments, the two-dimensional analogs of the input molecule may undergo further screening to ensure sufficient similarity with the three- dimensional pharmacophore-shape profile of the input molecule. As noted, in some cases, one or more molecular libraries may be searched for two-dimensional analogs of the input molecule. The two-dimensional analogs of the input molecule are those molecules in the one or more molecular libraries whose molecular fingerprint exhibits sufficient similarity to the molecular fingerprint of the input molecule. For example, in some cases, a two-dimensional analog of the input molecule may be a molecule whose two-dimensional similarity metric, which quantifies the similarity between the molecular fingerprint of the molecule and the molecular fingerprint of the input molecule, satisfies one or more criteria. In some cases, a molecule from the one or more molecular libraries identified as the two-dimensional analog of the input molecule may be synthesized, purified, and / or analyzed, for example, using the one or more laboratory equipment 130 shown in FIG.1. For instance, the input molecule 101 may be a known binder of a biological target. Accordingly, the two-dimensional analog of the input molecule 101 may undergo synthesis, which may include purification, at least because the two-dimensional analog of the input molecule 101 may be more likely to be a binder of the same biological target. In some cases, the molecules resulting from the synthesis of the two-dimensional analog of the input molecule 101 may undergo various analysis, for example, to determine one or more properties of the two-dimensional analog of the input molecule 101. In some cases, theAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 synthesized molecules may be analyzed for various drug-like properties, such as binding affinity, specificity, biological activity, developability, and / or the like.
[0069] While two-dimensional comparisons may be performed quickly and with little computational resources, it is possible for a two-dimensional analog of the input molecule identified in this manner to lack sufficient similarity to the input molecule in three-dimensional space. Accordingly, in some cases, the two-dimensional analogs of the input molecule may undergo further screening that includes a comparison of respective three-dimensional pharmacophore-shape profiles. For instance, in some cases, the two-dimensional analogs of the input molecules may undergo further screening in which the pharmacophore-shape profile of each two-dimensional analog with the pharmacophore-shape profile of the input molecule before a three-dimensional similarity metric is computed to quantify the similarity therebetween. In some cases, a molecule that is identified as a two-dimensional analog of the input molecule 101 may be synthesized, purified, and / or analyzed upon further determining that the three- dimensional similarity metric quantifying the three-dimensional structural similarity between the molecule and the input molecule 101 satisfies one or more thresholds.
[0070] In some cases, the three-dimensional similarity metric for two molecules may be a hybrid similarity metric, such as a Tanimoto Combo score, that measures the steric similarity (or the similarity between the arrangement of atoms in three-dimensional space) between the two molecules as well as the spatial overlap in the functional groups that are present in each molecule. Comparing two molecules in three-dimensional space, which includes computing the aforementioned three-dimensional similarity metric, may consume significant time and computational resources. As such, various embodiments of the present disclosure reduces the number of three-dimensional comparisons to the two-dimensional analogs of theAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 input molecule, which can be identified quickly and with minimal computational resources through two-dimensional comparisons (e.g., similarity metric between two molecular fingerprints).
[0071] As noted, in some example embodiments, a molecule may be represented in a one-dimensional linear form (e.g., as a simplified molecular-input line-entry system (SMILES) string, a self-referencing embedded string (SELFIES), and / or the like), as a two-dimensional graph, or as densities across a three-dimensional grid of voxels. Moreover, in some cases, molecules having a similar three-dimensional pharmacophore-shape profile as an input molecule may be identified by searching one or more molecular libraries for two-dimensional analogs of the input molecule before these two-dimensional analogs are screened for three-dimensional conformity with the input molecule. For example, as shown in FIG.2, the one or more molecular libraries may be searched using two-dimensional representations (e.g., circular molecular fingerprints such as Morgan fingerprints, path-based fingerprints, RDKit fingerprints, and / or the like) before the three-dimensional voxelized representations of two-dimensional analogs are compared to verify three-dimensional conformity.
[0072] To further illustrate, FIG.3A depicts an example of a two-dimensional molecular graph 300 representative of a molecule while the corresponding three-dimensional shape profile 310 and pharmacophore-shape profile 350 are shown in FIGS.3B and 3C, respectively. In some cases, the one or more molecular libraries may be searched using the two- dimensional molecular graph 300 while the two-dimensional analogs identified therefrom may be further screened using a combination of the three-dimensional shape profile 310 and the pharmacophore-shape profile 350. As shown in FIGS.3B-C, the three-dimensional shape profile 310 and the pharmacophore-shape profile 350 are voxelized representation of the molecule,Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 which depicts the molecule as densities (e.g., Gaussian-like densities) across a three-dimensional grid of voxels. In the three-dimensional shape profile 310, the densities are centered around the location of each constituent atom in three-dimensional space. In the pharmacophore-shape profile 350, the densities are centered around the location of each pharmacophore in the molecule.
[0073] In some example embodiments, the voxelized representation of the molecule may include multiple channels including, for example, one channel depicting the locations of the individual atoms (e.g., as in the three-dimensional shape profile 310) and one channel depicting the location of each type of pharmacophore. In the example shown in FIG.3C, the pharmacophore-shape profile 350 includes six different types of pharmacophores: hydrogen bond donor, hydrogen bond acceptor, cation, anion, aromatic ring, and hydrophobe. Accordingly, in some cases, the voxelized representation of the molecule may include a first channel depicting the locations of individual atoms, a second channel depicting the locations of hydrogen bond donors, a third channel depicting the locations of hydrogen bond acceptors, a fourth channel depicting the locations of anions, a fifth channel depicting the locations of cations, a sixth channel depicting the locations of aromatic rings, and a seventh channel depicting the locations of hydrophobe.
[0074] Referring again to FIGS.3B-C, in some cases, the voxelized representation ofthe molecule may a voxel grid containing ^^ × ^^ × ^^ voxels (e.g., 32 × 32 × 32 voxels,64 × 64 × 64 voxels, and / or the like). In some cases, each voxel in the voxel grid may beassociated with a value indicative of the density at the corresponding location. For example, in the three-dimensional shape profile 310 shown in FIG.3B, a first voxel associated with a higher density value may be more likely to be a portion of an atom than a second voxel associated withAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 a lower density value. Meanwhile, in the pharmacophore-shape profile 350 shown in FIG.3C, a first voxel associated with a higher density value may be more likely to be a portion of a pharmacophore than a second voxel associated with a lower density value. It should be appreciated that the volume of an individual atom or pharmacophore may span one or multiple voxels. That the densities are centered around the atoms or pharmacophores present in a molecule means that densities values will peak at the center of an atom or pharmacophore and decrease incrementally at locations away from the center. For instance, a voxel having a density value of 0 may be far away from the center of any atoms or pharmacophores in the molecule whereas a voxel having a density value of 1 may be at the center of an atom or pharmacophore in the molecule.
[0075] In some example embodiments, the voxelized pharmacophore-shape profile of a molecule may be generated based on the conformation of the molecule. For example, in some cases, the location of one or more pharmacophores present in an input molecule may be computed given an input molecular conformation, resulting in a point cloud in which the locations of the points corresponds to the locations of each type of pharmacophore (e.g., hydrogen bond donors, hydrogen bond acceptors, aromatic rings, cations, anions, hydrophobic moieties, and / or the like). In some cases, a shape point cloud may also be computed for the input molecular conformation. In the shape point cloud, the locations of the constituent points may correspond to the locations of the atoms in the input molecule. In some cases, the input molecule may be associated with a set of point clouds. For instance, where there are six different types ofpharmacophores, the input molecule may be associated with a set of 7 point clouds ^^ = ^^^^^^^ୀ^ .The ^^-th point cloud may ே^be given by ^^^ = ^^^^,^^^ୀ^ , wherein ^^^,^ denotes the ^^-thpoint cloud ^^^and ^^^denotes the number of points in the point cloud ^^^.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0076] The pharmacophore and shape point clouds may be voxelized such that the constituent pharmacophores and atoms in the input molecule are represented as densities (e.g., Gaussian-like densities) across a three-dimensional voxel grid. For example, in some cases, the voxelization may include converting each point ^^^,^into a three-dimensional density (e.g., athree-dimensional Gaussian-like density) as shown in Equation (1).మ^^^^,^^^^, ^^^ = exp ^− ௗ^.ଽଷ∙^^మ^ (1)whereinof ^^ fromthe center of the volume. The radius of all points ^^ may be ^^ = 1. Accordingly, Equation (2)shows that the voxelization of pharmacophore and shape point clouds may include computingthe occupancy Occ of each voxel in the voxel grid at each grid point (^^, ^^, ^^) in each channel ^^.Occ^,^,^,^ = 1 − ∏ே^^ୀ^ ^1 − ^^^^,^൫ฮ^^^,^,^ − ^^^^,^ฮ, ^^൯^ (2)whereinofthe corresponding point ^^^,^in the point cloud. It should be appreciated that ^^ indexes both channels and point clouds, meaning that points in the point cloud ^^^are assigned to channel ^^. Moreover, each voxel in the voxel grid may be associated with a density value between 0 and 1, with a voxel having a density value of 0 being located away from the center of any atoms or pharmacophores and a voxel having a density value of 1 being located at the center of an atom or a pharmacophore present in the molecule.
[0077] FIG.4 depicts a flowchart illustrating an example of a process 400 for generating a molecule, in accordance with some example embodiments. Referring to FIGS.1 and 4, the process 400 may be performed by the molecule design engine 110. As described in more details below, in some cases, the process 400 may include the molecule design engine 110Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 applying the molecule design computation model 115 to generate one or more molecular variants 103 of the input molecule 101 before the query controller 117 searches the one or more molecular libraries 125 for two-dimensional analogs of the input molecule 101 that exhibit a threshold degree of similarity (e.g., as quantified by a two-dimensional similarity metric) relative to the molecular fingerprint of the input molecule 101. In some cases, the output molecule 105 may be a two-dimensional analog of the input molecule 101 that further screening identified as having a sufficiently similar pharmacophore-shape profile (e.g., as quantified by a three- dimensional similarity metric) as the input molecule 101.
[0078] Referring again to FIG.4, in some cases, the process 400 may include a molecular design computation model being applied to operate on the pharmacophore-shape profile ^^ of an input molecule ^^. In some cases, the input molecule ^^ may be a ligand known to bind to a biological target. In some cases, the pharmacophore-shape profile ^^ of the input molecule ^^ may be a voxelized representation in which the locations of the pharmacophores present in the input molecule ^^ are represented as densities (e.g., Gaussian-like densities) across a three-dimensional voxel grid. In some cases, the molecular design computation model may be trained to predict the structures of molecules that can also exhibit the pharmacophore shape profile ^^ of the input molecule ^^. In doing so, the molecular design computation model may generate an ^^^number of molecular variants of the input molecule ^^, each of which being amolecule that can exhibit the same or similar pharmacophore-shape profile as the input molecule^^ and is thus likely to bind to the same biological target as the input molecule ^^. For example,in some cases, the molecular design computation model may generate, based at least on the pharmacophore ^^ of the input molecule ^^, the linear representation of each one of the ^^^number of molecular variants. It should be appreciated that two molecules with different typesAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 of atoms, bonds, rings, aromaticity, branching, stereochemistry, and isotopes may still exhibit the same or similar pharmacophore-shape profiles. Accordingly, the ^^^number of molecular variants may form the basis for a more diverse pool of binders for further drug development efforts.
[0079] In some example embodiments, a search may be performed in order to identify, in one or more molecular libraries, database molecules with sufficient two-dimensional similarity (e.g., as quantified by a two-dimensional similarity metric) to be considered a two- dimensional analog of the input molecule ^^. For example, FIG.4 shows that a molecular fingerprint (e.g., a circulate fingerprint such as a Morgan fingerprint, a path-based fingerprint, a RDKit fingerprint, and / or the like) may be computed for each one of the ^^^number of molecular variants. Furthermore, for every ^^-th molecular variant of the input molecule ^^, the molecular fingerprint of that molecular variant may be compared to the molecular fingerprint of at least a portion of database molecules in the one or more molecular libraries. For instance, where there are an ^^ number of database molecules, the molecular fingerprint of the ^^-th molecular variant may be compared to the molecular fingerprint of each one of the ^^ number of database molecules. In some cases, this comparison may include computing a two-dimensional similarity metric (e.g., a Tanimoto index and / or the like) in order to quantify the two-dimensional similarity between the ^^-th molecular variant and the ^^-th database molecule. In some cases, this comparison may be performed for every one of the ^^ number of database molecules and repeated for every one of the ^^^number of molecular variants.
[0080] In some example embodiments, a set of two-dimensional analogs of the input molecule ^^ present in the one or more molecular libraries may be identified based on the results of comparing the molecular fingerprints of each one of the ^^^number of molecular variantsAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 against the molecular fingerprints of every one of the ^^ number of database molecules. For example, FIG.4 shows one example in which an ^^^number of database molecules with the highest two-dimensional similarity (e.g., Jaccard similarity) relative to one of the ^^^number of molecular variants are identified two-dimensional analogs of the input molecule ^^. That is, those ^^^number of database molecules may be identified as being most similar to the input molecule ^^ and may thus undergo further screening to confirm three-dimensional conformity with the input molecule ^^. For instance, FIG.4 shows a three-dimensional similarity metric (e.g., a 3D Tanimoto Combo score) being computed for each one of the ^^^number of database molecules identified as being a two-dimensional analog of the input molecule ^^. In some cases, this three-dimensional similarity metric may quantify the similarity between the three- dimensional pharmacophore-shape profile of a two-dimensional analog and the pharmacophore shape profile ^^ of the input molecule ^^. In some cases, a two-dimensional analog of the input molecule ^^ whose pharmacophore-shape profile exhibits a sufficient degree of similarity (e.g., as quantified by the 3D Tanimoto Combo score) may be added to a binder pool 450 for further development. In some cases, for example, one or more molecules from the binder pool 450 may undergo synthesis, purification, and / or analysis, for example, using the one or more laboratory equipment 130 shown in FIG.1.
[0081] It should be appreciated that the molecules added to the binder pool 450 are likely to bind to the same biological target as the input molecule ^^ due to the similarity in pharmacophore-shape profile. That is, the binder pool 450 is generated to include other possible binders of the same biological target as the input molecule ^^. Moreover, due to the molecule design computation model being applied to generate molecular variants with a similar pharmacophore-shape profile as the input molecule ^^, the molecules in the final binder pool 450Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 remain diverse in terms of chemical substructures and patterns. Such diversity is advantageous in the context of drug development at least because it increases the likelihood that at least one of these molecules will bind to the biological target but exhibit minimal labilities (e.g., chemical or physical instabilities, poor pharmacokinetics, immunogenicity, and / or the like). Various embodiments of the present disclosure increases the diversity of the binder pool 450 with significantly less computational overhead than conventional approaches that rely on brute force three-dimensional searches across molecular databases. For example, the number of three-dimensional comparisons is reduced from ^^(^^) to ^^(^^^ × ^^^), a difference that can be severalorders of magnitude in size depending on the number ^^ of database molecules. In instanceswhere the molecular libraries contain hundreds of millions to billions of compounds, ^^^ × ^^^may be on the order of 1000 instead. Thus, where a brute force search against a full molecular library is intractable, especially where the number ^^^of molecular variants is in the hundreds or thousands, a search according to various embodiments of the present disclosure is significantly faster and cheaper computationally.
[0082] In some example embodiments, the molecule design computation model may be an encoder-decoder model that has been trained to ingest the voxelized pharmacophore-shape profile ^^ of the input molecule ^^. In some cases, the encoder may compress the voxelized pharmacophore-shape profile ^^ into a vector representation that is then decoded into a linear representation (e.g., a simplified molecular-input line-entry system (SMILES) string, a self- referencing embedded string (SELFIES), and / or the like) of a molecular variant exhibiting the pharmacophore-shape profile ^^ as the input molecule ^^ but different chemical substructures and patterns. FIG.5A depicts a block diagram illustrating the architecture of an example of an encoder 500 for implementing the molecule design computation model. In the example shown inAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 FIG.5A, the encoder 500 may have a three-dimensional convolutional neural network (CNN) based architecture. In some cases, the three-dimensional convolutional block in the encoder 500 may be applied to the voxels in the voxelized pharmacophore-shape profile ^^ an ^^ number of times before undergoing pooling (e.g., three-dimensional average pooling) to generate the vector representation. In some cases, the vector representation of the voxelized pharmacophore-shape profile ^^ may be used to initialize the hidden state of a long-short term memory (LSTM) decoder. The same vector representation of the voxelized pharmacophore-shape profile ^^ may be passed to the LSTM decoder at every timestep to be concatenated with the embedded token from that timestep. An example of a decoder 550 with a long-short term memory (LSTM) neural network based architecture is shown in FIG.5B.
[0083] In some example embodiments, the training of the molecule design computation model with an encoder-decoder architecture include training the encoder to generate a vector representation of the voxelized pharmacophore-shape profile ^^ of a training molecule ^^ that the decoder can then decode into the ground truth linear representation (e.g., simplified molecular input line entry system (SMILES) string, self-referencing embedded string(SELFIES), and / or the like) of the training molecule ^^. For example, given a training molecule^^ from a training dataset ^^ = ^^^^^ே^ୀ^ with an ^^ number of molecules, the tokenized linearrepresentation of the training molecule ^^ may be a string sequence u = (^^(^),⋯ ,^^(்)), where ^^denotes the length of the sequence and each element of u is an integer ^^ ∈ ℤ: 1 ≤ ^^ ≤ ^^representative of a token from the vocabulary ^^ = ^^^^^ோ^ୀ^ of size ^^. Each vocabulary element^^^ may be a sequence of characters. Theshape profile ^^ of the training molecule^^ may be computed, using a deterministic non-differentiable function ^^క, as ^^ = ^^క(^^) ∈Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 ℝ^×ௗ×ௗ×ௗ, with the number of channels ^^ = 7 and each side of the voxel grid having aof ^^ (or ^^ number of voxels).
[0084] The objective for training the molecule design computation model may be to predict the ground truth linear representation (e.g., SMILES string, SELFIES, and / or the like) sequence u from the voxelized pharmacophore-shape profile ^^ of the training molecule ^^. In some cases, the task may be to increase (or maximize) the likelihood of data with respect to the parameters ^^ of the molecular design computation model. With an autoregressive model, each datapoint may be factorized in accordance with Equation (3).(௧ା^) (^:௧) ^^(^^|^^; ^^) = ∏௧்ୀ^ ^^ ൫^^ ห^^ , ^^; ^^൯ (3)ℒ(^^|^^; ^^) = − log ^^(^^|^^;^^) = ∑ log^^൫^^(௧ା^)ห^^(^:௧)௧்ୀ^ , ^^;^^൯ (4)
[0086] According to Equation (4), in some cases, the molecule design computation model may be trained using teacher forcing. This training paradigm includes feeding the ground truth (actual) linear representation (e.g., SMILES string, SELFIES, and / or the like) sequence u as input for the next time step instead of the model's own previous output. Doing so may help the molecule design computation model learn the correct sequence of output tokens more efficiently as well as stabilize the training process. For each, at each timestep ^^, the molecule design computation model may use the softmax function to output a distribution ()భ:^ ^^୮(^൫^^ ,క;ఏ൯ )(௧ା^) (^:௧) ^ ∗ ^×ௗ×ௗ×ௗ ோ( ) = ^^ , = , in which ^^ . ; ^^ : (ℝ × ℝ ) → ℝ is a^^×ௗ×ௗ×ௗvoxel grid ℝ (with ^^ channels and a dimensionality of ^^) to a logit for each element in () the vocabulary ^^ and ^^ . ; ^^ denotes the logit that is output by this function for ^^ .^Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0087] In some cases, each training sample may undergo augmentation. In some cases, each training sample may undergo random augmentation. For example, in some cases, this augmentation may include a translation (e.g., a random translation) by a value uniformlysampled from [−1,1] for each axis, and a random rotation by a value uniformly sampled from[0, 2^^) for each of the three Euler angles. This augmentation may be applied to the point cloudsrepresentative of the atoms in each training molecule ^^ prior to voxelization.
[0088] At generation time, the trained molecule design computation model may ingest, as input, the voxelized pharmacophore-shape profile ^^ of an input molecule ^^. Sampling (e.g., ancestral sampling) may be used to generate linear representations s (e.g., SMILES strings, SELFIES, and / or the like) representative of molecular variants of the input molecule ^^, with adistribution ^^^ over the vocabulary ^^ that is annleaed by a temperature parameter ^^ ∈ ℝ: ^^ ≥ 1,in accordance with Equation (5). ()^^ ~^^^൫^^ ^) = ^^ หs(^:௧), ^^;^^ ^^୮(^൫^ భ:^(௧ା^) (௧ା^ ൯ = ,క;ఏ൯^⁄ ఛ )(భ:^) ⁄ (5)
[0090] Various embodiments of the molecule design computation model described herein may be deployed in different generative workflows. One example generative workflow is shown in FIGS.2 and 4, and includes a streamlined search of one or more molecular libraries that relies primarily on two-dimensional comparisons (e.g., of two-dimensional molecular fingerprints) for molecular variants generated by the molecule design computation model that have a similar pharmacophore-shape profile as an input molecule. Another example generative workflow is a de novo generative workflow in which the molecule design computation model is applied to generate molecular variants of an input molecule. Experiments were carried out to evaluate the performance of the molecule design computation model in both workflows.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0091] De Novo Generation from Pharmacophore-Shape Profile Input
[0092] Experiments were performed to assess the ability of the molecule design computation model described herein to generate a diverse set of molecules matching an input pharmacophore-shape profile for the purpose of ligand-based drug discovery (LBDD). In this context, a Tanimoto Combo score is used to quantify and evaluate the overlap between thepharmacophore-shape profiles of two molecules. This metric is abstracted as a functionTC(∙,∙):ℳ × ℳ → [0,2], wherein ℳ denotes the space of all possible molecules with theirconformations. The function TC ingests two molecules and outputs a score between 0 and 2, with 0 indicating no overlap between the pharmacophore-shape profiles of the two molecules and 2 indicating a perfect overlap.
[0093] This evaluation was carried out on two datasets: Geometric Ensemble Of Molecules (GEOM) drugs and the subset of ChEMBL used for machine learning. The GEOM drugs include approximately 430,000 drug-like molecules with up to 181 heavy atoms of the type {carbon (C), oxygen (O), nitrogen (N), fluorine (F), sulfur (S), chlorine (Cl), bromine (Br), phosphorus (P), iodine (I), and boron (B)}. The median number of heavy atoms per molecule is 45 and more than 99% of molecules have less than 80 heavy atoms. The GEOM dataset includes multiple conformations for each molecule. After experimenting with different resolutions and grid sizes, and the performance of the molecule design computation model on the GEOM dataset is assessed using a grid of dimension 48ଷwith a resolution of 0.35 Å. This volume covers over 99.8% of all compounds in the GEOM dataset. After preprocessing, the GEOM dataset is split 1.1M / 146K / 146K train / validation / test.
[0094] The ChEMBL (subset) dataset includes approximately 1,600,000 drug-like molecules with up to 88 heavy atoms of type {carbon (C), oxygen (O), nitrogen (N), fluorine (F),Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 sulfur (S), chlorine (Cl), bromine (Br), phosphorus (P), iodine (I), boron (B), selenium (Se), and silicon (Si)}. The median number of heavy atoms per molecule is 27. As the ChEMBL dataset is given as simplified molecular-input line-entry system (SMILES) string s, RdKit was used to generate the three-dimensional conformation for each molecule by optimizing the initial coordinates of the atoms of each molecule. The performance of the molecule design computation model on the ChEMBL dataset is assessed using a grid of dimension 64ଷwith a resolution of 0.35 Å. The ChEMBL is split 1.3M / 80k / 240k train / validation / test.
[0095] For each dataset, the performance of the molecule design computation model is assessed on 100 randomly selected input molecules from the test set. A three-dimensional conformation (e.g., a three-dimensional pharmacophore-shape profile) is computed with each input molecule. As a baseline, for each input molecule, 1000 molecules from the test set were randomly sampled. Multiple three-dimensional conformations (e.g., three-dimensional pharmacophore-shape profiles) were obtained for each sampled molecule and overlayed with that of the input molecule to compute a Tanimoto Combo score. For each sampled molecule, the three-dimensional conformations with the highest Tanimoto Combo score relative to the input molecule were retained. A hit in this case was defined as a sampled molecule ^^^^୬with aTanimoto Combo score TC(^^^^୬,^^୧୬୮^^) ≥ 1.2 with respect to the input molecule ^^୧୬୮^^.Four metrics were considered: number of hits per input molecule, number of hits per input molecule containing a unique scaffold, max Tanimoto Score per input molecule, and number of input molecules with at least one hit.
[0096] The performance of the molecule design computation model is evaluated against that of the Pharmacophore Guided Molecular Generation (PGMG) model based on a conditional variational autoencoder architecture. At training time, the PGMG model learns toAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 denoise a corrupted SMILES string conditioned on a pharmacophore point cloud that is computed from the uncorrupted SMILES string. At inference time, the PGMG model generates SMILES strings corresponding to a given pharmacophore point cloud. The point cloud is encoded using a Gated Graph Convolutional Neural Network that is equivariant to rotation before a transformer architecture is used to process the encoded representation along with the SMILES string.
[0097] The pharmacophore-shape profile ^^ was computed for each candidate molecule and passed as input to the molecule design computation model. Several SMILESstrings were sampled from the distribution ^^(s|^^;^^, ^^) using different values of the temperatureparameter ^^ as well as top-^^ sampling. For the PGMG model, the SMILES string of the candidate molecule was passed to the model along with a pharmacophore-shape profile computed from the SMILES string before several SMILES strings of molecular variants were sampled. Invalid and duplicate molecules generated by either model were removed prior to evaluation. Generated molecular variants that are identical to the input molecule were also removed because the purpose of this task is to generate novel, diverse molecules with a similar pharmacophore-shape profile as the candidate molecule and not to reproduce the candidate molecule. For each model, the first one thousand valid, non-duplicate generated molecule variants of each candidate molecule were selected. The four aforementioned metrics were determined for the three-dimensional pharmacophore-shape profiles of these generated molecular variants. In this way, we compare approaches using the same budget for three- dimensional evaluation, which is the time-consuming part of the process, compared to relatively cheaper forward passes through neural networks.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0098] The validation sets were used for early stopping and model selection based on the total number of unique scaffold hits in a small set of 100 generated samples for each of 10 randomly selected validation set query molecules.
[0099] Results for the GEOM-drugs and ChEMBL subset are given in Table 1 while Table 2 shows the mean and standard deviation summary statistics. Metrics corresponding to the dataset baseline are very low, showing that two randomly selected molecules from the same marginal distribution are expected to have low overlap. The molecule design computation model (MDCM) described herein outperformed PGMG by large margins across all metrics on both datasets. In contrast to PGMG, the molecule design computation model (MDCM) produced at least one hit for almost every query molecule in each test set. Of note, along with the high number of hits, is the high number of unique scaffold hits, showing that the molecule design computation model (MDCM) produced a diverse set of hypotheses with matching pharmacophore-shape profiles, providing diverse avenues for ligand exploration corresponding to a given input molecule. The top score among all generated ligands, averaged over all input molecules, was also higher for the molecule design computation model (MDCM) than for PGMG across both datasets. The molecule design computation model (MDCM) described herein was shown as being able to generate hits for input molecules drawn from the distribution on which it was trained. This in-distribution evaluation relates to drug-discovery campaigns where lead compounds are similar to library compounds that the campaign has access to, or has produced before, which have been used to train the model.
[0100] Table 1: De-novo generation where the input molecules are drawn from a test set with the same distribution as the training set. Values for hits, unique scaffold hits, and max score are medians across the 100 input molecules.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 Method Dataset Hits Unique Scaffold Hits Max # Inputs with ≥ 1 hit Dataset baseline 0 0 1.16 43 GEOM-dru sst set with the same distribution as the training set. Values for hits, unique scaffold hits and max score are means ± standard deviations across the 100 input molecules. Method Dataset Hits Unique Scaffold Hits Max Dataset baseline 1.37 ± 2.88 1.33 ± 2.75 1.18 ± 0.17computation model (MDCM) are comparable in all cases, showing that the superior performance of the molecule design computation model (MDCM) is not merely attributable to an extremely high contribution to the metrics from a small number of input molecules and low contribution from the rest, but is instead due to high contributions across many input molecules. The standard deviation across the 100 queries is, however, high for both the hits and unique scaffold hits metrics. Bringing this number down would indicate more predictable performance for a given input molecule.
[0103] Table 3 and Table 4 shows the results of measuring out-of-distribution performance. For this task, the models trained on GEOM-drugs are evaluated on input molecules from the ChEMBL subset, and vice-versa. The molecule design computation model (MDCM) once again outperformed PGMG across all metrics by large margins. Accordingly, theAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 molecule design computation model (MDCM) is able to generate hits for input molecules drawn from a distribution that is different to the one on which it was trained. This relates to drug- discovery campaigns where lead compounds are significantly different from the compounds that the campaign has access to, or has produced before. Results for models trained on ChEMBL are better than the corresponding models trained on GEOM-drugs, possibly because of the larger number of molecules in the ChEMBL subset.
[0104] Table 3: De-novo generation where the input molecules are drawn from a test set with a different distribution than the training set. Values for hits, unique scaffold hits and max score are medians across the 100 input molecules. Method Training / Test Dataset Hits Unique Scaffold Hits Max # Inputs with ≥ 1 hit PGMGChEMBL / GEOM-dr128 1.40 73st set with a different distribution than the training set. Values for hits, unique scaffold hits and max score are means ± standard deviations across the 100 input molecules. Method Training / Test Dataset Hits Unique Scaffold Hits Max PGMG5121 90042773 4386 139 030
[0107] The accuracy of a generative workflow in which molecular variants generated by the molecule design computation model are matched to molecules in one or more molecular libraries. In this particular generative workflow, one or more molecular libraries are searched for two-dimensional analogs of an input molecule. By limiting three-dimensional comparisons to aAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 smaller number of two-dimensional analogs with the highest two-dimensional similarity to the input molecule, this generative workflow imposes significantly less computational overhead than conventional brute force three-dimensional searches, which can be intractable for larger libraries such as the Enamine Real and Enamine Diverse databases.
[0108] To assess the accuracy of the fast molecular library search based generative workflow, 11 input molecules were randomly selected from the ChEMBL test set. For a brute force comparison, a three-dimensional Rapid Overlay of Chemical Structures (ROCS) based search was performed against every molecule in the union of the train and validation sets excluding the input molecule. Meanwhile, the molecule design computation model (MDCM) described herein was applied to generate several molecular variants given the input molecule’spharmacophore-shape profile before finding (i) ^^^ = 5 two-dimensional analogs for each of upto ^^^ = 100 unique, valid molecular variants generated by the molecule design computationmodel, or (ii) ^^^ = 1 two-dimensional analogs for each of up to ^^^ = 500 molecular variantsgenerated by the molecule design computation model. The ROCS scores of these two- dimensional analogs were computed against the input molecule. Moreover, the number of hits in this set of analogs were computed and compared against the number of hits in the set of top 500 molecules by ROCS score from the brute force search performed for each query molecule. The results are shown in Table 5.
[0109] Table 5: Fast database search results. Unique stands for unique scaffold hits. ROCS comparisons / query are given including the constant ^^, which is the average number of conformers per database molecule. Values for hits, unique scaffold hits, and max score are medians across the 11 input molecules. Method ^^^^^^Hits Unique Max Inputs w / ≥ 1 hit ROCS comparisons / inputAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 MDCM 100 5 10 9 1.58 10 α×500 MDCM 500 1 10 8 1.58 11 α×500
[0110] As the results in Table 5 library search basedgenerative workflow yielded significant numbers of hits and unique scaffold hits; with the two different settings of ^^^and ^^^yielding similar results. The number of ROCS comparisons perinput molecule (up to a constant, ^^, the average number of conformers per database molecule) is^^^ × ^^^. This is several orders of magnitude below that of a brute force search, which is the sizeof the database. The difference in ROCS comparisons would be even more pronounced in a real- world drug discovery application for large molecular libraries such as Enamine Real and Enamine Diverse. At the same time, a greater number of two-dimensional analogs and hits from the molecular variants generated by the molecule design computation model (MDCM) should also be found owing to the larger size of real-world molecular libraries. Accordingly, the fast molecular library search based generative workflow described eliminates a significant bottleneck in the current drug design pipeline by enabling a quick search of large-scale databases for sets of diverse hits with minimal computational overhead.
[0111] While the molecule design computation model identified fewer hits than conventional brute force three-dimensional searches, this may be due to the two-dimensional similarity between two molecules being generally low. This phenomenon was observed in the highest two-dimensional Tanimoto index for any given molecular variant being generally low,with a median top-1 similarity of 0.38 for the ^^^ = 500 and ^^^ = 1 setting. Consequently,many of the molecular variants generated by the molecule design computation model do not have close two-dimensional matches in the dataset. Graph 600 in FIG.6 shows a moderate correlation(Pearson ^^ = 0.4) between two-dimensional and three-dimensional similarities. The lack ofclose two-dimensional matches for many of the molecular variants generated by the moleculeAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 design computation model could reduce the number of hits. An additional contributor may be the larger number of duplicate two-dimensional hits per input molecule, where on average 2.4molecular variants are mapped to the same database molecule for the ^^^ = 500 and ^^^ = 1setting, thereby decreasing the number of hits that can be found during subsequent three- dimensional comparisons.
[0112] FIG.7 depicts a block diagram illustrating an example of a computing system 700, in accordance with some example embodiments. Referring to FIGS.1-7, the computing system 700 may be used to implement the molecule design engine 110, the data store 120, the one or more laboratory equipment 130, the client device 140, and / or any components therein.
[0113] As shown in FIG.7, the computing system 700 can include a processor 710, a memory 720, a storage device 730, and input / output devices 740. The processor 710, the memory 720, the storage device 730, and the input / output devices 740 can be interconnected via a system bus 750. The processor 710 is capable of processing instructions for execution within the computing system 700. Such executed instructions can implement one or more components of, for example, the molecule design engine 110, the data store 120, the one or more laboratory equipment 130, the client device 140, and / or the like. In some example embodiments, the processor 710 can be a single-threaded processor. Alternately, the processor 710 can be a multi- threaded processor. The processor 710 is capable of processing instructions stored in the memory 720 and / or on the storage device 730 to display graphical information for a user interface provided via the input / output device 740.
[0114] The memory 720 is a computer readable medium such as volatile or non- volatile that stores information within the computing system 700. The memory 720 can store data structures representing configuration object databases, for example. The storage device 730Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 is capable of providing persistent storage for the computing system 700. The storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 740 provides input / output operations for the computing system 700. In some example embodiments, the input / output device 740 includes a keyboard and / or pointing device. In various implementations, the input / output device 740 includes a display unit for displaying graphical user interfaces.
[0115] According to some example embodiments, the input / output device 740 can provide input / output operations for a network device. For example, the input / output device 740 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0116] In some example embodiments, the computing system 700 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, the computing system 700 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via the input / output device 740. The user interface can be generated and presented to a user by the computing system 700 (e.g., on a computer screen monitor, etc.).Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2
[0117] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0118] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object- oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.
[0119] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0120] In the descriptions above and in the claims, phrases such as “at least one of” or “one or more of” may occur followed by a conjunctive list of elements or features. The term “and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at least one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to meanAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 “A alone, B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.
[0121] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desired results. Other implementations may be within the scope of the following claims.
Claims
Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 CLAIMS What is claimed is:
1. A computer-implemented method, comprising: receiving a three-dimensional representation of an input molecule, wherein the three- dimensional representation of the input molecule specifies a location of one or more pharmacophores present in the input molecule; generating a first molecular variant of the input molecule by at least applying a molecule design computation model that has been trained to operate on the three-dimensional representation of the input molecule and output a linear representation of the first molecular variant; translating the linear representation of the first molecular variant into a two-dimensional representation of the first molecular variant; identifying a first molecule from a molecular library as a first two-dimensional analog of the input molecule based at least on a two-dimensional representation of the first molecule exhibiting a threshold degree of two-dimensional similarity to the two-dimensional representation of the first molecular variant of the input molecule; and generating, as an output molecule, the first two dimensional analog of the input molecule based at least on a three-dimensional representation of the first two-dimensional analog exhibiting a threshold degree of three-dimensional similarity to the three-dimensional representation of the input molecule.
2. The method of claim 1, further comprising:Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 determining a two-dimensional similarity metric quantifying a similarity between the two-dimensional representation of the first molecule and the two-dimensional representation of the first molecular variant of the input molecule; and determining, based at least on the two-dimensional similarity metric, that the two- dimensional representation of the first molecule exhibits the threshold degree of two-dimensional similarity to the two-dimensional representation of the first molecular variant of the input molecule.
3. The method of claim 2, wherein the two-dimensional representation of the first molecule is determined to exhibit the threshold degree of two-dimensional similarity to the two- dimensional representation of the first molecular variant of the input molecule based at least on the two-dimensional similarity metric satisfying one or more thresholds.
4. The method of any of claims 2 to 3, wherein the two-dimensional similarity metric comprises a Tanimoto index, a Jaccard coefficient, a cosine coefficient, a Dice metric, a Euclidean metric, a city-block metric, a Hamming index, and / or a Tversky index.
5. The method of any of claims 1 to 4, further comprising: translating the two-dimensional representation of the first two-dimensional analog of the input molecule into the three-dimensional representation of the first two-dimensional analog; determining a three-dimensional similarity metric quantifying a similarity between the three-dimensional representation of the first two-dimensional analog of the input molecule and the three-dimensional representation of the input molecule; determining, based at least on the three-dimensional similarity metric, that the three- dimensional representation of the first two-dimensional analog exhibits the threshold degree of similarity to the three-dimensional representation of the input molecule.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 6. The method of claim 5, wherein the determining the three-dimensional similarity metric includes aligning the three-dimensional representation of the first two-dimensional analog of the input molecule and the three-dimensional representation of the input molecule, determining an overlap between the three-dimensional representation of the first two- dimensional analog aligned with the three-dimensional representation of the input molecule, and determining, based at least on the overlap, the three-dimensional similarity metric quantifying the similarity between the three-dimensional representation of the first two- dimensional analog of the input molecule and the three-dimensional representation of the input molecule.
7. The method of any of claims 5 to 6, wherein the three-dimensional similarity metric comprises a Tanimoto Combo score.
8. The method of any of claims 5 to 7, wherein the three-dimensional similarity metric quantifies a similarity in an arrangement of atoms in three-dimensional space and a spatial overlap in one or more functional groups.
9. The method of any of claims 1 to 8, further comprising: identifying, within a same molecular library or a different molecular library, a second molecule as a second two-dimensional analog of the input molecule based at least on a two- dimensional representation of the second molecule exhibiting the threshold degree of two- dimensional similarity to the two-dimensional representation of the first molecular variant of the input molecule; and generating, as another output molecule, the second two dimensional analog of the input molecule based at least on a three-dimensional representation of the second two-dimensionalAttorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 analog exhibiting the threshold degree of three-dimensional similarity to the three-dimensional representation of the input molecule.
10. The method of any of claims 1 to 9, further comprising: applying the molecule design computation model to generate, based at least on the three- dimensional representation of the input molecule, a plurality of molecular variants including the first molecular variant and a second molecular variant.
11. The method of claim 10, further comprising: translating, into a two-dimensional representation of the second molecular variant, a linear representation of the second molecular variant output by the molecule design computation model; identifying, within a same molecular library or a different molecular library, a second molecule as a second two-dimensional analog of the input molecule based at least on a two- dimensional representation of the second molecule exhibiting the threshold degree of two- dimensional similarity to the two-dimensional representation of the second molecular variant of the input molecule; and generating, as another output molecule, the second two dimensional analog of the input molecule based at least on a three-dimensional representation of the second two-dimensional analog exhibiting the threshold degree of three-dimensional similarity to the three-dimensional representation of the input molecule.
12. The method of any of claims 1 to 11, wherein the linear representation of the first molecular variant comprises a simplified molecular input line entry system (SMILES) string or a self-referencing embedded string (SELFIES).Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 13. The method of any of claims 1 to 12, wherein the three-dimensional representation of the input molecule specifies the location of one or more pharmacophores as one or more densities across a three-dimensional voxel grid, and wherein each density is centered around a location of a corresponding pharmacophore.
14. The method of any of claims 1 to 13, wherein the three-dimensional representation of the input molecule further specifies a location of one or more atoms present in the input molecule.
15. The method of claim 14, wherein the three-dimensional representation of the input molecule specifies the location of the one or more atoms as one or more densities across a three-dimensional voxel grid, and wherein each density is centered around a location of a corresponding atom.
16. The method of any of claims 1 to 15, wherein the three-dimensional representation of the input molecule includes a plurality of channels, and wherein each channel of the plurality of channels corresponds to a different type of pharmacophore.
17. The method of claim 16, wherein the plurality of channels further includes a channel specifying a location of each atom in the input molecule.
18. The method of any of claims 16 to 17, wherein each channel of the plurality of channels corresponds to a different one of a hydrogen bond donor, a hydrogen bond acceptor, an aromatic ring, a cation, an anion, and a hydrophobic moiety.
19. The method of any of claims 1 to 18, wherein each of the two-dimensional representation of the first molecule and the two-dimensional representation of the first molecular variant of the input molecule comprise a molecular fingerprint.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 20. The method of claim 19, wherein the molecular fingerprint comprises a binary code, and wherein each bit in the binary code is set to a value indicating a presence or an absence of a chemical substructure.
21. The method of any of claims 19 to 20, wherein the molecular fingerprint comprises a circular fingerprint, a path-based fingerprint, or a RDKit fingerprint.
22. The method of any of claims 19 to 21, wherein the molecular fingerprint comprises a Morgan fingerprint or a RDKit fingerprint.
23. The method of any of claims 1 to 22, wherein the molecular design computation model includes an encoder that has been trained to compress, into a vector representation, the three-dimensional representation of the input molecule.
24. The method of claim 23, wherein the molecular design computation model further includes a decoder that has been trained to decode the vector representation of the input molecule to generate the linear representation of the first molecular variant.
25. The method of any of claims 1 to 24, further comprising: training, based at least on a training dataset, the molecule design computation model.
26. The method of claim 25, wherein the training dataset includes a plurality of training molecules, and wherein the training dataset includes, for each training molecule, a three- dimensional representation of the training molecule and a ground truth linear representation of the training molecule.
27. The method of claim 26, further comprising: augmenting the training dataset by at least translating and / or rotating the three- dimensional representation of one or more training molecules.Attorney Ref.: 14786-077-228 (103963-228077) / P39682-WO-2 28. The method of any of claims 26 to 27, wherein the training of the molecule design computation model includes applying the molecule design computation model to generate, for each training molecule in the training dataset, a linear representation of the training molecule; and adjusting one or more parameters of the molecule design computation model to reduce a difference between the ground truth linear representation of the training molecule and the linear representation of the training molecule generated by the molecule design computation model.
29. The method of any of claims 1 to 28, further comprising: synthesizing the output molecule to generate one or more synthesized molecules.
30. The method of claim 29, further comprising: analyzing the one or more synthesized molecules to determine one or more properties of the output molecule.
31. A system, comprising: at least one data processor; and at least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 30.
32. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1 to 30.