Three-dimensional molecule generation by de-noising voxel mesh
By using a molecular design computational model to denoise and update the voxelized representation of the input molecule, the difficulty of generating specific three-dimensional molecules in existing technologies is solved, and molecules that meet the property requirements are generated more efficiently and accurately.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to efficiently generate molecules exhibiting specific three-dimensional structures, especially in the development of macromolecular drugs. Conventional methods cannot fully explore molecular space and cannot effectively capture the three-dimensional conformation and long-range dependence of molecules.
A molecular design computational model is used to denoise and update the input molecule through voxelization representation. The model is trained to approximate the data distribution that represents the desired properties, and a voxelized representation of the output molecule is generated, thus avoiding the limitations of conventional three-dimensional representation.
It increases the likelihood that generated molecules will exhibit the desired properties, better captures the three-dimensional conformation and long-range dependence of molecules, expands the feasible molecular space for generation, and improves generation efficiency and accuracy.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 502,529, entitled "Three-Dimensional Molecular Generation by Voxel Mesh Denoising," filed May 16, 2023; U.S. Provisional Application No. 63 / 586,263, entitled "Three-Dimensional Molecular Generation by Voxel Mesh Denoising," filed September 28, 2023; and U.S. Provisional Application No. 63 / 623,062, entitled "Three-Dimensional Molecular Generation by Voxel Mesh Denoising," filed January 19, 2024, the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0003] The topics described in this paper generally relate to generative artificial intelligence, and more specifically to machine learning-enabled techniques for generating representations of three-dimensional molecules in discrete and potentially voxelized spaces. Background Technology
[0004] A molecule is a group of two or more atoms linked together by chemical bonds. Molecules form the smallest identifiable unit, and a pure substance can be broken down into such units while still retaining its composition and chemical properties. An example of a molecule is a small molecule, which is a low-weight compound with a molecular weight between approximately 100 and 1000 Daltons. Due to its many compelling advantages, small molecule therapeutics that modulate biochemical processes to diagnose, treat, and prevent a range of diseases have become a cornerstone of modern pharmacology. For example, small molecule drugs can penetrate cell membranes to reach intracellular targets. Furthermore, small molecule drugs are suitable for a wide variety of therapeutic applications. For example, small molecule drugs can be formulated as pills and capsules, intravenous or subcutaneous injections, inhaled medications, or suppositories. The development of small molecule drugs can be further extended to tailoring various pharmacokinetic properties, including release, absorption, distribution, metabolism, potency, efficacy, phenotypic effects, and excretion.
[0005] In contrast, macromolecules (also known as biopharmaceuticals, biologicals, or biologics) can range in molecular weight from approximately 3,000 Daltons to 150,000 Daltons. Macromolecular drugs are typically derivatives of natural human proteins that regulate many important cellular functions, such as enzymatic reactions, molecular transport, regulation and execution of numerous biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, and intercellular communication. A single macromolecule typically has more than 1,300 amino acid residues linked by peptide bonds to form one or more polypeptides. Due to their size and complexity, macromolecular drugs are generated through engineered cellular recombination rather than chemical synthesis, as is the case with most small molecule drugs. Furthermore, because oral administration is ineffective, macromolecular therapeutics are typically delivered by injection or infusion. The development of macromolecular drugs may require designing one or more sequences of amino acid residues capable of binding to targets (e.g., proteins, nucleic acids, etc.) that possess sufficient specificity and are free from undesirable properties such as immunogenicity, self-association, and instability. Summary of the Invention
[0006] Systems, methods, and articles (including computer program articles) for generating three-dimensional molecules in voxelized space are provided. In one aspect, a system for enabling machine learning-based three-dimensional molecule generation is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that provides operation when executed by the at least one processor. Operation may include: identifying an input molecule; generating a voxelized representation of the input molecule; applying a molecular design computational model to update the voxelized representation of the input molecule, wherein the molecular design computational model has been trained to approximate a data distribution of molecules exhibiting one or more desired properties by taking up a damaged voxelized representation of a sample molecule exhibiting one or more desired properties as input and recovering the voxelized representation of the sample molecule from the damaged voxelized representation of the sample molecule, and wherein the molecular design computational model updates the voxelized representation of the input molecule to increase the likelihood that the resulting updated voxelized representation is within the data distribution; and generating a voxelized representation of an output molecule based at least on the updated voxelized representation.
[0007] On the other hand, a method for enabling machine learning in the generation of three-dimensional molecules is provided. The method may include: identifying an input molecule; generating a voxelized representation of the input molecule; applying a molecular design computational model to update the voxelized representation of the input molecule, wherein the molecular design computational model has been trained to approximate a data distribution of molecules exhibiting one or more desired properties by taking up a damaged voxelized representation of a sample molecule exhibiting one or more desired properties as input and recovering the voxelized representation of the sample molecule from the damaged voxelized representation of the sample molecule; and wherein the molecular design computational model updates the voxelized representation of the input molecule to increase the likelihood that the resulting updated voxelized representation is within the data distribution; and generating a voxelized representation of an output molecule based at least on the updated voxelized representation.
[0008] On the other hand, a computer program product for enabling machine learning to generate three-dimensional molecules is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations. The operations may include: identifying an input molecule; generating a voxelized representation of the input molecule; applying a molecular design computational model to update the voxelized representation of the input molecule, wherein the molecular design computational model has been trained to approximate a data distribution of molecules exhibiting one or more desired properties by taking up a damaged voxelized representation of a sample molecule exhibiting one or more desired properties as input and recovering the voxelized representation of the sample molecule from the damaged voxelized representation of the sample molecule, and wherein the molecular design computational model updates the voxelized representation of the input molecule to increase the likelihood that the resulting updated voxelized representation is within the data distribution; and generating a voxelized representation of an output molecule based at least on the updated voxelized representation.
[0009] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.
[0010] In some variations, the voxelized representation of a molecule can include multiple voxels organized into a three-dimensional voxel grid. Each atom in the molecule can be represented as a continuous density of one or more voxels across the three-dimensional voxel grid.
[0011] In some variations, the continuous density of each atom in the molecule can be centered on the center of each atom. A first voxel located further away from any atom in the molecule can be associated with a lower atomic density value compared to a second voxel located closer to the center of the atoms in the molecule.
[0012] In some variations, each voxel in a three-dimensional voxel grid can be associated with a value indicating the atomic density at the corresponding location.
[0013] In some variations, the voxelized representation of a molecule may include one or more channels. Each channel may correspond to a type of atom present in the molecule.
[0014] In some variations, applying a molecular design computational model to update the voxelized representation of the input molecule may include updating the voxelized representation of the input molecule as a function of the probability values of the updated voxelized representation produced by the output indication within the data distribution.
[0015] In some variations, the function can be parameterized using multiple parameters of the molecular design computational model.
[0016] In some variations, the function can be a scoring function. The value output by the function can be a score indicating a local change in the density of the data distribution at the location of the updated voxelized representation.
[0017] In some variations, the molecular design computational model can update the voxelized representation of the input molecule at least by: updating the voxelized representation of the input molecule to generate a first updated voxelized representation; updating the voxelized representation of the input molecule to generate a second updated voxelized representation; applying a function parameterized by the molecular design computational model to determine a first value indicating a first local change in the density of the data distribution at a first location occupied by the first updated voxelized representation; applying the same function to determine a second value indicating a second local change in the density of the data distribution at a second location occupied by the second updated voxelized representation; and further updating the first updated voxelized representation instead of the second updated voxelized representation when the first and second values indicate that the density of the data distribution at the first location is higher than the density of the data distribution at the second location.
[0018] In some variations, molecular design computational models can be applied to further update the first updated voxelized representation until one or more criteria are met.
[0019] In some variations, one or more criteria may include at least one of the following: (i) an updated iteration of the voxelized representation of the input molecule with a threshold amount performed, (ii) a first value of the first updated voxelized representation satisfying one or more thresholds, and (iii) an output molecule with a threshold amount generated.
[0020] In some variations, molecular design computational models can be applied to further modify the first updated voxelized representation rather than the second updated voxelized representation, based at least on the first and second values indicating that the first updated voxelized representation has a higher probability of being within the data distribution compared to the second updated voxelized representation.
[0021] In some variations, molecular design computational models can be applied to further modify the first updated voxelized representation rather than the second updated voxelized representation by sampling from a higher-density region of the data distribution, based at least on the first and second values indicating that the first updated voxelized representation is sampled from a higher-density region of the data distribution compared to the second updated voxelized representation.
[0022] In some variations, the data distribution can be a noisy data distribution filled with noisy voxelized representations of molecules exhibiting one or more desired properties. The voxelized representation of the output molecule can be generated by denoising a first updated voxelized representation so as to map the first updated voxelized representation from the noisy data distribution to the true data distribution of molecules exhibiting one or more desired properties.
[0023] In some variations, the voxelized representation of the output molecule can be transformed into different representations of the output molecule.
[0024] In some variations, different representations of the output molecule may include a one-dimensional representation of the output molecule and / or a two-dimensional representation of the output molecule.
[0025] In some variations, the voxelized representation of the output molecule is transformed at least by determining the positions of one or more atoms in the output molecule by detecting one or more peaks among a plurality of atomic density values contained in the voxelized representation of the output molecule, and determining one or more interconnecting bonds at least based on the positions of one or more atoms.
[0026] Systems, methods, and articles (including computer program products) for generating three-dimensional molecules in voxelized space are provided. In one aspect, a system for enabling machine learning-based three-dimensional molecule generation is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that provides operation when executed by the at least one processor. Operation may include: identifying sample molecules exhibiting one or more desired properties; generating a noisy voxelized representation of the sample molecules; adding noise to the noisy voxelized representation of the sample molecules to generate a damaged voxelized representation of the sample molecules; training a molecular design computational model to approximate a data distribution of molecules exhibiting one or more desired properties, wherein training includes applying the molecular design computational model to recover the noisy voxelized representation of the sample molecules from the damaged voxelized representation of the sample molecules; and optionally, generating a voxelized representation of an output molecule by at least applying the molecular design computational model to denoise the voxelized representation of the input molecule.
[0027] On the other hand, a method for enabling machine learning to generate three-dimensional molecules is provided. This method may include: identifying sample molecules exhibiting one or more desired properties; generating a noisy voxelized representation of the sample molecules; adding noise to the noisy voxelized representation of the sample molecules to generate a damaged voxelized representation of the sample molecules; training a molecular design computational model to approximate a data distribution of molecules exhibiting one or more desired properties, wherein training includes applying the molecular design computational model to recover the noisy voxelized representation of the sample molecules from the damaged voxelized representation of the sample molecules; and optionally, generating a voxelized representation of the output molecule by at least applying the molecular design computational model to denoise the voxelized representation of the input molecule.
[0028] On the other hand, a computer program product for enabling machine learning in the generation of three-dimensional molecules is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations. The operations may include: identifying sample molecules exhibiting one or more desired properties; generating a noisy voxelized representation of the sample molecules; adding noise to the noisy voxelized representation of the sample molecules to generate a damaged voxelized representation of the sample molecules; training a molecular design computational model to approximate a data distribution of molecules exhibiting one or more desired properties, wherein training includes applying the molecular design computational model to recover the noisy voxelized representation of the sample molecules from the damaged voxelized representation of the sample molecules; and optionally, generating a voxelized representation of the output molecule by at least applying the molecular design computational model to denoise the voxelized representation of the input molecule.
[0029] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.
[0030] In some variations, the noisy voxelization representation of a sample molecule can include multiple voxels organized into a three-dimensional voxel grid. Each atom in the sample molecule can be represented as a continuous density of one or more voxels across the three-dimensional voxel grid.
[0031] In some variations, the continuous density of each atom in the sample molecule can be centered at the center of each atom.
[0032] In some variations, each voxel in a three-dimensional voxel grid can be associated with a value indicating the atomic density at the corresponding location.
[0033] In some variations, a first voxel located far from any atom in the sample molecule can be associated with a lower atomic density value compared to a second voxel located closer to the center of an atom in the sample molecule.
[0034] In some variations, the noisy voxelized representation of a sample molecule may include one or more channels. Each channel may correspond to the type of atom present in the sample molecule.
[0035] In some variations, the noisy voxelization representation of a sample molecule can collectively represent the type and position of one or more atoms present in the sample molecule.
[0036] In some variations, training the molecular design computation model may include adjusting multiple parameters of the molecular design computation model to reduce the difference between the recovered voxelized representation of the sample molecule generated by the molecular design computation model and the noisy voxelized representation of the sample molecule.
[0037] In some variations, multiple parameters of the molecular design computational model can be parameterized to the function. The values of these parameters can be adjusted so that the function output indicates local variations in the density of the data distribution of molecules exhibiting one or more desired properties.
[0038] In some variations, the molecular design computational model can denoise the voxelized representation of the input molecule by updating the atomic density of one or more voxels in at least one channel of the voxelized representation of the input molecule.
[0039] In some variations, updating the atomic density of one or more voxels in at least one channel of the voxelization of the input molecule may correspond to updating at least one of the types and / or positions of one or more atoms present in the input molecule.
[0040] In some variations, the molecular design computational model can undergo multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling to denoise the voxelized representation of the input molecule until one or more criteria are met.
[0041] In some variations, one or more criteria may include at least one of the following: (i) an iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling of the threshold amount has been performed, (ii) a voxelized representation of the output molecule sampled from a region having a threshold density, and (iii) an output molecule of the threshold amount has been generated.
[0042] In some variations, the molecular design computational model generates a voxelized representation of the output molecule at least by: applying a first update to the voxelized representation of the input molecule to generate a first updated voxelized representation; applying a second update to the voxelized representation of the input molecule to generate a second updated voxelized representation; and further updating the first updated voxelized representation after determining that the first updated voxelized representation is sampled from a higher density region of the data distribution compared to the second updated voxelized representation.
[0043] In some variations, the data distribution can be a noisy data distribution filled with noisy voxelized representations of molecules exhibiting one or more desired properties. The voxelized representation of the output molecule can be further generated by denoising the first updated voxelized representation so as to map the first updated voxelized representation from the noisy data distribution to the true data distribution of molecules exhibiting one or more desired properties.
[0044] In some variations, the voxelized representation of the output molecule can be transformed into different representations of the output molecule.
[0045] In some variations, different representations of the output molecule may include a one-dimensional representation of the output molecule and / or a two-dimensional representation of the output molecule.
[0046] In some variations, training the molecular design computation model may include applying a molecular design computation model with a first adjustment to denoise a damaged voxelized representation of a sample molecule and generate a first recovered voxelized representation of the sample molecule; determining a first mean square error (MSE) of a first difference between the first recovered voxelized representation of the sample molecule and the noisy voxelized representation; applying a molecular design computation model with a second adjustment to denoise the damaged voxelized representation of the sample molecule and generate a second recovered voxelized representation of the sample molecule; determining a second mean square error (MSE) of a second difference between the second recovered voxelized representation of the sample molecule and the noisy voxelized representation; and further adjusting the molecular design computation model with the first adjustment but not the second adjustment after determining that the first mean square error (MSE) is less than the second mean square error (MSE).
[0047] Implementations of the present subject matter may include, but are not limited to, methods consistent with the descriptions provided herein, and articles of art comprising a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to perform operations implementing one or more of the described features. Similarly, computer systems comprising one or more processors and one or more memories coupled to the one or more processors are also described. Memory that may include a non-transitory computer-readable or machine-readable storage medium may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implementations consistent with one or more implementations of the present subject matter may be implemented by one or more data processors existing in a single computing system or multiple computing systems. Such multiple computing systems may be interconnected and may exchange data and / or commands or other instructions via one or more connections, including, for example, direct connections between one or more of the multiple computing systems via a network (e.g., the Internet, wireless wide area network, local area network, wide area network, wired network, etc.).
[0048] Details of one or more variations of the subject matter described herein are set forth in the appended options and the following description. Other features and advantages of the subject matter described herein will become apparent from the description, the appended options, and the claims. While certain features of the subject matter currently disclosed are described for purposes of illustrating the computational design of molecules, including drug molecules, it should be readily understood that these features are not intended to constitute limitation. The claims following this disclosure are intended to define the scope of the protected subject matter.
[0049] Selection Instructions
[0050] The selections described herein, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the specification, help explain some principles associated with the disclosed embodiments. In the selections,
[0051] Figure 1A depicts a system diagram illustrating an example of a molecular design system according to some exemplary embodiments;
[0052] Figure 1B depicts a system diagram illustrating another example of a molecular design system according to some exemplary embodiments;
[0053] Figure 2 depicts a flowchart illustrating an example of a process for enabling machine learning to generate three-dimensional molecules in a voxelized space, according to some exemplary embodiments.
[0054] Figure 3A depicts a flowchart illustrating an example of a process for training a computational model for molecular design to generate three-dimensional molecules in voxelized space, according to some exemplary embodiments.
[0055] Figure 3B depicts a flowchart illustrating an example of a process for applying a molecular design computational model to generate three-dimensional molecules in voxelized space, according to some exemplary embodiments.
[0056] Figure 3C depicts a flowchart illustrating an example of a process for applying a molecular design computational model to generate three-dimensional molecules in voxelized space, according to some exemplary embodiments.
[0057] Figure 4 illustrates an example of a voxelized representation of a molecule according to some exemplary embodiments;
[0058] Figure 5A depicts a schematic diagram illustrating an example of a process for training a denoising engine to denoise a noisy voxelized representation of a molecule, according to some exemplary embodiments.
[0059] Figure 5B depicts a schematic diagram illustrating an example of a walking jump sampling scheme according to some exemplary embodiments;
[0060] Figure 5C depicts a schematic diagram illustrating an example of a computational model for molecular design that generates three-dimensional molecules by denoising a voxelized molecular representation according to some exemplary embodiments.
[0061] Figure 5D depicts a schematic diagram illustrating an example of a process for generating other molecular representations from a voxelized representation of a molecule, according to some exemplary embodiments.
[0062] Figure 6 illustrates an example of a process by which a computational model for molecular design generates a voxelized representation of a molecule by operating in a noisy latent voxelization space, according to some exemplary embodiments.
[0063] Figure 7 depicts a graph illustrating the effect of noise levels on the generative performance of molecular design computational models according to some exemplary embodiments;
[0064] Figure 8 illustrates a schematic diagram showing the effect of the number of sampling iterations in Markov chain Monte Carlo (MCMC) sampling on the generative performance of a molecular design computational model, according to some exemplary embodiments.
[0065] Figure 9A illustrates an example of a voxelized representation of a molecule generated by a molecular design computational model trained on the QM9 molecular dataset, according to some exemplary embodiments.
[0066] Figure 9B illustrates an example of a voxelized representation of a molecule generated by a computational model of molecular design trained on the Geometric Ensemble of Molecules (GEOM) drug dataset, according to some exemplary embodiments.
[0067] Figure 10A depicts a graph of the cumulative distribution function (CDF) of strain energy for molecules in the QM9 molecular dataset according to some exemplary embodiments, molecules generated by a conventional generative model, and molecules generated by a molecular design computational model trained on the QM9 molecular dataset.
[0068] Figure 10B depicts a graph comparing the empirical distribution of the number of atoms in each molecule in the QM9 molecular dataset according to some exemplary embodiments with the empirical distribution of the number of atoms in molecules generated by a molecular design computational model trained on the QM9 molecular dataset.
[0069] Figure 11A depicts the cumulative distribution function (CDF) of strain energy for molecules in the Geometry of Molecular Omnibus (GEOM) drug dataset according to some exemplary embodiments, molecules generated by a conventional generative model, and molecules generated by a molecular design computational model trained on the GEOM drug dataset.
[0070] Figure 11B depicts a graph comparing the empirical distribution of the number of atoms in each molecule in the Geometry of Molecular Ensemble (GEOM) drug dataset according to some exemplary embodiments with the empirical distribution of the number of atoms in molecules generated by a molecular design computational model trained on the GEOM drug dataset; and
[0071] Figure 12A depicts a schematic diagram showing a comparison of seed generation performance for molecular geometry ensemble (GEOM) drugs in discrete voxelization space and potential voxelization space according to some exemplary embodiments.
[0072] Figure 12B depicts a schematic diagram illustrating a comparison of seed generation performance for PubChem drugs in discrete voxelization space and potential voxelization space according to some exemplary embodiments.
[0073] Figure 12C depicts molecular diagrams of other examples of molecules generated by seeding a real drug in a potential voxelized space.
[0074] Figure 12D depicts molecular diagrams of other examples of molecules generated by de novo generation of real drugs in a potential voxelization space.
[0075] Figure 13 depicts a comparison of seed generation performance for molecular geometry ensemble (GEOM) drugs in discrete voxelization space and potential voxelization space according to some exemplary embodiments;
[0076] Figure 14 depicts a block diagram illustrating an example of a computing system according to some exemplary embodiments.
[0077] In practical applications, similar reference numerals indicate similar structures, features, or elements. Detailed Implementation
[0078] Generating new molecules with desired properties is a crucial task in chemistry, with applications spanning numerous scientific fields. In the context of drug discovery, conventional computational techniques for generating molecules with drug-like properties require searching the molecular space (or chemical space) occupied by each possible chemical compound (e.g., every possible combination of atoms of two or more chemical elements). For example, some search-based methods may include scoring and ranking different molecules in the molecular space based on one or more drug-like properties, such as affinity, specificity, bioactivity, and fertility. However, the aforementioned molecular space (estimated to contain 10...) 60 The molecular space (the number of possible chemical compounds) is extremely large and expands exponentially with molecular size (e.g., the number of constituent atoms). Even a tiny fraction of molecular space can contain billions to trillions of molecules. With state-of-the-art computational resources, conventional search-based methods can only explore a small fraction of molecular space, such as small regions selected based on prior domain knowledge. This limitation in search scope means that conventional search-based methods may overlook molecules with superior properties. Furthermore, conventional search-based methods do not explore molecular space in a principled manner, preventing the generation process from being conditionalized to specific properties.
[0079] Furthermore, whether a molecule exhibits certain desired properties may depend on its conformation (or three-dimensional structure). For example, the binding affinity between a drug molecule and a target molecule (e.g., a protein, nucleic acid, etc.) may depend on the ability of the drug molecule to adopt a conformation complementary to that of the target molecule. Moreover, molecules are flexible, meaning that a single molecule can present one of many possible conformations (or three-dimensional structures). In some cases, a group of the same molecule may exist as a collection of many different conformations in equilibrium with each other, but not every possible conformation is associated with the desired property. For example, in the context of binding affinity, the bioactive conformation of a molecule can be one or more of the conformations exhibited by the molecule in solution or a novel conformation induced by interaction with the target molecule. However, one-dimensional representations of molecules (e.g., simplified linear input canonical (SMILES) strings) or two-dimensional representations (e.g., molecular diagrams) cannot adequately capture the conformation (or three-dimensional structure) of a molecule. Therefore, when a molecular design computational model operates on a one-dimensional or two-dimensional representation of the input molecule, the resulting output molecule may not exhibit a conformation (or three-dimensional structure) associated with one or more desired properties.
[0080] Various exemplary embodiments of this disclosure can improve state-of-the-art computing resources by providing molecular design computational models that generate output molecules by exploring molecular (or chemical) space in a principled manner rather than by indiscriminately searching a finite portion of molecular space. For example, in some cases, the molecular design computational model can be trained to approximate a data distribution of molecules that exhibit one or more desired properties (e.g., pharmaceutical properties such as affinity, specificity, bioactivity, exploitability, etc.). Training the molecular design computational model may include determining the parameters of a function (e.g., a scoring function) such that the output of the function is a value indicating a density variation across the data distribution. In some cases, the molecular design computational model can sample the data distribution to generate output molecules that also exhibit one or more desired properties. For example, in some cases, the molecular design computational model can sample the data distribution by denoising the input molecules (such as a voxelized representation of the input molecules) through multiple sampling iterations. During each sampling iteration, the molecular design computational model can update the input molecules to remove a portion of the noise present in the input molecules. Doing so generates updated molecules (e.g., updated voxelized representations of the molecules) that constitute a sample selected from the data distribution. As described in more detail below, sampling can be guided by a function such that each successive sample (or updated molecule) is selected from regions of increasing density in the data distribution, regions that are more likely to be occupied by molecules exhibiting one or more properties.
[0081] In some exemplary embodiments, operating a molecular design computational model on a three-dimensional representation of the input molecule can increase (or maximize) the likelihood that the output molecule will exhibit one or more desired properties. For example, in some cases, the molecular design computational model can generate the output molecule at least by denoising the three-dimensional representation of the input molecule, for example, through multiple sampling iterations. In other cases, the molecular design computational model can generate the output molecule by denoising a voxelized representation of the input molecule instead of its conventional three-dimensional representation. A conventional three-dimensional representation of the input molecule, such as a point cloud representation, can specify the conformation (or three-dimensional structure) of the input molecule at least by specifying the coordinates of the constituent atoms (e.g., in Euclidean space). However, a conventional three-dimensional representation of the input molecule can impose many limitations on the generation process. For example, in order for the molecular design computational model to operate on a conventional three-dimensional representation of the input molecule, the number of atoms in the resulting output molecule must be known a priori. Denoising the conventional 3D representation of the input molecule may require workarounds to allow the molecular design computational model to approximate the distribution of atom types in the output molecule, resulting in a discrete distribution, while the positions of atoms in the output molecule (e.g., atomic coordinates in Euclidean space) form a continuous distribution. Furthermore, the conventional 3D representation of the input molecule may not adequately capture long-range dependencies existing across multiple atoms, especially as the number of constituent atoms increases.
[0082] In some exemplary embodiments, the voxelized representation of a molecule (e.g., an input molecule) can overcome the aforementioned limitations by representing the input molecule as a continuous distribution of atomic density across a voxel grid centered on the atomic coordinates of each individual atom present in the molecule. For example, in a graphical network representation of a molecule, the dependency between two adjacent atoms can be represented by interconnecting edges. However, these edges may not adequately capture longer-range dependencies, such as dependencies between non-adjacent atoms. In contrast, the voxelized representation of a molecule can better capture long-range dependencies between distant atoms, even when the input molecule contains a large number of atoms. Furthermore, a molecular design computational model can operate on the voxelized representation of the input molecule to generate an output molecule without any prior knowledge of the number of atoms present in the output molecule. This is because the molecular design computational model can freely add or remove different types of atoms by updating the atomic density distribution across the voxel grid. The voxelized representation of the input molecule also collectively represents the types and positions of atoms in the input molecule, thus avoiding workarounds for reconciling two different types of data distributions (e.g., a discrete distribution of atom types and a continuous distribution of atom positions).
[0083] In some exemplary embodiments, a voxelized molecular representation of a molecule (such as an input molecule) can represent each atom in the molecule (e.g., the input molecule) as a continuous (e.g., Gaussian-like) density of one or more voxels across a voxel grid. In this case, the voxel grid is a three-dimensional grid of voxels organized into continuous layers of rows and columns. Various examples of voxel grids described herein can contain multiple voxels, each voxel being a volume element (e.g., a three-dimensional cube) at the intersection of rows and columns. Each volume element can have a predetermined size, which may be the same or may not be the same for all voxels in the voxel grid. In the case where the input molecule is a drug molecule, the voxelized representation of the input molecule may include containing voxels (e.g., voxels, A voxel grid (e.g., voxels). In some cases, each voxel in the voxel grid can be associated with a value indicating the atomic density at the corresponding location. For example, a first voxel associated with a higher atomic density may be more likely to be part of an atom than a second voxel associated with a lower atomic density. It should be understood that the volume of a single atom can span one or more voxels. In some cases, the atomic density can also be centered on the atom present in the input molecule, meaning that the atomic density of a single atom spanning multiple voxels can be centered on the voxel that constitutes the center of that atom. A voxel with an atomic density of 0 can be far from any atom in the input molecule, while a voxel with an atomic density of 1 can be at the center of an atom in the input molecule. Furthermore, in some cases, the voxelized representation of a molecule (e.g., the input molecule) can include multiple channels, each corresponding to a type of atom that may be present in the input molecule. "Atom type" can refer to the single chemical element to which the atom belongs. The voxelized representation can include multiple channels, one channel for each type of atom or one channel for at least each type of heavy atom present in the molecule. For example, in some cases, the voxelized representation of a molecule (e.g., an input molecule) may include a first channel corresponding to a first atomic type (e.g., carbon (C) atoms) that may be present in the input molecule and a second channel corresponding to a second atomic type (e.g., nitrogen (N) atoms) that may be present in the input molecule. Each voxel in the first channel may be associated with a value indicating the density of atoms of the first atomic type at the corresponding position, and each voxel in the second channel may be associated with a value indicating the density of atoms of the second atomic type at the corresponding position. Thus, as described in more detail below, a molecular design computational model can denoise the voxelized representation of the input molecule at least by updating the atomic density of one or more voxels in at least one channel of the voxelized representation of the input molecule, for example, through multiple sampling iterations. That is, in some cases, the term "denoising" refers to updating the voxelized representation of the input molecule, which may include updating the atomic density of at least one voxel in the voxelized representation of the input molecule. In some cases, updating the atomic density of a voxel in a channel of the voxelized representation of the input molecule may change the probability that the voxel is part of the atomic type associated with that channel.
[0084] In some exemplary embodiments, the molecular design computational model may denoise the input molecule (e.g., a voxelized representation of the input molecule) through multiple sampling iterations, wherein each sampling iteration generates an updated voxelized representation different from the voxelized representation of the input molecule. In some cases, each updated voxelized representation may include a sample of a data distribution (e.g., a voxelized representation of a molecule) exhibiting one or more desired properties. In this context, the term "data distribution" may refer to a population of different molecular compositions and conformations (or three-dimensional structures). Molecules exhibiting one or more desired properties may cluster in higher-density regions of the data distribution, meaning the molecular design computational model should sample each updated voxelized representation from these higher-density regions. However, such a data distribution may be too high-dimensional to be directly approximated. For example, calculating the probability density function (PDF) characterizing the probabilities of different molecules in the data distribution requires normalization of constants. In the case of molecular design, normalization of constants may correspond to the total number of molecules in the data distribution, which may be impossible to estimate. Therefore, in some cases, molecular design computational models can be trained to approximate the data distribution by determining a function (such as a scoring function) that estimates the gradient (or density change) across the data distribution. As described in more detail below, molecular design computational models can use functions to guide sampling from the data distribution with updated voxelized representations, such that each successive sample is selected from regions of gradually increasing density in the data distribution.
[0085] As described, in some cases, denoising the input molecule through multiple sampling iterations (including continuous updates to the voxelized representation of the input molecule) can be equivalent to selecting successive samples from the data distribution of the molecule (e.g., the data distribution of the voxelized representation of the molecule), where each sample corresponds to an updated voxelized representation different from the voxelized representation of the input molecule. For example, in some cases, the voxelized representation of the input molecule can be denoised at least by updating the atomic density of one or more voxels in at least one channel of the voxel grid, thereby forming the voxelized representation of the input molecule. In some cases, molecules in higher-density regions of the data distribution can exhibit one or more desired properties, including, for example, pharmaceutical properties such as affinity, specificity, bioactivity, exploitability, etc. Operating on the three-dimensional representation of the input molecule (such as the voxelized representation of the input molecule) can increase the likelihood that the conformation (or three-dimensional structure) of the resulting output molecule is selected from higher-density regions of the data distribution and thus exhibits one or more desired properties.
[0086] In some cases, molecular design computational models can be trained to approximate a data distribution using training datasets of known molecules exhibiting one or more desired properties (e.g., the PubChem dataset, the QM9 molecular dataset, the Geometry of Molecular Ensemble (GEOM) drug dataset, etc.). For example, in some cases, a molecular design computational model can be trained to approximate a data distribution at least by determining, for example, a function (e.g., a scoring function) that approximates different densities across the data distribution using Bayesian inference. In some cases, the function can be parameterized by the molecular design computational model, meaning that the parameters of the function (e.g., the scoring function) are parameters of the molecular design computational model, which are adjusted as the model has been trained to approximate the data distribution. In some cases, high-density regions of the data distribution can be filled with molecules similar to known molecules exhibiting one or more desired properties, while low-density regions can be filled with molecules that are not similar to known molecules exhibiting one or more desired properties. The scoring function of the data distribution can indicate transitions between different density regions of the data distribution, including, for example, transitions between higher-density and lower-density regions of the data distribution. Therefore, once trained, the molecular design computational model can sample the data distribution based on the scoring function, so that each consecutive sample (or molecule) is selected from regions of gradually increasing density in the data distribution.
[0087] In some exemplary embodiments, a molecular design computational model can be trained to denoise damaged 3D representations of known molecules from a training dataset and recover the original 3D representation of the known molecules. For example, in some cases, a damaged 3D representation of a known molecule can be generated by corrupting the 3D representation of the known molecule with noise (e.g., Gaussian noise, such as isotropic Gaussian noise). Training the molecular design computational model may include adjusting one or more parameters of the molecular design computational model (e.g., weights, biases, etc.) to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the recovered 3D representation of the known molecule and the original 3D representation of the known molecule.
[0088] In some exemplary embodiments, to avoid overfitting the molecular design computation model to known molecules in the training dataset, the molecular design computation model can be trained to recover a noisy version of the 3D representation of the known molecules in the training dataset instead of the original 3D representation. That is, the 3D representation of each known molecule in the training dataset may be incorporated with additional noise, but this noise should not be confused with the noise that the molecular design computation model has been trained to remove from the damaged 3D representation of each known molecule in the training dataset. In other words, in some cases, the molecular design computation model can be trained based on a training dataset that includes noisy 3D representations of known molecules and their damaged versions. As described in more detail below, in some cases, the noisy 3D representation of a known molecule can be generated by incorporating a first amount of noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) into the 3D representation of the known molecule (e.g., a voxelized representation) to smooth the density of the data distribution of the known molecules, while still retaining at least a portion of the conformation (e.g., the 3D structure) of the known molecules, thereby obtaining a noisy representation of the known molecules. Then, a second amount of noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) can be used to further corrupt the noisy 3D representation of the known molecule to generate a damaged 3D representation. In some cases, the molecular design computational model can be trained to denoise the damaged 3D representation of the known molecule, for example by removing the second amount of noise, and recover the noisy 3D representation of the known molecule (which still includes the first amount of noise). Furthermore, in some cases, the training of the molecular design computational model can include gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.), in which the parameters of the molecular design computational model are adjusted through successive sampling iterations to increase the similarity between the 3D representation of each known molecule recovered by the molecular design computational model from the corresponding damaged 3D representations of known molecules in the training dataset and the noisy 3D representation of sample molecules in the training dataset (e.g., reducing the mean squared error (MSE)). The scoring function derived in this way can capture data distributions with smoother density transitions, which mitigates the pattern collapse phenomenon where molecular design computational models are less robust and can only generate a limited selection of output molecules (e.g., those in the immediate vicinity of molecules known in the data distribution).
[0089] As described in more detail below, during inference, a trained molecular design computational model generates one or more output molecules by applying noise reduction to the 3D representation of the input molecule. In some cases, the input molecule can be a random molecule (e.g., a molecule with randomly selected atom types and / or positions) or a known molecule with one or more unwanted properties. This means that the 3D representation of the input molecule can include at least some noise that needs to be removed, such that the resulting 3D representation of the output molecule is consistent with the 3D representation of the molecule exhibiting one or more desired properties. The molecular design computational model can do this by traversing a smoothed density of a noisy data distribution of noisy 3D representations of the molecules, for example, through one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo, etc.) toward regions of gradually increasing density in the data distribution. Each iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling can include updating the 3D representation of the input molecule, which is equivalent to selecting noisy 3D representations of one or more molecules from different positions in the noisy data distribution. Although selected from noisy data distributions, molecules corresponding to these noisy 3D representations may have less distortion compared to the original 3D representation of the input molecule. Such molecules corresponding to these noisy 3D representations may be more consistent with molecules exhibiting one or more desired properties compared to the input molecule. In some cases, the noisy 3D representations of molecules selected from noisy data distributions can undergo further denoising to recover the corresponding molecule by mapping the noisy 3D representation of each molecule from the noisy data distribution to the corresponding clean 3D representation of the molecule in the real data distribution exhibiting one or more desired properties. It should be understood that sampling from noisy data distributions can provide many advantages over sampling from the real data distribution of molecules exhibiting one or more desired properties. For example, sampling from a noisy data distribution exhibiting a smoother density transition in molecules exhibiting one or more desired properties may be less susceptible to mode collapse than sampling from the real data distribution. In some cases, this may be because steep or drastic gradients are less common in noisy data distributions than in the real data distribution, where steep gradients restrict sampling to the immediate vicinity of known molecules characterizing the real data distribution. In other words, given that molecular design computational models may be able to generate outputs with limited variation when sampling from real-world data distributions (e.g., the aforementioned phenomenon known as "mode collapse"), sampling from noisy distributions can increase the variability of the model output. Furthermore, sampling from noisy data distributions filled with noisy voxelized representations of molecules can offer many advantages over sampling from noisy data distributions filled with noisy conventional 3D representations of molecules, such as point cloud representations of molecules.For example, in some cases, computational molecular design models can be trained to operate on voxelized molecular representations and generate large drug-like molecules with greater ease of use, effectiveness, expressiveness, and scalability. Unlike operating on conventional three-dimensional molecular representations (e.g., point cloud representations), operating on voxelized molecular representations allows publicly available computational molecular design models to function without specifying the number of atoms present in the output molecule and without requiring workarounds to reconcile the discrete distribution of atom types with the continuous distribution of atom positions associated with each molecule.
[0090] Despite the advantages mentioned above, operating on voxelized molecular representations can impose a significant computational burden, which can scale exponentially with molecular size (e.g., the number of constituent atoms). For example, while a small molecule containing 10 heavy atoms already requires a [32×32×32] voxel grid with 32,000 features (or atomic density values) per molecule, larger, more realistic drug molecules may require voxel grids at least twice the size, with an exponentially larger number of points (e.g., a [64×64×64] voxel grid with 260,000 features (or atomic density values) per molecule). In some cases, applying molecular design computational models to voxelized representations of larger, more realistic drug molecule classes can be a challenging task, as can training such models on large training datasets (e.g., training datasets containing voxelized representations of millions of known molecules) to learn a more diverse molecular (or chemical) space. Furthermore, in practice, most candidate molecules generated by molecular design computational models may not be successfully synthesized in the laboratory, even if the candidate molecules are real and valid. Therefore, it may be necessary to apply molecular design computational models to generate tens of thousands or even millions of candidate molecules. The computational burden associated with generating molecules, especially larger molecules or larger quantities of molecules, can be reduced by using molecular design computational models that operate on lower-dimensional embeddings of voxelized molecular representations. For example, in some cases, molecular design computational models can be trained to generate output molecules by denoising the embeddings of three-dimensional representations of input molecules. As described in more detail below, in some cases, molecular design computational models can be trained based on training datasets that include damaged embeddings of sample molecules exhibiting one or more desired properties, each generated by encoding noisy three-dimensional representations (e.g., voxelized representations) of known molecules exhibiting one or more desired properties, and then corrupted by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.).
[0091] In some exemplary embodiments, encoding a three-dimensional representation of a molecule (such as a voxelized representation of the molecule) can project the three-dimensional representation of the molecule (e.g., a voxelized molecular representation) from a high-dimensional discrete space filled with the three-dimensional representation of the molecule (e.g., a discrete voxelized space filled with the voxelized molecular representation) to a lower-dimensional representation in a lower-dimensional latent space filled with the corresponding molecular embedding. In other words, encoding a three-dimensional representation of a molecule (such as a voxelized representation of the molecule) can take the three-dimensional representation of the molecule (e.g., a voxelized molecular representation) in a filled high-dimensional discrete space as input and produce a lower-dimensional representation of the molecule (i.e., a molecular embedding corresponding to the input three-dimensional representation) as output in a lower-dimensional latent space. Encoding a three-dimensional representation of a molecule, such as a voxelized representation of the molecule, can be performed using a machine learning model trained to recognize the latent space representation from which the three-dimensional representation of the input molecule can be recovered. In some cases, each embedding in the latent space can be the latent space representation of the corresponding voxelized representation of the molecule. Therefore, the embedding of the voxelized representation of the molecule can have a different dimension or feature quantity than the voxelized representation of the molecule. For example, in some cases, encoding a voxelized representation of a molecule can reduce the dimensionality or number of features present in that voxelized representation. Therefore, embedding the voxelized representation of a molecule into a molecular design computational model, rather than directly manipulating it, can reduce the computational burden of denoising the voxelized representation to generate one or more output molecules, at least because the embedding contains fewer features. Furthermore, in some cases, a molecular design computational model can be trained to approximate a noisy data distribution of molecules that exhibits one or more desired properties, allowing one or more output molecules to be generated by sampling from it. This noisy data distribution, potentially filled with noisy embeddings of the voxelized representation of the molecule, can exhibit a smoother density transition than the corresponding true data distribution. Therefore, noisy data distributions can support more efficient sampling (e.g., via gradient-based Markov chain Monte Carlo (MCMC) sampling, such as Langevin Markov chain Monte Carlo), at least because noisy data distributions can exhibit fewer steep gradient changes that would prevent the molecular design computational model from fully exploring the data distribution when sampling from it.
[0092] Figures 1A-1B depict system diagrams illustrating different examples of a molecular design system 100 according to some exemplary embodiments. Referring to Figures 1A-1B, in some cases, the molecular design system 100 may include a molecular design engine 110, a training engine 120, and a client device 130. In the example of the molecular design system 100 shown in Figures 1A-1B, the molecular design engine 110, the training engine 120, and the client device 130 may be communicatively coupled via a network 140. The client device 130 may be a processor-based device, including, for example, a workstation, desktop computer, laptop computer, smartphone, tablet computer, wearable device, etc. The network 140 may be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. Figure 1A-1B In the example shown, the molecular design computational model 115 may include a denoising model 117 trained to generate an output molecule 162 by at least denoising the input molecule 152. The denoising model is a machine learning model trained to take a damaged 3D representation of a molecule or its lower-dimensional embedding as input, where the damaged 3D representation is a 3D representation of a molecule with added noise, and to use training data containing 3D representations of multiple known molecules and their corresponding damaged 3D representations to produce a corresponding denoised 3D representation of the molecule or its lower-dimensional embedding as output. The known molecules in the training data may include multiple molecules exhibiting one or more desired properties. The denoising model may be an artificial neural network (ANN). The denoising model may be a deep learning model. The denoising model may be an encoder-decoder 3D convolutional neural network (CNN). For example, in some cases, the molecular design computational model 115 may apply the denoising model 117 to denoise the 3D representation of the input molecule 152, or alternatively, to denoise the embedding 154 of the 3D representation of the input molecule 152 to generate the output molecule 162. In some cases, the denoising model 117 can denoise the three-dimensional representation (or its embedding 154) of the input molecule 152 through multiple consecutive sampling iterations, wherein a portion of the noise present in the three-dimensional representation (or its embedding 154) of the input molecule 152 is removed at each sampling iteration.
[0093] As described in more detail below, denoising the three-dimensional representation of the input molecule 152 can alter the composition and / or conformation (or three-dimensional structure) of the input molecule 152 so that the composition and conformation (or three-dimensional structure) of the resulting output molecule 162 is consistent with the composition and conformation of a molecule exhibiting one or more desired properties. For example, in the case that the output molecule 162 is a drug molecule, one or more desired properties may include drug-like properties such as affinity, specificity, bioactivity, and exploitability. In some cases, whether the output molecule 162 exhibits certain desired properties may depend on whether the output molecule 162 presents a corresponding conformation (or three-dimensional structure). Therefore, in some cases, applying the denoising model 117 to the three-dimensional representation (or its embedding 154) of the input molecule 152, rather than the one-dimensional or two-dimensional representation of the input molecule 152, increases the likelihood that the resulting output molecule 162 exhibits a conformation (or three-dimensional structure) consistent with one or more desired properties.
[0094] In some exemplary embodiments, a molecular design computational model 115 (including a denoised model 117) can be trained to learn or approximate a data distribution of molecules exhibiting one or more desired properties (e.g., drug-like properties such as affinity, specificity, bioactivity, and exploitability). For example, in some cases, the molecular design computational model 115 can be trained to approximate a data distribution of molecules exhibiting one or more desired properties based on a training dataset of known molecules exhibiting one or more desired properties (e.g., the PubChem dataset, the QM9 molecular dataset, the Geometry of Molecular Ensemble (GEOM) drug dataset, etc.). As described in more detail below, the molecular design computational model 115 can be trained to approximate a noisy data distribution filled with noisy three-dimensional representations of known molecules exhibiting one or more desired properties. Furthermore, the molecular design computational model 115 can be trained to operate in a discrete space filled with three-dimensional representations of molecules exhibiting one or more desired properties (e.g., voxelized representations), or alternatively in a latent space filled with embeddings of three-dimensional representations of molecules as lower-dimensional representations of molecules (e.g., voxelized representation embeddings).
[0095] To further illustrate, Figure 1A depicts an example of the molecular design engine 110 of Figure 1A, where the molecular design computational model 115 is trained to operate in a discrete space (three-dimensional voxelized representation), while the example of the molecular design engine 110 shown in Figure 1B includes a molecular design computational model 115 that can be trained to operate in a latent space. The latent space can include continuous values (embeddings) of multiple features. In either example of the molecular design engine 110, the molecular design computational model 115 can be trained to approximate a noisy data distribution, for example, by training based on a noisy three-dimensional representation of a molecule exhibiting one or more desired properties. For example, the example of the molecular design computational model 115 shown in Figure 1A can be trained to approximate a noisy discrete distribution filled with noisy three-dimensional representations of molecules, while the example of the molecular design computational model 115 shown in Figure 1B can be trained to approximate a noisy latent distribution filled with noisy embeddings of the three-dimensional representation of molecules.
[0096] Referring first to Figure 1A, in some cases, the molecular design engine 110 may include a molecular design computational model 115 and a recovery model 118. Alternatively, Figure 1B depicts another example of the molecular design engine 110, which may include an encoder 111 and a decoder 119 in addition to the molecular design computational model 115 and the recovery model 118. In some exemplary embodiments, the molecular design engine 110 may apply the molecular design computational model 115 to generate an output molecule 162 based at least on an input molecule 152. For example, in the example of the molecular design engine 110 shown in Figure 1A, the molecular design computational model 115 may operate on a three-dimensional representation of the input molecule 152, which in some cases may be a voxelized representation of the input molecule 152. In doing so, the molecular design computational model 115 may operate in a discrete voxelized space filled with noisy voxelized representations of different molecules, such as molecules exhibiting one or more desired properties. Alternatively, in the example of the molecular design engine 110 shown in Figure 1B, the molecular design computational model 115 can operate on the embedding 154 of the input molecule 152. As described in more detail below, the embedding 154 of the input molecule 152 can be generated by encoder 111 encoding a three-dimensional representation (e.g., a voxelized representation) of the input molecule 152. In this variant of the molecular design engine 110, the molecular design computational model 115 can operate in a latent space filled with noisy embeddings of three-dimensional representations (e.g., voxelized representations) of different molecules.
[0097] In some exemplary embodiments, the molecular design computational model 115 may include a denoising model 117 trained to denoise the three-dimensional representation of the input molecule 152 based on a function 175, such that the resulting three-dimensional representation of the output molecule 162 is sampled from a higher-density region of the data distribution of molecules exhibiting one or more desired properties. Figure 1A An example of a denoising model 117 is depicted, trained to denoise a three-dimensional representation of input molecule 152, while Figure 1B depicts another example of a denoising model 117 trained to denoise an embedding 154 of the three-dimensional representation of input molecule 152. In some cases, the denoising model 117 may denoise the three-dimensional representation of input molecule 152 (e.g., a voxelized representation of input molecule 152) over multiple time steps, or alternatively, denoise its embedding 154. In some cases, the denoising performed at each time step may be equivalent to selecting one or more samples (e.g., intermediate molecules) from different locations in the data distribution. In some cases, function 175 may be a scoring function that outputs a value (e.g., a score) indicating a local density change at a specific location within the data distribution (e.g., a location occupied by a molecule). Therefore, denoising of input molecule 152 can be performed based on the output of the scoring function, such that each successive sample (or molecule) is selected from regions of gradually increasing density in the data distribution.
[0098] In some exemplary embodiments, denoising the input molecule 152 may include updating the three-dimensional representation (e.g., a voxelized representation) (or its embedding 154) of the input molecule 152, which may represent the composition and conformation (or three-dimensional structure) of the input molecule 152, to increase the likelihood that the resulting output molecule 162 will exhibit one or more of the desired properties in a data distribution. In some cases, the molecular design computational model 115 may apply the denoising model 117 through one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Markov chain Monte Carlo (MCMC) with Langevin dynamics, etc.) to modify the three-dimensional representation (or its embedding 154) of the input molecule 152. For example, in some cases, each iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling may include selecting samples (or molecules) from the data distribution that include one or more modifications to the three-dimensional representation (or its embedding 154) of the input molecule 152. As described, in some cases, sampling from the data distribution may be guided by a function 175 (e.g., a scoring function). For example, when function 175 is a scoring function, for each sample (or molecule) selected from the data distribution, function 175 can output a value (e.g., a score) corresponding to the density change observed at the location in the data distribution occupied by the sample (or molecule). Therefore, in some cases, sampling from the data distribution can be guided by function 175 such that each successive sample is selected from a region of gradually increasing density in the data distribution.
[0099] To further illustrate, in some cases, the denoising model 117 can be applied to update the three-dimensional representation (e.g., voxelized representation) (or its embedding 154) of the input molecule 152 at least by, for example, selecting a first sample and a second sample from the data distribution. It should be understood that each of the first and second samples may correspond to a modified three-dimensional representation of the input molecule 152 in the example shown in FIG1A, or, in the case of FIG1B, to a modified embedding of the three-dimensional representation of the input molecule 152. When function 175 is a scoring function, function 175 may assign a first value (first score) to the first sample to indicate a more aggressive local change (e.g., an increase or a smaller decrease) in the density of the data distribution at a first location of the first sample, and assign a second value to the second sample to indicate a less aggressive local change (e.g., a smaller increase or decrease) in the density of the data distribution at a second location of the second sample. In some cases, the molecular design computational model 115 can apply a denoising model 117 to select a third sample from the data distribution by further modifying the first sample to sample from a higher density region of the data distribution compared to the first and second samples (e.g., another modified three-dimensional representation or another modified embedding).
[0100] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to denoise the voxelized representation (or its embedding 154) of the input molecule 152 instead of its conventional three-dimensional representation (such as a point cloud representation of the input molecule 152). For example, in some cases, the voxelized representation of the input molecule 152 may represent the types and positions of atoms present in the input molecule 152 as continuous (e.g., Gaussian-like) densities across a three-dimensional voxel grid. To indicate the positions of atoms present in the input molecule 152, each voxel in the voxel grid may be associated with a value indicating the atomic density at the corresponding position. In some cases, the atomic density associated with a particular voxel in the voxel grid may correspond to the probability that the voxel is part of an atom at that position. For example, a first voxel with a higher atomic density may be more likely to be part of an atom forming the input molecule 152 than a second voxel with a lower atomic density. Therefore, the voxelization representation of input molecule 152 can represent the positions of the atoms of input molecule 152, which distinguishes voxels in the voxel grid that form part of the atoms in input molecule 152 from voxels in the voxel grid that do not form part of the atoms in input molecule 152 based on the atomic density associated with each voxel in the voxel grid. In some cases, the atoms forming input molecule 152 can be arranged at the positions of those voxels that are associated with atomic densities that satisfy one or more thresholds.
[0101] In some exemplary embodiments, the voxelized representation of input molecule 152 may include one or more channels, each corresponding to a type of atom that can be present in input molecule 152. For example, in some cases, the voxelized representation of input molecule 152 may include a separate channel for each type of heavy atom that can be present in input molecule 152. Including different channels for different types of atoms in the voxelized representation of input molecule 152 can eliminate the discrete distribution typically associated with atom types found in conventional three-dimensional representations (e.g., point cloud representations, etc.). Instead, the voxelized representation of input molecule 152 may represent the types and positions of atoms in input molecule 152 as one or more continuous (e.g., Gaussian-like) densities across the aforementioned three-dimensional voxel grid. For example, the voxelized representation of input molecule 152 may include a first channel representing a first type of atom (e.g., carbon C1 atoms) that can be present in input molecule 152. The presence of the first type of atom in input molecule 152 and their respective positions may be represented by a first continuous (e.g., Gaussian-like) density across the first channel in the voxelized representation of input molecule 152. Note that the density is continuous; in this sense, the value associated with each voxel can take continuous values (e.g., values within a continuous, bounded, or unbounded distribution). In some cases, the voxelized representation of input molecule 152 can further include a second channel representing a second type of atom (e.g., nitrogen (N) atom) that may be present in input molecule 152. The presence of the second type of atom and its respective position can be represented by a second continuous (e.g., Gaussian-like) density of the second channel in the voxelized representation of the input molecule.
[0102] Unlike the conventional three-dimensional representation (e.g., point cloud representation, etc.) of the input molecule 152, which represents the types and positions of atoms in the input molecule 152 as two different types of distributions (e.g., a discrete distribution of atom types and a continuous distribution of atom positions), the voxelized representation of the input molecule 152 can represent the types and positions of atoms in the input molecule 152 together as one or more continuous (e.g., Gaussian-like) distributions in the manner described above. Therefore, the molecular design computational model 115 can apply the denoising model 117 to operate on the voxelized representation of the input molecule 152 without requiring a workaround to reconcile the two different types of distributions, which is necessary for the conventional three-dimensional representation of the input molecule 152. The voxelized representation of the input molecule 152 may also be more representative of the conformation (or three-dimensional structure) of the input molecule 152 compared to its conventional three-dimensional representation. For example, the voxelized representation of the input molecule 152 can capture long-range dependencies between distant atoms, even when the input molecule 152 contains a large number of atoms. Furthermore, the molecular design computational model 115 can apply a denoising model 117 to denoise the voxelized representation of the input molecule 152 and generate the output molecule 162 without any prior knowledge of the number of molecules present in the output molecule 162.
[0103] In some exemplary embodiments, training engine 120 can be trained to generate one or more training samples to be included in a training dataset. Figure 1A depicts an example of training engine 120, where damage engine 121 generates each training sample by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to a noisy three-dimensional representation 182 of a sample molecule (e.g., a known molecule exhibiting one or more desired properties), thereby generating a damaged three-dimensional representation 184 of the sample molecule. It should be understood that damage engine 121 can add noise to the already noisy three-dimensional representation 182 of the sample molecule in order to train molecular design computation model 115 to approximate a noisy data distribution of molecules exhibiting one or more desired properties, having a smoother density transition than the true data distribution of molecules exhibiting one or more desired properties. Thus, in some cases, molecular design computation model 115 (including denoising model 117) can be trained to denoise the damaged three-dimensional representation of the sample molecule, and for each training sample, the corresponding noisy three-dimensional representation of the sample molecule is recovered instead of the clean (or original) three-dimensional representation of the sample molecule. This is equivalent to sampling from a noisy data distribution of molecules exhibiting one or more desired properties, rather than from the true distribution of the molecules.
[0104] Alternatively, Figure 1B depicts another example of the training engine 120, where encoder 111 first encodes a noisy 3D representation 182 of a sample molecule to generate its embedding 186, and then corruption engine 121 adds noise to generate a corrupted embedding 188 of the noisy 3D representation 182 of the sample molecule. In this variation of the molecular design engine 110, the molecular design computational model 115 can be trained to approximate a noisy latent distribution of the noisy embeddings of the 3D representation of molecules exhibiting one or more desired properties, rather than the molecular design computational model 115 in the example shown in Figure 1A being trained to approximate a noisy but discrete distribution. To achieve this result, training engine 120 can generate each training sample in the training dataset to include the corrupted embedding 188. For example, Figure 1B shows that the training engine may further include an encoder 111, which may first encode a noisy three-dimensional representation 182 of a sample molecule (e.g., a known molecule exhibiting one or more desired properties) to generate an embedding 186, and then a corruption engine 121 adds noise to it to generate a corrupted embedding 188. In some cases, a molecular design computational model 115 (including a denoising model 117) may be trained to denoise the corrupted embedding 188 and recover the embedding 186 therefrom. The denoising model 117, the encoder 111, and the corresponding decoder may be trained together.
[0105] In some exemplary embodiments, training the molecular design computational model 115 may include adjusting one or more parameters (e.g., weights, biases, etc.) of the denoising model 117. In the example shown in FIG1A, the parameters of the denoising model 117 may be adjusted such that the denoising model 117 is able to recover the corresponding noisy three-dimensional representation 182 of the sample molecule from the damaged three-dimensional representation 184 of the sample molecule. For example, as described in more detail below, one or more parameters of the denoising model 117 may be adjusted over multiple iterations to increase (or maximize) the similarity between the noisy three-dimensional representation 182 of the sample molecule recovered by the denoising model 117 from the damaged three-dimensional representation 184 of the sample molecule and the original noisy three-dimensional representation 182 of the sample molecule (e.g., reducing (or minimizing) the mean squared error (MSE)). Alternatively, the example in Figure 1B shows that the parameters of the denoising model 117 can be tuned so that the denoising model 117 can recover the embedding 186 generated by encoding the noisy three-dimensional representation 182 of the sample molecule from the damaged embedding 188. For example, in the example shown in Figure 1B, the parameters of the denoising model 117 can be tuned to increase (or maximize) the similarity between the embedding 186 of the three-dimensional representation 182 of the sample molecule recovered by the denoising model 117 from the damaged embedding 188 and the original undamaged embedding 186 (e.g., reduce (or minimize) any loss function, such as mean squared error (MSE), that quantifies the difference between the embedding 186 of the three-dimensional representation 182 of the sample molecule recovered by the denoising model 117 from the damaged embedding 188 and the original undamaged embedding 186).
[0106] In some exemplary embodiments, training the molecular design computational model 115 may include determining a function 175, which can be parameterized by parameters of the denoising model 117 (e.g., weights, biases, etc.). For example, in some cases, training the molecular design computational model 115 (including adjusting one or more parameters of the denoising model 117) may determine the function 175 (e.g., a scoring function) at least by adjusting the corresponding parameters of the function 175. In some cases, the function 175 may approximate different densities of molecules across a data distribution of one or more desired properties, which molecules are more likely to occupy higher density regions of the data distribution. In the example shown in FIG1A, this data distribution may be a noisy data distribution filled with noisy three-dimensional representations of molecules. Alternatively, in the example of the molecular design engine 110 shown in FIG1B, this data distribution may be a noisy latent distribution filled with noisy embeddings of three-dimensional representations of molecules.
[0107] As described, in some exemplary embodiments, training the denoising model 117 to approximate a noisy data distribution filled with noisy three-dimensional representations of molecules (such as noisy voxelized representations) can prevent the molecular design computation model 115 (including the denoising model 117) from overfitting to known molecules in the training dataset. Figure 1A shows such an example where the molecular design computation model 115 has been trained to recover a noisy three-dimensional representation 182 of the sample molecule instead of a clean (or original) three-dimensional representation of the sample molecule. As described in more detail below, once trained, the molecular design computation model 115 can generate a three-dimensional representation of the output molecule 158 by traversing the smooth density of the noisy data distribution (i.e., iteratively sampling different regions of it) to sample at least one updated three-dimensional representation 160. The noisy data distribution can be filled with noisy three-dimensional representations of molecules exhibiting one or more desired properties. Therefore, the updated three-dimensional representation 160 can include at least some noise that can be removed by applying the recovery model 118 to denoise the updated three-dimensional representation 160. The recovery model 118 can be a machine learning model that has been trained to take a noisy three-dimensional representation 160 of a molecule as input and produce a corresponding denoised three-dimensional representation as output. This generates a three-dimensional representation of the output molecule 152 that occupies the real data distribution of clean three-dimensional representations of molecules exhibiting one or more desired properties.
[0108] Alternatively, Figure 1B depicts another example of the molecular design engine 110, where the molecular design computational model 115 is trained to approximate a noisy latent distribution filled with embeddings of noisy 3D representations of molecules exhibiting one or more desired properties. In this variation of the molecular design computational model 115, a denoising engine 117 can be trained to recover embeddings 186 of the noisy 3D representation 182 of the sample molecule from the damaged embeddings 188. Once trained, the molecular design computational model 115 can apply the denoising model 117 to denoise the embeddings 154 generated by the encoder 111 encoding the 3D representation of the input molecule 152. Denoising may include the denoising model 117 updating the embeddings 154 over multiple successive sampling iterations to generate at least one updated embedding 156 during each sampling iteration. Doing so is equivalent to the molecular design computational model 115 selecting samples from the noisy latent distribution, and the molecular design computational model 115 can continue selecting samples from it until one or more criteria are met. In some cases, for example, the updated embedding 156 can be decoded by decoder 119, and then the resulting noisy 3D representation 158 can be denoised by recovery model 118 to generate a 3D representation of output molecule 162. As described in more detail below, in this example of molecular design engine 110, molecular design computation 115 can operate in a noisy latent space filled with noisy embeddings of the molecule's 3D representation rather than noisy 3D representations found in discrete voxelized space. Furthermore, it should be understood that although denoising model 117 and recovery model 118 may share the same architecture (e.g., artificial neural network (ANN) etc.) in some cases, the two models are trained to remove different noises. For example, denoising model 117 can be trained to denoise the 3D representation of input molecule 152 or its embedding 154 such that the resulting 3D representation of output molecule 162 is consistent with the composition and / or conformation of a molecule exhibiting one or more desired properties (e.g., drug-like properties). In contrast, recovery model 118 can be trained to remove noise, which is added to smooth the density of known molecules that can be used to train molecular design computation model 115.
[0109] Referring again to Figure 1B, the embedding 154 of the three-dimensional representation of the input molecule 152 can be generated by encoder 111 encoding the three-dimensional representation of the input molecule 152. In some cases, the embedding 154 can be a lower-dimensional representation of the three-dimensional representation of the input molecule 152 generated by encoder 111 reducing the dimension of the three-dimensional representation of the input molecule 152. For example, in some cases, encoder 111 and decoder 119 can form an autoencoder, including, for example, a variational autoencoder (VAE), such as a vector quantization variational autoencoder (VQ-VAE). Encoder 111 can generate the embedding 154 at least by reducing the dimension of the three-dimensional representation of the input molecule 152 (e.g., a voxelized representation). In this context, reducing the dimensionality of the three-dimensional representation of the input molecule 152 can include downsampling, compressing, or reducing the dimensionality of the three-dimensional representation of the input molecule 152, for example, by condensing at least some of the features (e.g., atomic density values) present in the three-dimensional representation of the input molecule 152, such that the resulting embedding 154 includes fewer features than the original three-dimensional representation of the input molecule 152, but those features still capture the same (or similar) information conveyed in the original three-dimensional representation of the input molecule 152. For a voxelized representation of the input molecule 152, each feature present can correspond to an atomic density value associated with each voxel included in the voxel grid representing the input molecule 152. For example, in the case where the voxelized representation of the input molecule 152 includes a [32×32×32] voxel grid, the voxelized representation of the input molecule 152 can include 32,000 features (or atomic density values). When generating the embedding 154, at least some of those 32,000 features can be condensed by the encoder 111. In doing so, embedding 154 can include fewer features, such as 4×4×4=64 features, so that the denoising model 117 can operate after generating the output molecule 162.
[0110] As described, the embedding 154 of the three-dimensional representation of the input molecule 152 shown in Figure 1B can include fewer features than the original three-dimensional representation of the input molecule 152. Furthermore, the embedding 154 can be generated by the encoder 111 downsampling, compressing, or reducing the dimensionality of the three-dimensional representation of the input molecule 152. Doing so is equivalent to the encoder 111 mapping the three-dimensional representation of the input molecule 152 from a high-dimensional discrete voxel space to a lower-dimensional latent space. In some cases, the molecular design computational model 115 (e.g., a denoising engine 117) can denoise the embedding 154, for example, through one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling. During each iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling, the molecular design computational model 115 (e.g., the denoising model 117) can sample at least one updated embedding 156 from this lower-dimensional latent space. The embedding 154 can include several orders of magnitude fewer features than the original three-dimensional representation of the input molecule 152. This means that the molecular design computational model 115 can apply the denoising model 117 to the embedding 154 and generate the output molecule 162 faster and with higher computational efficiency, while achieving comparable or better performance in both qualitative and quantitative aspects. This can be particularly advantageous in applications that require generating a large number of candidate molecules in a short timeframe, such as computational drug design. In some cases, reducing the dimensionality of the original three-dimensional representation of the input molecule 152 can enable the denoising model 117 to operate on and generate larger molecules (e.g., molecules containing more than 200 atoms) and a greater number of molecules.
[0111] In some cases, the compactness of embedding 154 relative to the three-dimensional representation of input molecule 152 also means that fewer computational resources are required when operating on embedding 154. For example, in order for denoising model 117 to directly address the three-dimensional representation of input molecule 152 (e.g., with 32,000 features)... The denoising model 117 can be implemented with a large number of trainable parameters (e.g., 100 million parameters) when operating on a voxel grid. In contrast, if the denoising model 117 is instead applied to operate on embedding 154, it can be implemented with far fewer trainable parameters. Implementing the denoising model 117 with fewer parameters can improve its performance because a larger number of parameters can reduce its generalization ability. For example, the denoising model 117 may be particularly prone to overfitting when it includes a large number of features but there are relatively few known molecules available to train it. When the denoising model 117 overfits to known molecules in the training dataset, meaning it has been overtrained on the training dataset (i.e. trained to learn a data distribution that is too tightly centered on molecules in the training dataset), it may fail to generalize. In this context, generalization refers to the ability to accurately denoise the input molecule 152 even when it is not one of the known molecules in the training dataset. Therefore, overfitting the denoising model 117 can prevent it from accurately denoising any molecule that is not one of the known molecules in the training data.
[0112] Referring again to Figure 1B, in the example of the molecular design engine 110 shown therein, the updated embedding 156 generated by sampling from the latent voxelized space can be decoded by decoder 119, and then the resulting noisy 3D representation 158 is denoised by recovery engine 118. While the decoding performed by decoder 119 maps the updated embedding 156 from the latent space to the discrete space, the subsequent denoising performed by recovery model 118 can constitute a jump from the noisy data distribution back to the true data distribution of the molecule. For example, during each round of gradient-based Markov chain Monte Carlo (e.g., Langevin Markov chain Monte Carlo, etc.), molecular design computational model 115 can apply denoising model 117 to sample from the noisy latent distribution of molecules exhibiting one or more desired properties. In this case, sampling from the noisy latent distribution may include updating the embedding 154 of the 3D representation (e.g., voxelized representation) of the input molecule 152. This is equivalent to selecting at least one updated embedding 156 from the data distribution, which is then decoded by decoder 119 and subsequently denoised by recovery model 118 to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162. As described in more detail below, molecular design computational model 115 can further update the noisy embedding 156, for example, through multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling, until one or more criteria are met.
[0113] Once one or more criteria are met, the molecular design engine 110 can recover the output molecule 162 based at least on its three-dimensional representation. For example, in some cases, the molecular design engine 110 can recover the positions (e.g., coordinates) of atoms present in the output molecule 162 and one or more bonds between them based at least on its three-dimensional representation. Recovering the positions of atoms present in the output molecule 162 can be performed by identifying the maximum value (e.g., peak) of the predicted density included in the three-dimensional representation of the output molecule 162. In doing so, the molecular design engine 110 can determine another representation of the output molecule 162, including, for example, a one-dimensional representation of the output molecule 162 (e.g., a simplified linear input specification (SMILES) string), a two-dimensional representation of the output molecule 162 (e.g., a molecular graph), etc. It should be understood that the output molecule 162 generated in this way may be more likely to exhibit one or more desired properties of the molecule in the data distribution. In particular, the denoising model 117 can generate an output molecule 152 to exhibit a composition and / or conformation (or three-dimensional structure) consistent with one or more desired properties. By operating on the embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152, the denoising model 117 can generate the output molecule 162 faster and with less computational burden.
[0114] As described, in some exemplary embodiments, the molecular design computational model 115 can be trained to recover a noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation of the sample molecule), which is generated by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) to the three-dimensional representation of the sample molecule. In the example shown in FIG1A, the molecular design computational model 115 can be trained to recover the noisy three-dimensional representation 182 of the sample molecule by at least modifying the damaged three-dimensional representation 184 of the sample molecule to denoise it. Alternatively, FIG1B shows an example of a molecular design engine 110 in which the molecular design computational model 115 is trained to recover the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule. As shown in Figure 1B, in some cases, the molecular design computational model 115 can denoise the damaged embedding 188 by at least modifying the damaged embedding 186 generated by adding noise (e.g., Gaussian noise, etc.) to the embedding 184 of the 3D representation 182 of the sample molecule by the damaged engine 121. As described, the embedding 184 can be generated by the encoder 111 by downsampling or reducing the dimensionality of the noisy 3D representation 182 of the sample molecule (e.g., a noisy voxelized representation). However, as described, the downsampling of the noisy 3D representation 182 of the sample molecule (e.g., a noisy voxelized representation) can be optional, which may be the case when the encoder 111 implements the identity function. Therefore, in some cases, the embedding 184 may include the same number of features as the original noisy 3D representation 182 of the sample molecule (e.g., a noisy voxelized representation). In those cases, the embedding 154 can capture the same information present in the original 3D representation (e.g., a voxelized representation) of the input molecule 152 without condensing the features present therein. In other words, the encoding of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 can be an optional operation, even in the case of including encoder 111 in the example of the molecular design system 100 shown in Figure 1B.
[0115] Figure 2 depicts a flowchart illustrating an example of a process 200 for enabling machine learning to generate three-dimensional molecules in voxelized space according to some exemplary embodiments. Referring to Figures 1A-1B and Figure 2, process 200 can be executed by a molecular design engine 110 to train and apply a molecular design computational model 115 to generate an output molecule 162, at least by denoising a three-dimensional representation (such as a voxelized representation) of an input molecule 152. For example, in some cases, the molecular design computational model 115 can be trained on a noisy three-dimensional representation 182 of a sample molecule, which is generated by adding noise to the original three-dimensional representation of the sample molecule, such that the molecular design computational model 115 is trained to approximate a noisy data distribution with a smoother density transition. Figure 1A shows a variation in which the molecular design computational model 115 is trained on a damaged three-dimensional representation 184 of a sample molecule, which can be generated by adding additional noise to the noisy three-dimensional representation 182 of the sample molecule without any downsampling or compression. Alternatively, Figure 1B shows another variation in which the molecular design computational model 115 is trained on a damaged embedding 186 of a noisy three-dimensional representation 182 of the sample molecule. This damaged embedding 188 can be generated by adding noise to the embedding 154 of the noisy three-dimensional representation 182 of the sample molecule, which is generated by encoder 111 downsampling (or compressing) features present in the noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation). In other words, it should be understood that the molecular design computational model 115 can be trained to operate in a noisy discrete voxelized space filled with the noisy three-dimensional representation of the molecule, or alternatively in a noisy latent voxelized space filled with the embeddings of the noisy three-dimensional representation of the molecule.
[0116] At 202, training engine 120 can generate a training dataset to include multiple damaged sample molecules. In some exemplary embodiments, generating the training dataset may include training engine 120 generating a training dataset to include multiple damaged sample molecules. This training dataset can then be used to train molecular design computational model 115 to approximate a data distribution of molecules exhibiting one or more desired properties (e.g., drug-like properties). In some cases, each damaged sample molecule may be a noisy three-dimensional representation of a known molecule, which is further corrupted by noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.). For example, Figure 1A shows such an example, where corruption engine 121 generates a damaged three-dimensional representation 184 of the sample molecule by adding additional noise to the noisy three-dimensional representation 182 of the sample molecule. As described in more detail below, molecular design computational model 115 (e.g., denoising model 117) can be trained to recover the noisy three-dimensional representation 182 of the sample molecule from the damaged three-dimensional representation 184 of the sample molecule to approximate a noisy data distribution with a smoother density transition. Alternatively, Figure 1B shows another example where training engine 120 generates a damaged embedding of each damaged sample molecule in the training dataset to include a damaged embedding of the noisy three-dimensional representation 182 of the sample molecule. For example, in some cases, the damage engine 121 can generate a damaged embedding 188 by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) to the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule (e.g., a noisy voxelized representation). In some cases, training engine 120 can further augment the training dataset by applying one or more transformations to the noisy three-dimensional representation 182 of the sample molecule (e.g., a voxelized representation), including, for example, translation (e.g., moving the center of the sample molecule in each of the three dimensions by sampling uniform shifts), rotation (e.g., sampling uniformly over the three Euler angles), reflection, etc.
[0117] Training a molecular design computational model based on a noisy 3D representation of sample molecules can mitigate overfitting and mode collapse, which often occurs when the molecular design computational model is trained to approximate high-dimensional data distributions (e.g., 10) based on too few known molecules (e.g., PubChem dataset, QM9 molecular dataset, Geometry of Molecular Omnibus (GEOM) drug dataset, etc.). 60 When considering the molecular space of a possible chemical compound.
[0118] In the example of the molecular design computational model 115 shown in Figures 1A-1B, the molecular design computational model 115 may include a denoising model 117. In this case, training the molecular design computational model 115 may include training the denoising model 117 to approximate a noisy data distribution, or in some cases, an approximate noisy latent distribution, either of which exhibits a smoother density transition and samples from it more efficiently than the true data distribution. In some cases, the denoising model 117 may be an artificial neural network (ANN), in which case training the denoising model 117 may include adjusting one or more parameters of the artificial neural network (ANN) (e.g., weights, biases, etc.). Doing so may also determine the parameters of function 175 such that function 175 outputs values indicating the probability that molecules exhibiting one or more desired properties are located at a specific position within the data distribution. For example, in some cases, function 175 may be a scoring function whose output is a value (e.g., a score) indicating the transition between different density regions of the data distribution, including, for example, a transition between higher density regions occupied by molecules more likely to exhibit one or more desired properties and lower density regions of the data distribution less likely to be occupied by molecules exhibiting one or more desired properties. In some cases, the denoising model 117 can be trained to recover a noisy three-dimensional representation 182 or embedding 186 of the sample molecule to avoid overfitting the denoising model 117 to, for example, a relatively small number of known molecules in the training dataset that can be used to characterize the data distribution.
[0119] To further illustrate, the use of This represents the true data distribution of voxelized representations of molecules exhibiting one or more desired properties, and uses... This represents the corresponding noisy data distribution, which may exhibit a smoother energy landscape and sample more efficiently from it than from an unknown data distribution. In some cases, the true data distribution... It may be unknown, which means that the denoising model 117 can be trained on a training dataset of known molecules from the real data distribution to approximate the real data distribution. To avoid overfitting the denoising model 117 to the training dataset, it is possible to train the denoising model 117 to approximate the distribution of noisy data. In some cases, noisy data distributions... It can be done by using a known covariance Gaussian kernel (e.g., isotropic Gaussian kernel) convolution real data distribution This can be obtained by sending data from a real data distribution. Voxelization of molecules Add noise To generate noisy voxelized molecular representations (For example, ,in , Given the above formula, there is a noisy voxelized molecule representation. It can be sampled from the following noisy data distribution. :
[0120]
[0121] Transform the distribution of real data in this way It can smooth the distribution of real data. The density, while still retaining the presence of a clean (or original) voxelized representation. Some of the structural information in the data, without any additional noise. If added to a clean voxelized molecule, it represents... Noise in If it is Gaussian (e.g., isotropic Gaussian), then the least squares estimator can be obtained by applying the following equation (1). Directly from the corresponding noisy voxelized molecules The restoration of clean voxelized molecules in the middle It should be understood that the least squares estimator It can act as a denoiser and remove noise present in voxelized molecules. Noise in To restore clean voxelized molecular representation .
[0122]
[0123] in, Corresponding to noisy data score function of distribution Equation (1) indicates that if there is noisy data distribution Based on the normalization constant (and the corresponding score function therefrom) If it is known, then it can be determined from its noisy counterpart. Estimate the clean voxelized molecular representation Similarly, noisy data distribution Scoring function It can also be based on real data distribution Least squares estimator To derive. As described in more detail below, in determining the scoring function. Alternatively, after determining the corresponding scoring function in some cases, the molecular design computational model 115 can apply the denoising model 117 based on the scoring function. (or the corresponding scoring function) from the distribution of noisy data Sampling.
[0124] As described in more detail below, once the denoising model 117 has been trained, the molecular design computational model 115 can apply the denoising model 117 to perform the "walking and jumping" generation process to generate data based on the real data distribution. The output molecule exhibits one or more desired properties of the molecule. For example, in some cases, the denoising model 117 can undergo multiple consecutive sampling iterations from the noisy data distribution. Mid-sampling, each sampling iteration includes denoising model 117 at least by performing noisy voxelization on the molecular representation. Denoising is achieved from the distribution of noisy data. Select at least one sample. In some cases, the data distribution is noisy. The sampling can be determined by the scoring function. Guidance ensures that the samples selected during a single sample iteration originate from a noisy data distribution. The location in the middle that differs from the sample selected during another sample iteration. This applies to noisy data distributions. The traversal is called the "walk" part of the generation process.
[0125] In some cases, the scoring function Can be based on conditions (For example, the gradient of a classifier) will molecule Sampling is limited to certain regions within the noisy data distribution. Instead of from the entire noisy data distribution Free sampling is used. Therefore, in some cases, each consecutive sample can be obtained from a noisy data distribution. Selected from regions where density gradually increases, such as by the scoring function. The output score is shown. Similarly, as described, it is derived from the scoring function. Guided traversal of noisy data distribution This can be considered a "walking" noisy data distribution. From the corresponding noisy voxelized molecule representation The restoration of clean voxelized molecules in the middle This may constitute a distribution from noisy data. Return to the real data distribution The "leap" in this context refers to the process of denoising data distributions. In some cases, this can be achieved by applying a denoiser (such as Denoising Engine 117). Return to the real data distribution The "jump" is used to remove the presence of noisy voxelized molecules. Noise in And restore the corresponding clean voxelized molecular representation. For example, in some cases, clean voxelized molecules represent... The least squares estimator can be obtained by the denoising engine 117. Application to noisy voxelized molecular representation To restore.
[0126] In some exemplary embodiments, training engine 120 can generate each damaged sample molecule in the training dataset by adding noise to the embedding of the noisy 3D representation of the sample molecule. For example, as shown in Figure 1B, in some cases, damage engine 121 can generate a damaged embedding 188 by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise, etc.) to the embedding 186 of the noisy 3D representation 182 of the sample molecule (e.g., a noisy voxelized representation) at least, rather than directly to the noisy 3D representation 182 of the sample molecule. Figure 1B further shows that embedding 186 can be generated by encoder 111 downsampling (or compressing) the noisy 3D representation 182 of the sample molecule. Downsampling (or compressing) the noisy 3D representation 182 of the sample molecule can reduce the dimensionality of the noisy 3D representation 182 of the sample molecule. For example, when the noisy 3D representation 182 of the sample molecule contains 32,000 features (or atomic density values), When using voxel meshes, downsampling (or compression) can produce a dataset containing 64 features (or atomic density values). Voxel grids. Therefore, downsampling (or compression) of the noisy 3D representation of a sample molecule can increase the overall speed and efficiency of the generation process. In some cases, large numbers of candidate molecules, such as tens of thousands or even millions, can be generated in a short period to support low-yield applications, such as computational drug design, where most candidate molecules fail to be successfully synthesized in the lab. Downsampling (or compression) of the 3D representation of a sample molecule can also enable the generation of larger molecules (e.g., molecules containing more than 200 atoms) that would be overly cumbersome to manipulate if the original 3D representation of the molecule were preserved without any downsampling (or compression).
[0127] At 204, the molecular design engine 110 may train the molecular design computation model 115 at least by applying the molecular design computation model 115 to recover the 3D representation of each damaged sample molecule in the training dataset from the damaged 3D representation of the sample molecule. In some exemplary embodiments, the training step 204 of the molecular design computation model 115 may include training a denoising model 117 at least based on the training dataset to approximate a noisy data distribution (or noisy latent distribution) of molecules that exhibit one or more desired properties. For example, in the example shown in FIG1A, the denoising model 117 may be trained to recover the noisy 3D representation 182 (e.g., voxelized representation) of the sample molecule from the damaged 3D representation 184 of the sample molecule. In the example shown in FIG1B, the denoising model 117 may be trained to recover the embedding 186 of the noisy 3D representation 182 of the sample molecule from the damaged embedding 188. It should be understood that, in either example, the denoising engine 117 may be trained on a noisy three-dimensional representation 182 of the sample molecules rather than a clean three-dimensional representation of the sample molecules, so that the denoising engine 117 approximates a noisy data distribution with a smoother density transition than the true data distribution.
[0128] In the example shown in Figure 1A, training the molecular design computational model 115 may include adjusting one or more parameters (e.g., weights, biases, etc.) of the denoising model 117 to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the noisy three-dimensional representation 182 of the sample molecule recovered by the denoising model 117 and the original noisy three-dimensional representation 182 of the sample molecule. Furthermore, training the denoising model 117 may include determining a function 175, which is parameterized by the parameters (e.g., weights, biases, etc.) of the denoising model 117. For example, in some cases, function 175 may be a scoring function that outputs values (e.g., scores) indicating local variations in the density (or gradient) of the data distribution. Thus, in some cases, function 175 may output a first value (e.g., a first score) indicating a first local variation in the density of the data distribution at a first location occupied by a first molecule and a second value (e.g., a second score) indicating a second local variation in the density of the data distribution at a second location occupied by a second molecule. In some cases, a first value (e.g., a first score) can indicate a more aggressive local change (e.g., an increase or a small decrease) in the density of the data distribution at a first position of the first molecule, while a second value (e.g., a second score) can indicate a less aggressive local change (e.g., a small increase or decrease) in the density of the data distribution at a second position of the second molecule. When function 175 is a scoring function, the sampling of the data distribution can be guided by the value (e.g., the score) output by function 175. As described in more detail below, the sampling of the data distribution can be guided by function 175 such that samples (or molecules) are selected from regions where the density of the data distribution gradually increases, regions more likely to be occupied by molecules exhibiting one or more desired properties.
[0129] In the example shown in Figure 1B, training the denoising model 117 may include adjusting one or more parameters of the denoising model 117 (e.g., weights, biases, etc.) to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the original undamaged embedding 186 and the embedding 186 of the noisy 3D representation 182 of the sample molecule recovered by the denoising model 117 from the damaged embedding 188. Doing so may also adjust the parameters of function 175. Similarly, if function 175 is a scoring function, the parameters of function 175 may be adjusted such that function 175 outputs a higher value (e.g., a higher score) for the first molecule occupying the first position in the data distribution, compared to a value for the second molecule occupying the second position in the data distribution that exhibits a less aggressive local change in the density of the data distribution, which exhibits a more aggressive local change in the density of the data distribution (e.g., a positive gradient indicating a transition from a lower density to a higher density region in the data distribution).
[0130] In some exemplary embodiments, the molecular design engine 110 may train a molecular design computational model 115 (including a denoised model 117) to approximate a function 175 by performing gradient-based Markov chain Monte Carlo (MCMC) sampling (such as Markov chain Monte Carlo (MCMC) sampling with Langevin dynamics). In some cases, function 175 may output a value (e.g., a score) indicating a transition between different density regions of the data distribution. For example, as a scoring function, the value (e.g., a score) output by function 175 for each molecule may indicate a local change in density (or gradient) at the corresponding location in the data distribution. In the example shown in Figure 1A, the gradient-based Markov chain Monte Carlo (MCMC) sampling determination function 175 may include undergoing multiple iterations to adjust the parameters of the denoising model 117 (e.g., weights, biases, etc.) and the parameters of function 175 (e.g., weights, biases, etc.) to increase (or maximize) the similarity between the noisy 3D representation 182 of the sample molecule recovered by the denoising model 117 and the original 3D representation 182 of the sample molecule (e.g., by reducing (or minimizing) the mean squared error (MSE)). For the example shown in Figure 1B, one or more parameters of the denoising model 117 and one or more parameters of the function 175 can be tuned through multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) to increase (or maximize) the similarity between the embedding 186 of the noisy 3D representation 182 of the sample molecule recovered by the denoising engine 117 from the damaged embedding 188 and the original undamaged embedding 186 (e.g., by reducing (or minimizing) the mean squared error (MSE)).
[0131] As described, the denoising model 117 can be trained to recover the noisy three-dimensional representation 182 of the sample molecules (e.g., a noisy voxelized representation) to avoid overfitting the denoising model 117 to known molecules that can be used to train the denoising model 117. When there are relatively few known molecules characterizing a high-dimensional data distribution, directly training the denoising model 117 based on known molecules can produce an overly jagged energy landscape, where there are sharp gradient changes between regions filled by known molecules. Sampling from the data distribution guided by steep gradients can prevent sufficient exploration of the data distribution, at least because the steepness of the gradients may limit sampling to regions within the immediate vicinity of known molecules. In contrast, training the denoising model 117 based on the noisy three-dimensional representation 182 of the sample molecules (e.g., a noisy voxelized representation) can produce smoother density transitions, where the gradient of function 175 is more asymptotic, thus enabling more efficient exploration of the data distribution when sampling from it.
[0132] At 206, the molecular design engine 110 can apply a trained molecular design computation model 115 to generate an output molecule, at least by denoising the voxelized representation of the input molecule. In some exemplary embodiments, the molecular design computation model 115 can use a denoising model 117 to generate an output molecule 162, at least by updating the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 under the guidance of function 175, or in some cases by updating its embedding 154. For example, the molecular design computation model 115 of FIG. 1A can directly update the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 without any downsampling or compression. Alternatively, the molecular design computation model 115 of FIG. 1B can update the embedding 154 of the three-dimensional representation of the input molecule 152, which can be generated by the encoder 111 by downsampling (or compressing) the three-dimensional representation (e.g., voxelized representation) of the input molecule 152. This reduces the dimensionality (or number of features) of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152, allowing the resulting embedding 154 to be more compact than the original (or uncompressed) three-dimensional representation of the input molecule 152. For example, while the original three-dimensional representation (e.g., voxelized representation) of the input molecule 152 could include 32,000 features (or atomic density values)... Voxel meshes, but embedding 154 can include 64 features (or atomic density values). Voxel grid. It should be understood that encoder 111 can be trained to downsample (or compress) the voxelized representation of input molecule 152 such that the resulting embedding 154 conveys the same (or similar) information as the voxelized representation of input molecule 152 in its original (or uncompressed) form. In some cases, encoder 111 may be part of an autoencoder (e.g., a variational autoencoder (VAE), such as a vector quantization variational autoencoder (VQ-VAE)) that also includes decoder 119. In some cases, encoder 111 can be trained to generate embedding 154 such that decoder 119 is able to recover the original voxelized representation of input molecule 152 at least by decoding embedding 154.
[0133] In some cases, the molecular design computational model 115 can denoise the input molecule 152 at least by updating the three-dimensional representation of the input molecule 152 or, alternatively, an embedding 154 of the three-dimensional representation of the input molecule 152. As described, in some cases, the three-dimensional representation of the input molecule 152 can be a voxelized representation of the input molecule 152, wherein the types and positions of atoms present in the input molecule 152 are represented as a continuous (e.g., Gaussian-like) atomic density centered on the atoms. For example, in some cases, the voxelized representation of the input molecule 152 may include... Amount of voxels [ A voxel grid, where each voxel is associated with a value indicating the atomic density at the corresponding location. In some cases, the atomic density associated with a single voxel can have values ranging from a range of values, such as... The atomic density at the lower end of this range indicates that the voxel is farther away from any atom in the input molecule 152, and the atomic density at the upper end of this range indicates that the voxel is closer to the center of the atoms in the input molecule 152. Furthermore, in some cases, the voxelized representation of the input molecule 152 may include multiple channels, where each channel corresponds to a different type of atom that may be present in the input molecule 152. Therefore, in some cases, the voxelized representation of the input molecule 152 can represent the types and positions of atoms present in the input molecule 152 as continuous (e.g., Gaussian-like) atomic densities across one or more channels.
[0134] In some exemplary embodiments, denoising of the input molecule 152 may include updating the three-dimensional representation of the input molecule 152, or alternatively, updating the embedding 154 of the input molecule 152. When the molecular design computational model 115 operates on the three-dimensional representation of the input molecule 152, or when the embedding 154 is generated without any downsampling (or compression) of the three-dimensional representation of the input molecule 152, denoising may include updating the atomic density of one or more voxels in at least one channel of the noisy voxelized representation of the input molecule 152. Doing so may be equivalent to adding, removing, and / or repositioning one or more atoms of different atom types in the input molecule 152. For example, increasing (or decreasing) the atomic density of one or more voxels in one channel of the noisy voxelized representation of the input molecule 152 may be equivalent to adding (or removing) atoms of the corresponding type to the input molecule 152. Alternatively and / or additionally, reducing the atomic density of the first voxel while increasing the atomic density of the second voxel can be equivalent to repositioning the atom from the first position of the first voxel to the second position of the second voxel.
[0135] Alternatively, when the molecular design computational model 115 operates on an embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152, denoising may include updating the values of voxels present in the embedding 154. As described, the embedding 154 may be generated by encoder 111 by condensing at least some of the features (e.g., atomic density values) present in the voxelized representation of the input molecule 152. The embedding 154 may include fewer features than the original voxelized representation of the input molecule 152, but still convey the same (or similar) information as the original three-dimensional representation of the input molecule 152. Therefore, denoising the embedding 154 may include updating one or more values present in the embedding 154, at least some of which may represent multiple features (or atomic density values) from the original voxelized representation of the input molecule 152.
[0136] Updating the three-dimensional representation of the input molecule 152 or its embedding 154 in the manner described above may include selecting samples (or updated molecules) from a noisy data distribution (or noisy latent distribution) of molecules exhibiting one or more desired properties. In the case of gradient-based Markov chain Monte Carlo (MCMC) sampling, the update may be guided by the output of function 175 (e.g., the score output by function 175) such that the samples (or updated molecules) selected during each successive sampling iteration originate from regions of gradually increasing density in the noisy data distribution, regions that are more likely to be filled by molecules exhibiting one or more desired properties.
[0137] To further illustrate, in some cases, the three-dimensional representation of the input molecule 152 or its embedding 154 may undergo a first update and a second update. Doing so is equivalent to selecting a first sample (or a first updated molecule) and a second sample (or a second updated molecule) from a noisy data distribution. In some cases, after selecting the first sample (or the first updated molecule) and the second sample (or the second updated molecule) from a noisy data distribution (or a noisy latent distribution), the molecular design computational model 115 may apply a function 175 to determine a value (e.g., a score) indicating the probability of each sample (or updated molecule) within the noisy data distribution (or noisy latent distribution). Where function 175 is a scoring function, for example, a higher value (e.g., a lower score) may indicate that the sample (or updated molecule) is selected from a noisy data distribution region exhibiting a more positive local change in density (e.g., an increase or a smaller decrease), or similarly, that the sample (or updated molecule) has a higher probability of being within the noisy data distribution. Therefore, in some cases, after selecting a first sample (or a first updated molecule) and a second sample (or a second updated molecule), the molecular design computational model 115 can apply a denoising model 117 to continue updating the three-dimensional representation or its embedding 154 of the input molecule 152 in order to select additional samples (or further updated molecules) from regions of gradually increasing density in the noisy data distribution, until, for example, a sample (or updated molecule) exhibiting a threshold probability within the noisy data distribution (or noisy latent distribution) is selected. For example, in some cases, if the three-dimensional representation or embedding 154 of the input molecule 152 with the first updated (or first updated molecule) is selected from a higher density region of the data distribution, the denoising model 117 can be applied to further modify the three-dimensional representation or embedding 154 of the input molecule 152 with the first updated (or first updated molecule) instead of the second updated (or second updated molecule). Doing so can be analogous to traversing the noisy data distribution (or noisy latent distribution) to sample from regions of gradually increasing density in the noisy data distribution. In the case where the denoising model 117 is modifying the embedding 154 of the three-dimensional representation of the input molecule 152, the denoising model 117 can operate in a noisy latent space, in which the distance between two or more embeddings can reflect the similarity (or dissimilarity) of the types and positions of atoms in different molecules. Dramatic shifts in density present in the real data distribution of molecules exhibiting one or more desired properties can be smoothed by adding noise.Since the denoising model 117 has been trained to approximate the data distribution of molecules exhibiting certain desired properties (e.g., drug-like properties), the updates to the embedding 154 made when denoising the input molecule 152 can be consistent with the types and positions of atoms found in molecules exhibiting one or more desired properties. Therefore, the same desired properties can also exist in the output molecule 162 generated by applying the denoising model 117 to denoise the input molecule 152 using the molecular design computational model 115.
[0138] Figure 3A depicts a flowchart illustrating an example of a process 300 for training a molecular design computational model 115 according to some exemplary embodiments. Referring to Figures 1-2 and 3A, process 300 may implement operation 204 of process 200 shown in Figure 2. In some cases, process 300 may be performed by molecular design engine 110 to train molecular design computational model 115 (including, for example, denoised model 117) to approximate a noisy three-dimensional representation (e.g., a noisy voxelized representation) of a molecule exhibiting one or more desired properties from a noisy data distribution. As described in more detail below, in some cases, molecular design computational model 115 (including denoised model 117) may be trained to approximate a noisy data distribution rather than a true data distribution to avoid overfitting molecular design computational model 115 to known molecules that can be used to train molecular design computational model 115. Furthermore, in some cases, the molecular design computational model 115 (including the denoised model 117) can be trained by gradient-based Markov chain Monte Carlo (MCMC) sampling (including, for example, Markov chain Monte Carlo (MCMC) sampling with Langevin dynamics).
[0139] At 302, the molecular design engine 110 may apply a molecular design computational model with a first adjustment to denoise the damaged sample molecule and generate a first updated molecule. In some exemplary embodiments, step 302 may include the molecular design engine 110 training a molecular design computational model 115 (including, for example, a denoising model 117) to approximate a data distribution of three-dimensional representations (e.g., voxelized representations) of molecules exhibiting one or more desired properties, such that candidate molecules exhibiting the same desired properties can be generated by sampling from it. In some cases, the molecular design computational model 115 may be trained to approximate the aforementioned data distribution based on a training dataset of damaged sample molecules, each damaged sample molecule being generated based on a noisy three-dimensional representation (e.g., a voxelized representation) of sample molecules (e.g., known molecules) from the data distribution. An example of such a model is shown in Figure 1A, where the molecular design computational model 115 (e.g., a denoising model 117) is trained to recover a noisy three-dimensional representation 182 of the sample molecule from a damaged three-dimensional representation 184 of the sample molecule generated by the damage engine 121. In some cases, the molecular design computational model 115 can be trained based on the damaged embeddings of those three-dimensional representations (e.g., voxelized representations) rather than being trained to directly recover the noisy three-dimensional representation (e.g., voxelized representation) of the sample molecule. This is shown in Figure 1B, where the molecular design computational model 115 (e.g., denoised model 117) is trained to recover the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule from the damaged embedding 188 generated by the damaged engine 121.
[0140] In some exemplary embodiments, training the molecular design computational model 115 may include applying a denoising model 117 to denoise the damaged 3D representation (e.g., voxelized representation) or alternatively its damaged embedding for each sample molecule. Figure 1A shows an example where training the molecular design computational model 115 includes adjusting the parameters of the denoising model 117 (e.g., weights, biases, etc.) to gradually reduce, for example, over multiple iterations, the difference (e.g., mean squared error (MSE) between the noisy 3D representation (e.g., voxelized representation) of the sample molecule and the noisy 3D representation recovered by the denoising model 117 from the damaged 3D representation of the sample molecule. Alternatively, in the example shown in Figure 1B, the molecular design computational model 115 may be trained by adjusting the parameters of the denoising model 117 (e.g., weights, biases, etc.) to gradually reduce, over multiple iterations, the difference (e.g., mean squared error) between the embedding of the noisy 3D representation of the sample molecule and the embedding recovered by the denoising model 117 from the corresponding damaged embedding. In some cases, the parameters of the denoising model 117 (e.g., weights, biases, etc.) may undergo different adjustments, followed by further adjustments that produce lower discrepancies (e.g., mean squared error (MSE)). For example, in some cases, the parameters of the denoising model 117 (e.g., weights, biases, etc.) may be first adjusted, and then the denoising model 117 with the first adjustment may be applied to denoise the damaged 3D representation of the sample molecule or the damaged embedding of the 3D representation of the sample molecule, generating at least a first updated molecule. In some cases, the first updated molecule may be an updated 3D representation of the first molecule (e.g., a voxelized representation), or alternatively, an updated embedding of the 3D representation of the first molecule (e.g., a voxelized representation).
[0141] In some cases, a damaged 3D representation of a sample molecule or a damaged embedding of that 3D representation can be denoised by updating one or more atomic density values representing the types and locations of atoms present in the sample molecule. In cases where a damaged embedding is generated by downsampling (or compressing) the 3D representation of the sample molecule, at least some of the updated values can condense multiple features (or atomic density values) from the original 3D representation (e.g., a voxelized representation) of the sample molecule. As described in more detail below, a denoising model 117 with a second adjustment (instead of a first adjustment) can be applied to denoise a damaged 3D representation (e.g., a damaged voxelized representation) or a damaged embedding of that 3D representation of the sample molecule, and generate at least a second updated molecule. The denoising model 117 with the first or second adjustment can be further adjusted. Doing so can train the denoising model 117 to approximate a noisy data distribution, or in some cases, a noisy latent distribution, exhibiting smoother density transitions to support more efficient sampling, since there are no steep gradient changes in regions that restrict sampling to the immediate vicinity of the sample molecules that form the basis of the training dataset.
[0142] In some exemplary embodiments, training the denoising model 117 may further include determining a function 175. As described, in some cases, function 175 may be a scoring function parameterized by the parameters of the denoising model 117 (e.g., weights, biases, etc.). Therefore, in some cases, training the molecular design computation model 115 (including adjusting the parameters of the denoising model 117) may also include adjusting the parameters of function 175. For example, in some cases, function 175 may be determined by performing gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.) to approximate the gradient of the noisy data distribution (or noisy latent distribution). Doing so may include adjusting the parameters of function 175 through one or more iterations such that function 175 outputs a value (e.g., a score) indicating local density changes in the noisy data distribution (or noisy latent distribution). When function 175 is a scoring function, the parameters of function 175 can be adjusted such that function 175 assigns higher values (e.g., higher scores) to samples from locations exhibiting more aggressive local changes in density (e.g., increases or smaller decreases) compared to samples from locations exhibiting less aggressive local changes in density (e.g., decreases or smaller decreases), such as 3D representations of molecules or their embeddings. Therefore, once the denoising model 117 has been trained, function 175 can output values (e.g., scores, etc.) that distinguish samples from higher-density regions of a noisy data distribution (or a noisy latent distribution) from samples sampled from lower-density regions of a noisy data distribution.
[0143] At 304, the molecular design engine 110 may apply a molecular design computational model with a second adjustment to denoise the damaged sample molecule and generate a second updated molecule. In some exemplary embodiments of step 304, after applying a denoising model 117 with a first adjustment to generate at least a first updated molecule, a denoising model 117 with a second adjustment may be applied to generate at least a second updated molecule, such as, for example, an updated three-dimensional representation (e.g., an updated voxelized representation of the second molecule or an updated embedding of the three-dimensional representation of the second molecule). It should be understood that the first and second adjustments may include different changes to the parameters of the denoising model 117 (e.g., weights, biases, etc.). Therefore, compared to applying a denoising model 117 with a first adjustment to denoise a damaged 3D representation 182 of a sample molecule or a damaged embedding 186 of a noisy 3D representation 182 of the same sample molecule, applying a denoising model 117 with a second adjustment to denoise a damaged 3D representation 184 of a sample molecule or a damaged embedding 186 of a noisy 3D representation 182 of a sample molecule can produce different updated molecules. As described in more detail below, training the denoising model 117 may include further adjusting the denoising model 117 with the first or second adjustment based on differences (e.g., mean squared error (MSE)) present in the noisy 3D representation 182 (FIG. 1A) or the embedding 184 (FIG. 1B) of the noisy 3D representation 182 of the sample molecule recovered by the denoising model 117.
[0144] At 306, the molecular design engine 110 may determine that the first updated molecule is more similar to the sample molecule than the second updated molecule. In some exemplary embodiments, step 306 may include selecting the denoising model 117 with the first adjustment rather than the denoising model 117 with the second adjustment for further adjustments during subsequent iterations if the first updated molecule generated by the denoising model 117 with the first adjustment is more similar to the noisy three-dimensional representation (or its embedding) of the sample molecule (e.g., exhibiting a lower mean square error (MSE)) compared to the second updated molecule generated by the denoising model 117 with the second adjustment. For example, in FIG1A, the first updated molecule may be an updated three-dimensional representation of a first molecule having a smaller difference relative to the noisy three-dimensional representation 182 of the sample molecule (e.g., a lower mean square error (MSE)). In Figure 1B, compared to the second updated molecule, the first updated molecule can be an updated embedding of the three-dimensional representation of the first molecule that has a smaller difference in embedding 186 relative to the noisy three-dimensional representation 182 of the sample molecule (e.g., a lower mean square error (MSE)).
[0145] The first updated molecule is more similar to the noisy three-dimensional representation 182 (or its embedding 186) of the sample molecule than the second updated molecule, which may indicate that the denoising model 117 with the first adjustment is better at recovering the noisy three-dimensional representation 182 (or its embedding 186) of the sample molecule than the denoising model 117 with the second adjustment. Therefore, the denoising model 117 with the first adjustment can better approximate the noisy data distribution (or noisy latent distribution) of molecules exhibiting one or more desired properties than the denoising model 117 with the second adjustment. Therefore, in some cases, the molecular design engine 110 may choose the denoising model 117 with the first adjustment instead of the denoising model 117 with the second adjustment to undergo one or more additional adjustment iterations.
[0146] At 308, the molecular design engine 110 may further adjust the molecular design computation model with the first adjustment instead of the second adjustment until one or more criteria are met. In some exemplary embodiments, if the first updated molecule generated by the denoised model 117 with the first adjustment is more similar to the noisy three-dimensional representation 182 (or its embedding 186) of the sample molecule (e.g., with a lower mean square error (MSE)) compared to the second updated molecule generated by the denoised model 117 with the second adjustment, the molecular design engine 110 may further adjust the denoised model 117 with the first adjustment instead of the denoised model 117 with the second adjustment. For example, during subsequent iterations of the adjustment, the molecular design engine 110 may further adjust the parameters (e.g., weights, biases, etc.) of the denoised model 117 with the first adjustment, and then apply the further adjusted denoised model 117 to generate one or more additional updated molecules. In some cases, the denoising model 117 can be further tuned to further increase the similarity (or reduce the mean squared error (MSE)) between the updated molecules generated by the denoising model 117 and the noisy 3D representations (or their embeddings) of sample molecules in the training dataset. In some cases, the molecular design engine 110 can continue to tune the denoising model 117 until one or more criteria are met. For example, in some cases, the molecular design engine 110 can continue to tune the parameters of the denoising model 117 (e.g., weights, biases, etc.) until the molecular design engine 110 has performed threshold-based tuning iterations. Alternatively and / or additionally, the molecular design engine 110 can continue to tune the parameters of the denoising model 117 (e.g., weights, biases, etc.) until the similarity (e.g., mean squared error (MSE)) between the updated molecules generated by the denoising model 117 and the noisy 3D representations (or their embeddings) of sample molecules in the training dataset meets one or more thresholds. In some cases, the molecular design engine 110 may continue to adjust the parameters of the denoising model 117 (e.g., weights, biases, etc.) until the updated molecules generated by the denoising model 117 exhibit a threshold probability in the data distribution of molecules in the training dataset that exhibit one or more desired properties.
[0147] As described in more detail below, once one or more criteria are met, the trained denoising model 117 can be applied to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162 by at least denoising the three-dimensional representation (e.g., a voxelized representation) of the input molecule 152. As shown in Figures 1A-1B, the trained denoising model 117 can generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162 by at least sampling from a noisy data distribution filled with noisy three-dimensional representations (e.g., voxelized representations) of molecules exhibiting one or more desired properties (e.g., drug-like properties) based on function 175. Sampling may include one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo, etc.), which may be guided by function 175 such that each sampling iteration includes selecting one or more samples (or molecules) from regions of gradually increasing density of a noisy data distribution (or a noisy latent distribution).
[0148] Figure 3B depicts a flowchart illustrating an example of a process 325 for applying a molecular design computational model to generate a three-dimensional molecule in voxelized space, according to some exemplary embodiments. Referring to Figures 1A, 1B, 2, and 3B, process 325 may implement operation 206 of process 200 shown in Figure 2. In some cases, process 325 may be performed by a molecular design engine 110. For example, in some cases, molecular design engine 110 may apply a molecular design computational model 115 (e.g., a denoising model 117) to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule, at least by denoising the three-dimensional representation (e.g., a voxelized representation) of the input molecule. In some cases, the input molecule may be a random molecule (e.g., a molecule with randomly selected atom types and / or positions) or a known molecule with one or more desired properties. Therefore, the three-dimensional representation (e.g., voxelized representation) of the input molecule may include noise that needs to be removed by the molecular design computational model 115 so that the resulting three-dimensional representation (e.g., voxelized representation) of the output molecule is consistent with a molecule exhibiting one or more desired properties (e.g., drug-like properties). The molecular design computational model 115 may denoise the three-dimensional representation (e.g., voxelized representation) of the input molecule at least by sampling from a noisy data distribution (or a noisy latent distribution) with higher sampling efficiency, because the smoother density transitions present therein allow for a full exploration of the data distribution. As described in more detail below, once the molecular design computational model 115 generates the three-dimensional representation (e.g., voxelized representation) of the output molecule, the molecular design engine 110 may further generate one or more other representations of the output molecule, including, for example, a one-dimensional representation of the output molecule, a two-dimensional representation of the output molecule, etc. The output molecule is generated by manipulating the three-dimensional representation (e.g., voxelization representation) of the input molecule to capture the conformation (or three-dimensional structure) of the input molecule. This means that the conformation (or three-dimensional) structure of the output molecule is more likely to be consistent with one or more desired properties (e.g., drug-like properties such as affinity, specificity, bioactivity, exploitability, etc.).
[0149] At 332, the molecular design engine 110 can update the three-dimensional representation of the input molecule to generate an updated three-dimensional representation. In some exemplary embodiments, updating the three-dimensional representation may include the molecular design engine 110 applying a molecular design computational model 115 to generate a three-dimensional representation (e.g., a voxelized representation) of the output molecule 162, at least by denoising the three-dimensional representation (e.g., a voxelized representation) of the input molecule 152. An example of this process is shown in Figure 1A, where a denoising engine 117 denoises the three-dimensional representation of the input molecule 152 to generate an updated three-dimensional representation 160. In some cases, the input molecule 152 may be a noisy molecule (e.g., a molecule with randomly selected atom types and / or positions) or a known molecule with one or more unwanted properties. This means that the three-dimensional representation of the input molecule 152 may include at least some noise, making it inconsistent with the three-dimensional representation of a molecule exhibiting one or more desired properties (e.g., drug-like properties). Therefore, in some cases, the denoising engine 117 can be trained to update the three-dimensional representation of the input molecule 152 such that the resulting updated three-dimensional representation 160 is consistent with the three-dimensional representation of the molecule exhibiting one or more desired properties.
[0150] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to update the three-dimensional representation of the input molecule 152 based on a function 175. In some cases, the function 175 may be a scoring function that, for each sample (or molecule) selected from a noisy data distribution, outputs a value (e.g., a score) indicating the likelihood of the sample (or molecule) within the noisy data distribution. For example, in some cases, the value output by function 175 for a particular sample (or molecule) may indicate a local variation in density at the location from which the sample (or molecule) was selected. The denoising model 117 may update the three-dimensional representation of the input molecule 152 over multiple successive sampling iterations, at least based on the value output by function 175. During each sampling iteration, the denoising model 117 may be applied to further update the three-dimensional representation of the input molecule 152, such that the resulting updated three-dimensional representation 160 is selected from a higher-density region of the noisy data distribution compared to the region selected during one or more previous sampling iterations.
[0151] In some exemplary embodiments, the molecular design computational model 115 may perform gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling) of a noisy data distribution, wherein the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 is updated over multiple successive sampling iterations. In some cases, each iteration may include the molecular design computational model 115 further updating the three-dimensional representation of the input molecule 152 to sample from regions with progressively increasing density of the noisy data distribution. Furthermore, in some cases, modifications to the three-dimensional representation of the input molecule 152 may accumulate over multiple successive iterations. For example, in some cases, the three-dimensional representation of the input molecule 152 may undergo a first update and a second update. The molecular design computational model 115 may apply a function 175 to determine a first value (e.g., a first score, etc.) of the three-dimensional representation of the input molecule 152 with the first update and a second value (e.g., a second score, etc.) of the three-dimensional representation of the input molecule 152 with the second update. During subsequent iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling, if the first and second values indicate that the three-dimensional representation of the input molecule 152 with the first update samples a higher-density region of the noisy data distribution and exhibits a higher probability of being within the noisy data distribution compared to the three-dimensional representation with the second update, then the denoising model 117 can be applied to further update the three-dimensional representation of the input molecule 152 with the first update.
[0152] In some cases, one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling may be performed, wherein the molecular design computational model 115 applies a denoising model 117 to further modify the three-dimensional representation of the input molecule 152 until one or more criteria are met. For example, in some cases, the molecular design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until a threshold amount of sampling iterations is performed. Alternatively and / or additionally, the molecular design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until the function 175 outputs a value (e.g., a score, etc.) that satisfies one or more thresholds for the updated three-dimensional representation 160. The fact that the value (e.g., a score, etc.) associated with the updated three-dimensional representation 160 satisfies one or more thresholds may indicate that the updated three-dimensional representation 160 is selected from a region with a sufficiently high density of noisy data distribution, and that the probability that the updated three-dimensional representation 160 satisfies one or more thresholds within the noisy data distribution. In some cases, one or more criteria may also include output molecules that have been generated in threshold quantities and exhibit one or more desired properties (e.g., at least one output molecule exhibiting one or more drug-like properties at the threshold level, such as affinity, specificity, bioactivity, extensibility, etc.).
[0153] At 336, the molecular design engine 110 can denoise the updated 3D representation to generate a 3D representation of the output molecule. In some exemplary embodiments of step 336, the molecular design computation model 115 can denoise the 3D representation of the input molecule 152 by sampling from a noisy data distribution occupied by noisy 3D representations of molecules exhibiting one or more desired properties. As described, the molecular design computation model 115 (including the denoising model 117) can be trained to approximate a noisy data distribution (rather than a true data distribution) at least by being trained to denoise a damaged 3D representation 184 of the sample molecule to recover a noisy 3D representation 182 of the sample molecule rather than a clean 3D representation of the sample molecule. Furthermore, such a noisy data distribution can exhibit smoother density transitions and is therefore more efficient in sampling. The updated 3D representation 160 being sampled from a noisy data distribution means that the updated 3D representation 160 can undergo additional denoising. For example, Figure 1A shows that the molecular design engine 110 can apply a recovery model 118 to denoise the updated 3D representation 160 and generate a 3D representation (e.g., a voxelized representation) of the output molecule 162. In some cases, the recovery model 118 can be trained to denoise the updated 3D representation 160 to map the updated 3D representation 160 from a noisy data distribution back to a true data distribution of molecules exhibiting one or more desired properties (e.g., drug-like properties). It should be understood that this denoising differs from the denoising performed by the denoising model 117, which involves updating the 3D representation of the input molecule 152 to sample from higher-density regions of the noisy data distribution that are more likely to be occupied by molecules exhibiting one or more desired properties.
[0154] At 338, the molecular design engine 110 can generate one or more other representations of the output molecule based at least on the three-dimensional representation of the output molecule. In some exemplary embodiments, the three-dimensional representation (e.g., a voxelized representation) of the output molecule 162 generated by the recovery model 118 denoising the updated three-dimensional representation 160 sampled from a noisy data distribution by the molecular computation model 115 can be further transformed into one or more other representations of the output molecule 162. For example, in some cases, the molecular design engine 110 can recover the positions (e.g., coordinates) of atoms present in the output molecule 162 and one or more bonds between them based at least on the three-dimensional representation (e.g., a voxelized representation) of the output molecule 162. In doing so, the molecular design engine 110 can determine another representation of the output molecule 162, including, for example, a one-dimensional representation of the output molecule 162 (e.g., a simplified molecular linear input specification (SMILES) string), a two-dimensional representation of the output molecule 162 (e.g., a molecular graph), etc. In some cases, the molecular design engine 110 can recover the positions of atoms present in the output molecule 162 by applying peak detection techniques. These peak detection techniques determine the positions (e.g., coordinates) of the atoms based on one or more peaks in the atomic density included in the three-dimensional representation (e.g., voxelized representation) of the output molecule 162, and then determine one or more interconnecting bonds based on these positions. Alternatively, the molecular design engine 110 can apply a machine learning model trained to transform the voxelized representation of the output molecule 162 into one or more other representations.
[0155] Figure 3C depicts a flowchart illustrating an example of a process 350 for applying a molecular design computational model to generate a three-dimensional molecule in voxelized space, according to some exemplary embodiments. Referring to Figures 1-2 and 3C, process 350 may implement operation 206 of process 200 shown in Figure 2. In some cases, process 350 may be performed by molecular design engine 110. For example, in some cases, molecular design engine 110 may apply molecular design computational model 115 (e.g., denoising model 117) to generate a three-dimensional representation of the output molecule, such as a voxelized representation of the output molecule, at least by denoising the three-dimensional representation (e.g., voxelized representation) of the input molecule. In some cases, the three-dimensional representation of the input molecule may be denoised at least by updating the embedding of the three-dimensional representation (e.g., voxelized representation) of the input molecule rather than by directly updating the three-dimensional representation of the input molecule, at least because the embedding may be more compact and computationally more efficient. In some cases, the embedding of the 3D representation of the input molecule can be generated by downsampling (or compressing) the 3D representation of the input molecule, although it is also possible to generate the embedding without any downsampling (or compression) of the 3D representation of the input molecule. In the former case, the embedding of the 3D representation of the input molecule can occupy a latent voxelization space, while in the latter case, the embedding of the 3D representation of the input molecule can remain in the same discrete voxelization space as the original 3D representation of the input molecule. It should be understood that the latent voxelization space can have a lower dimension than the discrete voxelization space, allowing operations on the embedding of the 3D representation of the input molecule to increase the speed and computational efficiency of the generation process, while achieving comparable or better generation performance.
[0156] It should be understood that the molecular design computational model 115 can denoise the embedding of the three-dimensional representation of the input molecule by sampling from a noisy latent distribution. That is, as described, the molecular design computational model 115 can be trained to approximate a noisy latent distribution rather than the true data distribution to at least avoid the steep density transitions present in the true data distribution. In other words, the updated embedding generated by updating the embedding of the three-dimensional representation of the input molecule by the molecular design computational model 115 may still occupy a noisy latent distribution. Sampling of such a noisy latent distribution may be more efficient because the smoother density transitions of the noisy latent distribution support a full exploration of the data distribution. As described in more detail below, the updated embedding can undergo decoding and further denoising to "jump" back to the true data distribution. Furthermore, in some cases, the molecular design engine 110 can generate one or more other representations of the output molecule based on the three-dimensional representation of the output molecule resulting from the decoding and denoising of the updated embedding, including, for example, a one-dimensional representation of the output molecule, a two-dimensional representation of the output molecule, etc. The output molecule is generated by manipulating the three-dimensional representation (e.g., voxelization representation) of the input molecule to capture the conformation (or three-dimensional structure) of the input molecule. This means that the conformation (or three-dimensional) structure of the output molecule is more likely to be consistent with one or more desired properties (e.g., drug-like properties such as affinity, specificity, bioactivity, exploitability, etc.).
[0157] At 352, the molecular design engine 110 can encode a three-dimensional representation of the input molecule to generate an embedding of the input molecule. In some exemplary embodiments, the encoder 111 can encode a three-dimensional representation (e.g., a voxelized representation) of the input molecule 152 to generate an embedding 154 of the input molecule 152. An example of this is shown in Figure 1B. In the case of “seed generation,” the input molecule 152 can be a known molecule (e.g., a molecule from a validation set derived from the PubChem dataset, the QM9 molecular dataset, the Geometry of Molecular Omnibus (GEOM) drug dataset, etc.). In some cases, the known molecule may exhibit one or more unwanted properties. When a known molecule is used as the input molecule 152, the generation process can be initialized with a voxel grid having an atomic density distribution corresponding to the types and positions of atoms expected to be found in the known molecule. Alternatively, the molecular design computational model 115 can perform de novo generation, in which case the input molecule 152 can be a noisy molecule whose atomic types and positions correspond to pure noise (e.g., uniform noise, etc.). When a noisy molecule is used as input molecule 152, the generation process can be initialized across the entire voxel grid without any expectation of atom type and / or position. In either case, the type and / or position of atoms in input molecule 152 may not be consistent with the type and / or position of molecules exhibiting one or more desired properties (e.g., drug-like properties). Therefore, molecular design computational model 115 can be applied to update the three-dimensional representation of input molecule 152 at least by updating embedding 154 and generating an updated embedding 156 such that the corresponding three-dimensional representation of output molecule 162 is more consistent with the corresponding three-dimensional representation of molecules exhibiting one or more desired properties.
[0158] In some exemplary embodiments, encoder 111 may encode the three-dimensional representation of input molecule 152 at least by downsampling or compressing it (e.g., voxelization). Doing so may include condensing at least some of the features present in the three-dimensional representation of input molecule 152, which reduces the dimensionality (or number of features) present in the three-dimensional representation of input molecule 152. For example, in a three-dimensional representation of input molecule 152 comprising 32,000 features (or atomic density values)... In the case of a voxel grid, encoder 111 can condense at least some of those 32,000 features (or atomic density values) to generate a dataset containing 64 features. The voxel grid is embedded in the input molecule 152 as 154.
[0159] In some exemplary embodiments, encoder 111 may generate embedding 154 of input molecule 152 with or without downsampling or compressing the three-dimensional representation (e.g., voxelized representation) of input molecule 152. In some cases, encoder 111 may implement an identity function, meaning that embedding 154 may include the same number of features (e.g., atomic density values) present in the three-dimensional representation of input molecule 152. Alternatively, in the case where embedding 154 is generated by downsampling the voxelized representation of input molecule 152, doing so projects the voxelized representation of input molecule 152 from a higher-dimensional discrete voxelized space to a lower-dimensional latent space. Sampling from the lower-dimensional latent space may impose less computational burden than sampling directly from the higher-dimensional discrete voxelized space. For example, where sampling from a discrete voxelization space is a resource-intensive task, such as when the input molecule 152 is large (e.g., containing between 80 and 200 atoms) or when a large number of candidate molecules are being generated from it, the molecular design engine 110 can sample from the latent voxelization space by applying an embedding 154 of the three-dimensional representation of the input molecule 152 to the molecular design computational model 115. It should be understood that even when the input molecule 152 is large (e.g., containing more than 200 atoms) or when a large number of candidate molecules are being generated, sampling from the latent voxelization space can impose a modest computational overhead.
[0160] In some exemplary embodiments, encoder 111 may be part of an autoencoder (e.g., a variational autoencoder (VAE), such as a vector quantization variational autoencoder (VQ-VAE)) together with decoder 119. In some cases, encoder 111 may be trained to encode a voxelized representation of input molecule 152, such that decoder 119 is able to recover a three-dimensional representation (e.g., a voxelized representation) of input molecule 152 from the resulting embedding 154. For further illustration, using... Voxelization representation of molecules, using Indicates encoder 111, using This indicates decoder 119, and uses... This indicates embedding 154. For example, voxelization of molecules represents... (Such as the voxelized molecular representation of input molecule 152) can be used with an encoder Encode the data to generate a continuous latent embedding according to equation (2) below. .
[0161] (2)
[0162] According to equation (3), continuous latent embedding Each of them can learn shared embeddings through nearest neighbor search. In the codebook One of the vectors is matched to quantize to a discrete latent embedding. .
[0163] ,in (3)
[0164] Quantified latent embeddings Through the decoder The original voxelized molecular representation is reconstructed according to equation (4) below. .
[0165] (4)
[0166] The latent embedding space can be represented as ,in It is the number of discrete latent vectors in the learned codebook, and This is the dimension of each potential embedding vector in the codebook. It should be understood that... and Hyperparameters can be selected experimentally.
[0167] In some cases, because the operation is non-differentiable, a gradient may not be defined for the nearest neighbor search in the codebook for each latent embedding. Instead, the nearest neighbor search in the codebook replaces each quantized latent embedding with one of the learned codebook embeddings of the same dimension. The stop gradient (sg) operation can transfer the gradient from the input to the decoder. Quantified latent embeddings Copying is performed by the encoder before quantization. Output continuous latent embedding The stopping gradient (sg) operation can act as a positive identity function by copying the variable without making any changes. However, when updating the encoder... During the backpropagation of the gradient, the Stop Gradient (sg) operation can prevent the gradient from flowing through the gradient update of the specific term to which the operation is applied, at least because the gradient cannot be computed for that term.
[0168] In some exemplary embodiments, the encoder forms an autoencoder (e.g., a variational autoencoder (VAE) or the like). and decoder Training can include adjusting the encoder. and decoder To reduce (or minimize) three separate losses or loss terms. The first loss term may include reconstruction loss (e.g., mean squared error (MSE) reconstruction loss), which corresponds to the loss caused by the encoder. Ingestion to generate embeddings Voxelization of molecules With the decoder Based on embedding The generated reconstruction The difference between them. The second loss term can be obtained by adjusting the embedding vector. Shift to encoder Output continuous latent embedding To force the embedding used for quantizing the latent space Learning the codebook. The third loss term can quantify the committed loss, which ensures the encoder... Commitment to Embedded Furthermore, its output will not grow arbitrarily. This third loss term may be related to the committed cost weight. Relatedly, the commitment cost weight may also be a hyperparameter set experimentally. Equation (5) below is used to train the encoder. and decoder Overall loss function Examples.
[0169] (5)
[0170] At 354, the molecular design engine 110 can generate an updated embedding at least by updating the embedding of the three-dimensional representation of the input molecule. In some exemplary embodiments, the molecular design engine 110 can apply a molecular design computation model 115 (e.g., a denoising model 117) to denoise the three-dimensional representation of the input molecule 152 at least by updating the embedding 154 of the three-dimensional representation of the input molecule 152 and generating an updated embedding 156. For example, in some cases, the three-dimensional representation of the input molecule 152 may include noise that causes inconsistencies between the types and / or positions of atoms present in the input molecule 152 and the types and / or positions of atoms in molecules exhibiting one or more desired properties (e.g., drug-like properties). In other words, the molecular design computation model 115 can update the embedding 154 of the three-dimensional representation of the input molecule 152 to increase the likelihood that the resulting output molecule 162 exhibits one or more desired properties. As described, the noise removed from the embedding 154 by the denoising model 117 should not be confused with the noise that projects the three-dimensional representation of the input molecule 152 from the real data distribution, which exhibits a jagged density transition, to a noisy data distribution exhibiting a smoother density transition for more efficient sampling (e.g., gradient-based Markov chain Monte Carlo (MCMC) sampling, etc.). As described in more detail below, by updating the embedding 154 of the three-dimensional representation of the input molecule 152, the molecular design computational model 115 (e.g., the denoising model 117) can traverse the smoother density of the noisy data distribution to sample from regions of gradually increasing density in the noisy data distribution using the updated embedding 156, and then “jump” back to the real data distribution when selecting a sample that exhibits a threshold probability within the noisy data distribution.
[0171] In some exemplary embodiments, the denoising model 117 may apply updates corresponding to changes in the type and / or position of atoms present in the input molecule 152 to the embedding 154 of the three-dimensional representation of the input molecule 152. Where the encoder 111 implements an identity function and generates the embedding 154 without any downsampling (or compression) of the underlying three-dimensional representation (e.g., voxelized representation) of the input molecule 152, the denoising model 117 may update the embedding 154 at least by updating the atomic density of one or more voxels in at least one channel of the embedding 154. Alternatively, where the generation of the embedding 154 includes downsampling (or compression) of the underlying three-dimensional representation (e.g., voxelized representation) of the input molecule 152, the denoising model 117 may update the embedding 154 at least by updating one or more values present in the embedding 154, at least some of which condense multiple atomic density values included in the three-dimensional representation (e.g., voxelized representation) of the input molecule 152.
[0172] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to update the embedding 154 of the input molecule 152 based on a function 175. In some cases, for each sample (or molecule) selected from the noisy data distribution, the function 175 may output a value (e.g., a score, etc.) indicating the likelihood of the sample (or molecule) in the noisy data distribution. For example, in some cases, the value output by the function 175 for a particular sample (or molecule) may indicate a local variation in density at the location from which the sample (or molecule) was selected. The denoising model 117 may update the embedding 154 at least based on the value output by the function 175 over multiple successive sampling iterations. During each sampling iteration, the denoising model 117 may be applied to further update the embedding 154, such that the resulting updated embedding 156 is selected from a higher-density region of the noisy data distribution compared to previous sampling iterations.
[0173] In some exemplary embodiments, the molecular design computational model 115 may perform gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling) of a noisy data distribution, wherein the embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 is updated through multiple successive sampling iterations, each iteration sampling from regions of gradually increasing density in the noisy data distribution to increase the likelihood of the resulting updated embedding 156 in the noisy data distribution. Furthermore, in some cases, updates to the embedding 154 of the input molecule 152 may accumulate over multiple successive iterations. To further illustrate, consider an example where the embedding 154 of the three-dimensional representation of the input molecule 152 undergoes a first update and a second update. The molecular design computational model 115 may apply a function 175 to determine a first value (e.g., a first score, etc.) of the embedding 154 with the first update and a second value (e.g., a second score, etc.) of the embedding 154 with the second update. During subsequent iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling, if the first and second values indicate that the embedding 154 with the first update samples from a higher-density region of the noisy data distribution and exhibits a higher probability of being within the noisy data distribution compared to the embedding 154 with the second update, then the denoising model 117 can be applied to further update the embedding 154 with the first update. In some cases, one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling can be performed, wherein the molecular design computational model 115 applies the denoising model 117 to further modify the embedding 154 of the three-dimensional representation (e.g., voxelized representation) of the input molecule 152 until one or more criteria are met. For example, in some cases, the molecular design computational model 115 can perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until a threshold amount of sampling iteration is performed. Alternatively and / or additionally, the molecular design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until the function 175 outputs a value (e.g., a score, etc.) that satisfies one or more thresholds for the updated embedding 156. The fact that the value (e.g., a score, etc.) associated with the updated embedding 156 satisfies one or more thresholds may indicate that the updated embedding 156 is selected from a region of noisy data distribution with a sufficiently high density, and that the probability that the updated embedding 156 is within the noisy data distribution satisfies one or more thresholds.In some cases, one or more criteria may also include output molecules that have been generated in threshold quantities and exhibit one or more desired properties (e.g., at least one output molecule exhibiting one or more drug-like properties at the threshold level, such as affinity, specificity, bioactivity, extensibility, etc.).
[0174] At 356, the molecular design computational model 115 can decode the updated embedding to generate a noisy three-dimensional representation of the output molecule. In some exemplary embodiments, the molecular design engine 110 can apply the decoder 119 to decode the updated embedding 156 and generate a noisy three-dimensional representation 158 of the output molecule 162 after the molecular design computational model 115 (e.g., denoising model 117) has been applied to update the embedding 154 of the three-dimensional representation of the input molecule 152 and generate the updated embedding 156. Decoding the updated embedding 156 can map the updated embedding 156 from a latent voxelized space filled with embeddings of the three-dimensional representations of various molecules to a latent discrete space. However, as described in more detail below, the latent discrete space can be a noisy latent space, which means that the noisy three-dimensional representation 158 generated by the decoder 119 decoding the updated embedding 156 may need further denoising in order to project the noisy three-dimensional representation 158 back to the true data distribution of the molecule exhibiting one or more desired properties.
[0175] In some exemplary embodiments, the decoder 119 of the molecular design engine 110 can generate a noisy 3D representation 158 at least by decoding an updated embedding 156 generated by the molecular design computational model 115 (e.g., a denoising engine 117). As described, in some cases, the decoder 119 may form part of an autoencoder (e.g., a variational autoencoder, such as a vector quantization variational autoencoder (VQ-VAE)) together with the encoder 111. In some cases, the encoder 111 and the decoder 119 may be trained in tandem, wherein the encoder 111 may be trained to generate embeddings of the 3D representation of the molecule (e.g., a voxelized representation), such as embedding 154 of the 3D representation of the input molecule 152, which enables the decoder 119 to recover the original 3D representation (e.g., the voxelized representation) from which. Thus, after generating the updated embedding 156, the decoder 119 may be applied to recover the noisy 3D representation 158 of the output molecule 162.
[0176] In some cases, decoding the updated embedding 156 may include upsampling (or decompressing) the updated embedding 156, which can project the updated embedding 156 from the latent voxelization space back to the discrete voxelization space. The noisy 3D representation 158 of the output molecule 162 (e.g., a noisy voxelization representation) may exhibit the same dimensions (or feature quantity) as the 3D representation (e.g., a voxelization representation) of the input molecule 152 taken up by the molecular design engine 110 at operation 352. For example, in some cases, the 3D representation of the input molecule 162 may include... A voxel grid, meaning that the three-dimensional representation of the input molecule 152 can include 32,000 features (or atomic density values). Meanwhile, each of the embeddings 154 operated by the molecular design computational model 115, and the resulting updated embedding 156, can include a voxel grid with 64 features. Voxel mesh. In some cases, decoder 119 can utilize the voxel mesh included in it. The voxel grid is upsampled (or decompressed) to decode the updated embedding 156 to generate a noisy 3D representation 158 (e.g., a noisy voxelized representation) for the output molecule 162. Voxel grid. It should be understood that this upsampling (or decompression) can recover 32,000 features (or atomic density values) in the noisy three-dimensional representation 158 (e.g., voxelized representation) of the output molecule 162. As described, these 32,000 features (or atomic density values) can indicate the positions of various atoms present in the output molecule 162. Furthermore, the 32,000 features (or atomic density values) can span one or more channels, each corresponding to a type of atom that may be present in the output molecule 162.
[0177] At 358, the molecular design engine 110 can denoise the noisy 3D representation of the output molecule to generate a 3D representation of the output molecule. In some exemplary embodiments, the denoising engine 117 of the molecular design computation model 115 can generate a 3D representation (e.g., a voxelized representation) of the output molecule 162 by at least denoising the noisy 3D representation 158 generated by decoding the updated embedding 156 by the decoder 119. As described, in some cases, the molecular design computation model 115 (e.g., the denoising model 117) can undergo one or more iterations of gradient-based Markov chain Monte Carlo (e.g., Langevin Markov chain Monte Carlo, etc.) to generate the updated embedding 156. In doing so, the molecular design computational model 115 can traverse the noisy latent distribution based at least on the output of function 175 (e.g., the score output by function 175) to sample updated embeddings 156 from higher-density regions of the noisy latent distribution filled with embeddings of 3D representations of molecules more likely to exhibit one or more desired properties (e.g., drug-like properties). However, decoding the updated embeddings 156 merely maps the updated embeddings 156 from the latent voxelization space to the discrete voxelization space, but the noisy 3D representations 158 still occupy the noisy data distribution rather than the true data distribution of molecules exhibiting one or more desired properties. Therefore, in some cases, the recovery model 118 can be applied to map the noisy 3D representations 158 from the noisy data distribution to the true data distribution. In some cases, this can constitute a “jump” back to the true data distribution, meaning that the 3D representation of the output molecule 162 generated by it occupies the true data distribution.
[0178] In some cases, the recovery model 118 may share the same architecture (e.g., an artificial neural network (ANN) or the like) as the denoising model 117, which is trained to traverse noisy latent distributions to denoise the embedding 154 of the three-dimensional representation of the input molecule 152 and generate an updated embedding 156. However, as described, the recovery model 118 can be trained to remove different types of noise. Therefore, in some cases, the recovery engine 118 may be trained based on a training dataset to denoise the noisy three-dimensional representation 182 of the sample molecule and recover the original three-dimensional representation 182 from it. In contrast, the denoising engine 117 may be trained to recover the embedding 186 of the noisy three-dimensional representation 182 of the sample molecule from the damaged embedding 188. In this context, training the denoising engine 117 may include adjusting one or more parameters of the denoising engine 117 (e.g., an artificial neural network (ANN) or the like) to reduce (or minimize) the difference (e.g., mean squared error (MSE)) between the original three-dimensional representation of the sample molecule and the three-dimensional representation of the sample molecule recovered by the denoising engine 117 from the noisy three-dimensional representation of the sample molecule.
[0179] To further illustrate, consider the quantized latent embeddings described in operation 252. As described, quantified latent embeddings It can be made by an encoder Encoding voxelization molecular representation To generate. In some cases, noise. (For example, Gaussian noise, such as isotropic Gaussian noise) can be added to the quantized latent embedding. For example, in some cases, there is a fixed high noise level. Noise in scaled unit covariance matrix (For example, Gaussian noise, such as isotropic Gaussian noise) can be added according to the following equation (6).
[0180] , (6)
[0181] It can be represented as a latent model The denoising engine 117 can be trained to work with latent embeddings. Perform denoising and restoration while reducing (or minimizing) the noise added (or not added). Previous original potential embedding With the denoised latent embedding generated by denoising engine 117 The reconstruction loss between (e.g., mean squared error (MSE) reconstruction loss). Derived from the latent model. Execution to generate denoised latent embeddings The denoising is shown in equation (7) below. Meanwhile, equation (8) shows the denoising used to train the latent model. loss function This includes reducing (or minimizing) noise in addition (or without addition). Previous original potential embedding With denoised potential embedding The difference between them (e.g., mean squared error (MSE)).
[0182] (7)
[0183] (8)
[0184] At 360°, the molecular design engine 110 can generate one or more other representations of the output molecule based at least on a three-dimensional representation of the output molecule. In some exemplary embodiments, the molecular design engine 110 can generate one or more other representations of the output molecule 162 based at least on a voxelized representation of the output molecule 162, including, for example, a one-dimensional representation of the output molecule 162 (e.g., a simplified linear input specification (SMILES) string), a two-dimensional representation of the output molecule 162 (e.g., a molecular graph), etc. For example, in some cases, the molecular design computational model 110 can recover the positions (e.g., coordinates) of atoms present in the output molecule 162 and the bonds between them from the voxelized representation of the output molecule 162. In some cases, the molecular design engine 110 can apply a peak detection technique that determines the positions (e.g., coordinates) of atoms present in the output molecule 162 based on one or more peaks of the atomic density included in the voxelized representation of the output molecule 162, and then determines one or more interconnecting bonds based on the positions of the atoms. Alternatively, the molecular design engine 110 may apply a machine learning model trained to transform the voxelized representation of the output molecule 162 into one or more other representations.
[0185] As described, in some exemplary embodiments, the molecular design computational model 115 (including the denoising model 117) may operate on a three-dimensional representation of the molecule rather than a one-dimensional or two-dimensional representation, at least because real and effective molecules exhibiting certain desired properties are more likely to be generated based on a molecular representation that captures both the composition (e.g., constituent atoms) and conformation (or three-dimensional structure) of the molecule. In some cases, the molecular design computational model 115 (including the denoising model 117) may operate on a voxelized representation of the molecule. Unlike conventional three-dimensional representations of molecules (e.g., point cloud representations, etc.), voxelized representations of molecules can represent both atom types and positions as one or more continuous (e.g., Gaussian-like) distributions across a voxel grid centered on the atomic coordinates of individual atoms. Therefore, unlike conventional three-dimensional representations of molecules (e.g., point cloud representations), the molecular design computational model 115 can apply a denoising model 117 to operate on the voxelized representation of the input molecule without requiring any workarounds to reconcile different types of data distributions (e.g., discrete distributions of atom types and continuous distributions of atom positions) and without requiring any prior knowledge of the number of atoms present in the resulting output molecule.
[0186] To further illustrate, Figure 4 depicts examples of voxelized representations and corresponding two-dimensional representations of different molecules according to some exemplary embodiments. For example, Figure 4 shows a voxelized representation 400 and a two-dimensional representation 450 of the molecule. In some exemplary embodiments, the voxelized representation 400 of the molecule can be generated by partitioning (or discretizing) the three-dimensional space surrounding the constituent atoms into a voxel grid 410, wherein each type of atom (or element) present in the molecule is represented by a different grid channel. This partitioning (or discretization) can generate Individualized molecules , , ,in This represents the length of each grid edge, and This indicates the number of channels in the dataset (e.g., the number of atoms (or elements) of different types).
[0187] In some cases, the voxel grid 410 can be a three-dimensional grid of voxels organized into continuous layers of rows and columns. Each voxel in the voxel grid 410 can be a volume element formed at the intersection of rows and columns, such as a three-dimensional cube. Furthermore, each voxel in the voxel grid 410 can be associated with a value indicating the atomic density at the corresponding location (e.g., having a value...). The voxelized representation of a single molecule can be a box around the center of the molecule, which is then divided into voxels. To generate a voxelized representation 400 of the molecule, each constituent atom can be converted into a three-dimensional continuous (e.g., Gaussian-like) density according to the following equation (9). For example, an example of the voxel grid 410 shown in Figure 4 may include a first atomic density 415a representing a first atom of a first type and a second atomic density 415b representing a second atom of a second type.
[0188] (9)
[0189] in Defined as the distance from the center of the atom It has a radius atoms The fraction of the volume occupied. Different types of atoms (or elements) can have different radii or the same radius (e.g., =.5Å). According to the following equation (10), the occupancy of each voxel in the voxel grid can be calculated by integrating the occupancy generated by each atom in the molecule. .
[0190] (10)
[0191] in Indicates the number of atoms in a molecule. for atom, Coordinates in a voxel grid ,and Represents atoms The coordinates of the center.
[0192] As mentioned, in some cases, the voxelization of a molecule can represent the atomic density in 400 ohms, centered on the atoms present in the molecule. Therefore, occupancy... The value can be maximized at the atom center (e.g., 1) and decreases to a minimum (e.g., 0) with increasing distance from the atom center. Each channel in the voxel grid can be independent. That is, channels do not interact or share volume contributions. In some cases, the size of the voxel grid 410 included in the voxelized representation 400 of the molecule can correspond to the size of the represented molecule (e.g., the amount of constituent atoms). For example, in some cases, if the molecule has fewer atoms (e.g., the QM9 molecular dataset), the voxel grid 410 can be a [32×32×32] voxel grid, or if the molecule has more atoms (e.g., the Geometry Ensemble of Molecular (GEOM) drug dataset), the voxel grid can be a [64×64×64] voxel grid. Furthermore, in some cases, the number of channels in the voxelized representation 400 of the molecule can correspond to the number of atom types (or elements) present in the molecule. For example, the voxelized representation of molecules in the QM9 molecular dataset can include five channels for the five types of atoms that make up those molecules (e.g., carbon (C), hydrogen (H), oxygen (O), nitrogen (N), and fluorine (F)). Similarly, the voxelized representation of molecules in the Geometry of Molecular Ensemble (GEOM) drug dataset can include eight channels for the eight types of atoms present in those molecules (e.g., carbon (C), hydrogen (H), oxygen (O), nitrogen (N), fluorine (F), sulfur (S), chlorine (Cl), and bromine (Br)). Therefore, the voxelized representation of each molecule in the QM9 molecular dataset can include... Voxel meshes, and the voxelized representation of each molecule in the Geometry of Molecular Omnibus (GEOM) drug dataset can include... Voxel grid.
[0193] As described, in some exemplary embodiments, the molecular design computational model 115 (including the denoising model 117) can be trained to approximate and subsequently sample from a noisy data distribution of a noisy voxelized representation of a molecule, or in some cases, to approximate and subsequently sample from a noisy embedding of a voxelized representation of a molecule, rather than from a true data distribution of a voxelized representation of a molecule that has not yet been subjected to any noise. Training the denoising model 117 to approximate a noisy data distribution of molecules, such as a noisy data distribution of a noisy voxelized representation of a molecule exhibiting certain desired properties (e.g., drug-like properties) or its noisy embedding, can include determining a function 175 such that, for each voxelized representation (or its noisy embedding) of a molecule sampled from the noisy data distribution, the function 175 outputs a value indicating the density at the corresponding location in the noisy data distribution. Where the function 175 is a scoring function, the function 175 can output a score corresponding to a local variation in the density (or gradient) of the noisy data distribution. Therefore, when function 175 is a scoring function, the score output by function 175 for the noisy voxelized representation of the molecule (or its noisy embedding) can indicate the local variation in density at the corresponding position in the noisy data distribution.
[0194] In some cases, the denoising engine 117 can be trained to denoise noisy voxelized representations of molecules, or in other cases, to denoise noisy embeddings of voxelized representations of molecules generated by the molecular design computational model 115 (e.g., the denoising model 117). To further illustrate, Figure 5A depicts a schematic diagram illustrating examples of training the denoising engine 117 to denoise noisy voxelized representations of molecules according to some exemplary embodiments. As shown in Figure 5A, a training dataset for training the denoising engine 117 can be generated to include multiple training samples, each corresponding to a sample molecule. For example, Figure 5A shows sample molecule 500, which can be a known molecule from the PubChem dataset, the QM9 molecular dataset, the Geometry of Molecular Ensemble (GEOM) drug dataset, etc. Sample molecule 500 can be presented in a one-dimensional representation (e.g., a simplified molecular linear input canonical (SMILES) string) or a two-dimensional representation (e.g., a molecular graph), neither of which adequately captures the conformation (or three-dimensional structure) of sample molecule 500. Therefore, in some cases, in order to generate training samples to be included in the training dataset, the one-dimensional or two-dimensional representation of sample molecule 500 can be transformed into a three-dimensional representation of sample molecule 500. For example, in some cases, the one-dimensional or two-dimensional representation of sample molecule 500 can be transformed into the voxelized representation shown in Figure 5A. Voxelized representation of sample molecule 500 The types and positions of atoms present in sample molecule 500 can be collectively represented as one of the more continuous (e.g., Gaussian-like) densities of a transvoxel grid centered on a single atom present in sample molecule 500.
[0195] Referring again to Figure 5A, in some cases, the voxelized representation of sample molecules 500 Noise can be mixed in. (For example, Gaussian noise, such as isotropic Gaussian noise, etc.), which can have a noise level. In order to generate noisy voxelized representations .noise The addition of voxelization can represent The true data distribution of filling is represented by clean (or pristine) voxels of molecules. Projected onto a noisy data distribution filled with noisy voxelized representations of molecules. As described, if the molecular design calculation model 115 directly applies data from the real data distribution... Clean (or pristine) voxelized representation of molecules (such as voxelized representation of sample molecule 500) The actual data distribution is then determined by performing the operation. The jagged energy landscape may prevent molecular design computational models115 from fully exploring the real data distribution when sampling from it. In contrast, noisy data distribution It can exhibit a smoother energy landscape with more gradual gradient changes, which means that the molecular design computational model115 can be derived from noisy data distributions. Mid-sampling is used to produce greater diversity in the resulting output molecules. Therefore, in some cases, the denoising engine 117 can be trained to handle noisy voxelized representations of molecules (such as noisy voxelized representations of sample molecules 500). Denoising is performed so that the denoising engine 117 can be applied to the noisy data distribution generated by the molecular design computational model 115. The noisy voxelized representation of the sampled molecules is denoised. As described in more detail below, in some cases, the voxelized representation of 500 sample molecules... You can add noise The previous downsampling (or compression) means that the denoising engine 117 can be trained to denoise the noisy embedding of the voxelized representation of the molecule rather than the noisy voxelized representation of the molecule shown in Figure 5A.
[0196] Referring again to Figure 5A, the denoising engine 117 can be trained to handle noisy voxelized representations. Denoising is performed. In some cases, the denoising engine 117 can be trained at least by denoising the voxelized representation. The corresponding clean voxelization representation is restored in the middle. This is used to denoise the noisy voxelized representation. For example, in some cases, the denoising engine 117 can be an encoder-decoder three-dimensional convolutional neural network (CNN) trained to denoise the noisy voxelized representation. The noisy voxels in the model are mapped to their corresponding clean voxels. In doing so, the denoising engine 117 can generate an approximately clean voxelized representation. Denoising voxel representation For example, in some cases, training the denoising engine 117 may include adjusting the parameters of the denoising engine 117 to reduce (or minimize) the denoised voxel representation. and the corresponding clean voxelization representation The differences between them (e.g., mean squared error (MSE)). In some cases, the addition of voxelization to the molecular representation is determined. noise The quantity of noise levels This can be set as a hyperparameter of the denoising engine 117. Furthermore, in some cases, the noise level... This can be kept constant (or fixed) during the training of the denoising engine 117, which reduces the complexity of the training process compared to the diffusion model. It should be understood that single-step denoising (as opposed to diffusion over multiple time steps) may be sufficient to reconstruct the original voxelized representation. This is because voxelization represents Its properties differ from natural images; it contains more structural information than texture information for 500 sample molecules.
[0197] In some exemplary embodiments, the molecular design computational model 115 may apply a denoising model 117 to generate a voxelized representation of the output molecule by denoising the noisy voxelized representation of the input molecule through at least one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling). In some cases, the denoising model 117 may be derived from noisy data distributions. Sampling, which includes noisy data distributions filled with molecules exhibiting one or more desired properties (e.g., drug-like properties). The density of the data gradually increases as the distribution of noisy data is traversed. To further illustrate, Figure 5B shows the distribution across noisy data. The traversal includes sampling iterations When selecting a sample (or molecule) During sampling iteration When selecting a sample (or molecule) and in sampling iteration When selecting a sample (or molecule) In some cases, noisy data distributions... The traversal can be guided by function 175, allowing the samples to... Sampling from and Compared to noisy data distribution The higher density region of the sample Sampling from and sample Compared to noisy data distribution Even higher density regions. In some cases, each iteration of gradient-based Markov chain Monte Carlo (MCMC) may include further modifications to the sample (or molecule) selected during previous iterations. Thus, as shown below, in the sampling iteration Period selected from noisy data distribution Sample It can be based on previous sampling iterations Samples selected during the period To generate. The following equation (10) represents the distribution of noisy data. The traversal.
[0198]
[0199] (10)
[0200] in express Standard Brownian motion in, and and These are hyperparameters (friction force and inverse mass, respectively). Discretization techniques (examples of which are shown in Algorithm 1 in Table 1 below) can be applied to generate samples. The sample includes a discretization step size. .
[0201] Referring again to Figure 5B, in some cases, when the denoising engine 117 selects from the noisy data distribution... The corresponding noisy voxelization representation During denoising, a voxelized representation of the molecules can be generated. As described, noisy voxelization representation Denoising can represent noisy voxels. Projecting back to the true data distribution For example, by applying least squares estimators This is projected. This constitutes the "jump" shown in Figure 5B. Furthermore, in the example shown in Figure 5B, the "jump" returns to the true data distribution. This can be executed at each sampling iteration, while simultaneously applying denoising model 117 to traverse the noisy data distribution. And select samples from them. For example, when sampling in iterations... Period selected from noisy data distribution Sample Denoise and project it back to the real data distribution When this happens, molecules can be generated. However, when sampling in subsequent iterations Period selected from noisy data distribution Sample Denoise and project it back to the real data distribution When this happens, molecules can be generated. .
[0202] Table 1
[0203]
[0204] In some exemplary embodiments, the denoising model 117 can continue to traverse the noisy data distribution. Samples are selected from these samples until one or more criteria are met. For example, denoising model 117 can continue to traverse the energy landscape of the noisy data distribution until sampling iterations. If a threshold sampling iteration is performed at that point, the denoising model 117 can alternatively and / or additionally continue to traverse the noisy data distribution. The energy landscape until the selection of samples If the sample It shows up in the distribution of noisy data The threshold probability in the data. To further illustrate, Figure 5C shows the threshold applied to the distribution of noisy data. A denoising model 117 is selected for multiple consecutive samples (including, for example, samples 510a to 510f). In the example shown in Figure 5C, sampling (e.g., gradient-based Markov chain Monte Carlo (MCMC) sampling) can begin by applying the denoising model 117 to the molecular design computational model 115 to start from the noisy data distribution. Select the first sample In some cases, the first sample The options may include denoising model 117 updating the noisy voxelized representation of the corresponding molecule (or its noisy embedding).
[0205] As shown in Figure 5C, the first sample can be... Denoising is performed to generate the corresponding voxelized representation. This denoising operation can be used to transform noisy data distributions. Return to the real data distribution The "jump" in this process. Each subsequent sampling iteration may include applying a denoised model 117 to further update the noisy voxelized representation of the molecule selected during the previous sampling iteration. In the example shown in Figure 5C, the molecular design computational model 115 can continue to apply the denoised model 117 until... The continuous samples have been distributed from the noisy data distribution Choose from. (Number) One sample Denoising can be performed, for example, by denoising engine 117 to generate the corresponding voxel representation. Doing so will allow the first One sample Distribution of noisy data Projecting back to the true data distribution It should be understood that The value can determine the number of sampling iterations and the distribution selected from the noisy data. The number of samples. Increase The value can increase the update performed on the initial input molecule (e.g., the "seed" molecule). Higher values can increase the difference between the initial input molecule (e.g., the "seed" molecule) and the final output molecule, as well as the novelty of the final output molecule.
[0206] In some exemplary embodiments, in the distribution of noisy data Select the first One sample And for the first One sample Denoising is performed to generate the corresponding voxelized representation. At that time, the molecular design engine 110 can be based on voxel representation. To generate one or more other representations. For example, in some cases, the molecular design engine 110 can be based at least on voxelized representations. To generate a one-dimensional representation (e.g., a simplified linear input specification (SMILES) string) and / or a two-dimensional representation (e.g., a molecular graph) of the corresponding molecule.
[0207] Figure 5D depicts a voxelized representation according to some exemplary embodiments. A schematic diagram illustrating examples of processes for generating other molecular representations. In the example shown in Figure 5D, the molecular design engine 110 can at least recognize voxelized representations. The peak values (e.g., atomic density values that satisfy one or more thresholds) in the molecular design engine are used to determine the atoms present in the corresponding molecule. Furthermore, the molecular design engine 110 can determine one or more bonds that interconnect the atoms present in the molecule. A one-dimensional or two-dimensional representation of the molecule can be generated based at least on atoms and interconnecting bonds. Alternatively, in some cases, the molecular design engine 110 can apply training to voxelize the representation. A machine learning model that transforms a molecule into one or more other representations of the corresponding molecule.
[0208] In some exemplary embodiments, the molecular design computational model 115 may operate in a noisy latent voxelization space rather than in a noisy discrete voxelization space, for example, as shown in Figures 5A to 5D. For example, in some cases, a denoising model 117 may be applied to the noisy embedding of the voxelization representation. Denoising is performed instead of molecular design computational model 115. A denoising model 117 is applied to the noisy voxelized representation. Denoising is performed. To further illustrate, Figure 6 depicts a schematic diagram illustrating an example of a process according to some exemplary embodiments, where a molecular design computational model 115 generates a voxelized representation of a molecule by operating in a noisy latent voxelized space. Referring to Figure 6, an input molecule 600, which may be presented in a one-dimensional or two-dimensional representation, can be transformed into a three-dimensional representation of an input molecule 152. In some cases, the three-dimensional representation of the input molecule 152 may be a voxelized representation of the input molecule 152, which collectively represents the type and position of atoms in one or more continuously distributed atomic densities across a voxel grid. In some cases, the encoder 111 may first generate an embedding 154 of the voxelized representation of the input molecule 152, and then noise... Instead of directly operating on the noisy voxelized representation of the input molecule 152, the embedding 154 is added to the voxelized representation of the input molecule 152 using the denoising model 117. The resulting noisy embedding 156 may undergo one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.). For example, each iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling may include the molecular design computation model 115 applying the denoising model 117 to denoise the noisy embedding 156 at least by updating the noisy embedding 156. As described, updating the noisy embedding 156 in this way can be equivalent to selecting one or more samples from a noisy data distribution filled with noisy embeddings of voxelized representations of molecules exhibiting one or more desired properties. Sampling can be guided by a function 175 (e.g., a scoring function, etc.) such that successive samples are selected from regions of increasing density in a noisy data distribution, regions that are more likely to be filled with noisy embeddings represented by voxelizations of molecules exhibiting one or more desired properties.
[0209] Referring again to Figure 6, the molecular design computational model 115 can generate an updated embedding 156 by updating the voxelized representation of the input molecule 152's embedding 154 through one or more iterations, for example, gradient-based Markov chain Monte Carlo (MCMC) sampling. As shown in Figure 6, the embedding 154 can be denoised, for example, by a denoising engine 117 to generate the updated embedding 156. Denoising the embedding 154 can include sampling the updated embedding 156 from a noisy latent distribution of molecules exhibiting one or more desired properties. Furthermore, as shown in Figure 6, the decoder 119 can decode the updated embedding 156 to generate a voxelized representation of the corresponding output molecule 162. Decoding the updated embedding 156 can project the updated embedding 156 from the latent voxelized space back to the discrete voxelized space. The resulting voxelized representation of the output molecule 162 can be further transformed into the reconstructed molecule 650. It should be understood that the reconstructed molecule 650 can correspond to a one-dimensional representation of the output molecule (e.g., a simplified linear input canonical (SMILES) string) or a two-dimensional representation (e.g., a molecular graph).
[0210] In some exemplary embodiments, the generative performance of the molecular design computational model 115 can be evaluated based on various metrics, some examples of which are described in Table 2 below.
[0211] Table 2
[0212]
[0213] In some exemplary embodiments, the generative performance of the molecular design computational model 115 may depend on one or more factors, including, for example, the noise level. Number of sampling iterations The differences and atomic density radii in the voxelized molecular representation. Figure 7 depicts the differences when different levels are used. noise (For example, Gaussian noise, such as isotropic Gaussian noise) when added to the voxelized representation of a molecule operated by the molecular design computational model 115, the noise level The stability and uniqueness of molecules generated by molecular design computational model 115 (Figure 7(a)), total atomic variation and total bond variation (Figure 7(b)), and valence were evaluated. and key angle (Figure 7(c)) is a diagram illustrating the effects. As described, unlike the diffusion model, the noise level, according to the various exemplary embodiments described herein, This can be fixed during training and sampling. Furthermore, it should be understood that the noise level... It is a hyperparameter that imposes a trade-off between the quality of sampling (e.g., gradient-based Markov chain Monte Carlo (MCMC) sampling) and denoising (e.g., denoising of empirical Bayesian frameworks). In some cases, the molecular design engine 110 can determine the noise level. The noise, corresponding to the voxelized representation of the molecules to be denoised, can still be learned by adding it to the denoising engine 117. The maximum amount. For example, in some cases, the molecular design computational model 115 and the denoising engine 117 can operate at different noise levels. The dataset was trained on the QM9 molecular dataset, while other hyperparameters remained constant. Figures 7(a), 7(b), and 7(c) show that while some metrics were trained at higher noise levels... There has been some improvement, but molecular stability and valence remain. With noise level The performance deteriorates with increasing noise. For the QM9 molecular dataset, the best overall performance across all metrics is achieved at a noise level of 0.9. The following implementation.
[0214] In some exemplary embodiments, the number of sampling iterations performed as part of a gradient-based Markov chain Monte Carlo (MCMC) algorithm. This may affect the novelty of molecules generated by the molecular design computational model 115. This phenomenon is shown in Figure 8, which depicts the molecular design computational model 115 (trained on the Geometry of Molecular Ensemble (GEOM) drug dataset) undergoing different sampling iterations. The output molecules are updated multiple times, starting with noisy molecules (for de novo generation) and known molecules (for seed generation). For example, Figure 8 shows the molecular design computation model 115 undergoing... The voxelized representation of the first molecule 810 generated by denoising the noisy molecules (for de novo generation) in the next sampling iteration, and the molecular design calculation model 115 undergoes... The voxelized representation of the first molecule 820 generated by denoising the noisy molecules (for de novo generation) in the next sampling iteration, and derived from the molecular design calculation model 115 through... The voxelized representation of the third molecule 830 generated by denoising the noise molecule (for de novo generation) in the next sampling iteration, etc.
[0215] In addition to the novelty of the molecules generated by the molecular design computational model 115, the sampling iteration is adjusted. The number of iterations may also affect other aspects of the generative performance of molecular design computational model 115. Table 3 below compares the performance of molecular design computational model 115 at different sampling iteration numbers. The generation performance of the sampled model is compared with that of the conventional generative model EDM, which performs 1,000 diffusion steps. The results in Table 3 show that the generation performance increases with the number of sampling iterations. With the increase in the number of sampling iterations, the molecular design computational model 115 performs better in some metrics. As expected, the average time (in seconds) to generate each molecule decreases with the number of sampling iterations. The speed increases linearly with the increase of . However, even at 500 sampling iterations, the molecular design computational model 115 is still faster than EDM. It is worth noting that in just 50 sampling iterations, the molecular design computational model 115 has already outperformed EDM in most metrics, while being an order of magnitude faster on average.
[0216] Table 3
[0217]
[0218] In some exemplary embodiments, the generative performance of the molecular design computational model 115 may also be affected by the size of the atomic radii in the voxelized representation operated by the molecular design computational model 115. It should be understood that the size of the atomic radii can be varied, while the resolution of the voxel grid remains fixed (e.g., at .25 Å). The generative performance of the molecular design computational model 115 (even with different hyperparameters) can peak at certain atomic radii. For example, when the molecular design computational model 115 is applied to operate on voxelized representations with atomic radii of .25, .5, .75, and 1.0, the fixed radius of .5 consistently outperforms the other values, even when the hyperparameters of the molecular design computational model 115 vary.
[0219] In some exemplary embodiments, the generative performance of the molecular design computational model 115 can be compared with existing generative models that operate on conventional three-dimensional molecular representations, such as the GSchNet point cloud autoregressive model and the EDM point cloud diffusion-based model. Each model is applied to generate 10,000 samples, which are then evaluated based on atomic stability, molecular stability, efficiency, uniqueness, total atomic variation (TV), total bond variation (TV), and valence. bond length and key angle These samples were evaluated. Table 4 below shows the results for samples generated by the Molecular Design Computational Model 115 (MDCM) trained on the QM9 molecular dataset, with the mean and standard deviation of three runs. Figure 9A depicts some examples of voxelized representations of molecules generated by the Molecular Design Computational Model 115 trained on the QM9 molecular dataset, along with the corresponding molecular plots. The cumulative distribution function (CDF) of strain energy of molecules generated by the Molecular Design Computational Model 115 trained on the QM9 molecular dataset is shown in Figure 1000, depicted in Figure 10A, compared to molecules in the QM9 molecular dataset and molecules generated by the conventional generative model EDM. Figure 10B depicts Figure 1050, showing a comparison between the empirical distribution of the number of atoms in each molecule in the QM9 molecular dataset and the empirical distribution of the number of atoms in molecules generated by the Molecular Design Computational Model 115 trained on the QM9 molecular dataset.
[0220] Table 4
[0221]
[0222] In some cases, the molecular design computational model 115 was also trained on the Geometry of Molecular Organelles (GEOM) drug dataset before being applied to generate 10,000 samples. These samples are compared with 10,000 samples generated by the conventional generative model EDM, as shown in Table 5 below, with the mean and standard deviation of three separate runs. Figure 9B depicts some examples of voxelized representations of molecules generated by the molecular design computational model 115 (MDCM) trained on the Geometry of Molecular Organelles (GEOM) drug dataset, along with corresponding molecular plots. The cumulative distribution function (CDF) of the strain energy of molecules generated by the molecular design computational model 115 (MDCM) trained on the Geometry of Molecular Organelles (GEOM) drug dataset, compared to molecules in the GEOM drug dataset and molecules generated by the conventional generative model EDM, is shown in Figure 1100, depicted in Figure 11A. Figure 11B depicts a comparison of the empirical distribution of the number of atoms in each molecule in the Geometry of Molecular (GEOM) drug dataset with the empirical distribution of the number of atoms in molecules generated by the Molecular Design Computational Model 115 (MDCM) trained on the GEOM drug dataset.
[0223] Table 5
[0224]
[0225] With the Molecular Design Computational Model 115 (MDCM) trained on the QM9 dataset, it exhibits generative performance comparable to the conventional generative model EDM. However, when the MDCM is trained on the Geometry of Molecular Ensemble (GEOM) Drugs dataset (a more challenging and realistic drug-like dataset than QM9), it outperforms EDM in eight out of nine metrics, with a significant lead. For example, molecules generated by the MDCM trained on the GEOM Drugs dataset show a significantly lower median strain energy compared to molecules generated by EDM. The results in Tables 3 and 4 also demonstrate that enhancing the training dataset through rotation and translation improves the generative performance of the MDCM (e.g., MDCM). 无旋转 Compared to MDCM). Overall, the molecular design computational model 115 is a more expressive model with better data scalability. In particular, the molecular design computational model 115 is better able to capture many patterns present in large-scale data distributions, such as the Geometry of Molecular Ensemble (GEOM) drug dataset.
[0226] Figure 12A depicts a schematic diagram comparing the generation of seeds for a Geometry of Molecular Omnibus (GEOM) drug in discrete voxelization space and potential voxelization space under different sampling iterations, according to some exemplary embodiments. Figure 1210 shows molecular graphs of molecules generated at steps (or sampling iterations) of 10, 20, 50, 100, and 200, where the molecular design computational model 115 operates in the potential voxelization space and updates the voxelized representations of seed molecules from the Geometry of Molecular Omnibus (GEOM) drug dataset. The corresponding voxelized representations of these molecules are shown in Figure 1220. Figure 1215 shows molecular graphs of molecules generated at steps (or sampling iterations) of 5, 10, 50, 100, and 200, where the molecular design computational model 115 operates in the discrete voxelization space and updates the voxelized representations of seed molecules from the Geometry of Molecular Omnibus (GEOM) drug dataset. The corresponding voxelized representations of these molecules are shown in Figure 1225. As shown in Figures 12A and 12B, the molecular design computation model 115 can generate stable, effective, and unique molecules, which are also very similar to the seed molecules from the Geometry of Molecular Ensemble (GEOM) drug dataset, whether operating in the latent voxelization space or the discrete voxelization space.
[0227] Table 6 below further illustrates the seed generation results for the Geometry of Molecular Ensemble (GEOM) drug dataset (average of 5 replicates).
[0228] Table 6
[0229]
[0230] Figure 12B illustrates a schematic comparison of seed generation for PubChem drugs in discrete voxelization space and potential voxelization space under different sampling iterations, according to some exemplary embodiments. Figure 1450 shows molecular graphs of molecules generated at steps (or sampling iterations) of 10, 20, 50, 100, and 200, where the molecular design computational model 115 operates in the potential voxelization space and updates the voxelized representations of seed molecules from the PubChem dataset. The corresponding voxelized representations of these molecules are shown in Figure 1260. Figure 1255 shows molecular graphs of molecules generated at steps (or sampling iterations) of 5, 10, 50, 100, and 200, where the molecular design computational model 115 operates in the discrete voxelization space and updates the voxelized representations of seed molecules from the PubChem dataset. The corresponding voxelized representations of these molecules are shown in Figure 1265. As shown in Figure 12B, the molecular design computation model 115 can generate stable, efficient and unique molecules, which are very similar to the seed molecules from the PubChem dataset, regardless of whether it operates in the latent voxelization space or the discrete voxelization space.
[0231] Table 7 below further illustrates the seed generation results for the PubChem dataset (average after 5 repetitions).
[0232] Table 7
[0233]
[0234] Figure 12C depicts molecular graphs of additional examples of molecules generated at steps (or sampling iterations) of 10, 20, 50, 100, and 200 by manipulating and updating the voxelized representations of two real drug seed molecules in a latent voxelized space using the molecular design computational model 115. Molecular graphs of some exemplary molecules generated at randomly selected steps (or sampling iterations) by manipulating and updating the embeddings of random molecules (e.g., molecules with randomly selected atom types and / or positions) in a latent voxelized space using the molecular design computational model are shown in Figure 12D.
[0235] Table 8 below depicts the seed generation results for five real drugs (average of five replicates).
[0236] Table 8
[0237]
[0238] Figure 1300 depicts a comparison of the number of stable, effective, and unique molecules generated over time by the molecular design computational model 115 operating in latent voxelization space, by the molecular design computational model 115 operating in discrete voxelization space, and by the state-of-the-art generative model. As shown in Figure 13, the molecular design computational model 115 is able to generate a significantly greater number of stable, effective, and unique molecules compared to the state-of-the-art generative model, regardless of whether it operates in latent or discrete voxelization space. Furthermore, the molecular design computational model 115 generates a greater number of stable, effective, and unique molecules when operating in latent voxelization space compared to when operating in discrete voxelization space.
[0239] Table 9 below shows the molecular design computational model 115 in the potential voxelization space (MDCM). 潜在 Execute de novo generation and molecular design computational models in Discrete Voxelized Space (MDCM) 115 离散 The performance of de novo generation and the state-of-the-art generation model EDM on the generation of drugs in the geometry ensemble of molecular geometry (GEOM) was compared (average of 10,000 molecules generated in 3 repetitions).
[0240] Table 9
[0241]
[0242] Table 10 below shows the molecular design computational model 115 in the potential voxelization space (MDCM). 潜在 Execute de novo generation and molecular design computational models in Discrete Voxelized Space (MDCM) 115 离散 The performance of de novo generation and the state-of-the-art generation models GSchNet and EDM on QM9 drugs was compared (average of 10,000 molecules generated in 3 repetitions).
[0243] Table 10
[0244]
[0245] Figure 14 depicts a block diagram illustrating an example of a computing system 1400 according to some exemplary embodiments. Referring to Figures 1 through 14, the computing system 1400 may be used to implement a molecular design engine 110, a training engine 120, a client device 130, and / or any of its components.
[0246] As shown in Figure 14, the computing system 1400 may include a processor 1410, a memory 1420, a storage device 1430, and an input / output device 1440. The processor 1410, memory 1420, storage device 1430, and input / output device 1440 may be interconnected via a system bus 1450. The processor 1410 is capable of processing instructions for execution within the computing system 1400. Such executed instructions may implement one or more components, such as a molecular design engine 110, an analysis engine 120, a client device 130, etc. In some exemplary embodiments, the processor 1410 may be a single-threaded processor. Alternatively, the processor 1410 may be a multi-threaded processor. The processor 1410 is capable of processing instructions stored on the memory 1420 and / or storage device 1430 to display graphical information for a user interface provided via the input / output device 1440.
[0247] Memory 1420 is a computer-readable medium, such as a volatile or non-volatile computer-readable medium, that stores information within computing system 1400. For example, memory 1420 may store a data structure representing a configuration object database. Storage device 1430 provides persistent storage for computing system 1400. Storage device 1430 may be a floppy disk device, hard disk device, optical disk device, magnetic tape device, or other suitable persistent storage device. Input / output device 1440 provides input / output operations for computing system 1400. In some exemplary embodiments, input / output device 1440 includes a keyboard and / or a pointing device. In various embodiments, input / output device 1440 includes a display unit for displaying a graphical user interface.
[0248] According to some exemplary embodiments, input / output device 1440 may provide input / output operations for network devices. For example, input / output device 1440 may include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).
[0249] In some exemplary embodiments, the computing system 1400 can be used to execute various interactive computer software applications that can be used to organize, analyze, and / or store data in various formats. Alternatively, the computing system 1400 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generating, managing, and editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications may include various additional functions or may be standalone computing products and / or functions. Once activated within the application, the functions can be used to generate a user interface provided via the input / output device 1440. The user interface can be generated by the computing system 1400 and presented to the user (e.g., on a computer screen monitor, etc.).
[0250] One or more aspects or features of the subject matter described herein can be implemented as digital electronic circuits, integrated circuits, specially designed ASICs, field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These aspects or features may be implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor (which may be dedicated or general-purpose, coupled to receive and send data and instructions), a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Typically, clients and servers are remotely configured to each other and generally interact via a communication network. Relationships between clients and servers arise from computer programs running on their respective computers and the client-server relationships between them.
[0251] These computer programs may also be referred to as programs, software, software applications, applications, components, or code, including machine instructions for a programmable processor, and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer product, apparatus, and / or device (such as, for example, a disk, optical disk, memory, and programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. Machine-readable media may (such as, for example, non-transitory solid-state memory or magnetic hard disk drive or any equivalent storage medium) store such machine instructions non-transitory. Machine-readable media may (such as, for example, a processor cache or other random access memory associated with one or more physical processor cores) optionally or additionally store such machine instructions transiently.
[0252] To provide interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or a light-emitting diode (LED) monitor for displaying information to the user) and a keyboard and pointing device (such as, for example, a mouse or trackball, through which the user can provide input to the computer). Other kinds of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; input from the user can be received in any form, including sound, speech, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, speech recognition hardware and software, optical scanners, optical indicators, digital image capture devices, and associated interpretation software, etc.
[0253] In the foregoing description and claims, phrases such as “at least one” or “one or more” may appear, followed by a list of combinations of elements or features. The term “and / or” may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to mean any element or feature listed alone, or any other recounted element or feature in combination with any other recounted element or feature. For example, the phrases “at least one of A and B”; “one or more of A and B”; and “A and / or B” are each intended to mean “A alone, B alone, or A and B together”. A similar interpretation applies to lists comprising three or more items. For example, the phrases “at least one of A, B, and C”; “one or more of A, B, and C”; and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together”. The use of the term "based on" in the above and claims is intended to mean "at least partially based on," so that undescribed features or elements are also permissible.
[0254] Depending on the desired construction, the subject matter described herein can be embodied in systems, apparatuses, methods, and / or articles of manufacture. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations may be provided in addition to those features and / or variations set forth herein. For example, the above embodiments may be provided for various combinations and sub-combinations of the disclosed features and / or for combinations and sub-combinations of several further features disclosed above. Furthermore, the logical flows depicted in the drawings and / or described herein do not necessarily require the specific order or sequential order shown to achieve the desired results. Other embodiments may be within the scope of the following claims.
Claims
1. A computer-implemented method for identifying molecules having one or more desired properties, the method comprising: Generate a voxelized representation of the input molecule; A molecular design computational model is applied to update the voxelized representation of the input molecule. The molecular design computational model has been trained to approximate the data distribution of molecules exhibiting one or more desired properties by taking up a damaged voxelized representation of a sample molecule exhibiting the one or more desired properties as input and recovering the voxelized representation of the sample molecule from the damaged voxelized representation of the sample molecule. The molecular design computational model updates the voxelized representation of the input molecule to increase the likelihood that the resulting updated voxelized representation is within the data distribution. as well as The voxelized representation of the output molecule is generated based at least on the updated voxelized representation.
2. The method of claim 1, wherein the voxelization of the molecule represents a plurality of voxels organized into a three-dimensional voxel grid, and wherein each atom in the molecule represents a continuous density across one or more voxels in the three-dimensional voxel grid.
3. The method of claim 2, wherein the continuous density of each atom in the molecule is centered on the center of each atom, and wherein a first voxel farther from any atom in the molecule is associated with a lower atomic density value compared to a second voxel closer to the center of the atom in the molecule.
4. The method according to any one of claims 2 to 3, wherein each voxel in the three-dimensional voxel grid is associated with a value indicating the atomic density at the corresponding location.
5. The method according to any one of claims 1 to 4, wherein the voxelization representation of the molecule comprises one or more channels, and wherein each channel corresponds to a type of atom present in the corresponding molecule.
6. The method according to any one of claims 1 to 5, wherein the voxelization representation of the input molecule collectively represents the type and position of one or more atoms present in the corresponding molecule.
7. The method according to any one of claims 1 to 6, wherein applying the molecular design computational model to update the voxelized representation of the input molecule comprises updating the voxelized representation of the input molecule based at least on a function whose output indicates the probability of the resulting updated voxelized representation within the data distribution.
8. The method of claim 7, further comprising parameterizing the function using a plurality of parameters of the molecular design computational model.
9. The method of any one of claims 7 to 8, wherein the function comprises a scoring function, and wherein the value output by the function comprises a score indicating a local change in the density of the data distribution at the location of the updated voxelized representation.
10. The method according to any one of claims 1 to 9, wherein the molecular design computational model updates the voxelized representation of the input molecule at least in the following manner: The voxelized representation of the input molecule is updated to generate a first updated voxelized representation. The voxelized representation of the input molecule is updated to generate a second updated voxelized representation. A first value is used, parameterized by the molecular design computational model, to determine a first local change in the density of the data distribution at a first location occupied by the first updated voxelized representation. The function is applied to determine a second value indicating a second local change in the density of the data distribution at the second location occupied by the second updated voxelized representation, and When the first value and the second value indicate that the density of the data distribution at the first location is higher than the density of the data distribution at the second location, the first updated voxelized representation is further updated instead of the second updated voxelized representation.
11. The method of claim 10, wherein the molecular design computational model is applied to further update the first updated voxelized representation until one or more criteria are met.
12. The method of claim 11, wherein the one or more criteria include at least one of: (i) an iteration of updating the voxelized representation of the input molecule by a threshold amount, (ii) the first value of the first updated voxelized representation satisfies one or more thresholds, and (iii) an output molecule by a threshold amount has been generated.
13. The method of any one of claims 10 to 12, wherein the molecular design computational model is applied to further modify the first updated voxelized representation rather than the second updated voxelized representation to indicate, at least based on the first and second values, that the first updated voxelized representation has a higher probability of being within the data distribution compared to the second updated voxelized representation.
14. The method of any one of claims 10 to 13, wherein the molecular design computational model is applied to further modify the first updated voxelized representation rather than the second updated voxelized representation by indicating, at least based on the first and second values, that the first updated voxelized representation is sampled from a higher-density region of the data distribution compared to the second updated voxelized representation.
15. The method according to any one of claims 1 to 14, wherein the data distribution is a noisy data distribution filled with noisy voxelized representations of the molecules exhibiting the one or more desired properties, and wherein the voxelized representations of the output molecules are generated by denoising the first updated voxelized representation in order to map the first updated voxelized representation from the noisy data distribution to a real data distribution of the molecules exhibiting the one or more desired properties.
16. The method according to any one of claims 1 to 15, further comprising: The voxelized representation of the output molecule is transformed into different representations of the output molecule.
17. The method of claim 16, wherein the different representations of the output molecule include a one-dimensional representation of the output molecule and / or a two-dimensional representation of the output molecule.
18. The method according to any one of claims 16 to 17, wherein the voxelized representation of the output molecule is converted at least by: The positions of one or more atoms in the output molecule are determined at least by detecting one or more peaks among a plurality of atomic density values contained in the voxelized representation of the output molecule, and One or more interconnecting bonds are determined based at least on the positions of the one or more atoms.
19. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 1 to 18.
20. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 1 to 18.
21. A computer-implemented method, comprising: Identify sample molecules exhibiting one or more desired properties; Generate a noisy voxelized representation of the sample molecules; Noise is added to the noisy voxelized representation of the sample molecule to generate a damaged voxelized representation of the sample molecule; The data distribution of molecules is used to train a computational model for molecular design to approximately represent one or more of the desired properties. The training includes: applying the molecular design computational model to recover the noisy voxelized representation of the sample molecule from the damaged voxelized representation of the sample molecule; as well as Optionally, at least the voxelized representation of the output molecule is generated by applying the molecular design computational model to denoise the voxelized representation of the input molecule.
22. The method of claim 21, wherein the noisy voxelization of the sample molecule represents a plurality of voxels organized into a three-dimensional voxel grid, and wherein each atom in the sample molecule represents a continuous density across one or more voxels in the three-dimensional voxel grid.
23. The method of claim 22, wherein the continuous density of each atom in the sample molecule is centered on the center of each atom.
24. The method according to any one of claims 22 to 23, wherein each voxel in the three-dimensional voxel grid is associated with a value indicating the atomic density at the corresponding location.
25. The method according to any one of claims 22 to 24, wherein a first voxel located further away from any atom in the sample molecule is associated with a lower atomic density value compared to a second voxel located near the center of an atom in the sample molecule.
26. The method according to any one of claims 21 to 25, wherein the noisy voxelization representation of the sample molecule comprises one or more channels, and wherein each channel corresponds to a type of atom present in the sample molecule.
27. The method according to any one of claims 21 to 26, wherein the noisy voxelization representation of the sample molecule collectively represents the type and position of one or more atoms present in the sample molecule.
28. The method of any one of claims 21 to 27, wherein the training of the molecular design computational model comprises adjusting a plurality of parameters of the molecular design computational model to reduce the difference between the recovered voxelized representation of the sample molecule generated by the molecular design computational model and the noisy voxelized representation of the sample molecule.
29. The method of claim 28, wherein the plurality of parameters of the molecular design computational model are parameterized to a function, and wherein the values of the plurality of parameters are adjusted such that the function outputs values indicating local variations in the density of the data distribution of molecules exhibiting the one or more desired properties.
30. The method according to any one of claims 21 to 29, wherein the molecular design computational model denoises the voxelized representation of the input molecule by updating the atomic density of one or more voxels in at least one channel of the voxelized representation of the input molecule.
31. The method of claim 30, wherein the update of the atomic density of the one or more voxels in the at least one channel represented by the voxelization of the input molecule corresponds to an update of at least one of the types and / or positions of the one or more atoms present in the input molecule.
32. The method according to any one of claims 1 to 31, wherein the molecular design computational model undergoes multiple iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling to denoise the voxelized representation of the input molecule until one or more criteria are met.
33. The method of claim 32, wherein the one or more criteria comprise at least one of: (i) an iteration of gradient-based Markov chain Monte Carlo (MCMC) sampling of the threshold amount has been performed, (ii) the voxelization of the output molecule represents sampling from a region having a threshold density, and (iii) an output molecule of the threshold amount has been generated.
34. The method according to any one of claims 21 to 33, wherein the molecular design computational model generates the voxelized representation of the output molecule at least by: The first update is applied to the voxelized representation of the input molecule to generate a first updated voxelized representation. The second update is applied to the voxelized representation of the input molecule to generate a second updated voxelized representation, and After determining that the first updated voxelized representation is sampled from a higher density region of the data distribution compared to the second updated voxelized representation, the first updated voxelized representation is further updated.
35. The method of claim 34, wherein the data distribution is a noisy data distribution filled with noisy voxelized representations of the molecules exhibiting the one or more desired properties, and wherein the voxelized representations of the output molecules are further generated by denoising the first updated voxelized representation in order to map the first updated voxelized representations from the noisy data distribution to the true data distribution of the molecules exhibiting the one or more desired properties.
36. The method according to any one of claims 21 to 35, further comprising: The voxelized representation of the output molecule is transformed into different representations of the output molecule.
37. The method of claim 36, wherein the different representations of the output molecule include a one-dimensional representation of the output molecule and / or a two-dimensional representation of the output molecule.
38. The method according to any one of claims 21 to 37, wherein the training of the molecular design computational model comprises: The molecular design computational model with a first adjustment is applied to denoise the damaged voxelized representation of the sample molecule and generate a first restored voxelized representation of the sample molecule. A first mean squared error (MSE) is determined, which quantifies a first difference between the first recovered voxelized representation of the sample molecule and the noisy voxelized representation. The molecular design computational model with a second adjustment is applied to denoise the damaged voxelized representation of the sample molecule and generate a second restored voxelized representation of the sample molecule. A second mean square error (MSE) is determined, which quantifies a second difference between the second recovered voxelized representation of the sample molecule and the noisy voxelized representation. After determining that the first mean square error (MSE) is less than the second mean square error (MSE), the molecular design calculation model with the first adjustment instead of the second adjustment is further adjusted.
39. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 21 to 38.
40. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 21 to 38.